Path planning method and device based on photon pulse reinforcement learning network
By employing a path planning method based on photonic pulse reinforcement learning networks and utilizing a combination of explicit MZI and DFB-SA lasers, the problems of slow computation speed and high power consumption in path planning of photonic computing are solved, achieving faster computation speed and lower energy consumption.
Patent Information
- Application Number
- CN202511373626.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-25
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2045-09-25
AI Technical Summary
Existing technologies for the integration of photonic computing and pulse reinforcement learning suffer from slow computation speed, high power consumption, and difficulty in configuring weights, making it difficult to meet the requirements of path planning tasks for fast decision-making and low power consumption.
A path planning method based on photonic pulse reinforcement learning network is adopted. Linear calculation is achieved using an explicit MZI network and nonlinear calculation is achieved using a DFB-SA laser. Combining the event-driven and sparse computing characteristics of pulse neural network, multi-level calculation is completed by adjusting the phase of the phase shifter to configure the weight matrix.
It achieves high-speed parallel computing, reduces energy consumption, simplifies weight configuration, improves inference speed and energy efficiency, and meets the rapid decision-making requirements of path planning tasks.
Smart Images

Figure CN120874931A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of photon pulse reinforcement learning networks, and specifically relates to a path planning method and apparatus based on photon pulse reinforcement learning networks. Background Technology
[0002] With the continuous development of intelligent technologies, the demand for path planning in various intelligent agents is increasing. Traditional electronic computing methods are increasingly facing bottlenecks in terms of speed and power consumption, especially in task scenarios requiring real-time response, making it difficult to simultaneously meet the requirements of high-efficiency computing and low energy consumption. Photonic computing, with its high-speed processing capabilities and extremely low energy consumption, has become an ideal solution to overcome this limitation.
[0003] Spiking Neural Networks (SNNs), by mimicking the spiking mechanism of biological neurons, possess excellent temporal awareness, event-driven mechanisms, and naturally sparse computational characteristics, enabling efficient learning and decision-making processes. In path planning tasks, SNNs can be combined with reinforcement learning to optimize strategies through continuous interaction with the environment, thereby endowing agents with stronger adaptability and long-term decision-making capabilities.
[0004] Most current solutions are based on spiking reinforcement learning networks and photonic reinforcement learning networks, with few solutions specifically for photonic spiking reinforcement learning networks. Among them, Patel D. et al.'s research applied a conversion method to spiking neural networks from standard neural networks in deep reinforcement learning, achieving the conversion of a complete deep Q-network to a spiking neural network, providing a new path for robust deep reinforcement learning applications based on SNNs. Liu et al. proposed a directly trained deep spiking Q-network, integrating leaky integral ignition neurons with a deep Q-network architecture and designing a dedicated learning algorithm, verifying the efficiency of directly trained SNNs in deep reinforcement learning tasks. Yang Z's team constructed an optical neural network based on a Mach-Zendel interferometer (MZI) and combined it with the deep Q-network framework in deep learning, proposing an optical deep Q-network algorithm. This algorithm achieved performance comparable to electronic neural networks in 2D path planning tasks, providing a new method for the efficient application of MZI photonic computing in reinforcement learning. The team then proposed an adjustable bias optical neural network, which significantly improves the network's representation capabilities by increasing the optical bias. Based on this, they designed an optical deep Q network algorithm, which improves the computation speed of 2D / 3D path planning tasks by 2.5 times and 4.5 times respectively compared to the traditional DQN (Deep Q Network), providing a new solution for accelerating photonic computing for complex tasks.
[0005] Current photonic computing is primarily used to accelerate the inference process of traditional deep neural networks, and it has not been combined with spiking neural networks, thus limiting the potential of photonic computing in low-power and high-efficiency inference. Furthermore, existing optical neural networks only achieve linear computation through MZI devices, and traditional MZI devices are difficult to map matrices, and cascading methods result in severe phase shifter error accumulation. In other words, current technologies have significant shortcomings in the integration of photonic computing and spiking reinforcement learning, making it difficult to meet the demands of tasks such as path planning for rapid decision-making and low-power inference.
[0006] Therefore, how to provide a high-speed, low-power, and easily configurable path planning method based on photonic pulse reinforcement learning networks has become an important issue. Summary of the Invention
[0007] To address the aforementioned problems in the prior art, this invention provides a path planning method and apparatus based on photon pulse reinforcement learning networks.
[0008] The technical problem to be solved by this invention is achieved through the following technical solution: In a first aspect, the present invention provides a path planning method based on a photon pulse reinforcement learning network, the path planning method comprising: A multi-level spiking neural network is determined based on multiple agents and the interactive environment; The phase of each phase shifter in the explicit MZI network is adjusted according to the pulse neural network to configure the weight matrix; the explicit MZI network includes multiple MZI units; each MZI unit includes a phase shifter; The input signal is obtained by applying a pulse signal to the light source using an optical modulator; Multiple linear calculation results are obtained using the weight matrix and the input signal through interference and coupling effects; The multiple linear calculation results are input into the DFB-SA laser. By adjusting the injection current in the gain region and the reverse bias voltage in the saturation region of the DFB-SA laser, a neuron-like response is generated to the multiple linear calculation results, resulting in nonlinear activation results. Using the nonlinear activation result as a new input signal, the process is repeated to obtain multiple linear calculation results by using the weight matrix and the input signal through interference and coupling effects, until the calculation of all layers in the spiking neural network is completed, and the path planning result is obtained.
[0009] Optionally, the spiking neural network includes an input layer, a first fully connected layer, a first LIF neuron, a second fully connected layer, a second LIF neuron, a third fully connected layer, and an output layer connected in sequence.
[0010] Optionally, an input signal is obtained by applying a pulse signal to the light source using an optical modulator, including: A multi-channel light source with wavelengths spaced 0.5 nm apart is used as the input signal source; The input signal is obtained by applying a pulse signal to the input signal source using an optical modulator.
[0011] Optionally, the weight values in the weight matrix can be calculated in the following ways: ; in, It refers to the phase of the phase shifter inside the MZI cell; It is the weight value represented by the current MZI cell; Represents the imaginary unit; It is the input value of the next MZI unit after the current MZI unit.
[0012] Optionally, the multiple linear calculation results are input into the DFB-SA laser, including: The multiple linear calculation results are input into the DFB-SA laser through channels; one channel corresponds to one linear calculation result.
[0013] Optionally, a multi-level spiking neural network is determined based on multiple agents and the interaction environment, including: A multi-level initial spiking neural network is determined based on multiple agents and the interactive environment; The initial spiking neural network is trained using a deep Q-learning strategy to obtain a multi-level spiking neural network.
[0014] Secondly, the present invention provides a path planning device based on a photon pulse reinforcement learning network, the path planning device comprising: The determination module is used to determine multi-level spiking neural networks based on multiple agents and the interaction environment; A configuration module is used to adjust the phase of each phase shifter in the explicit MZI network according to the pulse neural network and configure the weight matrix; the explicit MZI network includes multiple MZI units; each MZI unit includes a phase shifter; The loading module is used to apply a pulse signal to the light source using an optical modulator to obtain the input signal. The module is used to obtain multiple linear calculation results using the weight matrix and the input signal through interference and coupling effects; The activation module is used to input the multiple linear calculation results into the DFB-SA laser, and by adjusting the injection current in the gain region and the reverse bias voltage in the saturation region of the DFB-SA laser, a neuron-like response is generated to the multiple linear calculation results to obtain nonlinear activation results; The path acquisition module is used to take the nonlinear activation result as a new input signal and return to re-execute the step of obtaining multiple linear calculation results by using the weight matrix and the input signal through interference and coupling effects until the calculation of all layers in the spiking neural network is completed, and the path planning result is obtained.
[0015] Thirdly, the present invention provides an electronic device, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; When a processor executes a computer program stored in memory, it implements the steps described in any of the above-mentioned path planning methods based on photon pulse reinforcement learning networks.
[0016] Fourthly, the present invention provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of the path planning method based on any of the above-described photon pulse reinforcement learning networks.
[0017] This invention provides a path planning method based on photonic pulse reinforcement learning networks. It utilizes an explicit MZI network for linear computation and a DFB-SA laser for nonlinear computation, achieving full-optical-domain computation. This fully leverages the high-speed parallel advantages of photonic computing, enabling the path planning task to have faster computation speed and lower energy consumption during inference.
[0018] Furthermore, by replacing the triangular or rectangular network structure used in traditional photonic computing with an explicit MZI network, the configuration process of the weight matrix in the photonic neural network is simplified, resulting in simpler weight configuration and faster response speed.
[0019] Meanwhile, by leveraging the event-driven and sparse computing characteristics of spiking neural networks, computational redundancy is reduced, and inference speed and energy efficiency are improved.
[0020] The present invention will now be described in further detail with reference to the accompanying drawings. Attached Figure Description
[0021] Figure 1 This is a flowchart illustrating a path planning method based on a photon pulse reinforcement learning network provided in an embodiment of the present invention. Figure 2This is a schematic diagram of the spiking neural network construction process provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the training framework of the spiking neural network provided in an embodiment of the present invention; Figure 4 This is a schematic diagram illustrating the growth curve of rewards as the number of game rounds increases for the agent; Figure 5 This is a schematic diagram illustrating the results of online reasoning and path planning by the intelligent agent. Figure 6 This is a schematic diagram of the structure of the pulse reinforcement learning network inference device based on explicit MZI and DFB-SA lasers provided in an embodiment of the present invention; Figure 7 This is a schematic flowchart of a simulation experiment using a path planning method based on a photon pulse reinforcement learning network provided in an embodiment of the present invention. Figure 8 This is a schematic diagram of the structure of a path planning device based on a photon pulse reinforcement learning network provided in an embodiment of the present invention; Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0022] The present invention will be further described in detail below with reference to specific embodiments, but the implementation of the present invention is not limited thereto.
[0023] To address the significant shortcomings of existing path planning methods in integrating photonic computing and pulse reinforcement learning, including slow computation speed, high power consumption, and difficulty in configuring weights, this invention provides a path planning method based on a photonic pulse reinforcement learning network. (See [link to relevant documentation]). Figure 1 , Figure 1 This is a flowchart illustrating a path planning method based on a photon pulse reinforcement learning network provided in an embodiment of the present invention, specifically including the following steps: Step S101: Determine a multi-level spiking neural network based on multiple agents and the interactive environment.
[0024] In this embodiment of the invention, the intelligent agent and the interaction environment are defined as follows: See Figure 2 , Figure 2 This is a schematic diagram of the spiking neural network construction process provided in an embodiment of the present invention. First, a two-dimensional grid path planning environment is customized based on the OpenAI Gym framework (an open-source toolkit for reinforcement learning), which is formally defined as a quadruple, as follows: ; in, The state space represents the agent's current position and offset relative to the target point, encoded as a 4-dimensional vector: This representation includes not only position coordinates. It also explicitly provides direction information to the target point. This helps to enhance the convergence of policy learning.
[0025] The action space contains four discrete actions: up ,Down ,Left ,right ,Right now: ; Each action will cause the agent to attempt to move one square in that direction; if that square is an obstacle or it crosses the boundary, it will remain in place.
[0026] Let be the state transition probability. The state transition function is constrained by the distribution of environmental obstacles. That is, although some actions are defined within the action space, the presence of obstacles will prevent the corresponding state transition and force a regression to the current position, exhibiting incompleteness and asymmetry. ; in, It is to perform an action The next state after that; Indicates the current state; Indicates the action to be performed; The reward function is calculated based on Manhattan distance: ; in, Indicates the Manhattan distance before the action; Indicates the Manhattan distance after the action; Indicates whether the collision occurred with a wall or an obstacle; Indicates whether the target point has been reached; , , and All of these represent hyperparameters used for regulation.
[0027] In this embodiment of the invention, there are two ways to set the initial state of the agent: one is to use a fixed initial position, in which case the agent is initialized at (0,0) and the target point is at (5,1); the other is to use randomized initialization, in which the agent and the target point are randomly selected in an obstacle-free area. The agent is the main body that performs the path planning task, and can specifically be a virtual vehicle or robot, etc. The target point refers to the specific location that needs to be reached in the path planning task.
[0028] The following examples illustrate the construction and initialization of the intelligent agent and the interaction environment in this embodiment of the invention: First, GPU (Graphics Processing Unit) devices are checked. PyTorch (an open-source deep learning framework for machine learning and deep learning) is used, and NVIDIA GPUs (graphics processing units) are required for computational acceleration during training. A random seed is set to ensure the reproducibility of the experiment.
[0029] Set grid size The two-dimensional grid path planning environment is as follows: ; obstacle location : ; Initial state of the agent: ; in, This represents the initial state of the agent; This represents the initial offset of the agent.
[0030] In this embodiment of the invention, to adapt to the needs of neuromorphic computing and support the deployment of neural chip inference, the spiking neural network provided in this embodiment is a three-layer spiking neural network (Spiking-DQN, Spiking Deep QNetwork). Based on the classic DQN architecture, this model introduces LIF (Leaky Integrate-and-Fire) neurons as nonlinear units, explicitly modeling the neuronal firing and time integration mechanisms.
[0031] In one implementation, the spiking neural network includes an input layer (Input), a first fully connected layer (Fully Connected1), a first LIF neuron (LIF1), a second fully connected layer (Fully Connected2), a second LIF neuron (LIF2), a third fully connected layer (Fully Connected3), and an output layer (Output) connected in sequence.
[0032] The input layer is used to input the agent's 4-dimensional state vector. ; The first fully connected layer, the second fully connected layer, and the third fully connected layer are used for state feature extraction. The first and second LIF neurons are used to accumulate input at each time step and fire a pulse when a threshold is reached.
[0033] The output layer is used to output the Q-value estimate representing each action.
[0034] In this embodiment of the invention, the layer-by-layer calculation process of the spiking neural network is as follows: The computation of the first level includes the computation of the first fully connected layer and the first LIF neuron: ; ; in, The first state characteristics of the agent; Indicates the first pulse sequence; This represents the computation of the first fully connected layer; Represents the computation of neurons; The second level of computation includes the computation of the second fully connected layer and the second LIF neuron: ; ; in, Indicates the characteristics of the second state; Indicates the second pulse sequence; This represents the computation of the second fully connected layer; The computation of the third level includes the computation of the third fully connected layer: ; in, The third-state feature represents the Q-value estimate for each action; This represents the computation of the third fully connected layer; In addition, considering that the fully connected layers need to be deployed on the MZI photonic chip, the parameters of each fully connected layer are restricted to positive numbers.
[0035] The LIF neuron provided in the embodiments of the present invention will be described below: In each hidden layer activation function, bio-inspired LIF neurons are introduced to simulate the accumulation of membrane potential and impulse firing process of biological neurons. The mathematical model is as follows: Each hidden layer is activated using LIF neurons: ; in, This represents the membrane time constant, used to control the rate of membrane potential decay; Indicates the time of LIF neurons The membrane potential; This represents the input current, indicating synaptic input from other neurons.
[0036] When there is an input current, the membrane potential rises; when there is no input, the membrane potential decays exponentially back to the resting state.
[0037] When the membrane potential exceeds the threshold At this time, the neuron generates a pulse: ; ; in, Indicates the pulse output sequence; It is a step function; After the pulse is delivered, the membrane potential is reset to a lower level: ; In the standard LIF model, It is usually set to 0 or a negative value.
[0038] To ensure greater accuracy in subsequent operations based on the spiking neural network, it is necessary to first train the spiking neural network. See [link to documentation]. Figure 3 , Figure 3 This is a schematic diagram of the training framework of the spiking neural network provided in an embodiment of the present invention.
[0039] In this embodiment of the invention, a classic deep Q-learning strategy can be adopted. An adaptive training mechanism is designed to address the temporal dynamic characteristics of spiking neural networks, thereby achieving a deep integration of neuromorphic architecture and reinforcement learning algorithms, as detailed below: First, spiking neural networks introduce a time dimension. To model the dynamic process of neuronal firing, refer to the aforementioned layer-by-layer calculation process of spiking neural networks. During the training phase, the input state needs to be expanded into a T-dimensional sequence along the time axis. The specific process is as follows: Original state Repeated expansion to The temporal input tensor is used to simulate state observations at multiple time steps; The network performs time-series computation layer by layer, with each LIF neuron in Accumulates membrane potential within a time step and generates a pulse sequence. , ; Output layer pairs The pulse signals at each time step are aggregated, and the final value is calculated using the time averaging method. The value, the formula is: ; This operation enables the network to capture the time dependence of state changes, which aligns with the dynamic information processing mechanism of biological nervous systems.
[0040] To clarify the form of the algorithm's loss function, this embodiment of the invention uses the MSE (mean squared error) loss function. The error is optimized. The MSE loss function is: ; in, It is the current Network (with parameters) (Indicates) the state Next action of Value estimate represents "the value currently perceived". The goal Value represents "the actual value it should have". It is a state Execute action The instant reward obtained afterward; It is a discount factor, with a value between 0 and 1, used to balance the weight of current and future rewards; It is to perform an action The next state after that; The goal Network All actions of Value estimation, Take the largest one The value represents the "expected future optimal value," a parameter. Regularly from the current Network replication updates are used to stabilize training.
[0041] Furthermore, to avoid temporal correlation between training samples and improve data utilization efficiency, an experience replay mechanism is employed. Let the experience pool be denoted as... The capacity is Its storage format is a 5-tuple: ; in, express The state at any given moment; express Actions at any given moment; express Momentary rewards; express The state at any given moment; express The experience of moments; Each training session from Updates are performed by randomly selecting batches, reducing time dependency.
[0042] The experience pool is set as a fixed-length FIFO queue, such as 100,000 entries. Each training session randomly samples a minibatch, such as 64 entries. Network update.
[0043] During the training phase, an ε-greedy strategy is adopted, where ε is a parameter between 0 and 1 that determines whether the agent explores or exploits when selecting actions. Exploration means that the agent randomly selects an action to discover new behaviors that may bring higher rewards; exploitation means that the agent selects the action it believes will yield the greatest reward based on its current knowledge.
[0044] See Figure 4 , Figure 4 This is a schematic diagram illustrating the growth curve of reward as the agent progresses through game rounds. The graph shows how the reward changes with the number of training rounds in a path planning task. The blue curve represents the initial reward for each round, fluctuating greatly in the early stages (instability of the strategy during the exploration phase) and gradually stabilizing later. The orange curve represents the rolling average reward over 100 rounds, rapidly increasing in the early stages (strategy optimization) and stabilizing at a higher value in the later stages (strategy maturation and convergence), reflecting the process of exploration, utilization, and gradual strategy optimization during training.
[0045] In this embodiment of the invention, after the spiking neural network is trained, it can be deployed in an online environment for real-time path planning. The inference phase no longer involves gradient calculation or parameter updates; the focus is on efficient action selection and low-latency forward propagation, while also considering chip deployment friendliness and inference accuracy.
[0046] In this embodiment of the invention, the agent adopts a greedy policy, that is, at each step, it selects the option with the maximum potential in the current state. Value actions: ; Compared to ε-greedy strategies during the training phase, greedy strategies are more suitable for deployment phases, as they can fully apply the learned strategic capabilities to actual decision-making.
[0047] In addition, to adapt to hardware, the embodiments of the present invention also employ pruning (clip weights) and weight quantization operations. Pruning can limit the weights of the fully connected layer to the range [0,1], while weight quantization can reduce storage requirements and improve computational efficiency.
[0048] See Figure 5 , Figure 5 This is a schematic diagram illustrating the results of online reasoning and path planning by the intelligent agent. Figure 5In (a) of the equation, the number of steps is 0, the single-step reward is 0.00, and the total reward is 0.00. Figure 5 In (b) of the equation, the number of steps is 4, the reward per step is 0.40, and the total reward is 1.60. Figure 5 In (c), the number of steps is 9, the single-step reward is 0.40, and the total reward is 3.60. Figure 5 In (d), the number of steps is 12, the reward per step is 0.40, and the total reward is 4.80. Figure 5 The demonstration showcased a trained agent (red section) stably executing a path planning task from "starting point (top left corner of the grid) → obstacle avoidance (black section) → destination (yellow section)" in a "grid map + static obstacles" scenario. Each step demonstrated the training results in "obstacle avoidance, path finding, and maximizing rewards." During this process, the single-step reward was 0.40, while the total reward gradually increased from 0.00 to 4.80. As the number of steps increased, the red square gradually approached the yellow square, and finally, the planning ability was verified by reaching the destination.
[0049] Step S102: Adjust the phase of each phase shifter in the explicit MZI network according to the spiking neural network and configure the weight matrix; the explicit MZI network includes multiple MZI units; each MZI unit includes a phase shifter.
[0050] Next, a spiking neural network is deployed on an explicit MZI network to utilize the path planning results output by the explicit MZI network, including: First, based on the fully connected layer weights obtained during training and limited to positive values, the phase of each phase shifter in the explicit MZI network is adjusted to achieve the mapping between the weights and the optical hardware, thereby completing the hardware deployment of the network weights in the MZI and configuring the weight matrix.
[0051] See Figure 6 , Figure 6 This is a schematic diagram of the pulse reinforcement learning network inference device based on an explicit MZI and a DFB-SA laser (two-stage semiconductor laser) provided in an embodiment of the present invention. In the explicit MZI network, each independent MZI unit represents a weight value. The weight matrix can be obtained by using the weight values corresponding to all MZI units in the explicit MZI network. The calculation method for the weight value corresponding to each MZI unit is as follows: ; in, It refers to the phase of the internal phase shifter of the MZI; It is the weight value represented by the current MZI cell; Represents the imaginary unit; It is the input value of the subsequent MZI element, that is, the MZI element following the current MZI element, which is adjusted... The size of the MZI is used to explicitly configure the weight value. Finally, the weight matrix is achieved through a cascaded explicit MZI network. Loading.
[0052] Step S103: Use an optical modulator to apply a pulse signal to the light source to obtain an input signal.
[0053] In this embodiment of the invention, an optical modulator (MZM) is used to modulate the light source. The input signal is obtained by applying a pulse signal, including: A multi-channel light source with wavelengths spaced 0.5 nm apart is used as the input signal source; The input signal is obtained by applying a pulse signal to the input signal source using an optical modulator.
[0054] in, This indicates the total number of light sources.
[0055] In this embodiment of the invention, the output wavelength of a multi-channel light source is configured as the input channel carrier. The 4-dimensional state vector of the input agent is then converted using an optical modulator. The signal is loaded onto a light source to achieve optical domain representation of the input pulse signal.
[0056] Step S104: Using the weight matrix and input signal, multiple linear calculation results are obtained through interference and coupling effects.
[0057] In this embodiment of the invention, an optical signal, i.e. an input signal, is injected into the input port of an explicit MZI network. Linear matrix calculations are performed through interference and coupling effects, and the corresponding linear calculation results are obtained at the output port.
[0058] Specifically, a light source loaded with a pulsed signal, i.e., the input signal, is injected into the input port of an explicit MZI network through an optical fiber. After the optical signal enters the explicit MZI network, the physical structure of the explicit MZI network encodes the weight matrix. Changes in light intensity are achieved through the interference effect within each MZI unit and the coupling effect of the directional coupler. The calculation yields a linear result at the output port. See Figure 6 In .
[0059] Step S105: Input multiple linear calculation results into the DFB-SA laser. By adjusting the injection current in the gain region and the reverse bias voltage in the saturation region of the DFB-SA laser, a neuron-like response is generated to the multiple linear calculation results to obtain nonlinear activation results.
[0060] In this embodiment of the invention, the purpose of step S105 is to replace the first LIF neuron in the multi-level spiking neural network with a DFB-SA device, thereby completing the function of nonlinear activation of the LIF neuron.
[0061] In one implementation, multiple linear calculation results are input into the DFB-SA laser, including: Multiple linear calculation results are input into the DFB-SA laser via channels; one channel corresponds to one linear calculation result.
[0062] Then, by precisely controlling the injection current in the gain region and the reverse bias voltage in the saturation region of the DFB-SA laser, its operating state is simulated to mimic the pulse firing characteristics of a LIF neuron, ultimately generating an output signal with a nonlinear response, i.e., the nonlinear activation result. (See [link to relevant documentation]). Figure 6 In .
[0063] Step S106: Using the nonlinear activation result as a new input signal, return to re-execute the step of obtaining multiple linear calculation results by using the weight matrix and input signal through interference and coupling effects, until the calculation of all layers in the spiking neural network is completed, and the path planning result is obtained.
[0064] Specifically, the spiking neural network provided in this embodiment of the invention is multi-layered, wherein each fully connected layer requires linear computation using an explicit MZI network, and each LIF neuron requires non-linear activation using a DFB-SA array.
[0065] Therefore, after the aforementioned steps are completed, the nonlinear activation result obtained in step S105 needs to be used as the next input signal and injected into the explicit MZI network again.
[0066] The computation process is then re-executed, utilizing the weight matrix and input signal to obtain multiple linear calculation results through interference and coupling effects. Finally, through a neural network model, cascading multiple layers of "MZI linear calculation" and "DFB-SA nonlinear activation" modules, the calculations for the second fully connected layer, the second LIF neuron, and the third fully connected layer are achieved, completing the full photon pulse deep learning network computation and obtaining the output layer results, thus completing the inference for the reinforcement learning task.
[0067] Specifically, the nonlinear activation result output in step S105 (corresponding to the processing result of the first LIF neuron) is used as a new input signal to be fed into the explicit MZI network (corresponding to the linear calculation of the second fully connected layer). Then, the linear calculation of the explicit MZI network and the nonlinear activation of DFB-SA are repeated, and so on, to gradually complete the calculation of all levels of "second fully connected layer → second LIF neuron → third fully connected layer", and finally obtain the output layer result, that is, the path planning result.
[0068] In this embodiment of the invention, an explicit MZI network is used to achieve linear computation, and a DFB-SA laser is used to achieve nonlinear computation, thereby realizing full-optical-domain computation. This fully leverages the high-speed parallel advantages of photonic computing, enabling the path planning task to have faster computation speed and lower energy consumption during the inference process.
[0069] Furthermore, by replacing the triangular or rectangular network structure used in traditional photonic computing with an explicit MZI network, the configuration process of the weight matrix in the photonic neural network is simplified, resulting in simpler weight configuration and faster response speed.
[0070] Meanwhile, by leveraging the event-driven and sparse computing characteristics of spiking neural networks, computational redundancy is reduced, and inference speed and energy efficiency are improved.
[0071] The simulation experiment of a path planning method based on a photonic pulse reinforcement learning network provided by the embodiments of the present invention is as follows: See Figure 7 , Figure 7 This is a schematic flowchart of a simulation experiment using a path planning method based on a photonic pulse reinforcement learning network provided in this embodiment of the invention. A tunable light source generates a multi-channel light source with wavelengths spaced at 0.5 nm intervals as the input signal source. A polarization controller adjusts the polarization state of the light source entering the electro-optic modulator. An electro-optic modulator array applies pulse signals to the light source to obtain the input signal. The host computer generates an input signal based on the initial state, and the signal is generated by an arbitrary waveform generator and then amplified by an electric amplifier.
[0072] A light source loaded with a pulse signal is injected into each input port of an explicit MZI network after passing through a polarization controller to complete linear calculations. The results of the linear calculations are then obtained at the output port. The linear calculation results are introduced into the corresponding input channels of the DFB-SA array via fiber coupler arrays and optical circulator arrays, so that each DFB-SA channel processes one linear calculation result. Simultaneously, an optical power meter is used to detect the magnitude of the injected power of the DFB-SA.
[0073] By maintaining a constant temperature for the DFB-SA laser through temperature control, and then adjusting the current in the gain region and the reverse bias voltage in the saturation region of the DFB-SA laser through current and voltage sources, the DFB-SA laser is set to a neuron-like state, enabling it to generate a neuron-like response to the input signal, thus obtaining a nonlinear activation result.
[0074] The nonlinear activation result obtained from the DFB-SA array is then passed through an optical circulator array and a fiber coupler array. One end of the output light is fed into a spectrometer to observe the spectrum of the DFB-SA, while the other end enters a photodetector. The data is then collected back to the host computer via an oscilloscope, which generates the next action based on the current action. Fully connected layers 1 and 1 (LIF1 and 2) and Fully connected layers 2 and 2 (LIF2) are sequentially implemented on a single-layer MZI+DFB-SA network through cyclic multiplexing. Finally, the MZI network is used to complete the Fully connected layer 3 for network inference.
[0075] Based on the same inventive concept, embodiments of the present invention also provide a path planning device based on a photonic pulse reinforcement learning network, see [link to relevant documentation]. Figure 8 , Figure 8 This is a schematic diagram of a path planning device based on a photon pulse reinforcement learning network provided in an embodiment of the present invention. The path planning device includes: The determination module 801 is used to determine a multi-level spiking neural network based on multiple agents and the interaction environment; The configuration module 802 is used to adjust the phase of each phase shifter in the explicit MZI network according to the spiking neural network and configure the weight matrix; the explicit MZI network includes multiple MZI units; each MZI unit includes a phase shifter; The loading module 803 is used to load a pulse signal onto the light source using a light modulator to obtain an input signal; Module 804 is used to obtain multiple linear calculation results by utilizing the weight matrix and input signal through interference and coupling effects; The activation module 805 is used to input multiple linear calculation results into the DFB-SA laser. By adjusting the injection current in the gain region and the reverse bias voltage in the saturation region of the DFB-SA laser, a neuron-like response is generated to the multiple linear calculation results to obtain nonlinear activation results. The path acquisition module 806 is used to take the nonlinear activation result as a new input signal and return to re-execute the steps of obtaining multiple linear calculation results by using the weight matrix and the input signal through interference and coupling effects until the calculation of all layers in the spiking neural network is completed, and the path planning result is obtained.
[0076] In this embodiment of the invention, an explicit MZI network is used to achieve linear computation, and a DFB-SA laser is used to achieve nonlinear computation, thereby realizing full-optical-domain computation. This fully leverages the high-speed parallel advantages of photonic computing, enabling the path planning task to have faster computation speed and lower energy consumption during the inference process.
[0077] Furthermore, by replacing the triangular or rectangular network structure used in traditional photonic computing with an explicit MZI network, the configuration process of the weight matrix in the photonic neural network is simplified, resulting in simpler weight configuration and faster response speed.
[0078] Meanwhile, by leveraging the event-driven and sparse computing characteristics of spiking neural networks, computational redundancy is reduced, and inference speed and energy efficiency are improved.
[0079] Optionally, the spiking neural network includes an input layer, a first fully connected layer, a first LIF neuron, a second fully connected layer, a second LIF neuron, a third fully connected layer, and an output layer connected in sequence.
[0080] Optionally, the loading module 803 is specifically used to use a multi-channel light source with a wavelength interval of 0.5nm as the input signal source; and to use an optical modulator to load a pulse signal onto the input signal source to obtain the input signal.
[0081] Optionally, the weight values in the weight matrix can be calculated in the following ways: ; in, It refers to the phase of the phase shifter inside the MZI cell; It is the weight value represented by the current MZI cell; Represents the imaginary unit; It is the input value of the next MZI cell after the current MZI cell.
[0082] Optionally, the activation module 805 inputs multiple linear calculation results into the DFB-SA laser, including: Multiple linear calculation results are input into the DFB-SA laser via channels; one channel corresponds to one linear calculation result.
[0083] Optionally, the determining module 801 is specifically used to determine a multi-level initial spiking neural network based on multiple agents and the interactive environment; and to train the initial spiking neural network using a deep Q-learning strategy to obtain a multi-level spiking neural network.
[0084] This invention also provides an electronic device, such as... Figure 9As shown, it includes a processor 901, a communication interface 902, a memory 903, and a communication bus 904, wherein the processor 901, the communication interface 902, and the memory 903 communicate with each other through the communication bus 904. Memory 903 is used to store computer programs; When the processor 901 executes the program stored in the memory 903, it implements the method steps of any of the above-mentioned path planning methods based on photon pulse reinforcement learning networks.
[0085] The communication bus mentioned in the above electronic devices can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of representation, only one thick line is used in the diagram, but this does not indicate that there is only one bus or one type of bus.
[0086] The communication interface is used for communication between the aforementioned electronic devices and other devices.
[0087] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.
[0088] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0089] The present invention also provides a computer-readable storage medium. A computer program is stored in the computer-readable storage medium, and when executed by a processor, the computer program implements the method steps of any of the above-described path planning methods based on photon pulse reinforcement learning networks.
[0090] Optionally, the computer-readable storage medium may be non-volatile memory (NVM), such as at least one disk storage device.
[0091] Optionally, the aforementioned computer-readable storage medium may also be at least one storage device located remotely from the aforementioned processor.
[0092] In another embodiment of the present invention, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute the steps of the method described in any of the above-described path planning methods based on photon pulse reinforcement learning networks.
[0093] It should be noted that the terms "first," "second," etc., are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the invention.
[0094] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Furthermore, those skilled in the art can combine and integrate the different embodiments or examples described in this specification.
[0095] Although the invention has been described herein in conjunction with various embodiments, those skilled in the art will understand and implement other variations of the disclosed embodiments by reviewing the accompanying drawings and the disclosure in carrying out the claimed invention. In the description of the invention, the word "comprising" does not exclude other components or steps, "a" or "an" does not exclude a plurality, and "a plurality" means two or more, unless otherwise explicitly specified. Furthermore, while different embodiments may describe certain measures, this does not mean that these measures cannot be combined to produce good results.
[0096] The method provided in this invention can be applied to electronic devices. Specifically, the electronic device can be a desktop computer, a portable computer, a smart mobile terminal, a server, etc. No limitation is made herein; any electronic device that can implement this invention falls within the protection scope of this invention.
[0097] For the embodiments of the device / electronic device / storage medium, since they are basically similar to the method embodiments, the description is relatively simple, and relevant parts can be referred to in the description of the method embodiments.
[0098] It should be noted that the device, electronic device, and storage medium in the embodiments of the present invention are respectively devices, electronic devices, and storage media that apply the above-mentioned path planning method based on photon pulse reinforcement learning network. Therefore, all embodiments of the above-mentioned path planning method based on photon pulse reinforcement learning network are applicable to the device, electronic device, and storage medium, and can achieve the same or similar beneficial effects.
[0099] The above description, in conjunction with specific preferred embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art, various simple deductions or substitutions can be made without departing from the concept of the present invention, and all such modifications and substitutions should be considered within the scope of protection of the present invention.
Claims
1. A path planning method based on photon pulse reinforcement learning networks, characterized in that, The path planning method includes: A multi-level spiking neural network is determined based on multiple agents and the interactive environment; The phase of each phase shifter in the explicit MZI network is adjusted according to the pulse neural network to configure the weight matrix; the explicit MZI network includes multiple MZI units; each MZI unit includes a phase shifter; The input signal is obtained by applying a pulse signal to the light source using an optical modulator; Multiple linear calculation results are obtained using the weight matrix and the input signal through interference and coupling effects; The multiple linear calculation results are input into the DFB-SA laser. By adjusting the injection current in the gain region and the reverse bias voltage in the saturation region of the DFB-SA laser, a neuron-like response is generated to the multiple linear calculation results, resulting in nonlinear activation results. Using the nonlinear activation result as a new input signal, the process is repeated to obtain multiple linear calculation results by using the weight matrix and the input signal through interference and coupling effects, until the calculation of all layers in the spiking neural network is completed, and the path planning result is obtained.
2. The path planning method according to claim 1, characterized in that, The spiking neural network comprises an input layer, a first fully connected layer, a first LIF neuron, a second fully connected layer, a second LIF neuron, a third fully connected layer, and an output layer connected in sequence.
3. The path planning method according to claim 1, characterized in that, The input signal is obtained by applying a pulse signal to the light source using an optical modulator, including: A multi-channel light source with wavelengths spaced 0.5 nm apart is used as the input signal source; The input signal is obtained by applying a pulse signal to the input signal source using an optical modulator.
4. The path planning method according to claim 1, characterized in that, The weight values in the weight matrix are calculated in the following ways: ; in, It refers to the phase of the phase shifter inside the MZI cell; It is the weight value represented by the current MZI cell; Represents the imaginary unit; It is the input value of the next MZI unit after the current MZI unit.
5. The path planning method according to claim 1, characterized in that, The multiple linear calculation results are input into the DFB-SA laser, including: The multiple linear calculation results are input into the DFB-SA laser through channels; one channel corresponds to one linear calculation result.
6. The path planning method according to claim 1, characterized in that, A multi-level spiking neural network is determined based on multiple agents and the interactive environment, including: A multi-level initial spiking neural network is determined based on multiple agents and the interactive environment; The initial spiking neural network is trained using a deep Q-learning strategy to obtain a multi-level spiking neural network.
7. A path planning device based on a photon pulse reinforcement learning network, characterized in that, The path planning device includes: The determination module is used to determine multi-level spiking neural networks based on multiple agents and the interaction environment; A configuration module is used to adjust the phase of each phase shifter in the explicit MZI network according to the pulse neural network and configure the weight matrix; the explicit MZI network includes multiple MZI units; each MZI unit includes a phase shifter; The loading module is used to apply a pulse signal to the light source using an optical modulator to obtain the input signal. The module is used to obtain multiple linear calculation results using the weight matrix and the input signal through interference and coupling effects; The activation module is used to input the multiple linear calculation results into the DFB-SA laser, and by adjusting the injection current in the gain region and the reverse bias voltage in the saturation region of the DFB-SA laser, a neuron-like response is generated to the multiple linear calculation results to obtain nonlinear activation results; The path acquisition module is used to take the nonlinear activation result as a new input signal and return to re-execute the step of obtaining multiple linear calculation results by using the weight matrix and the input signal through interference and coupling effects until the calculation of all layers in the spiking neural network is completed, and the path planning result is obtained.
8. The path planning device according to claim 7, characterized in that, The spiking neural network comprises an input layer, a first fully connected layer, a first LIF neuron, a second fully connected layer, a second LIF neuron, a third fully connected layer, and an output layer connected in sequence.
9. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, when executing a computer program stored in memory, implements the path planning method based on a photon pulse reinforcement learning network as described in any one of claims 1-6.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the path planning method based on a photon pulse reinforcement learning network as described in any one of claims 1-6.
Citation Information
Patent Citations
Optical pulse neural network implementation device based on MZI array and FP-SA
CN115496195A
Method for simultaneously realizing pulse activation and synaptic weight through monolithic integration DFB-SA
CN116451759A
Hybrid photon neural network device and classification method
CN119647539A
Cited By
Photon neuromorphic autonomous navigation method and system
CN121655519A
A photonic neuromorphic autonomous navigation method and system
CN121655519B
Multi-modal physical neural network system and method based on gradient neurons
CN121809564A