Photoelectric hybrid reinforcement learning control method for continuous action task
In the photoelectric hybrid reinforcement learning control method, the optical neural network is used as the main body of the policy network, combined with the actor-judge architecture and the optical coherent filtering system, the problem of difficult application of optical neural networks in dynamic environments is solved, and efficient continuous motion control is achieved, which is suitable for complex dynamic environments.
Patent Information
- Application Number
- CN202510260448.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-06
AI Technical Summary
Existing optical neural networks are rarely used in the field of classical control, especially in dynamic environments or unknown systems. Due to their dependence on a large number of labeled data, it is difficult to effectively apply to continuous action tasks.
A photoelectric hybrid reinforcement learning control method for continuous action tasks is proposed. By using the optical neural network as the main body of the policy network, combining the actor-judgment architecture for strategy learning, and using digital micromirror devices and optical coherent filtering system for parameter updates, the coordination of optical and electronic computing is achieved.
It has successfully got rid of the dependence on data labels, significantly expanded the scope of application of optical neural networks in the control field, improved the control accuracy and robustness of continuous action objects, and is suitable for real-time responses in complex dynamic environments.
Smart Images

Figure CN120106162A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of optical computing and control engineering technology, and in particular to an optoelectronic hybrid reinforcement learning control method for continuous action tasks. Background Art
[0002] Deep Neural Networks (DNN) have shown excellent performance in processing various nonlinear tasks. As the data scale and model complexity continue to increase, traditional electronic computing methods face bottlenecks in energy consumption and computing speed. Optical Neural Networks (ONN) are seen as a potential solution to break through the limitations of traditional computing because they utilize the high-speed propagation and parallel computing characteristics of light. At present, research on optical neural networks is mainly focused on image recognition, reliable communication transmission and other fields, and there is relatively little research and application in the field of classical control.
[0003] In addition, most of the current training of optical neural networks uses supervised learning methods, which requires a large amount of labeled data, and updates the network parameters by calculating the loss function and using the back-propagation algorithm. In complex control tasks, it is often challenging to obtain effective labeled data, especially in dynamic environments or unknown systems, which limits the application of optical neural networks in the control field.
[0004] Comparative document "Decision-making and control with diffractive optical networks" (Jumin Qiu, Shuyuan Xiao, Lujun Huang, et al . Adv. Photon. Nexus 3(4) 046003 (30 May 2024)) Based on spatial light modulators, diffractive optical neural networks are applied to decision-making in a variety of games, verifying the feasibility of optical neural networks in tasks of interacting with the environment. However, this method completes the training of optical neural networks through the idea of imitation learning. The network only focuses on short-term behavior and is difficult to learn long-term optimal strategies. It is not suitable for the regulation of continuous action space.
[0005] Reinforcement learning, as a training method that does not require pre-labeled data, enables intelligent agents to autonomously learn optimal strategies by interacting with the environment, providing a new approach to solving training problems in control tasks. However, combining reinforcement learning with optical neural networks still faces many challenges, such as how to efficiently update strategies in optical systems, how to cope with the control requirements of continuous action spaces, and how to achieve efficient collaboration between optics and electronic computing.
[0006] Therefore, there is an urgent need for a new method that can apply optical neural networks to continuous control tasks, give full play to the high speed and parallelism advantages of optical computing, and overcome the limitations of traditional methods in training and control accuracy. This will help promote the application of optical neural networks in fields such as end-to-end autonomous driving. Summary of the invention
[0007] In order to solve the problems existing in the background technology, the present invention provides an optoelectronic hybrid reinforcement learning control method for continuous motion tasks. The method successfully applies optical neural networks to the task of continuous motion space control, which helps to promote the application of optical neural networks in continuous motion control scenarios such as end-to-end autonomous driving.
[0008] The present invention discloses a photoelectric hybrid reinforcement learning control method for continuous motion tasks, the method comprising the following steps:
[0009] S1: Data loading: real-time collection of agent status information , and with the historical hidden sequence According to a certain relationship to form a two-dimensional matrix Loaded onto a digital micromirror device (DMD), the input light field is generated by collimating and expanding the laser beam ;
[0010] S2: Forward propagation: The input light field is globally mixed through the optical neural network, and the output light field of the network is photoelectrically converted and normalized using a charge-coupled device (CCD) to obtain a historical hidden sequence containing the current state. , then perform nonlinear mapping, full connection and other operations to obtain the continuous actions output by the strategy network , thereby controlling the agent. The agent gets rewards after interacting with the environment and the next moment status information , and combined with get ;
[0011] S3: Neural network parameter update: The value network evaluates the output of the policy network and uses the policy gradient method to update the neural network parameters. At the same time, multiple groups of random sampling are performed in the experience replay pool. , calculate the time difference target to get the loss, thus completing the parameter update of the value network;
[0012] S4: Optical neural network parameter update: Encode the policy gradient and load it on the digital micromirror device for forward propagation. According to the reciprocity of the light field, obtain the gradient light field on the conjugate surface of the optical neural network, and use the optimizer to update the optical neural network parameters.
[0013] S5: Reasoning stage: After the training is completed, the value network is discarded, and only the optical neural network and a small amount of electronic calculations are retained to complete the continuous action task control of the intelligent agent.
[0014] Preferably, a two-dimensional matrix By the same size status information Hidden sequence with history The matrix is composed of a certain ratio; the elements in the matrix are integers ranging from 0 to 255; the historical hidden sequence Initialized to all zeros during the first forward propagation; the state information When it is one-dimensional data, its encoding is mapped into a two-dimensional state matrix so that the state information is evenly distributed in the state matrix.
[0015] Preferably, the proportional relationship can be described by the following formula:
[0016]
[0017] in, Measuring current state information Hidden sequence with history The weight relationship, The value of Whether it can fully reflect the current status is related.
[0018] Preferably, the optical neural network is composed of an optical coherence filtering system, which performs global feature mixing of the input light field as a dynamic convolution kernel in the spatial domain, and the optical coherence filtering system is composed of two lenses and a phase-type spatial light modulator (SLM).
[0019] Among them, the focal lengths of the two lenses are equal, and the rear focal plane of the first lens coincides with the front focal plane of the second lens. The input light field is located at the front focal plane of the first lens, and the spectrum information of the input light field is obtained at the rear focal plane of the lens; the spatial light modulator is loaded with optical neural network parameters and is located between the two lenses. As part of the strategy network, it phase modulates the spectrum of the input light field and obtains the output light field of the optical neural network at the rear focal plane of the second lens, and uses a charge-coupled device to complete the photoelectric conversion of the output light field.
[0020] Preferably, the charge-coupled device needs to select a suitable bit depth according to the structure of the strategy network. If the strategy network is composed of only the optical neural network part, the charge-coupled device needs to be at least Mono10 to ensure high-precision data acquisition and signal quality.
[0021] Preferably, the value network evaluates the actions generated by the policy network and provides optimization feedback to guide the update of the policy network. At the same time, the value network randomly samples multiple groups of state quadruplets from the experience replay pool. , calculate the loss function, and use the double Q truncation mechanism and delayed update strategy to optimize its own neural network parameters.
[0022] Preferably, the input of the strategy network is the incident light field modulated by the digital micromirror device The light field passes through the optical neural network, layer normalization, nonlinear activation layer and fully connected layer in sequence to generate continuous actions to control the intelligent agent. The policy network calculates the loss function and updates its parameters based on the action evaluation provided by the value network to optimize the control strategy.
[0023] Among them, the fully connected layer can be implemented electronically or optically. When the implementation method is optical, a combination of digital micromirror devices and lenses is used to realize an optical fully connected layer without bias terms, thereby completing the linear transformation of the input light field and the weight matrix.
[0024] As described above, the optoelectronic hybrid reinforcement learning control method for continuous motion tasks applied by the present invention has the following beneficial effects:
[0025] a. Using optical neural networks as the main body of the policy network in reinforcement learning and performing policy learning based on the actor-judge architecture successfully gets rid of the dependence of optical neural networks on data labels, especially in dynamic environments or unknown systems, thereby significantly expanding the application scope of optical neural networks in the control field.
[0026] b. Compared with the existing methods that directly phase modulate the current state, this method introduces a historical hidden sequence into the input light field and uses a dynamic convolution kernel composed of a coherent filter system to process it in the spatial domain. This enhances the global feature mixing ability of the policy network and effectively improves the control accuracy and robustness of objects with continuous motion, thereby significantly improving the adaptability and control performance of the system in complex dynamic environments.
[0027] c. This method ultimately retains only the strategy network part, and completes the efficient control of continuous motion tasks through a hybrid architecture of optical neural networks combined with a small number of electronic computing units. This design makes full use of the parallelism and low energy consumption characteristics of optical computing, effectively reduces the computational burden of the system, and ensures real-time response capabilities to complex dynamic environments. It is suitable for continuous motion control fields such as robots and autonomous driving. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Other objects and results of the present invention will become more apparent and easily understood by referring to the following description in conjunction with the accompanying drawings, and as the present invention is more fully understood, in which:
[0029] Figure 1 This is an architecture diagram of an optoelectronic hybrid reinforcement learning control method for continuous motion tasks according to an embodiment of the present invention;
[0030] Figure 2 Schematic diagram of an optoelectronic hybrid control system for a rotating inverted pendulum according to an embodiment of the present invention; DETAILED DESCRIPTION
[0031] In order to enable those skilled in the art to better understand the technical solution and advantages of the present invention, the present application is described in detail below in conjunction with the accompanying drawings, but it is not intended to limit the protection scope of the present invention.
[0032] The prior art has the following problems:
[0033] 1. Existing research on optical neural networks is mainly focused on fields such as image recognition and communication transmission, but is less used in classical control tasks, especially in dynamic environments or unknown systems, where it is often challenging to obtain effective labeled data because a large amount of labeled data is required for supervised training. This dependence significantly limits the scope of application of optical neural networks in real-time control and continuous motion tasks.
[0034] 2. Some studies have attempted to apply optical neural networks to decision-making and control tasks, such as training networks through imitation learning. However, such methods usually only focus on short-term behavior, cannot learn long-term optimal strategies, and have difficulty adapting to the regulation needs of continuous action spaces. Therefore, in complex dynamic environments that require fine control, existing methods have poor results and lack applicability and robustness.
[0035] In view of this, the present invention provides an embodiment of an optoelectronic hybrid reinforcement learning control method for continuous motion tasks:
[0036] The rotating inverted pendulum is a nonlinear control experimental device with nonlinear, unstable and under-actuated characteristics. It consists of a rotary arm and a swing arm, and is used to verify the control algorithm in complex dynamic systems. The rotary arm is driven by a motor and can rotate around a fixed axis. One end of the swing arm is hinged to the rotary arm, and the other end is free to move. It tends to droop when there is no control. Figure 2 As shown, the system goal is to use the method of the present invention to apply an appropriate control strategy to the swing arm so that the swing arm gradually transitions from an initial unbalanced state to a vertical upward equilibrium position while maintaining stability.
[0037] The state information of the rotating inverted pendulum is collected in real time through the motor encoder and angle sensor ,in The spiral arm angles are contained in the trigonometric form and , angular velocity of the swing arm , the swing arm angle in trigonometric form and , swing arm angular velocity , and encode it into a matrix of 64x64 pixels, and the state information is evenly distributed in the state matrix, where the elements in the matrix are integers ranging from 0 to 255. Then the matrix is superimposed with the historical hidden sequence of the same size (the initial state is an all-zero matrix) according to the following relationship to obtain a two-dimensional matrix , and loaded onto a digital micromirror device (DMD), which generates an input light field by collimating and expanding the laser beam .
[0038]
[0039] in, Measuring current state information Hidden sequence with history The weight relationship, The value of Whether it can fully reflect the current state is related to this implementation case. Take 0.8.
[0040] Input light field Pass in sequence Figure 2 The lens 2, polarizer 1, spatial light modulator, polarizer 2, lens 3 shown in the figure finally obtain the historical hidden sequence including the current state on the charge coupled device Among them, lens 2, spatial light modulator, and lens 3 form an optical coherence filter system to perform global feature mixing on the input light field. Input light field Located at the front focal plane of lens 2, the input light field spectrum information is obtained at the rear focal plane of the lens. The spatial light modulator is loaded with optical neural network parameters, and the spectrum of the input light field is phase modulated as part of the strategy network. The output light field of the optical neural network is obtained at the rear focal plane of the second lens. The output light field of the network is photoelectrically converted and normalized using a charge coupled device (CCD) to obtain a historical hidden sequence containing the current state. , then perform nonlinear mapping, full connection and other operations to obtain the continuous actions output by the strategy network , and Mapped to the voltage value of the control motor, thereby controlling the rotation of the inverted pendulum. Get the state of the inverted pendulum after the action is performed at 5 millisecond intervals , combined with get , and the reward is calculated according to the following relationship , get the state quaternary , and stored in the experience replay pool.
[0041]
[0042] On the one hand, the value network randomly samples multiple groups of state quadruplets from the experience replay pool , calculate the time difference target to get the loss, and use the double Q truncation mechanism and delayed update strategy to optimize its own neural network parameters; on the other hand, by evaluating the actions generated by the policy network, it provides optimization feedback to guide the update of the policy network.
[0043] After obtaining the policy gradient using the value network, it is encoded and loaded on the digital micromirror device for forward propagation. According to the reciprocity of the light field, the gradient light field is obtained on the conjugate surface of the optical neural network, and the optimizer is used to update the optical neural network parameters.
[0044] After training, the value network is discarded and only the Figure 2 The optical neural network shown and a small amount of electronic calculations can complete the control of the rotating inverted pendulum.
[0045] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. An optoelectronic hybrid reinforcement learning control method for continuous motion tasks, characterized in that: The continuous action task refers to a task in which the control variable range of the intelligent agent is continuous; the optoelectronic hybrid reinforcement learning includes a value network implemented by an electronic neural network and a strategy network based on an optical neural network, and uses the actor-judge architecture of reinforcement learning to perform strategy learning, and finally only retains the strategy network part to realize the continuous action task control; the method includes: S1: Data loading: real-time collection of agent status information , and with the historical hidden sequence According to a certain relationship to form a two-dimensional matrix Loaded onto a digital micromirror device (DMD), the input light field is generated by collimating and expanding the laser beam ; S2: Forward propagation: The input light field is globally mixed through the optical neural network, and the output light field of the network is photoelectrically converted and normalized using a charge-coupled device (CCD) to obtain a historical hidden sequence containing the current state. , then perform nonlinear mapping, full connection and other operations to obtain the continuous actions output by the strategy network , thereby controlling the agent. The agent gets rewards after interacting with the environment and the next moment status information , and combined with get ; S3: Neural network parameter update: The value network evaluates the output of the policy network and uses the policy gradient method to update the neural network parameters. At the same time, multiple groups of random sampling are performed in the experience replay pool. , calculate the time difference target to get the loss, thus completing the parameter update of the value network; S4: Optical neural network parameter update: Encode the policy gradient and load it on the digital micromirror device for forward propagation. According to the reciprocity of the light field, obtain the gradient light field on the conjugate surface of the optical neural network, and use the optimizer to update the optical neural network parameters. S5: Reasoning stage: After the training is completed, the value network is discarded, and only the optical neural network and a small amount of electronic calculations are retained to complete the continuous action task control of the intelligent agent.
2. The optoelectronic hybrid reinforcement learning control method for continuous motion tasks according to claim 1, characterized in that: The two-dimensional matrix By the same size status information Hidden sequence with history The matrix is composed of a certain ratio; the elements in the matrix are integers ranging from 0 to 255; the historical hidden sequence Initialized to all zeros during the first forward propagation; the state information When it is one-dimensional data, its encoding is mapped into a two-dimensional state matrix so that the state information is evenly distributed in the state matrix.
3. The optoelectronic hybrid reinforcement learning control method for continuous motion tasks according to claim 2, characterized in that: The proportional relationship can be described by the following formula:
4. Among them, Measuring current state information Hidden sequence with history The weight relationship, The value of Whether it can fully reflect the current status is related.
5. The optoelectronic hybrid reinforcement learning control method for continuous motion tasks according to claim 1, characterized in that: The optical neural network is composed of an optical coherence filter system, which performs global feature mixing of the input light field as a dynamic convolution kernel in the spatial domain. The optical coherence filter system is composed of two lenses and a phase-type spatial light modulator (SLM).
6. Among them, The focal lengths of the two lenses are equal, and the rear focal plane of the first lens coincides with the front focal plane of the second lens. The input light field is located at the front focal plane of the first lens, and the spectrum information of the input light field is obtained at the rear focal plane of the lens. The spatial light modulator is loaded with optical neural network parameters and is located between the two lenses. As part of the strategy network, it phase modulates the spectrum of the input light field and obtains the output light field of the optical neural network at the rear focal plane of the second lens, and uses a charge-coupled device to perform photoelectric conversion on the output light field.
7. The optoelectronic hybrid reinforcement learning control method for continuous motion tasks according to claim 4, characterized in that: The charge-coupled device needs to select a suitable bit depth according to the structure of the strategy network. If the strategy network is only composed of the optical neural network part, the charge-coupled device needs to be at least Mono10 to ensure high-precision data acquisition and signal quality.
8. The optoelectronic hybrid reinforcement learning control method for continuous motion tasks according to claim 1, characterized in that: The value network evaluates the actions generated by the policy network and provides optimization feedback to guide the update of the policy network. At the same time, the value network randomly samples multiple sets of data quads from the experience replay pool. , calculate the loss function, and use the double Q truncation mechanism and delayed update strategy to optimize its own neural network parameters.
9. The optoelectronic hybrid reinforcement learning control method for continuous motion tasks according to claim 1, characterized in that: The input of the strategy network is the incident light field modulated by the digital micromirror device. The light field passes through the optical neural network, layer normalization, nonlinear activation layer and fully connected layer in sequence to generate continuous actions to control the intelligent agent. . The policy network calculates the loss function and updates its parameters based on the action evaluation provided by the value network to optimize the control strategy.
10. Among them, the fully connected layer can be implemented electronically or optically. When the implementation method is optical, a combination of a digital micromirror device and a lens is used to implement an optical fully connected layer without a bias term, thereby completing the linear transformation of the input light field and the weight matrix.