Beamline station parameter optimization method based on extended Kalman filter and reinforcement learning
By combining extended Kalman filtering and reinforcement learning in beamline station parameter optimization, a state transfer model is established and precise state estimation is carried out, the state estimation deviation problem caused by device error is solved, the accuracy and stability of strategy learning is improved, and more efficient beamline station parameter optimization is achieved.
Patent Information
- Application Number
- CN202510053213.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-14
AI Technical Summary
During the optimization of beamline station parameters, equipment errors lead to deviations between the state estimate and the actual equipment status, affecting the learning effect of reinforcement learning strategies. Especially in sparse reward scenarios, the accumulation of errors makes strategy optimization more difficult.
Using an extended Kalman filtering (EKF) and reinforcement learning method, a state transfer model is established by training a probability neural network, and the state is accurately estimated in combination with EKF to reduce the impact of device error on state estimation. Use the DDPG algorithm to optimize the current strategy and gradually obtain new strategies through multiple rounds of iteration.
It significantly alleviates the impact of device noise on state estimation, improves the accuracy and stability of strategy learning, ensures that reinforcement learning can still learn more accurate strategies under device noise and error interference, and improves experimental accuracy and efficiency.
Smart Images

Figure CN119474620B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of reinforcement learning, and in particular to a beamline station parameter optimization method based on extended Kalman filtering and reinforcement learning. Background Art
[0002] The beamline station is a key component of the synchrotron radiation source device, responsible for guiding the synchrotron radiation light generated by the high-speed electrons in the storage ring to a specific experimental station. Synchrotron radiation source is an extremely important scientific research tool that can provide high-brightness, high-resolution light sources and is widely used in many fields such as materials science, life science, chemistry and physics. In the beamline station, the light beam is regulated by a series of optical elements, such as focusing, monochromatization and collimation, to meet the needs of different experiments. Each beamline station is usually designed for a specific experimental technique or research field, such as X-ray absorption spectroscopy, X-ray diffraction, photoelectron spectroscopy, etc. In implementation, the beamline station parameters need to be optimized to achieve high-precision regulation of the optical elements in the beamline station, so as to ensure that the beam characteristics can meet the experimental requirements. Therefore, how to optimize the beamline station parameters to improve the experimental accuracy and efficiency is one of the current research focuses.
[0003] At present, optimization methods such as reinforcement learning, Bayesian optimization algorithm and particle swarm algorithm have been widely used in the parameter optimization of beamline stations. The above methods can optimize the parameter combination of beamline stations through continuous trials and adjustments to achieve the best experimental results. However, in actual operation, there are errors in the equipment, which will cause deviations between the state estimation and the actual state of the actual equipment, making it more difficult for reinforcement learning strategies to learn effective strategies in sparse reward scenarios. Especially in areas close to physical boundaries, due to the increase in the error of state estimation, reinforcement learning is difficult to obtain high-precision strategies. This deviation is gradually amplified during the strategy optimization process. Specifically, reinforcement learning relies on accurate perception of the device state to evaluate the correspondence between actions and rewards. When the state estimation is inaccurate, the strategy may tend to optimize the wrong state. This deviation affects the update direction of the strategy, making the actual learned strategy unable to accurately reflect the actual situation of the device, which can easily lead to inaccurate strategy learning. Summary of the invention
[0004] This application provides a beamline station parameter optimization method based on extended Kalman filtering and reinforcement learning, which uses a combination of Kalman filtering and reinforcement learning to alleviate the impact of system errors in the process of equipment parameter tuning and improve the accuracy of state estimation, thereby making the strategy learning more accurate. This application provides the following technical solutions:
[0005] In a first aspect, the present application provides a beamline station parameter optimization method based on extended Kalman filtering and reinforcement learning, the method comprising:
[0006] Based on the initial strategy and the preset target state, several initial states are randomly selected from the environment and sampled to collect multiple trajectory data consisting of continuous experience quadruplets;
[0007] In the first round of sampling, the probabilistic neural network is trained using the collected trajectory data to obtain the state transition model;
[0008] For each trajectory data, an extended Kalman filter is performed in combination with the state transfer model, and the filtered next moment state is used to replace the experience quadruple of each trajectory data and the new experience quadruple is saved in the experience playback pool;
[0009] The DDPG algorithm is used to randomly sample experience quadruplets from the experience replay pool and learn and update the current strategy to obtain a new strategy, and this cycle is repeated until the strategy learning is completed.
[0010] In a specific implementation scheme, the use of the collected trajectory data to train a probabilistic neural network to obtain a state transition model includes:
[0011] After the first round of sampling is completed and multiple trajectory data are collected, the preset probabilistic neural network is trained using the collected multiple trajectory data to obtain a state transition model. The state transition model is as follows:
[0012] ;
[0013] in, and Represents the current state and current action respectively. are the model parameters of the probabilistic neural network, is the mean vector, is the variance vector.
[0014] In a specific implementation scheme, for each trajectory data, the extended Kalman filter is performed in combination with the state transfer model, and the filtered next moment state is used to replace the experience quadruple of each trajectory data and the new experience quadruple is saved in the experience replay pool, including:
[0015] The state at the current moment is predicted through the state transition model and the error covariance matrix of the state is calculated;
[0016] Observation values and Kalman gain are introduced to correct the predicted state value and error covariance matrix;
[0017] After the correction is completed, the error covariance matrix is updated, and the corrected predicted state value is replaced into the original experience quadruple as the state value at the next moment to form a new experience quadruple.
[0018] In a specific implementation scheme, predicting the state at the current moment through the state transition model and calculating the error covariance matrix of the state includes:
[0019] Use the state transition model to predict the current moment The status is as follows:
[0020] ;
[0021] For the current moment The prediction formula of the error covariance is as follows:
[0022] ;
[0023] in, It is a state transition model about the state The Jacobian matrix at the current time, is the covariance matrix, is the error covariance matrix of the last update phase, is the covariance matrix of the current prediction stage, is the process noise covariance matrix, Represents a transpose operation.
[0024] In a specific implementation scheme, the predicting the state at the current moment by the state transition model and calculating the error covariance matrix of the state also includes:
[0025] Using N consecutive experience quadruplets in the current trajectory And the state transfer model calculates the error of each empirical quadruple as follows:
[0026] ;
[0027] ;
[0028] in, , calculate the average error of all empirical quadruplets as follows:
[0029] ;
[0030] Use the following formula to calculate :
[0031] ;
[0032] in, represents the transpose operation, Represents the trajectory data Four-tuple samples, and Represent the previous moment and the current moment in the experience quadruple respectively.
[0033] In a specific implementation scheme, the introduction of observation values and Kalman gains to correct the predicted state values and the error covariance matrix includes:
[0034] Calculate the difference between observed and predicted values as follows:
[0035] ;
[0036] The Kalman gain is calculated using the following formula :
[0037] ;
[0038] in, It is a diagonal matrix calculated and constructed by using the continuous trajectory data in the experience replay pool using dynamic estimation. It uses N continuous experience samples sampled in the current trajectory to take , assuming there is Dimensional components, calculate the standard deviation of each component , The calculation formula is as follows:
[0039] ;
[0040] The correction formula for the predicted state value is as follows:
[0041] ;
[0042] To predict the corrected state value; the correction formula of the error covariance matrix is as follows:
[0043] ;
[0044] is the corrected error covariance matrix.
[0045] In a specific implementation scheme, the predicting the state at the current moment by the state transition model and calculating the error covariance matrix of the state also includes:
[0046] The error covariance needs to be initialized urgently, using the trajectory data of N consecutive empirical quadruplets in the current trajectory , assuming have The dimensions of , calculate N The standard deviation of each component , the initial value of the error covariance is calculated as follows:
[0047] ;
[0048] ;
[0049] in, represents the mean of the first component, Indicates Track data status The value of the first component of , calculate the covariance matrix of the observed value and the predicted value of each empirical quadruple respectively, and take the average value as the initial value of the error covariance in the prediction stage .
[0050] In the second aspect, the present application provides a beamline station parameter optimization system based on extended Kalman filtering and reinforcement learning, which adopts the following technical solutions:
[0051] A beamline station parameter optimization system based on extended Kalman filtering and reinforcement learning, comprising:
[0052] The trajectory data collection module is used to randomly select several initial states from the environment and perform sampling based on the initial strategy and the preset target state, and collect multiple trajectory data consisting of continuous experience quadruplets;
[0053] A state transition model generation module is used to train a probabilistic neural network using the collected trajectory data to obtain a state transition model in the first round of sampling;
[0054] An extended Kalman filter module is used to perform an extended Kalman filter on each trajectory data in combination with the state transfer model, replace the filtered next moment state into the experience quadruple of each trajectory data and save the new experience quadruple into the experience playback pool;
[0055] The strategy learning module is used to use the DDPG algorithm to randomly sample experience quadruplets from the experience replay pool and learn and update the current strategy to obtain a new strategy, and repeat this cycle until the strategy learning is completed.
[0056] In a third aspect, the present application provides an electronic device comprising a processor and a memory; a program is stored in the memory, and the program is loaded and executed by the processor to implement a beamline station parameter optimization method based on extended Kalman filtering and reinforcement learning as described in the first aspect.
[0057] In a fourth aspect, the present application provides a computer-readable storage medium, wherein the storage medium stores a program, and when the program is executed by a processor, the program is used to implement a beamline station parameter optimization method based on extended Kalman filtering and reinforcement learning as described in the first aspect.
[0058] In summary, the beneficial effects of this application include at least:
[0059] 1) During the optimization of beamline station parameters, equipment errors will cause deviations between state estimates and actual equipment states. This error will affect the learning effect of the reinforcement learning strategy, especially in a sparse reward environment. The accumulation of errors will make strategy optimization more difficult. This application introduces an extended Kalman filter (EKF) to predict and update the state, effectively reducing the impact of equipment noise on state estimation. When predicting the state, the extended Kalman filter can correct the previous estimate based on the current observation information, making the state estimate more accurate. Through this process, the error is significantly alleviated, ensuring that reinforcement learning can still learn more accurate strategies under the interference of equipment noise and errors, further improving the effectiveness and stability of reinforcement learning in practical applications.
[0060] 2) In reinforcement learning, the experience replay pool stores multiple historical experience quadruplets, which are usually used for subsequent policy updates. However, due to inaccurate state estimation in traditional methods, the data in the replay pool may contain noise, resulting in limited contribution to policy updates. This application combines extended Kalman filtering to accurately estimate and correct the state in each trajectory data, reducing the impact of noise. The corrected state estimation enhances the validity of the data, making the data sampled from the experience replay pool more valuable for policy optimization. Accurate state estimation accelerates policy convergence in the reinforcement learning process, improves data utilization efficiency, shortens learning time, and improves overall optimization performance.
[0061] 3) Reinforcement learning algorithms usually need to show good adaptability and stability in dynamic and complex environments. However, the inevitable noise and uncertainty in the environment may cause the algorithm to perform unstable. By introducing the extended Kalman filter, not only the accuracy of state estimation is improved, but also the algorithm is more adaptable to equipment errors and external disturbances. In a complex physical environment, the algorithm can reduce the impact of noise interference on the optimization process by continuously updating the predicted state and correcting its error. This enables this method to maintain stable performance in an environment with high uncertainty, improves the robustness of the algorithm, and can run efficiently in a wider range of application scenarios, ensuring the robustness and reliability of the optimization process.
[0062] By training a probabilistic neural network to establish a state transition model and combining it with an extended Kalman filter to accurately estimate the state, the impact of device errors on state estimation can be reduced, and the accuracy and stability of policy learning can be improved. Finally, the DDPG algorithm is used to optimize the current policy, and a new policy is gradually obtained through multiple rounds of iterations. The combination of Kalman filtering and reinforcement learning is used to alleviate the impact of system errors in the process of device parameter tuning, improve the accuracy of state estimation, and make policy learning more accurate.
[0063] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application and implement it according to the contents of the specification, the following is a detailed description of the preferred embodiments of the present application in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 It is a flow chart of a beamline station parameter optimization method based on extended Kalman filtering and reinforcement learning in an embodiment of the present application.
[0065] Figure 2 It is a schematic diagram of the overall process of the beamline station parameter optimization method based on extended Kalman filtering and reinforcement learning in an embodiment of the present application.
[0066] Figure 3 It is a structural block diagram of a beamline station parameter optimization system based on extended Kalman filtering and reinforcement learning in an embodiment of the present application.
[0067] Figure 4 It is a block diagram of an electronic device for beamline station parameter optimization based on extended Kalman filtering and reinforcement learning in an embodiment of the present application. DETAILED DESCRIPTION
[0068] The specific implementation methods of the present application are further described in detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present application but are not intended to limit the scope of the present application.
[0069] Optionally, the present application takes the beamline station parameter optimization method based on extended Kalman filtering and reinforcement learning provided in each embodiment as an example of being used in an electronic device, where the electronic device is a terminal or a server, and the terminal can be a computer, a tablet computer, etc. This embodiment does not limit the type of electronic device.
[0070] Reference Figure 1 , is a flow chart of a beamline station parameter optimization method based on extended Kalman filtering and reinforcement learning provided by an embodiment of the present application, the method comprising at least the following steps:
[0071] Step S101: Based on the initial strategy and the preset target state, several initial states are randomly selected from the environment and sampled to collect multiple trajectory data consisting of continuous experience quadruplets.
[0072] Specifically, first based on the preset target state Random Selection Initial state , Then based on the initial strategy Sampling in the environment, collecting Each trajectory consists of multiple consecutive and ordered empirical quadruplets. composition, and Represent the previous moment and the current moment in the experience quadruple, represent The state of the moment, represent The action of the moment, represent The rewards of the moment, represent The state of the moment.
[0073] The environment refers to the environment where the beamline station system is located.
[0074] Step S102: In the first round of sampling, the collected trajectory data is used to train a probabilistic neural network to obtain a state transition model.
[0075] In the implementation, after the first round of sampling is completed and multiple trajectory data are collected, the preset probabilistic neural network is trained using the collected multiple trajectory data to obtain a state transition model. The state transition model is as follows:
[0076] ;
[0077] in, are the model parameters of the probabilistic neural network, and Representing the current state and current action respectively, the state transition model has two outputs, namely, is the mean vector, which is also the expected state vector at the next moment. is the variance vector, which represents the randomness of existence. The input of the probabilistic neural network is the current state and actions , in implementation, the goal of the probabilistic neural network is to learn the current state and actions With the next state The probability distribution relationship between them, the model requires the next state It obeys Gaussian distribution in all dimensions and outputs the mean vector , the variance vector , , in a Gaussian distribution, the mean vector It represents the center position of the distribution and is the most likely value in the distribution. Therefore, after the model training is completed, the mean vector output by the final state transition model is used as the predicted value of the next state of the system. The variance vector represents uncertainty and is used to reflect the credibility of the prediction result. For example, when the variance is small, it means that the model has a high confidence in the predicted value.
[0078] It should be noted that the probabilistic neural network is an existing model pre-designed according to a specific reinforcement learning task, and its goal is to predict the next state and its uncertainty by inputting the current state and action. The training process is completed based on the trajectory data collected in the first round of sampling, which contains the state of the device, the action taken, and the result state. During training, the trajectory data is divided into input and target output using the standard method of the neural network, and the network parameters are adjusted by optimizing the loss function so that the model can accurately predict the state transition relationship. This process ensures that the probabilistic neural network can efficiently represent the complex relationship between the state and action of the device, and provide reliable state prediction for subsequent steps. In implementation, due to the high cost of training the probabilistic neural network, it is impossible to repeatedly train the state transition model in each round of policy update. At the same time, since the initial strategy is randomly generated, its exploration ability is strong and it will fully traverse each state and action. Therefore, this application only trains the state transition model by increasing the amount of sampled data after the first round of sampling, and the parameters of the state transition model applied after each subsequent round of sampling are fixed. Although the subsequent strategy may be more convergent, the state transition model is sufficient to describe the entire optimization process and does not need to be updated repeatedly.
[0079] In summary, by training the probabilistic neural network, step S102 successfully establishes the state transition model of the system. The mean vector is used as the predicted value of the next state to simplify subsequent calculations while maintaining the description of the uncertainty of state transition. After fixing the model parameters, the subsequent steps can more efficiently use the model for policy optimization, thereby reducing the overall computational cost.
[0080] Step S103: For each trajectory data, an extended Kalman filter is performed in combination with a state transfer model, and the filtered next moment state is used to replace the experience quadruple of each trajectory data and the new experience quadruple is saved in the experience playback pool.
[0081] In step S103, the state at the current moment is predicted by the state transition model and the error covariance matrix of the state is calculated. , where the error covariance matrix is used to predict the uncertainty of the state. Then the observation value and Kalman gain are introduced to correct the predicted state value and the error covariance matrix to obtain a more accurate predicted state value. After the correction is completed, the error covariance matrix is updated, and the corrected predicted state value is replaced into the original experience quadruple as the state value at the next moment to form a new experience quadruple, thereby ensuring that the state estimation of each trajectory data is more accurate and improving the efficiency and quality of subsequent training.
[0082] Specifically, first use the state transition model trained in step S102 to predict the current moment The status is as follows:
[0083] ;
[0084] For the current moment It should be noted that although the state transition model outputs two values, the variance is not applied.
[0085] In the implementation, the error covariance needs to be initialized. Specifically, the trajectory data of N consecutive empirical quadruplets in the current trajectory are used. , where it is assumed have The dimensions of , calculate N The standard deviation of each component , the initial value of the error covariance is calculated as follows:
[0086] ;
[0087] ;
[0088] in, represents the mean of the first component, Indicates Track data status The value of the first component of , calculate the covariance matrix of each empirical quadruple observation and prediction value, and then take the average value as the initial value of the error covariance in the prediction stage .
[0089] In implementation, the prediction formula for the error covariance is as follows:
[0090] ;
[0091] in, It is a state transition model about the state At the present moment The Jacobian matrix is used to linearize the nonlinear process model and propagate the state estimation error of the previous moment to the current moment, reflecting the influence of the state of the previous moment on the state of the current moment. is the error covariance matrix of the last update phase, It is the covariance matrix of the current prediction stage, which indicates the uncertainty of the state after prediction. is the process noise covariance matrix, which is used to describe the additional uncertainty introduced in the prediction process, Represents a transpose operation.
[0092] Optionally, the present application uses an experience replay pool and adopts an empirical statistical estimation method to calculate Specifically, we use N consecutive empirical quadruple pairs in the current trajectory And the state transfer model calculates the error of each empirical quadruple as follows:
[0093] ;
[0094] ;
[0095] in, . Then calculate the average error of all empirical quadruple as follows:
[0096] ;
[0097] Finally, the following formula is used to calculate :
[0098] ;
[0099] in, represents the transpose operation, Represents the trajectory data Four-tuple samples, and Represent the previous moment and the current moment in the experience quadruple respectively.
[0100] In the implementation, after the state and covariance matrix are predicted, the observation value and Kalman gain are introduced to correct the predicted state value and error covariance to obtain a more accurate current state prediction value and covariance matrix. Specifically, the difference between the observation value and the predicted value is first calculated. as follows:
[0101] ;
[0102] The Kalman gain is then calculated using the following formula :
[0103] ;
[0104] in, It is a diagonal matrix calculated and constructed by using the continuous trajectory data in the experience replay pool using dynamic estimation. The calculation method is the same as the initial value of the error covariance Similarly, using N consecutive experience samples in the current trajectory, take , assuming there is Dimensional components, calculate the standard deviation of each component , The calculation formula is as follows:
[0105] ;
[0106] In practice, the correction formula for the predicted state value is as follows:
[0107] ;
[0108] To predict the corrected state value; the correction formula of the error covariance matrix is as follows:
[0109] ;
[0110] is the corrected error covariance matrix; after the correction is completed, the corrected predicted state value and error covariance matrix of each trajectory data are replaced into the corresponding experience quadruple to form a new experience quadruple and put into the experience replay pool.
[0111] In step S103, through the extended Kalman filter processing, the state estimation of the trajectory data is more accurate and of higher quality, and the training efficiency and effect of reinforcement learning are significantly improved. At the same time, through the dynamic estimation and correction mechanism, the robustness and adaptability of the system are improved, so that the algorithm can better cope with complex physical environments and random challenges.
[0112] Step S104: Use the DDPG algorithm to randomly sample experience quadruplets from the experience replay pool and learn and update the current strategy to obtain a new strategy, and repeat the process until the strategy learning is completed.
[0113] In the implementation, the deep deterministic policy gradient algorithm (DDPG algorithm) is used to optimize the current policy. The DDPG algorithm is a technology that combines the policy gradient method and deep reinforcement learning, which can achieve efficient learning in the continuous action space. Specifically, a number of experience quadruplets are randomly sampled from the experience replay pool, the gradient of the policy network is calculated using the sampled data, and the policy network parameters are updated through back propagation. At the same time, the value network parameters are updated according to the target value. Through multiple rounds of iterations, the current policy is gradually optimized to obtain a better policy as the new policy.
[0114] In summary, combined with Figure 2 , this application provides a beamline station parameter optimization method based on extended Kalman filtering and reinforcement learning, which aims to optimize beamline station parameters and solve the problem that reinforcement learning strategies are difficult to learn effective strategies in sparse reward scenarios when there are errors in the equipment. By training a probabilistic neural network to establish a state transition model and combining the extended Kalman filter to accurately estimate the state, it is possible to reduce the impact of equipment errors on state estimation and improve the accuracy and stability of strategy learning. Finally, the DDPG algorithm is used to optimize the current strategy, and a new strategy is gradually obtained through multiple rounds of iterations. The combination of Kalman filtering and reinforcement learning is used to alleviate the impact of system errors in the process of equipment parameter tuning, improve the accuracy of state estimation, and make strategy learning more accurate.
[0115] Figure 3 This is a structural block diagram of a beamline station parameter optimization system based on extended Kalman filtering and reinforcement learning provided by an embodiment of the present application. The system includes at least the following modules:
[0116] The trajectory data collection module is used to randomly select several initial states from the environment and perform sampling based on the initial strategy and the preset target state, and collect multiple trajectory data consisting of continuous experience quadruplets;
[0117] A state transition model generation module is used to train a probabilistic neural network using the collected trajectory data to obtain a state transition model in the first round of sampling;
[0118] The extended Kalman filter module is used to perform extended Kalman filtering on each trajectory data in combination with the state transfer model, replace the filtered next moment state into the experience quadruple of each trajectory data and save the new experience quadruple into the experience playback pool;
[0119] The strategy learning module is used to use the DDPG algorithm to randomly sample experience quadruplets from the experience replay pool and learn and update the current strategy to obtain a new strategy, and repeat this cycle until the strategy learning is completed.
[0120] For relevant details, refer to the above method embodiment.
[0121] Figure 4 4 is a block diagram of an electronic device provided by an embodiment of the present application. The device at least includes a processor 401 and a memory 402.
[0122] The processor 401 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 401 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 401 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 401 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 401 may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.
[0123] The memory 402 may include one or more computer-readable storage media, which may be non-transitory. The memory 402 may also include a high-speed random access memory, and a non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 402 is used to store at least one instruction, which is used to be executed by the processor 401 to implement the beamline station parameter optimization method based on extended Kalman filtering and reinforcement learning provided in the method embodiment of the present application.
[0124] In some embodiments, the electronic device may further optionally include: a peripheral device interface and at least one peripheral device. The processor 401, the memory 402 and the peripheral device interface may be connected via a bus or a signal line. Each peripheral device may be connected to the peripheral device interface via a bus, a signal line or a circuit board. Schematically, the peripheral devices include but are not limited to: a radio frequency circuit, a touch display screen, an audio circuit, and a power supply.
[0125] Of course, the electronic device may also include fewer or more components, which is not limited in this embodiment.
[0126] Optionally, the present application also provides a computer-readable storage medium, in which a program is stored, and the program is loaded and executed by a processor to implement the beamline station parameter optimization method based on extended Kalman filtering and reinforcement learning of the above method embodiment.
[0127] Optionally, the present application also provides a computer product, which includes a computer-readable storage medium, in which a program is stored, and the program is loaded and executed by a processor to implement the beamline station parameter optimization method based on extended Kalman filtering and reinforcement learning of the above-mentioned method embodiment.
[0128] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0129] The above embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.
Claims
1. A beamline station parameter optimization method based on extended Kalman filtering and reinforcement learning, characterized in that: The method comprises: Based on the initial strategy and the preset target state, several initial states are randomly selected from the environment where the beamline station system is located and sampled to collect multiple trajectory data consisting of continuous experience quadruplets; In the first round of sampling, the probabilistic neural network is trained using the collected trajectory data to obtain the state transition model; For each trajectory data, an extended Kalman filter is performed in combination with the state transfer model, and the filtered next moment state is used to replace the experience quadruple of each trajectory data and the new experience quadruple is saved in the experience playback pool; The DDPG algorithm is used to randomly sample experience quadruplets from the experience replay pool and learn and update the current strategy to obtain a new strategy, and this cycle is repeated until the strategy learning is completed.
2. The beamline station parameter optimization method based on extended Kalman filtering and reinforcement learning according to claim 1, characterized in that: The method of using the collected trajectory data to train a probabilistic neural network to obtain a state transition model includes: After the first round of sampling is completed and multiple trajectory data are collected, the preset probabilistic neural network is trained using the collected multiple trajectory data to obtain a state transition model. The state transition model is as follows: f(s,a;θ)→μ,σ 2 ; Among them, s and a represent the current state and current action respectively, θ is the model parameter of the probabilistic neural network, μ is the mean vector, σ 2 is the variance vector.
3. The beamline station parameter optimization method based on extended Kalman filtering and reinforcement learning according to claim 2, characterized in that: For each trajectory data, the extended Kalman filter is performed in combination with the state transfer model, and the filtered next moment state is used to replace the experience quadruple of each trajectory data and the new experience quadruple is saved in the experience playback pool, including: The state at the current moment is predicted through the state transition model and the error covariance matrix of the state is calculated; Introduce observation values and Kalman gain to correct the predicted state value and error covariance matrix; After the correction is completed, the error covariance matrix is updated, and the corrected predicted state value is replaced into the original experience quadruple as the state value at the next moment to form a new experience quadruple.
4. The beamline station parameter optimization method based on extended Kalman filtering and reinforcement learning according to claim 3 is characterized in that: The method of predicting the state at the current moment through the state transition model and calculating the error covariance matrix of the state includes: Use the state transition model to predict the current time s t The status is as follows: is the current time s t The prediction formula of the error covariance is as follows: Among them, F t+1 is the Jacobian matrix of the state transition model with respect to the state s at the current moment, P is the covariance matrix, P t is the error covariance matrix of the last update phase, P t+1 - is the covariance matrix of the current prediction stage, Q t+1 is the process noise covariance matrix, and T represents the transpose operation.
5. The beamline station parameter optimization method based on extended Kalman filtering and reinforcement learning according to claim 4, characterized in that: The method of predicting the state at the current moment by the state transition model and calculating the error covariance matrix of the state also includes: Using N consecutive empirical quadruple sets (s t ,a t ,r t ,s t+1 ) and the state transition model calculates the error w of each empirical quadruple i t+1 as follows: w i t+1 =s i t+1 -m i t+1 ; Among them, i∈[1,N], calculate the average error of all empirical quadruple as follows: Q is calculated using the following formula t+1 : Among them, T represents the transposition operation, i represents the i-th quadruple sample in the trajectory data, and t and t+1 represent the previous moment and the current moment in the experience quadruple, respectively.
6. The beamline station parameter optimization method based on extended Kalman filtering and reinforcement learning according to claim 4, characterized in that: The introduction of observation values and Kalman gains to correct the predicted state values and the error covariance matrix includes: Calculate the difference y between the observed and predicted values t+1 as follows: The Kalman gain K is calculated using the following formula t+1 : K t+1 =P t+1 - (P t+1 - +R) -1 ; Among them, R is a diagonal matrix calculated and constructed by dynamic estimation using continuous trajectory data in the experience replay pool, and N continuous experience samples are sampled in the current trajectory, and the Assuming there are n-dimensional components, calculate the standard deviation σ of each component. The calculation formula of R is as follows: R=diag(σ1 2 ,σ2 2 ,σ3 2 ,…,s n 2 ); The correction formula for the predicted state value is as follows: To predict the corrected state value; the correction formula of the error covariance matrix is as follows: P t+1 =(1-K t+1 )P t+1 - ; P t+1 is the corrected error covariance matrix.
7. The beamline station parameter optimization method based on extended Kalman filtering and reinforcement learning according to claim 3, characterized in that: The method of predicting the state at the current moment by the state transition model and calculating the error covariance matrix of the state also includes: The error covariance needs to be initialized urgently, using the trajectory data of N consecutive empirical quadruplets in the current trajectory (s t ,a t ,r t ,s t+1 ), assuming that st has n-dimensional components, s t =(s t1 ,s t2 ,s t3 ,…,s tn ), calculate N s t The standard deviation σ of each component of , and the initial value of the error covariance are calculated as follows: P0=diag(σ1 2 ,σ2 2 ,σ3 2 ,…,s n 2 ); in, represents the mean of the first component, s i1 Represents the state s in the i-th trajectory data t The value of the first component of is used to calculate the covariance matrix of the observed and predicted values of each empirical quadruple, and the average value is taken as the initial value P0 of the error covariance in the prediction stage.
8. A beamline station parameter optimization system based on extended Kalman filtering and reinforcement learning, characterized in that: include: The trajectory data acquisition module is used to randomly select several initial states from the environment where the beamline station system is located and sample them based on the initial strategy and the preset target state, and collect multiple trajectory data consisting of continuous experience quadruplets; A state transition model generation module is used to train a probabilistic neural network using the collected trajectory data to obtain a state transition model in the first round of sampling; An extended Kalman filter module is used to perform an extended Kalman filter on each trajectory data in combination with the state transfer model, replace the filtered state into the experience quadruple of each trajectory data and save the new experience quadruple into the experience playback pool; The strategy learning module is used to use the DDPG algorithm to randomly sample experience quadruplets from the experience replay pool and learn and update the current strategy to obtain a new strategy, and repeat this cycle until the strategy learning is completed.
9. An electronic device, characterized in that: The device includes a processor and a memory; a program is stored in the memory, and the program is loaded and executed by the processor to implement a beamline station parameter optimization method based on extended Kalman filtering and reinforcement learning as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The storage medium stores a program, and when the program is executed by the processor, it is used to implement a beamline station parameter optimization method based on extended Kalman filtering and reinforcement learning as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Automatic assembly line work efficiency optimization system and method based on reinforcement learning
CN111223141A
Model-enhanced unmanned aerial vehicle flight path reinforcement learning optimization method
CN114879738A