Position precision compensation method for radiation-resistant absolute photoelectric encoder
By introducing a dual sensing module and Kalman filter reinforcement learning algorithm into the photoelectric encoder, the position information of the main sensor is predicted and compensated, which solves the position error and communication delay problems of the photoelectric encoder in a nuclear radiation environment and improves the stability and accuracy of the robot in a nuclear radiation environment.
Patent Information
- Application Number
- CN202511136468.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-14
- Publication Date
- 2025-10-03
AI Technical Summary
The position error and communication delay problems of photoelectric encoders in nuclear radiation environments cause unstable robot operation, which is difficult to effectively solve with existing technologies.
The Kalman filter and reinforcement learning algorithm of the dual-sensor module are used to predict and compensate the position information of the main sensor through the secondary sensor, thereby optimizing the radiation resistance of the photoelectric encoder.
The sensing accuracy and service life of the photoelectric encoder in a radiation environment are improved, the problem of position loss when the main sensor is damaged is solved, and the stability of the robot in a nuclear radiation environment is enhanced.
Smart Images

Figure CN120740657A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of nuclear radiation robot development, and in particular to a position accuracy compensation method for a radiation-resistant absolute photoelectric encoder. Background Art
[0002] As nuclear industry technology matures, the safe operation of nuclear power plants has attracted much attention, and nuclear emergency rescue equipment and technology have become a key research area. As an efficient and safe nuclear emergency and maintenance equipment, nuclear environment operation robots can effectively replace manual labor to perform work tasks in high-risk radiation environments. In the development of nuclear robots, the radiation damage effect of the circuit system has become a key technical bottleneck. The reason is that the performance of electronic components degrades and even fails under the action of cumulative radiation doses, eventually leading to the paralysis of the circuit module. As an angle displacement sensor, the photoelectric encoder is an important circuit system of the robot's articulated arm. Radiation damage to the photoelectric encoder will cause position errors, communication delays, and communication interruptions, which seriously interfere with the operation of the robot. Therefore, a method is needed to improve the service life and sensing accuracy of the photoelectric encoder in a radiation environment.
[0003] Existing solutions to radiation damage to robot angular displacement sensors can be divided into three main categories: one is to reinforce the outer shell shielding material and add materials such as lead and tungsten to protect the internal circuit from damage by gamma rays; however, the high-density materials added by this method will significantly increase the weight of the robot and reduce its flexibility. The second is to replace the photoelectric encoder with a rotary transformer. Since the rotary transformer has no electronic components inside, it exhibits high radiation resistance in a radiation environment; however, the rotary transformer is large in size and has lower accuracy than the photoelectric encoder. The third is to improve the radiation resistance of the photoelectric encoder circuit by reinforcing the circuit with radiation resistance, and to reduce the accuracy loss caused by radiation damage through software algorithms; however, the difficulty of this method lies in the radiation resistance circuit design and algorithm design. Summary of the Invention
[0004] The purpose of the present invention is to use the position information of the secondary sensor to predict and compensate the position information of the main sensor, so as to optimize the radiation resistance performance of the photoelectric encoder without signal interruption.
[0005] The present invention discloses a method for compensating the position accuracy of an absolute photoelectric encoder resistant to radiation, wherein the absolute photoelectric encoder comprises a main sensing module and a secondary sensing module;
[0006] The method for compensating the position accuracy of the radiation-resistant absolute photoelectric encoder comprises the following steps:
[0007] During the power-on initialization phase of the absolute photoelectric encoder, the main sensor module is powered on by an external power supply, while the secondary sensor module is kept off through a switch circuit to reduce radiation damage.
[0008] When the encoder is working in an ionizing radiation environment, as the cumulative absorbed dose of gamma radiation reaches the preset threshold, it detects communication delays and errors in the main sensor, starts powering the secondary sensor, and simultaneously reads the position data of the main and secondary sensors;
[0009] The optimized and compensated position information is obtained through the dual-sensor Kalman compensation algorithm.
[0010] Furthermore, the dual-sensor Kalman compensation algorithm uses dual-sensor readings, including absolute position information from the primary and secondary sensors, which are fed into the Kalman filter prediction module and the DDQN reinforcement learning module.
[0011] The Kalman filter prediction module uses the main sensor reading as the state at the previous moment and calculates the optimal estimate of the main sensor through the secondary sensor reading;
[0012] At the same time, the DDQN reinforcement learning module inputs the dual sensor readings and the optimal estimated parameters of the Kalman filter as the current state, and obtains the optimal action, which is the optimized Kalman filter parameters;
[0013] The Kalman filter obtains the final compensated Kalman filter value through the primary and secondary sensor readings at the next moment and the optimized Kalman filter parameters;
[0014] The training process of the DDQN reinforcement learning module is:
[0015] The Kalman filter module calculates the compensated Kalman filter value using the new filter parameters and calculates the reward function based on the difference between the new filter parameters and the reference value. The reward function and the next-moment primary and secondary sensor readings are fed into the DDQN reinforcement learning module to calculate the cross-entropy loss and backpropagate the updated network. The new round of optimized Kalman filter parameters is then output to enter the training loop.
[0016] The prediction process of the DDQN reinforcement learning module is:
[0017] The compensated position information after Kalman filtering is no longer returned to the DDQN reinforcement learning module for updating, but is directly output to the external device as the final position information.
[0018] Furthermore, the optimal estimate of the primary sensor is:
[0019]
[0020] Where x i is the position information after compensation, x i-1 is the position information after compensation at the previous moment, α is the optimal Kalman gain parameter output by DDQN reinforcement learning, P(i-1) is the radiation damage period difference ratio of the primary and secondary sensors, is the displacement output difference between the main sensor at the i-th moment and the previous moment, is the displacement output difference between the auxiliary sensor at the i-th moment and the previous moment.
[0021] Furthermore, the radiation damage period difference ratio of the primary and secondary sensors is:
[0022]
[0023] i represents the current moment, ΔT represents the sensor rated position update period, ΔV1 i-1 Indicates the displacement output difference between the main sensor at the i-1th moment and the previous moment, ΔV2 i-1 Indicates the displacement output difference between the auxiliary sensor at the i-1th moment and the previous moment.
[0024] Furthermore, the DDQN reinforcement learning module first inputs the difference between the displacement outputs of the two sensors and The current Kalman filter parameter α is also used as the current state s; the optimal action is obtained by maximizing the Q value, that is, the optimized Kalman filter parameter α′, which is calculated as follows:
[0025] α′=arg max Q online (s,a)
[0026] Where a represents all possible actions in the current state, that is, the possible increase or decrease of α. Substitute the obtained α′ into the dual-sensor Kalman filter to obtain the optimized compensated position output x i , the output position is interpolated with the reference true value to convert it into a reward coefficient r. The sensor output at the next moment and the optimized Kalman filter parameter α′ are combined as the next state s′, and then DDQN updates the parameters by the following formula:
[0027] y=r+γQ target (s′,max Q online (s′,a′))
[0028] l=(yQ online (s,a))
[0029] Where y is the Q value of the target, r is the current reward, γ is the discount factor, and Q target With Q online are the Q values of the target network and the current network respectively, and a′ is the possible action in the next state;
[0030] Calculate all possible a' values, select the optimal action based on the action with the largest Q value, and substitute the value into Q target Calculate the y value and compare it with Q onlineThe loss is calculated by the difference between (s, a), and then the network is updated by backpropagation of the loss and put into the experience replay pool as an experience reference.
[0031] Furthermore, we calculate all possible a' values, select the optimal action based on the action with the largest Q value, and substitute the value into Q target Calculate the y value and compare it with Q online (s, a) is subtracted to calculate the loss, and then the network is updated through loss backpropagation and put into the experience replay pool as an experience reference.
[0032] The beneficial effects achieved by the present invention are:
[0033] This invention addresses the issue of radiation damage to absolute photoelectric encoders operating in radiation-prone environments by designing a dual-sensor enhanced Kalman compensation algorithm to address the issue of reduced accuracy caused by radiation damage. The dual-sensor module of the photoelectric encoder activates the secondary sensor if radiation damage damages the primary sensor, improving the encoder's overall radiation resistance.
[0034] This paper uses the position data obtained by the secondary sensor as parameters in the state estimation equation of the Kalman filter method to compensate for the position error of the primary sensor caused by radiation damage. This paper also uses reinforcement learning to optimize the hyperparameters in the Kalman filter algorithm, addressing the state estimation bias caused by increasing cumulative radiation dose.
[0035] The present invention proposes a method for compensating the position accuracy of a radiation-resistant absolute photoelectric encoder, which can improve the radiation resistance of the absolute photoelectric encoder, compensate for the accuracy drop caused by radiation damage, and solve the problem of position loss when switching from main sensing to auxiliary sensing. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 Schematic diagram of dual-sensor radiation damage of the absolute photoelectric encoder of the present invention;
[0037] Figure 2 This is a flowchart of the dual-sensor optimization of the absolute encoder of the present invention;
[0038] Figure 3 The specific process of the dual-sensor Kalman compensation algorithm of the present invention is
[0039] Figure 4 This is the experimental effect of the present invention. DETAILED DESCRIPTION
[0040] The present invention will be further described below with reference to specific embodiments, and the advantages and features of the present invention will become clearer as the description proceeds. However, these embodiments are merely exemplary and do not constitute any limitation to the scope of the present invention. It should be understood by those skilled in the art that the details and forms of the technical solutions of the present invention may be modified or replaced without departing from the spirit and scope of the present invention, and such modifications and replacements fall within the scope of protection of the present invention.
[0041] The present invention provides a method for compensating the position accuracy of an absolute photoelectric encoder that is resistant to radiation. The absolute photoelectric encoder outputs a communication signal containing absolute position information by reading position information on a grating code disk.
[0042] Radiation can damage absolute photoelectric encoders. Damage to the primary sensor can cause position errors and increased communication intervals. However, the secondary sensor, due to its late startup, suffers less radiation damage and maintains normal signals. The increase in communication interval caused by radiation damage is difficult to predict. This invention estimates the state change of the primary sensor by using the period difference between the secondary and primary sensors.
[0043] Figure 1 As shown in the figure, ΔV1 is the main sensor that has position error and increased communication interval after damage, and ΔV2 is the secondary sensor that has less radiation damage due to late startup, so the signal is in a normal state.
[0044] Calculate the radiation damage change rate of the primary and secondary sensors, that is, the period difference ratio P(i-1):
[0045]
[0046] i represents the current moment, ΔT represents the sensor rated position update period, ΔV1 i-1 Indicates the displacement output difference between the main sensor at the i-1th moment and the previous moment, ΔV2 i-1 Indicates the displacement output difference between the auxiliary sensor at the i-1th moment and the previous moment.
[0047] like Figure 2 As shown in the figure, during the power-on initialization phase of an absolute photoelectric encoder, the main sensor module is directly connected to a +5V DC power supply via an external power supply and starts up, while the secondary sensor module is kept off by a switching circuit to reduce radiation damage. At this point, the encoder defaults to using the absolute position information output by the main sensor module, such as the BISS or EnDat protocol encoded signal, as the only position information.
[0048] When the encoder operates in an ionizing radiation environment, as the cumulative absorbed gamma radiation dose reaches a preset threshold, the main sensor module gradually exhibits performance degradation, including increased position signal transmission delay and error. At this point, external control logic intervenes to trigger a redundant switching protocol, initiating dual-channel data acquisition and inputting the dual-sensor Kalman compensation algorithm.
[0049] Figure 3 As shown in the figure, the specific process of the dual-sensor Kalman compensation algorithm is shown. The dual-sensor readings include the absolute position information transmitted by the main and auxiliary sensors ( and Representing the binary position information of the main sensor and the auxiliary sensor respectively), the information is passed to the Kalman filter prediction module and the DDQN reinforcement learning module.
[0050] The Kalman filter prediction module uses the primary sensor reading as the state at the previous moment and calculates the optimal estimate of the primary sensor through the secondary sensor reading:
[0051]
[0052] Where x i is the position information after compensation, x i-1 is the position information after compensation at the previous moment, α is the optimal Kalman gain parameter output by DDQN reinforcement learning, P(i-1) is the radiation damage period difference ratio of the primary and secondary sensors, represents the displacement output difference between the main sensor at the i-th moment and the previous moment, represents the displacement output difference between the auxiliary sensor at the i-th moment and the previous moment.
[0053] The DDQN reinforcement learning module first inputs the dual sensor displacement output difference and The current Kalman filter parameter α is also used as the current state s. The optimal action is obtained by maximizing the Q value, that is, the optimized Kalman filter parameter α′, which is calculated as follows:
[0054] α′=arg max Q online (s,a)
[0055] Where a represents all possible actions in the current state, that is, the possible increase or decrease of the value of α. Substitute the obtained α′ into the dual-sensor Kalman filter to obtain the optimized compensated position output x i , the output position is interpolated with the reference true value to convert it into a reward coefficient r. The sensor output at the next moment and the optimized Kalman filter parameter α′ are combined as the next state s′, and then DDQN updates the parameters by the following formula:
[0056] y=r+γQ target(s′,max Q online (s′,a′))
[0057] l=(yQ online (s,a)) where y is the Q value of the target, r is the current reward, γ is the discount factor, Q target With Q online are the Q values of the target network and the current network respectively, and a′ is the possible action for the next state.
[0058] Calculate all possible a' values, select the optimal action based on the action with the largest Q value, and substitute the value into Q target Calculate the y value and compare it with Q online The loss is calculated by the difference between (s, a), and then the network is updated by backpropagation of the loss and put into the experience replay pool as an experience reference.
[0059] In summary, during the training phase, the dual sensors output the current position information and Kalman filter parameters to the DDQN reinforcement learning module as the current state s. The DDQN reinforcement learning module calculates the maximum Q value to obtain the current optimal action, i.e., the optimal Kalman filter parameter α, which is then passed back to the Kalman filter module. The Kalman filter module calculates the compensated position output x using the new filter parameters. i , and the difference from the reference value yields the reward function r. This reward function r and the next state s′ are fed into the DDQN reinforcement learning module to calculate the cross-entropy loss and backpropagate to update the network. The new Kalman filter parameters α are then output to the training loop. Prediction phase: The compensated position after the dual-sensor Kalman filter is directly output to the external device as the final position information. The reward function r is no longer calculated to update the DDQN reinforcement learning module.
[0060] Figure 4 The compensation algorithm of the present invention is tested by using the absolute photoelectric encoder data obtained from the γ radiation experiment. It can be seen from the figure that the error after compensation is significantly smaller than the original error, proving the effectiveness of the present invention.
[0061] The above are only specific steps of the present invention and do not constitute any limitation to the scope of protection of the present invention; any technical solutions formed by equivalent transformation or equivalent replacement fall within the scope of protection of the present invention; the parts not elaborated in detail in the present invention belong to the common knowledge of those skilled in the art.
Claims
1. A method for compensating the position accuracy of an absolute photoelectric encoder resistant to radiation, characterized in that: The absolute photoelectric encoder includes a main sensing module and a secondary sensing module; The method for compensating the position accuracy of the radiation-resistant absolute photoelectric encoder comprises the following steps: During the power-on initialization phase of the absolute photoelectric encoder, the main sensor module is powered on by an external power supply, while the secondary sensor module is kept off through a switch circuit to reduce radiation damage. When the encoder is working in an ionizing radiation environment, as the cumulative absorbed dose of gamma radiation reaches the preset threshold, it detects communication delays and errors in the main sensor, starts powering the secondary sensor, and simultaneously reads the position data of the main and secondary sensors; The optimized and compensated position information is obtained through the dual-sensor Kalman compensation algorithm.
2. The method for compensating the position accuracy of an absolute photoelectric encoder according to claim 1, characterized in that: The dual-sensor Kalman compensation algorithm uses dual-sensor readings, including absolute position information from the primary and secondary sensors. This absolute position information is then fed into the Kalman filter prediction module and the DDQN reinforcement learning module. The Kalman filter prediction module uses the main sensor reading as the state at the previous moment and calculates the optimal estimate of the main sensor through the secondary sensor reading; At the same time, the DDQN reinforcement learning module inputs the dual sensor readings and the optimal estimated parameters of the Kalman filter as the current state, and obtains the optimal action, which is the optimized Kalman filter parameters; The Kalman filter obtains the final compensated Kalman filter value through the primary and secondary sensor readings at the next moment and the optimized Kalman filter parameters; The training process of the DDQN reinforcement learning module is: The Kalman filter module calculates the compensated Kalman filter value using the new filter parameters and calculates the reward function based on the difference between the new filter parameters and the reference value. The reward function and the next-moment primary and secondary sensor readings are fed into the DDQN reinforcement learning module to calculate the cross-entropy loss and backpropagate the updated network. The new round of optimized Kalman filter parameters is then output to enter the training loop. The prediction process of the DDQN reinforcement learning module is: The compensated position information after Kalman filtering is no longer returned to the DDQN reinforcement learning module for updating, but is directly output to the external device as the final position information.
3. The method for compensating the position accuracy of an absolute photoelectric encoder according to claim 2, wherein: The optimal estimate of the primary sensor is: Where x i is the position information after compensation, x i-1 is the position information after compensation at the previous moment, α is the optimal Kalman gain parameter output by DDQN reinforcement learning, P(i-1) is the radiation damage period difference ratio of the primary and secondary sensors, is the displacement output difference between the main sensor at the i-th moment and the previous moment, is the displacement output difference between the auxiliary sensor at the i-th moment and the previous moment.
4. The method for compensating the position accuracy of an absolute photoelectric encoder according to claim 3, wherein: The radiation damage period difference ratio of the main and auxiliary sensors is: i represents the current moment, ΔT represents the sensor rated position update period, ΔV1 i-1 Indicates the displacement output difference between the main sensor at the i-1th moment and the previous moment, ΔV2 i-1 Indicates the displacement output difference between the auxiliary sensor at the i-1th moment and the previous moment.
5. The method for compensating position accuracy of an absolute photoelectric encoder for radiation resistance according to claim 3, characterized in that: The DDQN reinforcement learning module first inputs the dual sensor displacement output difference and The current Kalman filter parameter α is also used as the current state s; The optimal action is obtained by maximizing the Q value, that is, the optimized Kalman filter parameter α′, which is calculated as follows: α′=arg max Q online (s,a) Where a represents all possible actions in the current state, that is, the possible increase or decrease of α. Substitute the obtained α′ into the dual-sensor Kalman filter to obtain the optimized compensated position output x i , the output position is interpolated with the reference true value to convert it into a reward coefficient r. The sensor output at the next moment and the optimized Kalman filter parameter α′ are combined as the next state s′, and then DDQN updates the parameters by the following formula: y=r+γQ target (s′,max Q online (s′,a′)) l=(y-Q online (s,a)) Where y is the Q value of the target, r is the current reward, γ is the discount factor, and Q target With Q online are the Q values of the target network and the current network respectively, and a′ is the possible action in the next state; Calculate all possible a' values, select the optimal action based on the action with the largest Q value, and substitute the value into Q target Calculate the y value and compare it with Q online The loss is calculated by the difference between (s, a), and then the network is updated by backpropagation of the loss and put into the experience replay pool as an experience reference.
6. The method for compensating the position accuracy of an absolute photoelectric encoder resistant to radiation according to claim 5, characterized in that: Calculate all possible a' values, select the optimal action based on the action with the largest Q value, and substitute the value into Q target Calculate the y value and compare it with Q online (s, a) is subtracted to calculate the loss, and then the network is updated through loss backpropagation and put into the experience replay pool as an experience reference.
Citation Information
Cited By
Grating signal acquisition and processing method and system
CN121877071A