Nuclear power maintenance robot control method and device

CN120901963BActive Publication Date: 2026-08-18CHINA NUCLEAR POWER TECH RES INST CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511219333.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-28
Publication Date
2026-08-18
Estimated Expiration
2045-08-28

AI Technical Summary

Technical Problem

然而,在高辐射环境下,辐射会导致传感器数据漂移和执行器响应异常,影响机器人感知和操作的精准性

Benefits of technology

[0055] The aforementioned nuclear power plant maintenance robot control method and device acquires measurement data from multimodal sensors and obtains environmental state data based on the measurements. Using a large language model, it generates task instructions for the robot based on the environmental state data and the target task. The large language model then infers the task instructions to obtain radiation interference prediction information. During the robot's execution of the target task, the task instructions are adjusted based on the radiation interference prediction information obtained at the current moment. The robot is then controlled to perform corresponding task operations according to the adjusted instructions. This allows for continuous optimization of the robot's execution strategy based on real-time environmental changes, precise robot control, and thus improved accuracy and reliability in task completion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120901963B_ABST
    Figure CN120901963B_ABST
Patent Text Reader

Abstract

The application relates to a nuclear power maintenance robot control method and device. The method comprises the following steps: obtaining measurement data of a multi-modal sensor, obtaining environment state data according to the measurement, generating a robot task instruction according to the environment state data and a target task through a large language model, reasoning the task instruction through the large language model to obtain irradiation interference prediction information, adjusting the task instruction at the current moment according to the irradiation interference prediction information obtained at the current moment in the process of executing the target task by the robot, and controlling the robot to execute corresponding task operations according to the adjusted task instruction. The method can accurately control a radiation-resistant nuclear power robot.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent control technology, and in particular to a control method and device for a nuclear power plant maintenance robot. Background Technology

[0002] Radiation-resistant nuclear power robots are widely used in nuclear power plant inspection, maintenance, and emergency response tasks to reduce human exposure to high-radiation environments. These robots typically rely on radiation-resistant hardware and shielding materials to extend their service life, while also being equipped with sensors to collect environmental data to assist in navigation and task execution.

[0003] In existing technologies, machine learning techniques are introduced to improve the adaptability and intelligence of control systems. However, in high-radiation environments, radiation can cause sensor data drift and actuator response anomalies, affecting the accuracy of robot perception and operation. Summary of the Invention

[0004] Therefore, it is necessary to provide a control method and device for a nuclear power maintenance robot that can accurately control a radiation-resistant nuclear power robot, addressing the aforementioned technical problems.

[0005] Firstly, this application provides a control method for a nuclear power plant maintenance robot, including:

[0006] The system acquires measurement data from multimodal sensors and obtains environmental status data based on the measurements. The multimodal sensors include multiple static sensors installed around nuclear power equipment in different radiation areas inside the nuclear power plant and multiple mobile sensors installed on the robot. The measurement data includes environmental parameters and robot status information.

[0007] Based on environmental state data and target tasks, a large language model is used to generate task instructions for the robot; the task instructions include task paths and operation instructions.

[0008] By reasoning about task instructions using a large language model, we can obtain irradiation interference prediction information.

[0009] During the robot's execution of the target task, the task instructions at the current moment are adjusted based on the radiation interference prediction information obtained at the current moment, and the robot is controlled to perform the corresponding task operations according to the adjusted task instructions.

[0010] In one embodiment, a static sensor is used to collect environmental parameters of the radiation area, and a motion sensor is used to collect the robot's state information; the step of acquiring environmental state data based on the measurement data includes:

[0011] The measurement data is dynamically corrected, and a multidimensional environmental dataset is constructed based on the dynamically corrected measurement data. The multidimensional environmental dataset includes time parameters, spatial location parameters, radiation level parameters, and equipment status parameters.

[0012] Environmental risk distribution information of nuclear power plants is obtained from multidimensional environmental datasets. The environmental risk distribution information includes the spatial gradient variation trend of environmental parameters, the boundary characteristics of the target radiation area, the variation trend of humidity anomaly areas, and the degree of deviation of equipment status parameters.

[0013] Based on the installation location of the multimodal sensors, the measurement data is mapped to the physical space model of the nuclear power plant;

[0014] In the physical space model, abnormal data in the measurement data are corrected, and the corrected measurement data is fused with environmental risk distribution information to obtain environmental status data.

[0015] In one embodiment, the step of correcting abnormal data in the measurement data includes:

[0016] Based on the spatial distribution of multimodal sensors and historical measurement data, a reference sensor whose radiation level meets the preset conditions is obtained;

[0017] Based on the measurement data of the reference sensor, the measurement data of the remaining sensors other than the reference sensor are corrected for deviation to obtain reference measurement data;

[0018] Based on reference measurement data, the trend of radiation level change in different radiation areas is obtained, and the current acquisition strategy of the multimodal sensor is adjusted according to the rate of change of the radiation level trend.

[0019] Based on the adjusted acquisition strategy, the updated measurement data of the multimodal sensor is reacquired. Based on the updated measurement data and the reference measurement data, the characteristics of radiation intensity change at different measurement times are obtained.

[0020] Based on the characteristics of radiation intensity changes and the operating status data of nuclear power equipment, the interfered data in the reference measurement data is obtained, and the interfered data is dynamically corrected based on the updated measurement data.

[0021] In one embodiment, the step of generating task instructions for the robot based on environmental state data and the target task using a large language model includes:

[0022] Based on environmental status data and target tasks, obtain multidimensional task input data and task context information for the large language model; the task context information includes radiation scene features, equipment status information, and task objectives.

[0023] Based on task context information, the multidimensional task input data is processed through a large language model to obtain the target execution strategy corresponding to the target task.

[0024] By using a multi-objective path planning algorithm, the objective execution strategy is optimized to obtain candidate task paths;

[0025] Conduct radiation impact assessments on candidate task paths, and obtain target task paths that meet preset safety conditions based on the radiation impact assessment results.

[0026] Generate robot task instructions based on the target task path.

[0027] In one embodiment, the step of optimizing the target execution strategy and obtaining candidate task paths using a multi-objective path planning algorithm includes:

[0028] Based on the three-dimensional environment model of the nuclear power plant, a path search space is established, and risk elements are labeled on the three-dimensional environment model within the path search space.

[0029] Based on the labeled risk factors and target tasks, obtain a multi-level path cost matrix;

[0030] In the process of optimizing the target execution strategy through multi-objective path planning algorithm, the cumulative radiation exposure on different task paths is obtained according to the current radiation distribution data, and the weights in the multi-level path cost matrix are adjusted based on the cumulative radiation exposure and the robot's radiation tolerance threshold.

[0031] Based on the adjusted multi-level path cost matrix, path evaluation is performed on each task path, and candidate task paths are obtained based on the path evaluation results.

[0032] In one embodiment, the step of reasoning about task instructions using a large language model to obtain irradiation interference prediction information includes:

[0033] Based on environmental state data and the robot's historical operation data, the sensor error change trend under different task stages is obtained, and the potential sensing error of the multimodal sensor is obtained according to the sensor error change trend.

[0034] By using a large language model to reason about potential perceptual errors and task instructions, we can obtain prediction data for the robot's execution errors.

[0035] Based on the error prediction data, irradiation interference prediction information is generated; the irradiation interference prediction information includes potential sensor anomaly types, execution response errors, potential control deviations, and execution compensation strategies.

[0036] In one embodiment, the step of adjusting the task instructions at the current moment based on the radiation interference prediction information obtained at the current moment includes:

[0037] Based on the radiation interference prediction information obtained at the current moment, obtain the sensor error information for the current mission stage, and correct the measurement data at the current moment based on the sensor error information;

[0038] Based on the corrected measurement data, the control deviation at the current moment is obtained, and the task instructions at the current moment are adjusted according to the control deviation.

[0039] In one embodiment, the step of adjusting the task instructions at the current moment based on the control deviation includes:

[0040] Based on the corrected measurement data, the robot's current pose data and current path deviation are obtained;

[0041] The environmental status data is updated based on the current pose data and the current path deviation;

[0042] Adjust the task instructions for the current moment based on the updated environmental status data and control deviations.

[0043] In one embodiment, the method further includes:

[0044] Based on the sensor error information, the temporal variation characteristics of the sensor signal are obtained, and based on the temporal variation characteristics, abrupt changes in the measurement data are obtained.

[0045] When the anomaly type of the mutation data is interference anomaly, the mutation data is smoothed to obtain updated measurement data;

[0046] In cases where the anomaly type of the mutation data is a loss anomaly, the mutation data is compensated to obtain updated measurement data.

[0047] Secondly, this application also provides a nuclear power plant maintenance robot control device, comprising:

[0048] The data acquisition module is used to acquire measurement data from multimodal sensors and obtain environmental status data based on the measurements. The multimodal sensors include multiple static sensors installed around nuclear power equipment in different radiation areas inside the nuclear power plant and multiple mobile sensors installed on the robot. The measurement data includes environmental parameters and robot status information.

[0049] The instruction generation module is used to generate task instructions for the robot based on environmental state data and target tasks using a large language model; the task instructions include task paths and operation instructions.

[0050] The interference acquisition module is used to reason about the task instructions through a large language model to obtain irradiation interference prediction information;

[0051] The robot control module is used to adjust the task instructions at the current moment based on the radiation interference prediction information obtained at the current moment during the robot's execution of the target task, and control the robot to perform the corresponding task operations according to the adjusted task instructions.

[0052] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method steps of any one of the first aspects.

[0053] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method steps of any one of the first aspects.

[0054] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the method steps of any one of the first aspects.

[0055] The aforementioned nuclear power plant maintenance robot control method and device acquires measurement data from multimodal sensors and obtains environmental state data based on the measurements. Using a large language model, it generates task instructions for the robot based on the environmental state data and the target task. The large language model then infers the task instructions to obtain radiation interference prediction information. During the robot's execution of the target task, the task instructions are adjusted based on the radiation interference prediction information obtained at the current moment. The robot is then controlled to perform corresponding task operations according to the adjusted instructions. This allows for continuous optimization of the robot's execution strategy based on real-time environmental changes, precise robot control, and thus improved accuracy and reliability in task completion. Attached Figure Description

[0056] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0057] Figure 1 This is a diagram illustrating the application environment of a nuclear power plant maintenance robot control method in one embodiment.

[0058] Figure 2 This is a flowchart illustrating a nuclear power plant maintenance robot control method in one embodiment;

[0059] Figure 3 This is a flowchart illustrating the control method for a nuclear power plant maintenance robot in another embodiment;

[0060] Figure 4 This is a structural block diagram of the nuclear power plant maintenance robot control device in one embodiment;

[0061] Figure 5 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0062] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0063] The nuclear power plant maintenance robot control method provided in this application embodiment can be applied to, for example... Figure 1 In the application environment shown, terminal 102 communicates with the robot's controller 104 via a network. Terminal 102 acquires measurement data from multimodal sensors and obtains environmental state data based on these measurements. Using a large language model, it generates task instructions for the robot based on the environmental state data and the target task. The large language model then infers the task instructions to obtain radiation interference prediction information. During the robot's execution of the target task, the terminal adjusts the current task instructions based on the obtained radiation interference prediction information and controls the robot to perform corresponding task operations according to the adjusted instructions. Terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, drones, low-altitude aircraft, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart vehicle devices, projection devices, etc. Portable wearable devices can include smartwatches, smart bracelets, head-mounted devices, etc. Head-mounted devices can be virtual reality (VR) devices, augmented reality (AR) devices, smart glasses, etc.

[0064] In one exemplary embodiment, such as Figure 2 As shown, a control method for a nuclear power plant maintenance robot is provided, which is applied to... Figure 1 Taking terminal 102 as an example, the explanation includes the following steps 202 to 208. Wherein:

[0065] S202: Acquire measurement data from multimodal sensors and obtain environmental status data based on the measurements; the multimodal sensors include multiple static sensors installed around nuclear power equipment in different radiation areas inside the nuclear power plant and multiple mobile sensors installed on the robot; the measurement data includes environmental parameters and robot status information.

[0066] Optionally, nuclear power plant robots are equipped with radiation-resistant multimodal sensors to ensure accurate acquisition of necessary environmental parameters and their own status information even in high-radiation environments. Multimodal sensors include, but are not limited to, radiation dose detectors, temperature sensors, humidity sensors, inertial measurement units (IMUs), lidar (LiDAR), and cameras. Radiation dose detectors measure the radiation intensity in the environment, temperature and humidity sensors monitor environmental temperature and humidity conditions, the IMU acquires the robot's attitude and motion status, and lidar and cameras construct environmental maps and identify obstacles or critical mission targets around the robot. Due to the high radiation levels in nuclear power plant environments, ordinary sensors may drift, increase noise, or fail due to prolonged exposure to high doses of radiation. Therefore, the multimodal sensors used must possess radiation-resistant characteristics. This can be achieved through the use of radiation-resistant material encapsulation, shielded circuit design, redundant sensor arrangements, or algorithm-based radiation noise filtering techniques. For example, radiation dose detectors can employ metal-oxide-semiconductor field-effect transistors (MOSFETs) or semiconductor PIN diodes to improve immunity to gamma-ray and neutron radiation.

[0067] Optionally, multi-sensor data fusion can be performed after data acquisition to improve the accuracy and stability of environmental perception. Fusion methods can employ Kalman filtering, particle filtering, or deep learning-based fusion algorithms to integrate data from various sensors and reduce potential errors from individual sensors. For example, the attitude data from the inertial measurement unit (IMU) may be affected by sensor drift, while the position information provided by LiDAR is relatively stable. Therefore, combining the two through extended Kalman filtering can yield more accurate robot attitude and position estimates. After data fusion, environmental state data is generated, including current radiation levels, temperature and humidity information, the robot's current position, attitude angles, and the distribution of surrounding obstacles. This ensures the robot can plan and execute subsequent tasks based on complete and accurate environmental information. This data can be stored locally and transmitted to a remote control center via wireless communication modules (such as Wi-Fi, 5G, or industrial-grade wireless protocols) for subsequent analysis and remote monitoring. Simultaneously, the environmental state data must meet real-time requirements to ensure the robot can respond quickly in dynamically changing environments. Typically, the update cycle for environmental state data can be set at the millisecond level to adapt to the complex operational needs of nuclear power plants.

[0068] S204: Generates robot task instructions based on environmental state data and target task using a large language model; task instructions include task path and operation instructions.

[0069] Optionally, a large language model trained on radiation environment data from nuclear power plants can be employed. This model not only understands task requirements but also optimizes robot movement strategies based on historical task experience and environmental perception information to improve task execution efficiency and stability. Specifically, the input to the large language model is environmental state data, including current radiation levels, ambient temperature, humidity, robot position, posture information, and the distribution of surrounding obstacles. This environmental state data serves as the foundation for robot task planning, ensuring that path planning avoids high-radiation areas, dynamically adjusts the travel route, and optimizes operational steps to reduce energy consumption and actuator workload. During task path generation, the large language model analyzes the specific requirements of inspection, maintenance, or emergency response tasks. For example, inspection tasks may involve the detection of specific pipelines, valves, or equipment areas; maintenance tasks may include component replacement or safety checks; and emergency response may involve rapid response to an accident area and the execution of specific interventions. Based on these task objectives, the model references historical task data to identify the optimal path and operational sequence, while also considering the impact of current environmental conditions on robot movement and task execution. For example, in inspection tasks, the model might prioritize traveling through low-radiation areas and plan the shortest path to reduce the robot's exposure time to radiation. In maintenance tasks, the model might optimize the robot arm's operational sequence to minimize unnecessary energy consumption and mechanical wear.

[0070] Furthermore, to ensure the rationality of the path, the large language model employs a reinforcement learning strategy for path optimization. By simulating the execution effects of different path schemes, the optimal scheme is selected to meet the task requirements. In this optimization process, the model considers not only the robot's own mobility capabilities, such as maximum speed, joint flexibility, and endurance, but also dynamic environmental factors, such as real-time changes in radiation dose, the movement trends of potential obstacles, and unforeseen circumstances that may occur during task execution. Through continuous iterative optimization, the model can generate more accurate, efficient, and safer task paths.

[0071] Beyond path planning, the large language model is also responsible for generating operational instructions to ensure the robot can accurately execute tasks. These instructions include the robot's dwell time at specific locations, the timing of sensor activation and deactivation, and the sequence of actuator movements. For example, in a valve inspection task, instructions might include steps such as aligning the camera with the target area, adjusting the robotic arm angle, and collecting data from the radiation detector. In emergencies, such as the detection of a high-radiation leak, instructions might adjust the robot's action strategy, such as rapidly evacuating the danger zone or switching to a backup control mode, to ensure the robot's safety. Throughout this process, the large language model's task planning and instruction generation capabilities enable the robot to autonomously adapt to different task requirements and optimize based on the actual environment, providing precise data support for subsequent task execution and control.

[0072] S206: Obtain irradiation interference prediction information by reasoning about task instructions through a large language model.

[0073] Optionally, the core task of the large language model is to perform in-depth analysis and reasoning on the robot's task path and operation instructions to predict potential sensor signal anomalies and actuator response delays or drifts that the robot may encounter during execution, thereby generating radiation interference prediction information. Specifically, the input data for the large language model consists of the task path and operation instructions, including the robot's movement path within a specific time period, the robotic arm's operation sequence, and sensor data acquisition plans. To ensure the robot can accurately execute tasks, the model analyzes the feasibility of these instructions in high-radiation environments and, combined with current environmental data, infers potential interference situations the robot may encounter. For example, in high-radiation areas, due to long-term exposure of electronic components to radiation, some sensor data may drift, increase noise, or even fail completely. Furthermore, radiation may affect the actuator's motor control system, leading to sluggish joint movements, decreased displacement accuracy, or increased execution deviations. Therefore, the large language model, based on existing environmental parameters and historical actuator data, predicts the potential risks these effects pose to task execution.

[0074] Furthermore, the large language model first performs pattern recognition using historical mission data from nuclear power plants to acquire robot behavior characteristics under historical radiation environments. For example, if the attitude angle measurement value of the inertial measurement unit deviates by a fixed amount over a long period at a certain dose level, the model can learn this deviation pattern and compensate in advance for it during future mission execution. In addition, the model can analyze the response characteristics of actuators under different radiation levels, such as the response hysteresis that motors may exhibit under high radiation conditions, thereby predicting potential motion errors during mission execution and adjusting control strategies in advance. The reasoning capability of the large language model is not limited to static data analysis; it can also perform dynamic evaluation by combining the robot's current real-time operating status. For example, if there is a significant deviation between the data fed back by sensors and the predicted values ​​during robot movement, the model can identify potential signal anomalies using anomaly detection algorithms and assess whether the anomaly will affect the execution of critical tasks in conjunction with the mission path. For example, if there is strong radiation interference in a certain area along the robot's path, causing large fluctuations in lidar data, the model can analyze the range of influence of this fluctuation and determine whether it is necessary to adjust the sensor usage strategy, such as switching to a vision sensor or using historical map data for path correction.

[0075] Optionally, the large language model can also analyze the potential impact of the external environment on actuators based on the task path. For example, when a robot needs to operate a valve, the model considers the impact of radiation on the joint motors of the robotic arm, infers the possible response lag of the motors, and, combined with the task execution time requirements, provides adjustment suggestions, such as increasing the operation redundancy time or optimizing the robotic arm's movement sequence, to ensure smooth task execution. Through the above analysis and reasoning, the large language model ultimately generates radiation interference prediction information, which includes the types of sensor anomalies that may be encountered at a specific task stage, the range of actuator response errors, possible control deviations, and suggested compensation strategies. This prediction information is used to adjust the robot's sensor signals and actuator control commands, compensate for the impact of the radiation environment, and ensure that the robot can accurately execute tasks and maintain long-term stable operation.

[0076] S208: During the process of the robot performing the target task, the task instructions at the current moment are adjusted according to the radiation interference prediction information obtained at the current moment, and the robot is controlled to perform the corresponding task operation according to the adjusted task instructions.

[0077] Optionally, based on the radiation interference prediction information generated by the large language model, the sensor signals and actuator control commands during task execution are dynamically corrected to ensure accurate task execution and stable operational performance even in high-radiation environments. Specifically, the raw sensor data is first corrected in real time to compensate for measurement errors caused by radiation interference. Since high-dose radiation may cause sensor signal drift, increased noise, or data loss, dynamic adjustments using data fusion and filtering algorithms are necessary. For example, if the inertial measurement unit experiences angular velocity shift in a strong radiation area, data fusion using extended Kalman filtering or particle filtering is performed, combining historical robot motion data and information from external environmental sensors, to generate more stable and reliable attitude information. Similarly, if the ranging error of the lidar data increases due to radiation, a deep neural network prediction model can be used to compare historical data with current measurements, identify error patterns, and compensate before the data is input into the control system.

[0078] After correcting the sensor data, the actuator control commands are adjusted to reduce the impact of radiation on the actuator. Radiation can cause time delays in motor control signals, changes in response curves, and even affect the stability of the drive system. Therefore, when correcting control commands, the radiation interference prediction information provided by the large language model is first analyzed to identify potential actuator response errors. For example, if a decrease in the response speed of the robotic arm at a specific joint is predicted, the amplitude of the control signal or the response curve can be adjusted to compensate for the hysteresis effect. Furthermore, during operation, the expected motion trajectory can be compared with the actual execution result, and the control gain can be adjusted using an adaptive control algorithm. Real-time feedback control ensures that the actuator can accurately complete the task.

[0079] Furthermore, to make task execution more robust, path planning will be dynamically adjusted to adapt to changes in the high-radiation environment. For example, when performing inspection or maintenance tasks, if sensor data correction indicates that the radiation level in a certain area is higher than expected, the robot's route will be adjusted to avoid high-risk areas without affecting the overall task objective. If an actuator response deviation is detected during operation that affects a critical task step, such as valve operation or equipment inspection, the task sequence can be adjusted or the execution instructions optimized to ensure the task can still be completed successfully. The corrected environmental state data and adjusted operation instructions will serve as new inputs, feeding back to the task execution system to guide the robot's subsequent perception and control decisions.

[0080] In the aforementioned nuclear power plant maintenance robot control method, measurement data from multimodal sensors is acquired, and environmental state data is obtained based on these measurements. A large language model is then used to generate task instructions for the robot based on the environmental state data and the target task. These instructions are then reasoned through using the large language model to obtain radiation interference prediction information. During the robot's execution of the target task, the task instructions are adjusted based on the radiation interference prediction information obtained at the current moment. The robot is then controlled to perform corresponding task operations according to the adjusted instructions. This method continuously optimizes the robot's execution strategy based on real-time environmental changes, precisely controlling the robot and thus improving the accuracy and reliability of task completion.

[0081] In an exemplary embodiment, a static sensor is used to collect environmental parameters of the radiation area, and a mobile sensor is used to collect the robot's state information. The step of obtaining environmental state data based on the measurement data includes: dynamically correcting the measurement data; constructing a multidimensional environmental data set based on the dynamically corrected measurement data; the multidimensional environmental data set includes time parameters, spatial location parameters, radiation level parameters, and equipment state parameters; obtaining environmental risk distribution information of the nuclear power plant based on the multidimensional environmental data set; the environmental risk distribution information includes the spatial gradient change trend of environmental parameters, the boundary characteristics of the target radiation area, the change trend of humidity anomaly areas, and the offset degree of equipment state parameters; mapping the measurement data to the physical space model of the nuclear power plant according to the installation location of the multimodal sensor; correcting abnormal data in the measurement data in the physical space model; and fusing the corrected measurement data with the environmental risk distribution information to obtain environmental state data.

[0082] Optionally, a hierarchical distributed sensor network architecture is adopted, deploying sensor clusters with differentiated protection levels in different functional areas of the nuclear power plant. This includes static sensors fixed around critical equipment and mobile sensors mounted on the robot itself, to cover different spatial ranges and ensure long-term data acquisition capabilities in high-radiation environments. Based on the sensor clusters, data on radiation intensity, temperature and humidity, and equipment status in the nuclear power plant's operating environment are collected. Combined with sensor location information, measurement timestamps, and historical measurement data, a multi-dimensional environmental data set encompassing time, space, radiation levels, and equipment status is constructed. The multi-dimensional environmental data set is analyzed to calculate the spatial gradient changes of environmental parameters, identify the boundary characteristics of high-radiation areas, the changing trends of abnormal temperature and humidity areas, and the offsets of equipment status parameters, generating environmental risk distribution information for the nuclear power plant. This environmental risk distribution information is then fused with real-time sensor data, integrating the risk levels of different areas, sensor measurement deviations, and equipment operating status to generate environmental perception data.

[0083] Specifically, due to significant differences in radiation intensity, temperature and humidity conditions, and equipment status across different areas within a nuclear power plant, sensor deployment requires differentiated protection levels. Static sensors are installed around critical equipment, fixed in specific locations such as pipe interfaces, electrical cabinets, and high-temperature pressure vessels, to monitor parameter changes in high-radiation environments over long periods. These static sensors are encapsulated in radiation-resistant materials and possess high anti-interference capabilities to ensure long-term stable operation. On the other hand, mobile sensors mounted on the robot body, including radiation detectors, infrared temperature sensors, humidity sensors, and inertial measurement units (IMUs), are used to collect real-time environmental data from different areas during inspections. These mobile sensors are mounted on the robot's external frame, the end effector of the robotic arm, or the motion chassis to ensure accurate acquisition of surrounding environmental data during robot movement and operation. The combination of these two types of sensors provides comprehensive coverage of both fixed and dynamically changing areas, maintaining stable data acquisition capabilities throughout extended operation.

[0084] During the data acquisition phase, all sensors acquire data at preset sampling intervals and transmit the data to the central control system via wired or wireless means. Each sensor's measurement data is accompanied by a corresponding timestamp, sensor type identifier, measurement unit, and sensor location information to ensure correct data source matching during subsequent data processing. To improve data accuracy, sensor data undergoes local preprocessing before transmission, including signal filtering, outlier detection, and error correction. For example, for sensors that may experience measurement drift due to radiation, historical data can be used for deviation compensation, and redundant sensor data can be used for cross-validation to eliminate measurements with significant errors.

[0085] Furthermore, the collected data, after local preprocessing by the sensors, is transmitted to the central control system to construct a multi-dimensional environmental dataset. This dataset includes time, spatial location, radiation level, and equipment status parameters, and all data undergoes format standardization for subsequent analysis. In the dataset, the time dimension records the changes in measurement values ​​of each sensor at different times; the spatial location dimension matches the physical area where the sensors are located; the radiation level dimension reflects the radiation dose in each area; and the equipment status parameter dimension correlates the relationship between equipment operating status and environmental changes. After obtaining the multi-dimensional environmental dataset, environmental parameter analysis is performed to calculate the environmental change trends and spatial gradients in different areas. First, spatial difference calculations are performed on the radiation intensity data of different areas to identify the boundary characteristics of high-radiation areas and analyze whether there are dynamic changes in radiation levels in these areas. For example, in areas with a high risk of pipeline leakage, an abnormally rising radiation level may indicate the presence of micro-cracks on the pipeline surface or damage to the shielding material. Temperature and humidity data are used for environmental anomaly detection to determine whether the operation of certain equipment is causing abnormal changes in the surrounding temperature and humidity. For example, during the operation of a high-temperature pressure vessel, the surrounding temperature curve should rise steadily. An abnormal, sudden drop may indicate a coolant leak. Furthermore, changes in equipment status parameters are also included in the analysis. For instance, abnormal changes in data such as motor current and hydraulic system pressure may reflect a decline in the operational stability of the equipment in a high-radiation environment. These analytical methods can generate risk distribution information for the nuclear power plant environment, which can be used for subsequent decision support.

[0086] After environmental risk distribution information is generated, it is fused with real-time sensor data to ensure data integrity and accuracy. During data fusion, data from different sources is time-aligned, synchronizing all sensor data according to timestamps to ensure correct correlation of data from the same moment. Then, spatial matching is performed, mapping data from different sensors to the physical space model of the nuclear power plant based on sensor location information to analyze discrepancies in data measured by different sensors within the same area. If significant deviations exist in the measurement data from multiple sensors within a certain area, the reliability of the data is further evaluated. For example, among multiple sensors near the same pipeline, if one sensor measures a radiation value significantly higher than others, it may indicate that the sensor has experienced temporary interference or has a hardware malfunction; in this case, data from surrounding sensors can be used to correct the reading.

[0087] To further improve data reliability, dynamic weighting is applied to data based on the risk level of different regions. In high-risk areas, such as those near reactors or regions with historically significant radiation fluctuations, the weight of sensor data is adjusted to be higher to ensure that measurements from these areas have a greater impact on the final decision. Conversely, the weight of data in low-risk areas can be reduced to minimize unnecessary computational burden. Furthermore, if a sensor's data remains abnormal for an extended period, its input can be temporarily disabled to avoid affecting the overall accuracy of the data.

[0088] In this embodiment, by dynamically correcting the measurement data, a multi-dimensional environmental data set is constructed based on the dynamically corrected measurement data. Environmental risk distribution information of the nuclear power plant is obtained based on the multi-dimensional environmental data set. According to the installation location of the multi-modal sensors, the measurement data is mapped to the physical space model of the nuclear power plant. In the physical space model, abnormal data in the measurement data is corrected. The corrected measurement data is fused with the environmental risk distribution information to obtain environmental state data. This can accurately obtain the real-time operating status of the nuclear power plant, improve the reliability of the input data, and thus improve the reliability of robot task control.

[0089] In an exemplary embodiment, the step of correcting abnormal data in the measurement data includes: obtaining a reference sensor whose radiation level meets preset conditions based on the spatial distribution of the multimodal sensors and historical measurement data; correcting the deviation of the measurement data of the remaining sensors other than the reference sensor based on the measurement data of the reference sensor to obtain reference measurement data; obtaining the radiation level change trend of different radiation areas based on the reference measurement data, and adjusting the current acquisition strategy of the multimodal sensors according to the rate of change of the radiation level change trend; re-acquiring updated measurement data of the multimodal sensors according to the adjusted acquisition strategy, and obtaining the radiation intensity change characteristics at different measurement times based on the updated measurement data and the reference measurement data; obtaining the interfered data in the reference measurement data based on the radiation intensity change characteristics and the operating status data of the nuclear power equipment, and dynamically correcting the interfered data based on the updated measurement data.

[0090] Optionally, by comparing and analyzing sensor data from different regions, reference sensors less affected by radiation can be selected. These sensors are typically located in well-shielded, low-radiation areas or have been specially designed to have higher radiation resistance. For example, multiple redundant sensors may be deployed in certain critical areas, and through long-term comparison, the sensor with the smallest error and fluctuation can be selected as the reference benchmark. Then, using the measurement values ​​of these reference sensors as a benchmark, the data of other sensors more susceptible to radiation interference are corrected. A dynamic deviation compensation method is used to calculate the degree of sensor drift under different radiation levels and correct its output data in real time, reducing the cumulative error of the radiation environment on the measurement results.

[0091] After data calibration, the radiation level variation trends in different regions are further calculated, and the data acquisition strategy is optimized based on the rate of change of these trends. By analyzing historical measurement data, the variation patterns in high-radiation areas are identified, such as whether the radiation level in a certain area exhibits periodic fluctuations or whether a sudden increase in radiation level is caused by changes in equipment operating status. In areas with drastic radiation changes, such as near reactors or around critical components affected by high temperatures, the sampling frequency of the sensors is automatically increased to ensure sufficient high-temporal-resolution data to reflect dynamic changes in radiation dose. In areas with stable radiation levels, such as shielded areas where multiple measurements have confirmed no significant changes, the data sampling frequency is appropriately reduced to reduce data redundancy and optimize energy consumption. For mobile sensors, the sampling strategy is dynamically adjusted in conjunction with the robot's task path, enabling more intensive data acquisition in high-risk areas and reduced measurements in low-risk areas to balance data quality and energy consumption.

[0092] Furthermore, after optimizing the data acquisition strategy, the adjusted dataset is compared over time to analyze the characteristics of radiation intensity changes in different time periods. By calculating the gradient of historical data changes, the normal fluctuation range of radiation levels in a certain area is determined, and compared with current measurements to detect any anomalies exceeding the normal range. Combined with equipment operating status information, further analysis is conducted to determine if external factors cause data deviations; for example, the start-up or shutdown of certain equipment may affect sensor measurements, and these changes should be correctly classified to avoid misjudgments. Pattern recognition technology is used to extract potentially interfered sensor data, and abnormal data points are further filtered out to prevent radiation-interferenced data from affecting the overall environmental perception results. To ensure the accuracy and stability of the final generated environmental data, abnormal data is dynamically corrected, and historical data is used for error compensation. If a sensor's measurement deviates significantly from historical trends or other sensor measurements, statistical regression or weighted data fusion methods can be used to adjust it, ensuring the stability of the output data. In addition, if some sensors exhibit long-term abnormal behavior, data interpolation or redundant sensor data replacement strategies can be used to reduce the impact of single-point failures on the environmental dataset. The filtered, optimized, and corrected environmental data sets can be used for robot task planning, path decision-making, and intelligent control to support automated inspection, equipment monitoring, and emergency response tasks within nuclear power plants.

[0093] In this embodiment, by correcting the deviation of the measurement data based on the reference sensor and obtaining the radiation level change trend in different radiation areas, the current acquisition strategy of the multimodal sensor is adjusted. Based on the radiation intensity change characteristics at different measurement times, the data subject to interference is dynamically corrected, which can improve the reliability of the measurement data and thus improve the reliability of robot task control.

[0094] In an exemplary embodiment, the step of generating robot task instructions based on environmental state data and target task using a large language model includes: obtaining multidimensional task input data and task context information of the large language model based on environmental state data and target task; the task context information includes radiation scene features, device state information, and task objectives; processing the multidimensional task input data using the large language model based on the task context information to obtain the target execution strategy corresponding to the target task; optimizing the target execution strategy using a multi-objective path planning algorithm to obtain candidate task paths; conducting radiation impact assessment on the candidate task paths, and obtaining target task paths that meet preset safety conditions based on the radiation impact assessment results; and generating robot task instructions based on the target task paths.

[0095] Optionally, to enable the large language model to be applied to the intelligent control of radiation-resistant nuclear power robots, the model undergoes specialized training for the radiation environment of nuclear power plants. This ensures that it can correctly analyze the characteristics of the radiation environment, predict possible interference situations, and provide reasonable task planning and control strategies for the robot. The training process begins by constructing a high-quality training dataset. This dataset should contain various information about the nuclear power plant environment, such as sensor data at different radiation levels, the robot's motion trajectory during task execution, actuator response, historical inspection and maintenance task execution records, sensor signal drift data caused by radiation interference, and robot operation logs under abnormal conditions. Data collection can be conducted in various ways, including extracting real task data from existing inspection and maintenance databases of nuclear power plants, obtaining simulated operation data from radiation-resistant robot experimental tests, and generating radiation impact data under different operating conditions through numerical simulation. Based on real data, data augmentation techniques can also be used to simulate the impact of changes in radiation intensity, temperature, humidity, and other factors on robot sensors and actuators, thereby improving the model's adaptability to environmental changes.

[0096] After data collection, the data is preprocessed, including noise removal, format conversion, and standardization, to ensure the consistency and usability of the data input to the model. Subsequently, a combination of supervised learning and reinforcement learning is used to train the DeepSeek model. In the supervised learning phase, labeled task execution data is used to train the model, enabling it to learn task planning, path optimization, and operation command generation capabilities. In the reinforcement learning phase, a simulation platform based on a nuclear power plant environment is built, allowing the model to perform autonomous decision-making training under different radiation scenarios, optimizing its ability to predict abnormal situations and its response strategies to emergencies. To improve the model's generalization ability, transfer learning techniques can be used to transfer the basic model trained in the general robot control domain to the specific scenario of a nuclear power plant. Fine-tuning is performed using a small amount of high-quality nuclear power plant data to enhance its environmental adaptability. Furthermore, adversarial training methods can be introduced to improve the model's robustness and anti-interference ability when facing the uncertainty brought by radiation interference. After training, the model is validated and optimized. First, in an offline testing environment, unseen nuclear power plant task data is used to evaluate the rationality of its task planning, the efficiency of path optimization, and the accuracy of anomaly prediction. The model was deployed on a robot in an experimental environment to simulate different radiation intensities, environmental changes, and unexpected situations, verifying its adaptability to actual task execution. After multiple rounds of optimization, the optimal model was selected for deployment, and a real-time update mechanism was provided to enable it to continuously learn new data and optimize task planning and control decisions during future nuclear power plant operations, adapting to ever-changing environmental and task requirements.

[0097] Optionally, when generating task instructions for the robot, the collected environmental state data is first transformed to serve as input to the model. Since the environmental parameters faced by the robot performing tasks in a nuclear power plant environment are complex and variable, a multi-dimensional task input data structure is constructed, including radiation scene characteristics, equipment status information, and task objectives. Radiation scene characteristics describe the radiation level of the current area, the trend of radiation dose changes, and the effectiveness of surrounding shielding facilities. Equipment status information includes the operating status of the equipment to be inspected or maintained, possible failure modes, and relevant historical data. The task objectives define the specific operations to be performed by the robot, such as inspection paths, maintenance tasks, or emergency response measures. Through this data transformation process, the model can accurately analyze the operational requirements in a high-radiation environment and form a reasonable basis for task planning.

[0098] After obtaining the task context, the model's task reasoning module is invoked to perform analogical analysis using historical task data to identify the optimal execution strategy. The model retrieves past task execution records, analyzes the paths, execution methods, and final task completion outcomes of the robot under similar environmental conditions, and thus extracts the optimal execution mode for the current task. Simultaneously, considering the dynamic changes in the radiation environment, the model also adaptively adjusts the original strategy based on the latest environmental data to ensure the robot can optimize its task path and execution method according to the current situation. For example, in maintenance tasks, if historical data indicates that a certain inspection method has a high failure rate at a specific radiation level, the robot's path will be automatically adjusted to avoid high-risk areas or optimize the operation sequence to improve the success rate of task completion.

[0099] After task inference is completed, the robot's task path is optimized to balance task efficiency, minimized radiation exposure, and equipment safety. First, using a 3D environmental model of the nuclear power plant and real-time radiation distribution data, the radiation exposure level of the robot on different paths is calculated. The robot's radiation resistance, energy consumption, and safety during movement are considered to optimize the task path. During optimization, the model executes a multi-objective path planning algorithm, scores different path schemes, and dynamically adjusts the path selection strategy based on task priority. For example, if a path can quickly reach the target area, but the cumulative radiation dose to the robot's critical electronic components may exceed the safety threshold, a slightly longer but lower radiation exposure path will be automatically selected to ensure the robot's stability during long-term operation. After candidate task paths are generated, the impact of radiation on the robot's hardware is further evaluated to select the optimal path that meets safety requirements. This evaluation process includes calculating the cumulative radiation dose of each key component of the robot and, combined with the radiation tolerance limit of the robot hardware, predicting the long-term impact of different paths on sensors, actuators, and computing units. Prioritize paths within the safety threshold range of the robot's electronic components, while avoiding situations that may cause sensor signal drift, actuator lag, or data errors in the computing unit due to prolonged exposure to high radiation. For example, when performing tasks in high-radiation areas, the robot may experience distortion in some sensor readings. Path selection should consider protecting these components to ensure that critical functions are not affected during task execution.

[0100] After the final path is determined, the low-level operation instructions required for the robot to execute the task are generated based on the task planning results. The model combines the robot's dynamics model, actuator control characteristics, and radiation protection strategies to refine the high-level task path into specific robot motion control instructions, including the robot's travel path, the robotic arm's trajectory, and the sequence of sensor data acquisition. Necessary compensation strategies are added to address sensor errors and actuator deviations that may occur in high-radiation environments. For example, if a segment of the task path may be affected by radiation interference, the sensor calibration parameters are pre-adjusted, or real-time error correction is introduced into the actuator control to ensure the robot can accurately execute the task. Throughout the task execution process, the robot's state is continuously monitored, and instructions are dynamically adjusted as needed to ensure the robot is always in the optimal task execution state.

[0101] In this embodiment, multi-dimensional task input data and task context information of a large language model are obtained based on environmental state data and target task. The target execution strategy corresponding to the target task is obtained based on the task context information. The target execution strategy is optimized through a multi-target path planning algorithm. The target task path that meets the preset safety conditions is obtained based on the radiation impact assessment results. Then, the robot's task instructions are generated, which can ensure that the target task path is the optimal path, thereby ensuring that the generated task instructions can accurately control the robot's execution state.

[0102] In an exemplary embodiment, the step of optimizing the target execution strategy and obtaining candidate task paths using a multi-objective path planning algorithm includes: establishing a path search space based on a three-dimensional environment model of a nuclear power plant; labeling risk elements on the three-dimensional environment model within the path search space; obtaining a multi-level path cost matrix based on the labeled risk elements and the target task; during the optimization of the target execution strategy using the multi-objective path planning algorithm, obtaining the cumulative radiation exposure on different task paths based on current radiation distribution data, and adjusting the weights in the multi-level path cost matrix based on the cumulative radiation exposure and the robot's radiation tolerance threshold; evaluating each task path based on the adjusted multi-level path cost matrix, and obtaining candidate task paths based on the path evaluation results.

[0103] Optionally, a complete path search space can be established using a 3D environmental model of the nuclear power plant, combined with the robot's current operating status and task requirements. Since the interior of a nuclear power plant contains various complex environmental factors, including high-radiation areas, physical obstacles, and dynamic risk points that may affect robot movement, these factors need to be labeled during path search. Therefore, a multi-level path cost matrix needs to be constructed to quantify the impact of different path choices on task execution. This matrix is ​​constructed by partitioning the nuclear power plant structure, stratifying each area according to radiation dose level, accessibility, and equipment distribution, and assigning different path cost coefficients to different levels of areas. For example, high-radiation areas have higher costs to encourage the robot to choose low-radiation paths, while areas close to critical equipment but beneficial to task completion have lower costs to ensure the robot can efficiently complete inspection, maintenance, or emergency tasks. Simultaneously, this path cost matrix also needs to consider the robot's own movement patterns, turning radius, accessibility, and other physical constraints to ensure that the planned path conforms to the limitations of the robot's movement capabilities.

[0104] During path search, real-time radiation distribution data is used to calculate the cumulative radiation exposure on different paths, and the results are compared with the robot's radiation tolerance threshold to adjust the cost weights of the path search algorithm. Since radiation levels inside nuclear power plants may fluctuate dynamically with changes in equipment operating status, a dynamic adjustment mechanism needs to be introduced during path planning to allow the robot to adapt to environmental changes at different task stages. For example, when the radiation level in a certain area exceeds the safety threshold, path planning needs to immediately recalculate possible alternative paths and reduce the feasibility score of paths in that area to reduce the time the robot is exposed to high radiation. Simultaneously, if certain paths can significantly improve task completion efficiency by traversing high-radiation areas in a short time, the robot's radiation tolerance capacity in a short period is comprehensively evaluated, and the feasibility score of these paths is appropriately increased within a safe range to ensure the task can be completed within the optimal time. During path evaluation, the radiation tolerance thresholds of key robot components are considered, such as the long-term radiation resistance of key components like the robot's computing unit, sensor modules, and actuators, to ensure that the robot does not experience functional degradation or failure due to excessive radiation exposure.

[0105] Furthermore, after path selection, the execution stability of candidate paths is evaluated, and the final task path is optimized in conjunction with equipment safety constraints. Since the robot's motion stability, actuator load distribution, and energy consumption directly affect task reliability when performing tasks in high-radiation environments, global optimization of different paths is necessary to ensure successful task completion. During path stability analysis, the robot's motion patterns on different paths are calculated, including parameters such as acceleration, turning radius, and travel smoothness, and the impact of these factors on the accuracy of robot sensor data acquisition and actuator control is evaluated. For example, complex paths may contain sharp turns and elevation changes, which can distort robot sensor data and affect the accuracy of environmental perception. Therefore, when optimizing paths, priority should be given to paths with smooth motion and reduced high-frequency vibrations. In addition, path optimization should be combined with task sequence to ensure that the robot can complete all task points with the least travel distance while ensuring optimal energy consumption to extend task execution time.

[0106] For example, for each path, calculate the waypoints. Local radiation exposure ,in, The intensity of the radiation source, the distance from the path point to the radiation source, and the time decay characteristics of the radiation source are determined by the following formula:

[0107]

[0108] in, Let be the initial intensity of radiation source j, expressed in Gy / s (gray per second), representing the radiation output power of the source at the initial moment; m is the total number of radiation sources, determined based on the known number of radiation sources in the nuclear power plant environment or the radiation sources detected in real time. path point The distance to radiation source j, in meters, can be measured by sensors or calculated based on known environmental data. The radiation attenuation index represents the rate at which radiation decreases with increasing distance, and its preferred value is 2. The time decay factor represents the rate at which the radiation intensity of a radiation source decays over time; the preferred value is [value missing]. It can be modified according to the physical characteristics of the radiation source; The cumulative effect time of the radiation source; t represents the current time.

[0109] Total radiation cost along the path The calculation formula is:

[0110]

[0111] Where n is the total number of path points on the path, that is, the number of all discrete path points traversed by the robot from the starting point to the target point. path point The task weight, which is dynamically adjusted based on the robot's radiation tolerance threshold at different task stages, is calculated using the following formula:

[0112]

[0113] in, This is the radiation tolerance threshold of a robot component, measured in Gy, and is generally determined by the tolerance characteristics of the robot hardware. To adjust the parameter, the preferred value is 0.2.

[0114] Total path length The sum of the Euclidean distances between adjacent path points on the path is calculated using the following formula:

[0115]

[0116] Where n is the total number of path points on the path, that is, the number of all discrete path points traversed by the robot from the starting point to the target point. path point Coordinates in a three-dimensional environment; path point Coordinates and path points in a 3D environment path point The next path point.

[0117] Overall task execution efficiency along the path The calculation formula is:

[0118]

[0119] Among them, the overall task execution efficiency of the path Used to measure the time efficiency of robot task execution on different paths; Representing path points arrive The Euclidean distance between them; For the robot at the path point arrive The expected speed of the path segment between the two is generally determined by the performance of the robot's power system, with the preferred value usually between 0 and 10, depending on the robot's ability to move on different terrains; For the robot at path points during task execution The expected dwell time; n is the total number of path points on the path, that is, the number of all discrete path points that the robot passes through from the starting point to the target point.

[0120] The optimal path is obtained using the following formula:

[0121]

[0122] in, The total radiation cost along the path; This represents the total path length. This is the path optimization coefficient, with a preferred value of 0.5.

[0123] In this embodiment, a path search space is established based on a three-dimensional environment model of a nuclear power plant. During the optimization of the target execution strategy through a multi-objective path planning algorithm, the weights in the multi-level path cost matrix are adjusted according to the current radiation distribution data, and path evaluation is performed on each task path. Candidate task paths are obtained based on the path evaluation results, which can accurately select the optimal robot task path and ensure that the robot can efficiently and safely complete inspection, maintenance or emergency handling tasks in a high-radiation environment.

[0124] In an exemplary embodiment, the step of obtaining radiation interference prediction information by reasoning about task instructions using a large language model includes: obtaining sensor error change trends at different task stages based on environmental state data and the robot's historical operation data; obtaining potential sensing errors of multimodal sensors based on sensor error change trends; performing reasoning analysis on potential sensing errors and task instructions using a large language model to obtain robot execution error prediction data; and generating radiation interference prediction information based on the error prediction data. The radiation interference prediction information includes potential sensor anomaly types, execution response errors, potential control deviations, and execution compensation strategies.

[0125] Optionally, based on environmental state data and historical robot operation data, the sensor signal characteristics under different radiation dose levels are analyzed. Since radiation can affect the measurement accuracy of sensors, leading to signal drift, increased noise, or data jumps, the error patterns of the sensors are modeled. Specifically, the measurement data of sensors under different radiation levels are compared with baseline data in low-radiation environments to extract error growth trends. For example, for inertial sensors, there may be cumulative drift in angular velocity and acceleration data, while temperature and humidity sensors may experience offset due to radiation aging. The patterns of these errors need to be obtained through long-term data statistics. Simultaneously, based on historical task data, the evolution patterns of sensor errors at different task stages are analyzed to predict potential sensing errors during upcoming tasks. For example, when the robot is about to enter a high-radiation area, the radiation change rate in that area is calculated, and combined with sensor error information from similar past tasks, the range of sensor errors during task execution is estimated. These predicted sensing error data will serve as input to the model to improve its ability to infer sensor anomalies in high-radiation environments.

[0126] Furthermore, after acquiring perception error information, the response characteristics of the actuator in a high-radiation environment are inferred by combining the task path and operation instructions. Since actuators typically include motors, joint mechanisms, and hydraulic or pneumatic drive systems, these components may experience problems such as lubricant degradation, material fatigue, and motor control delays due to radiation. Therefore, an execution error prediction model is established. The dynamic characteristics of the actuator under different radiation doses are analyzed. For example, a high-radiation environment may lead to increased joint friction, thus affecting the accuracy of reaching the target position. Based on the robot's historical execution data, the model analyzes the motion response characteristics of the actuator under different radiation levels, infers possible control deviations, and calculates the execution error at different path points. For example, if a task path point requires the robot to complete a turning operation with a specific angular velocity, and the motor response delay in a high-radiation environment leads to a decrease in the actual angular velocity, the error range is calculated, and it is determined whether this deviation will affect the task completion accuracy. In addition, for dynamic tasks, such as when a robot needs to grasp an object while moving, the control deviation of the actuator may lead to an increase in the error of the grasping position. Therefore, combined with task requirements, the possible execution error range is calculated, and the operation strategy is adjusted in advance if necessary.

[0127] After completing sensor error prediction and execution error inference, comprehensive radiation interference prediction information is further generated. This information is based on the aforementioned error data and, combined with task requirements, classifies and quantifies key error factors affecting task execution. First, based on sensor error data, the types of sensor anomalies that may occur at different task stages are identified, such as signal attenuation of optical sensors, offset of temperature and humidity sensors, and angular velocity drift of inertial sensors. Second, based on actuator error prediction data, the actuator response error range at each path point is calculated, such as speed errors and joint angle deviations that may occur in a high-radiation environment. Furthermore, by analyzing the combined impact of sensor and execution errors, the overall control deviation of the robot is calculated, such as path tracking error, task completion time deviation, and stability changes.

[0128] Optionally, compensation strategies include optimizing sensor data by filtering data from high-error sensors, increasing redundant data fusion, or adjusting the sampling frequency to reduce the impact of high noise; adjusting actuator control by adjusting control parameters, such as adding feedforward control or correcting torque compensation, to improve the accuracy of task execution; and optimizing task paths by adjusting path planning to avoid high-error areas for specific task path points, or optimizing the execution sequence of key task points to reduce the impact of error accumulation.

[0129] In this embodiment, potential sensing errors of multimodal sensors are obtained based on environmental state data and the robot's historical operation data. Execution error prediction data of the robot is obtained through a large language model, and radiation interference prediction information is generated. This enables accurate prediction of deviations during the robot's task execution, thereby timely optimization of the control precision of task execution.

[0130] In an exemplary embodiment, the step of adjusting the task instructions at the current moment based on the radiation interference prediction information obtained at the current moment includes: obtaining sensor error information of the current task stage based on the radiation interference prediction information obtained at the current moment; correcting the measurement data at the current moment based on the sensor error information; obtaining the control deviation at the current moment based on the corrected measurement data; and adjusting the task instructions at the current moment based on the control deviation.

[0131] Optionally, based on sensor data comparison and correction, it is also necessary to combine historical data from the robot's task execution process to analyze the sensor error variation trends at different task stages. Since the robot experiences different environmental conditions during task execution, such as changes in radiation intensity, fluctuations in temperature and humidity, and mechanical vibrations, the sensor error patterns may change at different task stages. Therefore, it is necessary to establish an error trend analysis model to extract the sensor error growth patterns during historical task execution and, combined with the current environmental conditions, predict potential perception errors during the upcoming task. For example, when the robot is performing an inspection task near a reactor area, historical data may indicate that the radiation level in that area is high, leading to an increase in the drift rate of the inertial sensor. Therefore, the drift amplitude of the sensor at this task stage can be calculated in advance, and corresponding compensation measures can be taken. Furthermore, in certain task stages, such as prolonged exposure to high-radiation environments, the sensor may experience increased measurement noise due to accumulated radiation dose. Historical data can be used to predict noise variation trends and preprocess the data during the data processing stage to improve the reliability of sensor data.

[0132] After predicting potential sensor errors, the sensor data processing strategy needs to be dynamically adjusted to ensure the accuracy of the robot's perceived data. For predictable error patterns, such as slow signal drift or increased noise, error compensation methods can be used, such as fitting an error curve using historical sensor data and correcting the current measurement. For unpredictable anomalies, such as sudden data jumps or short-term data loss, data replacement strategies can be used, such as supplementing with data from redundant sensors or using time-series interpolation to fill in missing data. Furthermore, to ensure the stability of the model's input data, various signal optimization methods are applied during the data processing stage, such as adaptive filtering or outlier removal, to reduce the magnitude of input data errors and improve the model's inference accuracy regarding sensor anomalies.

[0133] In this embodiment, by correcting the measurement data at the current moment based on the radiation interference prediction information and sensor error information obtained at the current moment, and adjusting the task instructions at the current moment, the stability of task execution can be ensured.

[0134] In an exemplary embodiment, the step of adjusting the task instructions at the current moment based on the control deviation includes: acquiring the robot's current pose data and current path deviation based on the corrected measurement data; updating the environmental state data based on the current pose data and current path deviation; and adjusting the task instructions at the current moment based on the updated environmental state data and control deviation.

[0135] Optionally, after the sensor data is corrected, the control deviation of the actuator is calculated based on the corrected data, and the actuator control signal is adjusted. Since components such as the actuator's power unit, mechanical joints, and drive circuits may be affected in high-radiation environments, leading to decreased control accuracy, it is necessary to retrospectively analyze the actuator's historical operating data and calculate the actuator's control error in conjunction with the current task requirements. First, the actuator's response characteristics under different radiation dose conditions are extracted, including joint motion hysteresis, torque changes, and trajectory deviation, and compared with the actuator's real-time motion state to calculate the current control error. For example, in a high-radiation environment, the actuator's motor drive may experience increased response time due to radiation damage to electronic components, causing the actual execution time of the execution command to lag behind the expected time, affecting the accuracy of the robot's motion. To address this, the actuator's control signal is adjusted based on error compensation parameters. For example, a feedforward control strategy is introduced during task execution, applying compensation signals in advance so that the actuator can pre-adjust before the delay occurs, reducing the hysteresis effect. Furthermore, the actuator's joint torque control is optimized, and the control gain is adjusted to maintain the actuator's motion stability in high-radiation environments. If the motion deviation of the actuator exceeds the allowable range of the task during the path execution, the execution path is recalculated based on the corrected sensor data, and the motion control parameters are adjusted to ensure that the robot can complete the task according to the expected trajectory.

[0136] Furthermore, after adjusting the actuator control signals, the environmental state of the robot is recalculated based on the corrected sensor data and the adjusted control signals, and the environmental state data is updated. The core of updating the environmental state data lies in using the corrected sensor measurements to calculate the robot's current position, posture, path execution error, and the status of surrounding equipment, ensuring the accuracy of environmental perception during task execution. First, the robot's posture is re-estimated using the corrected inertial sensor data, and combined with the robot's dynamics model, the deviation between the actual path and the predetermined path is compared to update the robot's current navigation state. For example, in an inspection task, if the robot's trajectory deviates slightly due to actuator response errors, the impact of this deviation on the subsequent path is calculated, and the deviation is recorded in the environmental state data for compensation in the next adjustment of operational instructions. In addition, if the corrected sensor data indicates abnormal environmental parameters, such as the actual radiation level in a high-radiation area being higher than the predicted value, the risk level of that area is updated in the environmental state data for optimization and adjustment during the task planning phase.

[0137] Based on updated environmental state data, the task path, execution order, and control parameters are adjusted to generate the final operation instructions. First, based on the latest changes in the environmental state data, the priority of task execution is calculated, and it is determined whether the path needs adjustment or the task execution order needs to be changed. For example, if the radiation level in a certain area exceeds a safe threshold, the task path may be adjusted so that the robot executes tasks in low-radiation areas first, and then returns to execute tasks in high-radiation areas after the radiation level decreases. Furthermore, if actuator control adjustments increase task execution time, the control strategy is optimized, such as increasing task parallelism so that multiple operation units execute tasks simultaneously to improve execution efficiency. During the adjustment of operation instructions, the robot's joint motion parameters are optimized in conjunction with the robot's dynamic characteristics, enabling the robot to complete task operations more smoothly and reducing motion instability caused by actuator adjustments.

[0138] Optionally, as the robot performs its tasks, environmental data is continuously updated, including changes in radiation intensity, corrections to sensor signals, actuator response adjustments, and deviations encountered during actual execution. This data is used not only for short-term real-time adjustments but also as learning samples input into a reinforcement learning framework to optimize future control decisions. The reinforcement learning algorithm constructs a reward mechanism, enabling the robot to gradually learn better control strategies from task execution feedback. For example, in an inspection task, if the robot successfully avoids high-radiation areas and completes the inspection of designated checkpoints, it receives a high reward value; if it fails to effectively adjust its path, resulting in unnecessary increased radiation exposure time, the reward value is reduced, prompting it to optimize path selection in future tasks.

[0139] In actuator control, reinforcement learning algorithms can optimize control parameters for robotic arms, wheeled, or tracked motion systems through iterative trials and feedback mechanisms. High-radiation environments can cause motor torque attenuation or increased joint friction, affecting actuator accuracy. Reinforcement learning continuously adjusts control parameters, such as optimizing drive current compensation or adjusting actuator acceleration limits, to ensure stability under different radiation environments. Furthermore, when the robot encounters task failures, such as failing to complete an operation within a specified time due to response lag, the reinforcement learning algorithm adjusts the control strategy, enabling it to predict the lag's impact in advance and take appropriate compensatory measures when performing similar tasks in the future.

[0140] In this embodiment, by adjusting the task instructions at the current moment using the corrected measurement data, it is possible to ensure that the robot can adapt to sensor errors and actuator deviations when performing tasks in a high-radiation environment, thereby improving the stability of task execution.

[0141] In an exemplary embodiment, the method further includes: obtaining the temporal variation characteristics of the sensor signal based on sensor error information; obtaining abrupt change data in the measurement data based on the temporal variation characteristics; smoothing the abrupt change data to obtain updated measurement data when the abrupt change data anomaly type is interference anomaly; and compensating for the abrupt change data to obtain updated measurement data when the abrupt change data anomaly type is loss anomaly.

[0142] Optionally, the radiation-induced drift amplitude is calculated by comparing the current sensor measurements with historical error models, and the error compensation algorithm is adjusted based on real-time environmental parameters. Since the error accumulation characteristics of sensors may vary with time and radiation dose, an adaptive error compensation method is employed to dynamically correct data from different sensors. For example, for inertial sensor drift, drift correction parameters are continuously updated during operation, and inertial data is compensated in real-time during each measurement cycle. For temperature and humidity sensor data drift, the compensation factor is dynamically adjusted based on the temperature change rate and cumulative radiation dose to keep the measured values ​​within a reasonable range. Furthermore, to prevent long-term deviations caused by error accumulation, a drift error correction threshold is set, and a data recalibration mechanism is triggered when the error exceeds the set range. During operation, the sensor's reference value is adjusted to ensure that the measurement data remains accurate and reliable even after long-term operation. For example, in inspection tasks, if the inertial sensor drift exceeds the acceptable range, the inertial data is recalibrated based on the position information of other sensors to reduce the impact of accumulated errors.

[0143] To further improve the stability of measurement data, the temporal variation characteristics of sensor signals are analyzed to detect and compensate for sudden signal jumps, increased noise, or short-term data loss. In practical applications, radiation may cause sudden anomalies in sensor data, such as sudden jumps in measurement values ​​or signal loss within a short period. If this is not handled, it may affect the robot's perception of the environmental state. Therefore, a temporal filtering method is used to detect and compensate for abnormal signals. For example, when a sudden change in the data of a certain sensor is detected, the historical measurement trend of that sensor is compared with the data of other redundant sensors to determine whether the jump is a real environmental change or an erroneous measurement caused by radiation interference. If it is determined to be an interference signal, a data interpolation method is used to smooth the abnormal data points to reduce the impact of signal jumps. At the same time, to improve the long-term reliability of sensor data, a sensor self-calibration mechanism is introduced. During the data compensation process, the measurement frequency and data weight of the sensors are adjusted according to environmental changes. For example, in high-radiation areas, the data weight of affected sensors is appropriately reduced, and the reliance on low-noise sensor data is increased, thereby improving the accuracy of the overall measurement data. In addition, during the execution of the task, the sampling interval of the sensor is dynamically adjusted to increase the sampling frequency in noisy environments to enhance the robustness of the signal, while the sampling frequency is reduced in stable environments to reduce unnecessary data processing overhead.

[0144] In this embodiment, by using the temporal variation characteristics of sensor signals to compensate for abrupt changes in the measurement data, the stability and reliability of the sensing data can be ensured, and accurate environmental status information can be provided for subsequent task path planning and actuator control.

[0145] In one exemplary embodiment, such as Figure 3 As shown, a control method for a nuclear power plant maintenance robot is provided, which includes the following steps:

[0146] (1) Dynamically correct measurement data: acquire measurement data from multimodal sensors, dynamically correct the measurement data, and construct a multidimensional environmental data set based on the dynamically corrected measurement data; the multidimensional environmental data set includes time parameters, spatial location parameters, radiation level parameters, and equipment status parameters; obtain environmental risk distribution information of the nuclear power plant based on the multidimensional environmental data set; the environmental risk distribution information includes the spatial gradient change trend of environmental parameters, the boundary characteristics of the target radiation area, the change trend of the humidity abnormal area, and the degree of deviation of equipment status parameters; map the measurement data to the physical space model of the nuclear power plant based on the installation location of the multimodal sensors; in the physical space model, obtain reference sensors whose radiation levels meet the preset conditions based on the spatial distribution of multimodal sensors and historical measurement data.

[0147] (2) Obtaining environmental status data: Based on the measurement data of the reference sensor, the measurement data of the remaining sensors other than the reference sensor are corrected for deviation to obtain reference measurement data; based on the reference measurement data, the radiation level change trend of different radiation areas is obtained, and the current acquisition strategy of the multimodal sensor is adjusted according to the rate of change of the radiation level change trend; the updated measurement data of the multimodal sensor is re-acquired according to the adjusted acquisition strategy, and the radiation intensity change characteristics at different measurement times are obtained according to the updated measurement data and the reference measurement data; based on the radiation intensity change characteristics and the operating status data of the nuclear power equipment, the interfered data in the reference measurement data is obtained, and the interfered data is dynamically corrected according to the updated measurement data; the corrected measurement data is fused with the environmental risk distribution information to obtain environmental status data.

[0148] (3) Large language model reasoning: Based on the environmental state data and the target task, obtain the multidimensional task input data and task context information of the large language model; the task context information includes radiation scene features, equipment state information and task target; based on the task context information, process the multidimensional task input data through the large language model to obtain the target execution strategy corresponding to the target task.

[0149] (4) Generating task instructions: Based on the three-dimensional environment model of the nuclear power plant, a path search space is established. Within the path search space, risk elements are labeled on the three-dimensional environment model. Based on the labeled risk elements and target tasks, a multi-level path cost matrix is ​​obtained. During the optimization of the target execution strategy through a multi-objective path planning algorithm, the cumulative radiation exposure on different task paths is obtained based on the current radiation distribution data. Based on the cumulative radiation exposure and the robot's radiation tolerance threshold, the weights in the multi-level path cost matrix are adjusted. Based on the adjusted multi-level path cost matrix, each task path is evaluated. Candidate task paths are obtained based on the path evaluation results. Radiation impact assessment is performed on the candidate task paths. Based on the radiation impact assessment results, target task paths that meet the preset safety conditions are obtained. Task instructions for the robot are generated based on the target task paths.

[0150] (5) Predicting irradiation interference information: Based on environmental state data and the robot's historical operation data, obtain the sensor error change trend under different task stages, and obtain the potential sensing error of multimodal sensors according to the sensor error change trend; use a large language model to reason and analyze the potential sensing error and task instructions to obtain the robot's execution error prediction data; generate irradiation interference prediction information according to the error prediction data; the irradiation interference prediction information includes potential sensor anomaly types, execution response errors, potential control deviations and execution compensation strategies.

[0151] (6) Sensor error compensation: Based on the sensor error information, the temporal variation characteristics of the sensor signal are obtained, and based on the temporal variation characteristics, the abrupt change data in the measurement data is obtained; when the abrupt change data is an interference abrupt change, the abrupt change data is smoothed to obtain updated measurement data; when the abrupt change data is a loss abrupt change, the abrupt change data is compensated to obtain updated measurement data.

[0152] (7) Dynamically adjust task instructions: During the process of the robot performing the target task, the sensor error information of the current task stage is obtained based on the radiation interference prediction information obtained at the current moment, and the measurement data at the current moment is corrected based on the sensor error information; the control deviation at the current moment is obtained based on the corrected measurement data, and the task instructions at the current moment are adjusted based on the control deviation; the robot is controlled to perform the corresponding task operation according to the adjusted task instructions.

[0153] In this embodiment, by acquiring measurement data from multimodal sensors and obtaining environmental state data based on the measurements, a large language model is used to generate task instructions for the robot based on the environmental state data and the target task. The large language model is then used to reason about the task instructions to obtain radiation interference prediction information. During the robot's execution of the target task, the task instructions at the current moment are adjusted based on the radiation interference prediction information obtained at the current moment, and the robot is controlled to perform corresponding task operations according to the adjusted task instructions. This allows for continuous optimization of the robot's execution strategy based on real-time environmental changes, precise control of the robot, and thus improved accuracy and reliability of task completion.

[0154] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.

[0155] Based on the same inventive concept, this application also provides a nuclear power plant maintenance robot control device for implementing the aforementioned nuclear power plant maintenance robot control method. The solution provided by this device is similar to the solution described in the above method; therefore, the specific limitations in one or more embodiments of the nuclear power plant maintenance robot control device provided below can be found in the limitations of the nuclear power plant maintenance robot control method described above, and will not be repeated here.

[0156] In one exemplary embodiment, such as Figure 4 As shown, a nuclear power plant maintenance robot control device is provided, comprising: a data acquisition module 10, an instruction generation module 20, an interference acquisition module 30, and a robot control module 40, wherein:

[0157] The data acquisition module 10 is used to acquire measurement data from multimodal sensors and obtain environmental state data based on the measurements. The multimodal sensors include multiple static sensors installed around nuclear power equipment in different radiation areas within the nuclear power plant and multiple mobile sensors installed on the robot. The measurement data includes environmental parameters and robot state information. The instruction generation module 20 is used to generate task instructions for the robot based on the environmental state data and the target task using a large language model. The task instructions include task paths and operation commands. The interference acquisition module 30 is used to reason about the task instructions using a large language model to obtain radiation interference prediction information. The robot control module 40 is used to adjust the task instructions at the current moment based on the radiation interference prediction information obtained at the current moment during the robot's execution of the target task, and control the robot to perform corresponding task operations according to the adjusted task instructions.

[0158] In an exemplary embodiment, a static sensor is used to collect environmental parameters of the radiation area, and a mobile sensor is used to collect the robot's state information. The data acquisition module 10 is also used to dynamically correct the measurement data and construct a multidimensional environmental data set based on the dynamically corrected measurement data. The multidimensional environmental data set includes time parameters, spatial location parameters, radiation level parameters, and equipment state parameters. The environmental risk distribution information of the nuclear power plant is obtained based on the multidimensional environmental data set. The environmental risk distribution information includes the spatial gradient change trend of environmental parameters, the boundary characteristics of the target radiation area, the change trend of the humidity anomaly area, and the offset degree of equipment state parameters. Based on the installation location of the multimodal sensor, the measurement data is mapped to the physical space model of the nuclear power plant. In the physical space model, abnormal data in the measurement data is corrected, and the corrected measurement data is fused with the environmental risk distribution information to obtain environmental state data.

[0159] In an exemplary embodiment, the data acquisition module 10 is further configured to: acquire a reference sensor whose radiation level meets preset conditions based on the spatial distribution of the multimodal sensors and historical measurement data; perform deviation correction on the measurement data of the remaining sensors other than the reference sensor based on the measurement data of the reference sensor to obtain reference measurement data; acquire the radiation level change trend of different radiation areas based on the reference measurement data, and adjust the current acquisition strategy of the multimodal sensors according to the rate of change of the radiation level change trend; reacquire updated measurement data of the multimodal sensors according to the adjusted acquisition strategy, and acquire the radiation intensity change characteristics at different measurement times based on the updated measurement data and the reference measurement data; acquire the interfered data in the reference measurement data based on the radiation intensity change characteristics and the operating status data of the nuclear power equipment, and dynamically correct the interfered data based on the updated measurement data.

[0160] In an exemplary embodiment, the instruction generation module 20 is further configured to obtain multi-dimensional task input data and task context information of a large language model based on environmental state data and the target task; the task context information includes radiation scene features, device state information, and task objectives; based on the task context information, the multi-dimensional task input data is processed by the large language model to obtain the target execution strategy corresponding to the target task; the target execution strategy is optimized by a multi-objective path planning algorithm to obtain candidate task paths; the radiation impact assessment of the candidate task paths is performed, and the target task path whose safety meets preset conditions is obtained based on the radiation impact assessment results; and the robot's task instructions are generated based on the target task path.

[0161] In an exemplary embodiment, the instruction generation module 20 is further configured to establish a path search space based on a three-dimensional environment model of a nuclear power plant, and within the path search space, label risk elements on the three-dimensional environment model; obtain a multi-level path cost matrix based on the labeled risk elements and the target task; during the optimization of the target execution strategy through a multi-objective path planning algorithm, obtain the cumulative radiation exposure on different task paths based on the current radiation distribution data, and adjust the weights in the multi-level path cost matrix based on the cumulative radiation exposure and the robot's radiation tolerance threshold; perform path evaluation on each task path based on the adjusted multi-level path cost matrix, and obtain candidate task paths based on the path evaluation results.

[0162] In an exemplary embodiment, the interference acquisition module 30 is further configured to acquire sensor error change trends under different task stages based on environmental state data and the robot's historical operation data; acquire potential sensing errors of multimodal sensors based on sensor error change trends; perform reasoning analysis on potential sensing errors and task instructions through a large language model to obtain robot execution error prediction data; and generate irradiation interference prediction information based on the error prediction data. The irradiation interference prediction information includes potential sensor anomaly types, execution response errors, potential control deviations, and execution compensation strategies.

[0163] In an exemplary embodiment, the robot control module 40 is further configured to obtain sensor error information for the current task stage based on the radiation interference prediction information obtained at the current moment, correct the measurement data at the current moment based on the sensor error information, obtain the control deviation at the current moment based on the corrected measurement data, and adjust the task instructions at the current moment based on the control deviation.

[0164] In an exemplary embodiment, the robot control module 40 is further configured to acquire the robot's current pose data and current path deviation based on the corrected measurement data; update the environmental state data according to the current pose data and current path deviation; and adjust the task instructions at the current moment according to the updated environmental state data and control deviation.

[0165] In an exemplary embodiment, the interference acquisition module 30 is further configured to acquire the temporal variation characteristics of the sensor signal based on the sensor error information, acquire abrupt change data in the measurement data based on the temporal variation characteristics, smooth the abrupt change data to obtain updated measurement data when the abrupt change data anomaly type is interference anomaly, and compensate the abrupt change data to obtain updated measurement data when the abrupt change data anomaly type is loss anomaly.

[0166] The various modules in the aforementioned nuclear power plant maintenance robot control device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of a computer device in hardware form or independent of it, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0167] In one exemplary embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 5As shown, the computer device includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interface. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The input / output interface is used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When the computer program is executed by the processor, it implements a control method for a nuclear power plant maintenance robot. The display unit is used to form a visually visible image and can be a display screen, projection device, or virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, a trackball, or a touchpad located on the computer device casing, or an external keyboard, touchpad, or mouse, etc. Those skilled in the art will understand that... Figure 5 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0168] In one exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor, when executing the computer program, implements the steps in the above-described method embodiments. In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps in the above-described method embodiments. In one embodiment, a computer program product is provided, including a computer program, which, when executed by a processor, implements the steps in the above-described method embodiments.

[0169] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one of relational databases and non-relational databases. Non-relational databases may include blockchain-based distributed databases, etc., and are not limited thereto. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited thereto. The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not contradict each other, it should be considered within the scope of this application. The above embodiments only illustrate several implementation methods of this application, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the appended claims.

Claims

1. A control method for a nuclear power plant maintenance robot, characterized in that, The method includes: The system acquires measurement data from multimodal sensors and dynamically corrects the data. Based on the dynamically corrected measurement data, a multidimensional environmental data set is constructed. This multidimensional environmental data set includes time parameters, spatial location parameters, radiation level parameters, and equipment status parameters. The multimodal sensors include multiple static sensors installed around nuclear power equipment in different radiation areas within the nuclear power plant and multiple mobile sensors installed on the robot. The measurement data includes environmental parameters and robot status information. The static sensors are used to collect environmental parameters from the radiation areas, and the mobile sensors are used to collect the robot's status information. The environmental risk distribution information of the nuclear power plant is obtained based on the multidimensional environmental data set; the environmental risk distribution information includes the spatial gradient variation trend of environmental parameters, the boundary characteristics of the target radiation area, the variation trend of the humidity anomaly area, and the degree of deviation of the equipment status parameters; Based on the installation location of the multimodal sensor, the measurement data is mapped to the physical space model of the nuclear power plant; In the physical space model, a reference sensor whose radiation level meets the preset conditions is obtained based on the spatial distribution of the multimodal sensors and historical measurement data; Based on the measurement data of the reference sensor, the measurement data of the remaining sensors other than the reference sensor are corrected for deviation to obtain reference measurement data; Based on the reference measurement data, the radiation level variation trend in different radiation areas is obtained, and the current acquisition strategy of the multimodal sensor is adjusted according to the rate of change of the radiation level variation trend. The updated measurement data of the multimodal sensor is reacquired according to the adjusted acquisition strategy. Based on the updated measurement data and the reference measurement data, the characteristics of radiation intensity change at different measurement times are obtained. Based on the radiation intensity change characteristics and the operating status data of the nuclear power equipment, the interfered data in the reference measurement data is obtained, the interfered data is dynamically corrected based on the updated measurement data, and the corrected measurement data is fused with the environmental risk distribution information to obtain environmental status data. Based on the environmental state data and the target task, the robot generates task instructions using a large language model; the task instructions include task paths and operation instructions. The task instructions are reasoned through the large language model to obtain radiation interference prediction information; During the process of the robot performing the target task, the task instructions at the current moment are adjusted according to the radiation interference prediction information obtained at the current moment, and the robot is controlled to perform the corresponding task operation according to the adjusted task instructions.

2. The method according to claim 1, characterized in that, The process of generating task instructions for the robot based on the environmental state data and the target task using a large language model includes: Based on the environmental state data and the target task, multidimensional task input data and task context information of the large language model are obtained; the task context information includes radiation scene features, equipment status information and task objectives. Based on the task context information, the multidimensional task input data is processed by the large language model to obtain the target execution strategy corresponding to the target task; The target execution strategy is optimized by using a multi-objective path planning algorithm to obtain candidate task paths; The candidate task paths are subjected to radiation impact assessment, and the target task paths that meet the preset safety conditions are obtained based on the radiation impact assessment results. The robot's task instructions are generated based on the target task path.

3. The method according to claim 2, characterized in that, The step of optimizing the target execution strategy using a multi-objective path planning algorithm to obtain candidate task paths includes: Based on a three-dimensional environmental model of a nuclear power plant, a path search space is established, and risk elements are labeled on the three-dimensional environmental model within the path search space. Based on the labeled risk factors and the target task, obtain a multi-level path cost matrix; In the process of optimizing the target execution strategy through a multi-objective path planning algorithm, the cumulative radiation exposure on different task paths is obtained based on the current radiation distribution data, and the weights in the multi-level path cost matrix are adjusted based on the cumulative radiation exposure and the robot's radiation tolerance threshold. Based on the adjusted multi-level path cost matrix, path evaluation is performed on each task path, and candidate task paths are obtained based on the path evaluation results.

4. The method according to claim 1, characterized in that, The step of reasoning about the task instructions using the large language model to obtain radiation interference prediction information includes: Based on the environmental state data and the robot's historical operation data, the sensor error change trend under different task stages is obtained, and the potential sensing error of the multimodal sensor is obtained according to the sensor error change trend. The robot's execution error prediction data is obtained by reasoning and analyzing the potential perception error and the task instructions using the large language model. Based on the error prediction data, irradiation interference prediction information is generated; the irradiation interference prediction information includes potential sensor anomaly types, execution response errors, potential control deviations, and execution compensation strategies.

5. The method according to claim 1, characterized in that, The step of adjusting the task instructions at the current moment based on the radiation interference prediction information obtained at the current moment includes: Based on the radiation interference prediction information obtained at the current moment, obtain the sensor error information for the current mission stage, and correct the measurement data at the current moment based on the sensor error information; Based on the corrected measurement data, the control deviation at the current moment is obtained, and the task instructions at the current moment are adjusted according to the control deviation.

6. The method according to claim 5, characterized in that, The adjustment of the task instructions at the current moment based on the control deviation includes: Based on the corrected measurement data, the current pose data and current path deviation of the robot are obtained; The environmental state data is updated based on the current pose data and the current path deviation; The task instructions at the current moment are adjusted based on the updated environmental status data and the control deviation.

7. The method according to claim 5, characterized in that, The method further includes: Based on the sensor error information, the temporal variation characteristics of the sensor signal are obtained, and based on the temporal variation characteristics, abrupt change data in the measurement data is obtained; If the anomaly type of the mutation data is interference anomaly, the mutation data is smoothed to obtain updated measurement data. If the anomaly type of the mutation data is a loss anomaly, the mutation data is compensated to obtain updated measurement data.

8. A control device for a nuclear power plant maintenance robot, characterized in that, The device is applied to the nuclear power plant maintenance robot control method according to any one of claims 1-7; the device comprises: The data acquisition module is used to acquire measurement data from multimodal sensors and acquire environmental status data based on the measurement data; the multimodal sensors include multiple static sensors installed around nuclear power equipment in different radiation areas inside the nuclear power plant and multiple mobile sensors installed on the robot; the measurement data includes environmental parameters and robot status information; The instruction generation module is used to generate task instructions for the robot based on the environmental state data and the target task using a large language model; the task instructions include task paths and operation instructions. The interference acquisition module is used to reason about the task instructions through the large language model to obtain irradiation interference prediction information; The robot control module is used to adjust the task instructions at the current moment based on the radiation interference prediction information obtained at the current moment during the process of the robot performing the target task, and control the robot to perform corresponding task operations according to the adjusted task instructions.

Citation Information

Patent Citations

  • Nuclear decommissioning operation master-slave manipulator control system and method

    CN119974008A

  • Nuclear industrial robot body intelligent system based on large model

    CN120228719A