Photovoltaic power station intelligent inspection task planning and scheduling method and system

By constructing an inspection task priority matrix and a regional collaborative decision-making network in photovoltaic power plants, and combining deep learning and multi-agent collaborative decision-making, intelligent allocation and dynamic optimization of inspection tasks in photovoltaic power plants are realized. This solves the problems of unreasonable allocation of inspection resources and information isolation in existing technologies, and improves inspection efficiency and system robustness.

CN121010174APending Publication Date: 2025-11-25DATANG KUNYU CLEAN ENERGY CO LTD +1
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202511186470.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-24
Publication Date
2025-11-25

AI Technical Summary

Technical Problem

Existing intelligent inspection systems for photovoltaic power plants lack the ability to differentiate equipment status assessment and cannot dynamically adjust inspection priorities, resulting in unreasonable allocation of inspection resources, blind spots or duplicate inspections, and failure to effectively utilize distributed edge computing for collaborative decision-making, making it difficult to achieve real-time perception of robot status and optimization of task execution.

Method used

By acquiring real-time monitoring data from edge computing nodes in photovoltaic power plants, an inspection task priority matrix is ​​constructed. A regional collaborative decision-making network is built using distributed edge computing nodes. Combined with deep learning and a multi-agent collaborative decision-making framework, the inspection task scheme is dynamically optimized, and the task priority is updated in real time, thereby realizing feedback on robot status and closed-loop optimization of tasks.

Benefits of technology

It improves the efficiency of inspection tasks and resource utilization, solves the communication pressure and computing bottleneck problems of traditional centralized scheduling, enhances the scalability and robustness of the system, and strengthens the intelligent operation and maintenance level and operating efficiency of photovoltaic power plants.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121010174A_ABST
    Figure CN121010174A_ABST
Patent Text Reader

Abstract

The invention provides a photovoltaic power station intelligent inspection task planning and scheduling method and system, and relates to the technical field of photovoltaic power stations, and the method comprises the steps: obtaining real-time monitoring data of edge calculation nodes to evaluate an equipment state, constructing an inspection task priority matrix, employing distributed edge calculation nodes to make a collaborative decision, and generating a global task distribution result, and dynamically optimizing an inspection task scheme based on robot state information, and continuously updating the priority matrix according to task execution state data. According to the invention, intelligent distribution and dynamic optimization of the inspection tasks of the photovoltaic power station are realized, and the inspection efficiency and the equipment operation and maintenance quality are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to photovoltaic power plant technology, and more particularly to a method and system for intelligent inspection task planning and scheduling of photovoltaic power plants. Background Technology

[0002] With the rapid development of photovoltaic (PV) power generation technology, PV power plants have been widely used globally as important clean energy facilities. PV power plants typically occupy large areas and have numerous pieces of equipment, requiring regular inspections to ensure normal operation and power generation efficiency. Traditional manual inspection methods are not only time-consuming and labor-intensive but also pose safety hazards in adverse weather conditions and complex terrain. With the development of artificial intelligence and robotics, intelligent inspection robots are increasingly being applied to the daily inspection work of PV power plants, significantly improving inspection efficiency and accuracy.

[0003] The intelligent inspection system for photovoltaic power plants mainly adopts a centralized task planning and scheduling method, which uniformly allocates inspection tasks to various robots through a central control system. However, this method has the following defects and shortcomings in practical applications: Centralized task planning lacks the ability to differentiate the status of equipment in different areas of the power plant and cannot dynamically adjust the inspection priority according to the real-time changes in equipment status, resulting in unreasonable allocation of inspection resources and key areas not being inspected in a timely manner.

[0004] Existing technologies fail to fully leverage the advantages of distributed edge computing for collaborative decision-making. Information is isolated between different inspection areas, making it difficult to form a globally optimal task allocation scheme. This can easily lead to blind spots or duplicate inspections in large photovoltaic power plants.

[0005] Existing inspection systems lack the ability to perceive the robot's status in real time and dynamically optimize the task execution process. They cannot adaptively adjust based on the robot's battery status, execution capability, and task execution effect, resulting in low efficiency of inspection tasks and difficulty in establishing an effective feedback optimization mechanism to continuously improve inspection strategies. Summary of the Invention

[0006] This invention provides a method and system for intelligent inspection task planning and scheduling of photovoltaic power plants, which can solve the problems in the prior art.

[0007] A first aspect of this invention provides a method for intelligent inspection task planning and scheduling of photovoltaic power plants, comprising: The system acquires real-time monitoring data collected by multiple edge computing nodes in a photovoltaic power station, evaluates the equipment status of each area within the photovoltaic power station based on the real-time monitoring data, and generates equipment status evaluation results. Based on the equipment status assessment results, an inspection task priority matrix is ​​constructed. A regional collaborative decision-making network is built through distributed edge computing nodes. Each edge computing node generates a local task allocation strategy according to the inspection task priority matrix of its own region. Edge computing nodes in adjacent regions interact with their respective local task allocation strategies and perform iterative optimization to generate a global inspection task allocation result. An initial inspection task plan is generated based on the global inspection task allocation result. The status information of the robot executing the inspection task is obtained in real time. The initial inspection task plan is dynamically optimized based on the status information. The priority value of each task is updated by inputting the robot's remaining power and execution capability into the inspection task priority matrix. The tasks are then redistributed based on the updated priority values ​​to generate an optimized inspection task plan. The optimized inspection task scheme is sent to the robot execution unit. During the execution process, task execution status data, including task execution efficiency and completion quality, is collected. The task execution status data is used as a new influencing factor to feed back to the edge computing node to update the inspection task priority matrix for continuous optimization of subsequent task allocation.

[0008] Based on the equipment status assessment results, an inspection task priority matrix is ​​constructed. A regional collaborative decision-making network is built through distributed edge computing nodes. Each edge computing node generates a local task allocation strategy based on the inspection task priority matrix of its region, including: The equipment status assessment results are constructed into a status assessment result set. The weights of equipment status features are determined based on equipment runtime, fault frequency, and maintenance records. Temporal correlation analysis and multi-dimensional feature fusion are performed on the equipment status features in the status assessment result set. A deep learning network is used to dynamically predict and identify abnormal patterns of the fused equipment status features. An inspection task priority matrix is ​​generated based on the prediction results and abnormal identification results. Based on the equipment status information in the inspection task priority matrix, a task execution time window and a task priority level are generated. The task execution time window and the task priority level are then input into the deep learning network for multi-objective constraint evaluation to obtain the optimal state action value. Based on the optimal state action value, calculate the spatial clustering coefficient and task correlation degree between devices. According to the spatial clustering coefficient and the task correlation degree, divide the photovoltaic power station into multiple regions with task collaboration characteristics, and deploy edge computing nodes in the divided regions. Based on the edge computing node, the inspection task priority matrix and historical task execution feedback data are input into the deep learning network. A task execution strategy is generated by combining online strategy optimization and offline experience learning. The task execution strategy is updated by gradient, and the updated gradient is used to guide the optimization direction of the strategy, and a local task allocation strategy is output.

[0009] Based on the optimal state action values, the spatial clustering coefficient and task correlation degree between devices are calculated. According to the spatial clustering coefficient and the task correlation degree, the photovoltaic power station is divided into multiple regions with task collaboration characteristics. Edge computing nodes are deployed in the divided regions, including: Based on the optimal state action value, the spatial clustering coefficient between devices is calculated, and the equipment tasks in the photovoltaic power station are divided into emergency tasks, periodic tasks, and daily tasks. Long Short-Term Memory Network is used to perform load prediction on the historical operation data to obtain the task load prediction values ​​of the emergency task, the periodic task and the daily task. Based on the task load prediction values, corresponding priority weights are set for different types of tasks. The task correlation degree between devices is calculated based on the priority weight. The spatial clustering coefficient and the task correlation degree are weighted and combined to obtain the task collaboration comprehensive evaluation index. A task collaboration execution queue is constructed based on the task collaboration comprehensive evaluation index. Tasks with similar comprehensive evaluation indicators for task collaboration in the task collaboration execution queue are clustered into the same region to obtain multiple regions with task collaboration characteristics. The spatial distribution centroid of the devices in each region is calculated, and the spatial distribution centroid is determined as the deployment location of the edge computing node. The task processing status of the edge computing nodes is monitored in real time. When a change is detected in the task collaborative execution queue, the region division results are dynamically updated based on the changed task collaborative execution queue, and the deployment location of the edge computing nodes is adjusted accordingly.

[0010] An initial inspection task plan is generated based on the global inspection task allocation result. The status information of the robot executing the inspection task is acquired in real time. The initial inspection task plan is dynamically optimized based on the status information, including: A task dependency graph is constructed based on the global inspection task allocation results. The task nodes in the task dependency graph are connected based on spatiotemporal constraints. The task nodes in the task dependency graph are used to extract features using a graph neural network. The neighborhood information of the task nodes is aggregated through the graph neural network to obtain task relevance features. An initial inspection task plan is generated based on the task relevance features. The state information of the robot performing the inspection task is collected and constructed into a robot state vector. The robot state vector is processed by a multi-agent collaborative decision-making framework, wherein the multi-agent collaborative decision-making framework calculates a state decision value based on a local reward value and realizes collaborative decision-making of multiple robots through the state decision value, generating a decision result that considers the collaboration of multiple robots. A task adjustment function is constructed based on the robot's state vector, and the initial inspection task scheme is dynamically optimized using the calculation results of the task adjustment function and the decision results of the multi-agent collaborative decision-making framework.

[0011] The robot's state vector is processed through a multi-agent cooperative decision-making framework, wherein the framework calculates state decision values ​​based on local reward values, and uses these state decision values ​​to achieve collaborative decision-making among multiple robots, generating decision results that consider multi-robot collaboration, including: A dual-Q learning network is constructed, which simultaneously maintains an online evaluation network and a target value network through a parallel structure. Based on the online evaluation network, gradient descent is used to extract features from the robot's state vector to obtain node feature vectors. The dynamic dependencies between robots are calculated based on the node feature vectors to obtain attention weight coefficients. Based on the dual Q learning network, message passing is constructed according to the attention weight coefficients, and the node feature vector is updated to obtain the updated node feature vector. The target value network is used to calculate the combination of task completion state, distance state and collaborative state according to the updated node feature vector to obtain the local reward value. Based on the dual Q learning network, the state action value is updated by minimizing the output difference between the online evaluation network and the target value network according to the local reward value. The policy parameters are then dynamically adjusted according to the state action value to generate a collaborative decision-making result for multiple robots.

[0012] By inputting the robot's remaining battery power and execution capability as influencing factors into the inspection task priority matrix to update the priority values ​​of each task, and then reallocating tasks based on the updated priority values, an optimized inspection task scheme is generated, including: The remaining battery power and execution capability of the robot performing the inspection task are obtained. The remaining battery power and execution capability are used to construct robot state influencing factors. The robot state influencing factors are weighted and calculated to obtain a comprehensive robot state index. The robot's overall state index is input into the task priority adjustment function, and the priority values ​​in the inspection task priority matrix are updated through the task priority adjustment function. Multi-objective constraints are constructed based on the updated inspection task priority matrix. The multi-objective constraints are input into the task priority adjustment function, and the constraint violation degree is calculated through the task priority adjustment function. The constraint violation degree and the updated inspection task priority matrix are combined to construct a task allocation matrix. Under the condition of satisfying the single task unique allocation constraint, the task allocation matrix is ​​solved to obtain the optimal task allocation scheme. Based on the task-robot matching relationship corresponding to the optimal task allocation scheme, and combined with the priority values ​​in the updated inspection task priority matrix, the inspection tasks are reassigned to generate an optimized inspection task scheme that takes into account robot state constraints.

[0013] The task execution status data is used as a new influencing factor to feed back to the edge computing nodes to update the inspection task priority matrix, which is used for continuous optimization of subsequent task allocation, including: The task execution status data is transmitted to the edge computing node as a new influencing factor. The priority value in the inspection task priority matrix is ​​updated based on the task execution status data in the edge computing node to obtain the updated inspection task priority matrix that reflects the current task execution status. Based on the priority values ​​in the updated inspection task priority matrix, subsequent inspection tasks to be executed are reallocated to achieve continuous optimization of the inspection task scheme.

[0014] A second aspect of the present invention provides an intelligent inspection task planning and scheduling system for photovoltaic power plants, comprising: The first unit is used to acquire real-time monitoring data collected by multiple edge computing nodes in the photovoltaic power station, evaluate the equipment status of each area in the photovoltaic power station based on the real-time monitoring data, and generate equipment status evaluation results. The second unit is used to construct an inspection task priority matrix based on the equipment status assessment results, and to construct a regional collaborative decision-making network through distributed edge computing nodes. Each edge computing node generates a local task allocation strategy according to the inspection task priority matrix of its own region. Edge computing nodes in adjacent regions interact with their respective local task allocation strategies and perform iterative optimization to generate a global inspection task allocation result. The third unit is used to generate an initial inspection task plan based on the global inspection task allocation result, obtain the status information of the robot executing the inspection task in real time, dynamically optimize the initial inspection task plan based on the status information, update the priority value of each task by inputting the robot's remaining power and execution capability into the inspection task priority matrix, and redistribute the tasks based on the updated priority values ​​to generate an optimized inspection task plan. The fourth unit is used to distribute the optimized inspection task plan to the robot execution unit. During the execution process, it collects task execution status data including task execution efficiency and completion quality, and feeds the task execution status data as a new influencing factor to the edge computing node to update the inspection task priority matrix for continuous optimization of subsequent task allocation.

[0015] A third aspect of the present invention provides an electronic device, comprising: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.

[0016] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0017] The beneficial effects of this application are as follows: The intelligent inspection task planning and scheduling method for photovoltaic power plants provided by this invention uses edge computing nodes to collect real-time monitoring data to evaluate equipment status and construct an inspection task priority matrix, thereby realizing intelligent task allocation and decision-making, and effectively improving the execution efficiency and resource utilization of inspection tasks.

[0018] This method uses distributed edge computing nodes to build a regional collaborative decision-making network, realizing task collaborative optimization between adjacent regions. It solves the communication pressure and computing bottleneck problems faced by traditional centralized task scheduling, and improves the scalability and robustness of the system.

[0019] By acquiring robot status information in real time for dynamic task optimization and using task execution status data as feedback factors to update the priority matrix, closed-loop optimization of inspection tasks is achieved, improving the accuracy and adaptability of task allocation and further enhancing the intelligence level and operational efficiency of photovoltaic power plant operation and maintenance. Attached Figure Description

[0020] Figure 1 This is a flowchart illustrating the intelligent inspection task planning and scheduling method for photovoltaic power plants according to an embodiment of the present invention. Figure 2 A bar chart showing the performance comparison analysis of edge computing node deployment methods according to embodiments of the present invention; Figure 3 This is a schematic diagram comparing the task completion rates of different task priority allocation methods in embodiments of the present invention. Detailed Implementation

[0021] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0022] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0023] Figure 1 This is a flowchart illustrating the intelligent inspection task planning and scheduling method for photovoltaic power plants according to an embodiment of the present invention. Figure 1 As shown, the method includes: The system acquires real-time monitoring data collected by multiple edge computing nodes in a photovoltaic power station, evaluates the equipment status of each area within the photovoltaic power station based on the real-time monitoring data, and generates equipment status evaluation results. Based on the equipment status assessment results, an inspection task priority matrix is ​​constructed. A regional collaborative decision-making network is built through distributed edge computing nodes. Each edge computing node generates a local task allocation strategy according to the inspection task priority matrix of its own region. Edge computing nodes in adjacent regions interact with their respective local task allocation strategies and perform iterative optimization to generate a global inspection task allocation result. An initial inspection task plan is generated based on the global inspection task allocation result. The status information of the robot executing the inspection task is obtained in real time. The initial inspection task plan is dynamically optimized based on the status information. The priority value of each task is updated by inputting the robot's remaining power and execution capability into the inspection task priority matrix. The tasks are then redistributed based on the updated priority values ​​to generate an optimized inspection task plan. The optimized inspection task scheme is sent to the robot execution unit. During the execution process, task execution status data, including task execution efficiency and completion quality, is collected. The task execution status data is used as a new influencing factor to feed back to the edge computing node to update the inspection task priority matrix for continuous optimization of subsequent task allocation.

[0024] In one optional implementation, an inspection task priority matrix is ​​constructed based on the equipment status assessment results, and a regional collaborative decision-making network is built through distributed edge computing nodes. Each edge computing node generates a local task allocation strategy based on the inspection task priority matrix of its region, including: The equipment status assessment results are constructed into a status assessment result set. The weights of equipment status features are determined based on equipment runtime, fault frequency, and maintenance records. Temporal correlation analysis and multi-dimensional feature fusion are performed on the equipment status features in the status assessment result set. A deep learning network is used to dynamically predict and identify abnormal patterns of the fused equipment status features. An inspection task priority matrix is ​​generated based on the prediction results and abnormal identification results. Based on the equipment status information in the inspection task priority matrix, a task execution time window and a task priority level are generated. The task execution time window and the task priority level are then input into the deep learning network for multi-objective constraint evaluation to obtain the optimal state action value. Based on the optimal state action value, calculate the spatial clustering coefficient and task correlation degree between devices. According to the spatial clustering coefficient and the task correlation degree, divide the photovoltaic power station into multiple regions with task collaboration characteristics, and deploy edge computing nodes in the divided regions. Based on the edge computing node, the inspection task priority matrix and historical task execution feedback data are input into the deep learning network. A task execution strategy is generated by combining online strategy optimization and offline experience learning. The task execution strategy is updated by gradient, and the updated gradient is used to guide the optimization direction of the strategy, and a local task allocation strategy is output.

[0025] When constructing a condition assessment result set from the equipment condition assessment results, for each piece of equipment in the photovoltaic power station, its operating parameters are collected and a condition feature vector is formed. Taking a 100MW photovoltaic power station as an example, it includes 5000 photovoltaic modules, 200 inverters, 50 transformer substations, and 10 step-up transformers. The condition features of each piece of equipment include: output power (kW), operating temperature (°C), voltage (V), current (A), vibration value (mm / s), noise value (dB), etc., totaling 15 dimensions of condition features. For example, the condition feature vector of inverter No. 1 is [45.2, 42.5, 380.2, 118.6, 2.5, 58.3...], indicating that the current output power is 45.2kW, the operating temperature is 42.5°C, etc.

[0026] The weights of equipment status characteristics are determined based on equipment runtime, failure frequency, and maintenance records. Runtime is divided into five levels according to the equipment's commissioning time: 0-1 year = 0.2, 1-3 years = 0.4, 3-5 years = 0.6, 5-8 years = 0.8, and over 8 years = 1.0. Failure frequency is divided according to the average number of failures per year: 0 failures = 0.2, 1-2 failures = 0.4, 3-5 failures = 0.6, 6-10 failures = 0.8, and over 10 failures = 1.0. Maintenance records are divided according to the maintenance cycle: scheduled maintenance = 0.3, occasional delayed maintenance = 0.6, and frequent skipped maintenance = 0.9. The weights of these three indicators are calculated by weighted averaging, with weight coefficients of 0.3, 0.5, and 0.2 respectively. For example, the state characteristic weight of an inverter that has been running for 4 years, with an average of 3 failures per year and a good maintenance record is 0.3×0.6+0.5×0.6+0.2×0.3=0.54.

[0027] Time-series correlation analysis and multi-dimensional feature fusion were performed on the equipment condition characteristics in the condition assessment result set. The time-series correlation analysis employed a sliding window method with a window size of 7 days and a step size of 1 day, calculating the trend and fluctuation amplitude of characteristic values. For power output characteristics, the average, standard deviation, maximum, minimum, and rate of change over 7 days were calculated to form a time-series description vector. For temperature characteristics, the deviation of the intraday temperature curve from the standard curve was calculated. For vibration and noise characteristics, the frequency and duration of outliers were calculated. Multi-dimensional feature fusion adopted a hierarchical approach. First, highly correlated features were locally fused, such as fusion of voltage, current, and power into electrical performance indicators; and fusion of temperature, vibration, and noise into physical condition indicators. Then, all indicators were globally fused according to equipment characteristics to obtain a comprehensive condition characterization.

[0028] A deep learning network is used to dynamically predict and identify abnormal patterns in the fused device status features. The deep learning network adopts a dual-stream structure: a prediction stream and an identification stream. The prediction stream uses a recurrent neural network, containing two LSTM layers with 128 units each, to predict the device's status changes over the next 7 days. The identification stream uses a convolutional neural network, containing three convolutional layers (kernel sizes of 3×3, 5×5, and 7×7, and filter numbers of 32, 64, and 128, respectively) and two fully connected layers (with 256 and 128 nodes, respectively), to identify abnormal patterns. The network is trained using one year's worth of historical device operation data, with a batch size of 64, a learning rate of 0.001, and an Adam optimizer. The prediction results include the trend of device status changes and the probability of failure. For example, the temperature prediction result for a certain inverter is that it will gradually increase by 2.3℃ over the next 7 days, with a failure probability of 0.15. The anomaly identification results classify the device status into 5 categories: normal (0), slightly abnormal (1), moderately abnormal (2), severely abnormal (3), and about to fail (4).

[0029] A priority matrix for inspection tasks is generated based on the prediction and anomaly identification results. Priority calculation comprehensively considers four factors: equipment importance, anomaly level, failure probability, and inspection cycle. Equipment importance is determined according to equipment type: 0.9 for step-up transformers, 0.8 for box-type transformers, 0.7 for inverters, and 0.5 for photovoltaic modules. The anomaly level is directly derived from the identification results of the deep learning network. The failure probability is the probability value output by the prediction stream. The inspection cycle is determined based on equipment type and historical inspection records, as the ratio of the time since the last inspection to the standard cycle. The weights of the four factors are 0.3, 0.3, 0.25, and 0.15, respectively. A weighted average is used to obtain the inspection task priority value, ranging from 0 to 1. Rows in the priority matrix represent equipment IDs, columns represent feature combinations, and matrix elements are the corresponding priority values. For example, the priority value for box-type transformer No. 3 is 0.3×0.8 + 0.3×2 / 4 + 0.25×0.35 + 0.15×1.2 = 0.601.

[0030] The task execution time window and task priority level are generated based on the equipment status information in the inspection task priority matrix. The time window is determined based on the predicted failure probability and anomaly level: for equipment with a failure probability > 0.8 or anomaly level 4, the time window is within 24 hours; for equipment with a failure probability between 0.5 and 0.8 or anomaly level 3, the time window is within 48 hours; for equipment with a failure probability between 0.3 and 0.5 or anomaly level 2, the time window is within 72 hours; for equipment with a failure probability between 0.1 and 0.3 or anomaly level 1, the time window is within one week; for equipment with a failure probability < 0.1 and anomaly level 0, the time window is within two weeks. The task priority level is divided into five levels according to the priority value: priority value > 0.8 is level 1 (highest), 0.6-0.8 is level 2, 0.4-0.6 is level 3, 0.2-0.4 is level 4, and < 0.2 is level 5 (lowest).

[0031] The task execution time window and task priority are input into a deep learning network for multi-objective constraint evaluation. Constraints include inspection resource limitations (number of available robots, working time), inspection path constraints (shortest path, optimal energy consumption), and inter-task dependencies (completion status of prerequisite tasks). The evaluation employs reinforcement learning. The state space represents the current set of devices to be inspected and the state of available resources, while the action space represents the selectable inspection order and resource allocation scheme. The reward function comprehensively considers task completion rate, resource utilization rate, and time window satisfaction. Iterative optimization using the Q-learning algorithm yields the optimal state-action value, representing the expected long-term reward of choosing a specific action in the current state.

[0032] Based on optimal state action values, spatial clustering coefficients and task correlations among devices are calculated. The spatial clustering coefficients consider the physical location and type similarity of the devices, calculated using Euclidean distance and type matching degree. Task correlations consider functional associations and fault correlations among devices, determined using historical fault propagation data and an expert knowledge base. The spatial clustering coefficients and task correlations are weighted and combined with weights of 0.4 and 0.6 respectively to obtain a comprehensive correlation coefficient. A spectral clustering algorithm is used to construct a similarity matrix based on the comprehensive correlation coefficients, dividing the photovoltaic power station into eight regions. Devices within each region exhibit high task collaboration characteristics. Edge computing nodes are deployed at the spatial center of each region, configured with an 8-core CPU, 16GB of RAM, and 256GB of storage.

[0033] Based on edge computing nodes, the priority matrix of inspection tasks and historical task execution feedback data are input into a deep learning network. Online policy optimization is achieved using a policy gradient algorithm, while an offline experience base is introduced to provide pre-training knowledge. The policy network structure is a four-layer fully connected network with 128, 256, and 128 hidden layer nodes, and the output layer represents the action probability distribution. During online optimization, after each inspection task is completed, the task execution time, resource consumption, and task completion quality are recorded, the reward value is calculated, and the policy network is updated. Offline experience learning uses three months of historical task execution data to construct an experience replay buffer with a size of 10,000 records. Each update randomly samples 128 records for batch learning. The policy gradient update uses the REINFORCE algorithm with a learning rate of 0.0005 and a discount factor of 0.95. Each edge computing node independently generates a local task allocation strategy for its region, including task execution order, resource allocation scheme, and execution time arrangement.

[0034] The above methods have enabled efficient allocation of inspection tasks based on regional collaborative decision-making networks. Compared with traditional centralized scheduling methods, task response time has been reduced by 42.3%, resource utilization has increased by 37.8%, and inspection coverage has increased by 18.6%, effectively improving the intelligent operation and maintenance level and efficiency of photovoltaic power plants.

[0035] In one optional implementation, the spatial clustering coefficient and task correlation degree between devices are calculated based on the optimal state action value. The photovoltaic power station is then divided into multiple regions with task collaboration characteristics based on the spatial clustering coefficient and the task correlation degree. Edge computing nodes are deployed in the divided regions, including: Based on the optimal state action value, the spatial clustering coefficient between devices is calculated, and the equipment tasks in the photovoltaic power station are divided into emergency tasks, periodic tasks, and daily tasks. Long Short-Term Memory Network is used to perform load prediction on the historical operation data to obtain the task load prediction values ​​of the emergency task, the periodic task and the daily task. Based on the task load prediction values, corresponding priority weights are set for different types of tasks. The task correlation degree between devices is calculated based on the priority weight. The spatial clustering coefficient and the task correlation degree are weighted and combined to obtain the task collaboration comprehensive evaluation index. A task collaboration execution queue is constructed based on the task collaboration comprehensive evaluation index. Tasks with similar comprehensive evaluation indicators for task collaboration in the task collaboration execution queue are clustered into the same region to obtain multiple regions with task collaboration characteristics. The spatial distribution centroid of the devices in each region is calculated, and the spatial distribution centroid is determined as the deployment location of the edge computing node. The task processing status of the edge computing nodes is monitored in real time. When a change is detected in the task collaborative execution queue, the region division results are dynamically updated based on the changed task collaborative execution queue, and the deployment location of the edge computing nodes is adjusted accordingly.

[0036] When calculating the spatial clustering coefficient between devices based on optimal state action values, it is necessary to obtain the spatial coordinates and operating status data of all devices within the photovoltaic power station. Spatial coordinates are obtained through equipment installation drawings or GPS positioning, accurate to the meter level; operating status data includes parameters such as device operating status codes, output power, temperature, and voltage. Taking a 100MW photovoltaic power station as an example, the station contains 5000 photovoltaic panels, 50 inverters, 20 transformer substations, 5 step-up substations, and 1 centralized control center. The specific method for calculating the spatial clustering coefficient is as follows: For any two devices i and j, calculate the Euclidean distance dij between them; set a distance threshold D (generally 100 meters); when dij is less than D, the two devices are considered to have a spatial clustering relationship; the spatial clustering coefficient Cij is defined as: 1 minus the ratio of the distance between devices to the threshold D, multiplied by the device importance coefficient. The device importance coefficient is determined according to the device type: 0.9 for inverters, 0.8 for transformer substations, and 0.5 for photovoltaic panels. For example, if the distance between inverter A and transformer B is 30 meters, then their spatial clustering coefficient is (1-30 / 100)×0.9×0.8=0.504.

[0037] The tasks within the photovoltaic power station are categorized into emergency tasks, periodic tasks, and daily tasks, based on factors such as response time requirements, execution frequency, and impact scope. Emergency tasks refer to anomaly handling tasks requiring immediate response, such as equipment failures and safety hazards, with a response time requirement of less than 5 minutes. Periodic tasks refer to regularly performed inspection and maintenance tasks, such as equipment parameter collection and performance evaluation, with execution cycles on an hourly or daily basis. Daily tasks refer to routine monitoring tasks, such as environmental data collection and cleanliness checks, with a high execution frequency but lower priority. In this power station, emergency tasks account for 15% of the total tasks, periodic tasks account for 45%, and daily tasks account for 40%.

[0038] A Long Short-Term Memory (LSTM) network was used to predict the computational and communication loads of different task types based on historical execution data. The LSM network structure consisted of an input layer, an LSTM layer (containing 64 LSTM units), a fully connected layer, and an output layer. Input features included the task execution time, data volume, CPU utilization, and memory usage over the past 10 days; the output was the predicted task load for the next 24 hours. The network was trained using the Adam optimizer with a learning rate of 0.001, a batch size of 32, and 100 training epochs. Historical data was sampled every 5 minutes, containing at least 30 days of execution records. Prediction results showed that the average computational load for urgent tasks was 250 MIPS (millions of instructions per second), and the communication load was 5 Mbps; the average computational load for periodic tasks was 120 MIPS, and the communication load was 3 Mbps; and the average computational load for routine tasks was 50 MIPS, and the communication load was 1 Mbps.

[0039] Priority weights are assigned to different task types based on task load prediction values. These weights consider three factors: task urgency, resource consumption, and business importance. Specifically, task urgency, resource consumption, and business importance are quantified as values ​​between 0 and 1. Urgent tasks have an urgency of 0.9, resource consumption of 0.7, and business importance of 0.9; periodic tasks have an urgency of 0.6, resource consumption of 0.5, and business importance of 0.7; and routine tasks have an urgency of 0.3, resource consumption of 0.2, and business importance of 0.4. The priority weight is a weighted average of these three factors, with weights of 0.5, 0.2, and 0.3 respectively. The calculated priority weight is 0.84 for urgent tasks, 0.62 for periodic tasks, and 0.31 for routine tasks.

[0040] The task correlation degree between devices is calculated based on priority weights. This degree reflects the interdependence of task execution between devices. The calculation method is as follows: For any two devices i and j, the task interaction frequency fij (times / hour) between them is recorded; according to the type of interactive task, the frequency is multiplied by the corresponding priority weight; the weighted frequencies of different task types are accumulated to obtain the task correlation degree Rij between the devices. For example, if the emergency task interaction frequency between inverter A and transformer B is 0.5 times / hour, the periodic task interaction frequency is 2 times / hour, and the daily task interaction frequency is 5 times / hour, then their task correlation degree is 0.5×0.84+2×0.62+5×0.31=3.17. The spatial clustering coefficient and task correlation degree are weighted and combined to obtain the comprehensive task collaboration evaluation index. The calculation method is: the comprehensive task collaboration evaluation index Iij equals the spatial clustering coefficient Cij multiplied by a weight of 0.4 plus the task correlation degree Rij multiplied by a weight of 0.6, then divided by the maximum value of the task correlation degree (used for normalization). For inverter A and transformer B, assuming the maximum task correlation value is 5, their task coordination comprehensive evaluation index is (0.504×0.4+3.17×0.6) / 5=0.4406.

[0041] A task collaboration execution queue is constructed based on a comprehensive evaluation index for task collaboration. Tasks in the queue are sorted from highest to lowest according to the evaluation index. The queue construction process is as follows: traverse all equipment pairs and their associated tasks; calculate the comprehensive evaluation index for task collaboration of each equipment pair; add equipment pairs and their associated tasks with an index value greater than a threshold of 0.3 to the queue; and sort the tasks in the queue in descending order of the evaluation index. In this power plant, the task collaboration execution queue contains approximately 500 tasks, including 75 emergency tasks, 225 periodic tasks, and 200 daily tasks.

[0042] Tasks with similar comprehensive evaluation indicators for task collaboration in the task collaboration execution queue are clustered into the same region using the K-means clustering algorithm. The value of K is determined based on the power plant scale and equipment distribution; in this power plant, K=8. Clustering features include the comprehensive evaluation indicators for task collaboration and the spatial coordinates of the equipment. The clustering process is as follows: K initial cluster centers are randomly selected; the distance from each task to each cluster center is calculated, and the task is assigned to the category of the nearest cluster center; the cluster centers for each category are recalculated (based on the average of the evaluation indicators and the average of the spatial coordinates); these two steps are repeated until the clustering results stabilize or the maximum number of iterations (50) is reached. After clustering, eight regions with task collaboration features are obtained, each containing approximately 625 devices and 60-70 tasks.

[0043] The spatial centroid of devices within each region is calculated, and this centroid is used to determine the deployment location of edge computing nodes. The centroid calculation method is as follows: obtain the spatial coordinates (xi, yi) and device importance weights wi of all devices within the region; calculate the weighted average coordinates as the centroid coordinates. For example, the centroid coordinates of region 1 are (356.8m, 421.5m). There is a transformer substation nearby, so an edge computing node can be deployed next to this substation. The hardware configuration of the edge computing nodes is determined based on the region's task characteristics: for regions with a high proportion of urgent tasks, high-performance processors (such as Intel i7-9700, 8 cores 3.0GHz) and large-capacity memory (32GB) are configured; for regions with mainly periodic and daily tasks, medium-performance processors (such as Intel i5-9400, 6 cores 2.9GHz) and adequate memory (16GB) are configured.

[0044] Real-time monitoring of edge computing node task processing, including CPU utilization, memory usage, task queue length, and average response time. Monitoring frequency is once per minute. When any indicator exceeds a preset threshold (CPU utilization > 85%, memory usage > 80%, queue length > 20, average response time > 500ms) for more than 5 minutes, an update of the task collaboration execution queue is triggered. The update method is: recalculating the task correlation of devices within the affected area; updating the comprehensive evaluation indicators for task collaboration; and reconstructing the task collaboration execution queue. When a change in the task collaboration execution queue is detected to exceed 20%, the area division results are dynamically updated based on the changed task collaboration execution queue, and the deployment locations of edge computing nodes are adjusted accordingly. For example, if monitoring reveals that the CPU utilization of the edge computing node in area 3 consistently exceeds 90%, triggering an update process, and after recalculation, it is found that the task volume in this area has increased by 35%, it is decided to divide this area into two sub-areas, 3a and 3b, and deploy edge computing nodes in these sub-areas respectively, with new node locations of (245.6m, 318.2m) and (312.3m, 378.9m).

[0045] Using the above method, tasks within the photovoltaic power station are rationally divided based on collaborative characteristics, and the deployment locations of edge computing nodes are optimized, effectively improving data processing efficiency and task execution efficiency while reducing communication latency and energy consumption. Practical application results show that, compared to traditional centralized computing, this method reduces the average task response time by 43.2%, network bandwidth usage by 37.5%, and computing resource utilization by 28.6%.

[0046] Figure 2This is a bar chart comparing the performance of edge computing node deployment methods according to embodiments of the present invention. The chart illustrates the optimization effects of three different deployment architectures (traditional centralized deployment, fixed feedback partitioning deployment, and dynamic collaborative feature deployment) on five key performance indicators. Traditional centralized deployment (white bars) represents the conventional centralized management mode, fixed feedback partitioning deployment (dark gray bars) reflects a semi-dynamic optimization scheme, and dynamic collaborative feature deployment (light gray grid bars) reflects an innovative adaptive architecture. From the data performance, the dynamic collaborative feature deployment scheme demonstrates significant advantages across all indicators: system stability improvement reaches 23.2%, far exceeding the traditional method's 8.2% and fixed partitioning's 14.7%; task processing efficiency improves to 18.8%, significantly better than the traditional method's 7.5% and fixed partitioning's 11.3%; communication latency reduction reaches 18.5%, exceeding the traditional method's 9.6% and fixed partitioning's 13.4%; resource utilization improves to 19.9%, significantly leading the traditional method's 8.9% and fixed partitioning's 14.8%; and it also achieves a good performance of 14.6% in task response time optimization, better than the traditional method's 6.8% and fixed partitioning's 9.9%. Overall, the data highlights the comprehensive advantages of the dynamic collaborative feature deployment scheme in system performance optimization, especially its outstanding performance in the two key indicators of improved system stability and resource utilization, fully demonstrating the practical value and technological advancement of this scheme in real-world application scenarios.

[0047] In one optional implementation, an initial inspection task plan is generated based on the global inspection task allocation result, and the status information of the robot executing the inspection task is acquired in real time. Dynamic optimization of the initial inspection task plan based on the status information includes: A task dependency graph is constructed based on the global inspection task allocation results. The task nodes in the task dependency graph are connected based on spatiotemporal constraints. The task nodes in the task dependency graph are used to extract features using a graph neural network. The neighborhood information of the task nodes is aggregated through the graph neural network to obtain task relevance features. An initial inspection task plan is generated based on the task relevance features. The state information of the robot performing the inspection task is collected and constructed into a robot state vector. The robot state vector is processed by a multi-agent collaborative decision-making framework, wherein the multi-agent collaborative decision-making framework calculates a state decision value based on a local reward value and realizes collaborative decision-making of multiple robots through the state decision value, generating a decision result that considers the collaboration of multiple robots. A task adjustment function is constructed based on the robot's state vector, and the initial inspection task scheme is dynamically optimized using the calculation results of the task adjustment function and the decision results of the multi-agent collaborative decision-making framework.

[0048] When constructing a task dependency graph based on the global inspection task allocation results, each inspection task is represented as a node in the graph. Node attributes include task ID, location coordinates, estimated execution time, priority, and task type. For example, in a photovoltaic power station inspection scenario with 10 inspection tasks, task T1 has the following attributes: ID 1, location coordinates (125.3m, 87.6m), estimated execution time 25 minutes, priority 0.85, and task type: component detection. Connections are established between task nodes in the task dependency graph based on spatiotemporal constraints. Specifically, the spatial distance and temporal correlation between tasks are calculated. When the spatial distance between two tasks is less than a preset threshold of 50 meters and their time windows overlap, a connection is established between the corresponding nodes. The connection weight is determined based on the distance and the degree of time overlap. The calculation formula is: the reciprocal of the distance is multiplied by 0.7, and the ratio of the overlap time to the total time window is multiplied by 0.3. The two are added together to obtain the final weight. For example, if the distance between tasks T1 and T3 is 32 meters and the time overlap is 0.6, then the connection weight is 0.7×(1 / 32)+0.3×0.6=0.4.

[0049] Feature extraction of task nodes in a task dependency graph is performed using a graph neural network. A two-layer graph convolutional structure is employed, with each layer involving feature propagation and feature transformation. In the feature propagation stage, for task node i, features from all its neighboring nodes j are aggregated using a weighted summation method, with weights determined by the connection weights between nodes. In the feature transformation stage, the aggregated features undergo a non-linear transformation using a fully connected layer and a ReLU activation function. The first graph convolutional layer has an input feature dimension of 5 (the number of task attributes) and an output dimension of 32; the second graph convolutional layer has an input dimension of 32 and an output dimension of 64. The graph convolutional kernel parameters are obtained through pre-training using historical inspection task data, with the loss function being the mean squared error between predicted and actual execution times. By aggregating neighborhood information of task nodes through the graph neural network, task-related features are obtained, reflecting the dependencies between tasks and the rationality of their execution order. For example, for task T1, the first 5 elements of the extracted relevance feature vector are [0.82, 0.56, 0.23, 0.91, 0.45], which represent the degree of relevance to other tasks.

[0050] An initial inspection task plan is generated based on task relevance characteristics, and a task scheduling algorithm is used to determine the optimal execution order. The specific steps are as follows: For each robot, based on the task set assigned to it, the total execution time and resource consumption of all task permutations are calculated. Execution time includes the execution time of the task itself and the travel time between tasks, with travel time calculated based on the robot's speed and the distance between tasks. Resource consumption mainly considers power consumption, estimated based on travel distance and task type. The total execution time and resource consumption are weighted and combined as evaluation indicators, with weights of 0.6 and 0.4 respectively. The permutation with the smallest evaluation indicator value is selected as the initial inspection task plan for that robot. For example, robot R1 is assigned tasks T1, T3, and T8, and the generated initial plan is the execution order T3--T1--T8, with an estimated total execution time of 95 minutes and power consumption of 35%.

[0051] The system collects status information of the robot performing inspection tasks, including real-time location, remaining battery power, movement speed, navigation status, and task progress. Data is collected every 5 seconds via sensors and monitoring modules on the robot. Location information is obtained via GPS module with an accuracy of ±0.5 meters; remaining battery power is read from the battery management unit with an accuracy of ±1%; movement speed is calculated using an odometer with an accuracy of ±0.05 meters per second; navigation status includes discrete states such as normal, obstacle avoidance, and out of control; task progress is calculated based on the ratio of completed inspection points to the total number of inspection points. The state information is constructed as a robot state vector. The vector dimension is 10 state parameters for each robot. For example, the state vector of robot R2 at a certain moment is [152.6, 98.3, 0.5, 73, 0.8, 0.2, 1, 4, 65, 0], which respectively represent x coordinate 152.6 meters, y coordinate 98.3 meters, z coordinate 0.5 meters, remaining battery 73%, current speed 0.8 m / s, acceleration 0.2 m / s², navigation status normal (1), current task ID is 4, task completion rate 65%, and abnormal status code 0.

[0052] The robot's state vector is processed through a multi-agent collaborative decision-making framework, which employs a deep reinforcement learning method. The specific implementation is as follows: A dual Q-learning network is constructed, comprising an online evaluation network and a target value network. The online evaluation network has the following structure: 10 nodes in the input layer, 64, 128, and 64 nodes in the three hidden layers respectively, and an output layer with an action space dimension of 8. The target value network has the same structure but a lower parameter update frequency. Training is performed using an Adam optimizer with a batch processing of 32 samples and a learning rate of 0.001. Dynamic dependencies between robots are calculated based on an attention mechanism. The attention weight coefficients are obtained by concatenating the feature vectors of any two robots using a single-layer neural network. A message passing mechanism is constructed to update node features. Each robot obtains weighted feature information from other robots based on its attention weight and then fuses it. Local reward values ​​are calculated, including three parts: task completion state (weight 0.5), distance state (weight 0.3), and collaborative state (weight 0.2). The output difference between the online evaluation network and the target value network is minimized to update the state action values. The loss function includes temporal difference error and L2 regularization. An ε-greedy strategy is used to dynamically adjust the policy parameters, with an initial ε value of 0.3 that linearly decays to 0.05. The decision-making results of the multi-agent framework include the optimal action selection for each robot, such as moving forward, turning, and adjusting speed.

[0053] A task adjustment function is constructed based on the robot's state vector. This function considers the robot's battery level change rate, task execution efficiency, and environmental interference factors. The battery level change rate is calculated by dividing the difference in remaining battery power between two consecutive samples by the time interval. The task execution efficiency is evaluated by the ratio of actual task progress to expected progress. Environmental interference factors are determined based on sensor data, such as light intensity and wind speed. The task adjustment function takes the robot's state vector and current task information as inputs and outputs a task adjustment coefficient, with a value range of [-1, 1]. Negative values ​​indicate task demotion or cancellation, while positive values ​​indicate task priority improvement. For example, in a certain calculation, robot R3's battery level change rate is -0.15% / second (higher than the expected -0.1% / second), the task execution efficiency is 0.85, and the environmental interference factor is 0.2. The combined calculation yields a task adjustment coefficient of -0.3.

[0054] The initial inspection task plan is dynamically optimized using the calculation results of the task adjustment function and the decision results of the multi-agent collaborative decision-making framework. The specific steps are as follows: obtain the task adjustment coefficient and decision result for each robot; reassign or cancel tasks with a task adjustment coefficient below the threshold of -0.5; adjust the task execution order according to the decision results, prioritizing tasks with higher decision scores; recalculate the task path and select the energy-optimal route; update task execution parameters, such as detection speed and sampling frequency; generate a new inspection task plan and distribute it to each robot for execution. Compared to the initial plan, the optimized task plan reduces the average task completion time by 15.8%, power consumption by 12.3%, and task coverage by 8.5%. For example, the initial plan T3--T1--T8 for robot R1 is optimized to T1--T3--T8, reducing the total execution time from 95 minutes to 82 minutes and power consumption from 35% to 31%.

[0055] In practical applications, this method performs dynamic optimization every 30 seconds to ensure that the inspection tasks can adapt to changes in robot status and environment. When a robot's battery is detected to be severely low (less than 1.5 times the power required to return to the charging station), an emergency task adjustment is triggered, reassigning the remaining tasks to other robots. When high-priority anomalies (such as component overheating or equipment failure) are detected, new inspection tasks are immediately generated and inserted into the current task queue. Through continuous dynamic optimization, this method can significantly improve the intelligence level and efficiency of photovoltaic power plant inspection work.

[0056] In one optional implementation, the robot state vector is processed by a multi-agent cooperative decision-making framework, wherein the multi-agent cooperative decision-making framework calculates state decision values ​​based on local reward values, and realizes cooperative decision-making of multiple robots through the state decision values, generating a decision result that considers multi-robot cooperation, including: A dual-Q learning network is constructed, which simultaneously maintains an online evaluation network and a target value network through a parallel structure. Based on the online evaluation network, gradient descent is used to extract features from the robot's state vector to obtain node feature vectors. The dynamic dependencies between robots are calculated based on the node feature vectors to obtain attention weight coefficients. Based on the dual Q learning network, message passing is constructed according to the attention weight coefficients, and the node feature vector is updated to obtain the updated node feature vector. The target value network is used to calculate the combination of task completion state, distance state and collaborative state according to the updated node feature vector to obtain the local reward value. Based on the dual Q learning network, the state action value is updated by minimizing the output difference between the online evaluation network and the target value network according to the local reward value. The policy parameters are then dynamically adjusted according to the state action value to generate a collaborative decision-making result for multiple robots.

[0057] The multi-agent collaborative decision-making framework processes robot state vectors that include multi-dimensional features such as position coordinates, velocity, remaining battery power, and current task completion rate. In a practical application scenario, a photovoltaic power station deploys five inspection robots. Each robot's state vector has 12 dimensions, including spatial position coordinates (x, y, z), velocity (vx, vy, vz), remaining battery percentage, CPU utilization, memory usage, current task ID, task completion percentage, and task priority. For example, the state vector of robot R1 at a certain moment is [25.4, 78.2, 0.5, 0.3, 0.4, 0, 85, 23, 15, 3, 47, 0.8], representing x-coordinate 25.4 meters, y-coordinate 78.2 meters, z-coordinate 0.5 meters, x-direction velocity 0.3 m / s, y-direction velocity 0.4 m / s, z-direction velocity 0 m / s, remaining battery power 85%, CPU utilization 23%, memory usage 15%, task ID 3, task completion rate 47%, and task priority 0.8.

[0058] When constructing the dual Q-learning network, two neural networks with identical structures but independent parameters are created: an online evaluation network and a target value network. The online evaluation network has the following structure: 12 nodes in the input layer, corresponding to the state vector dimension; 64 nodes in the first hidden layer using the ReLU activation function; 128 nodes in the second hidden layer using the ReLU activation function; 64 nodes in the third hidden layer using the ReLU activation function; and 8 nodes in the output layer, corresponding to 8 combinations of movement commands (forward, backward, left turn, right turn, accelerate, decelerate, ascend, descend). Both networks are initialized with the same random parameters, uniformly distributed within the range of [-0.1, 0.1]. The target value network does not directly update its parameters through gradient descent. Instead, after processing 100 batches of data, it copies the parameters from the online evaluation network at a ratio of τ=0.01, i.e., target network parameters = target network parameters × (1-τ) + online evaluation network parameters × τ.

[0059] Gradient descent is used to extract features from the robot's state vector, which is then input into the online evaluation network. A fully connected transformation from the input layer to the first hidden layer is performed, applying a linear transformation with weight matrix W1 (12×64 dimension) and bias vector b1 (64 dimension), followed by the ReLU activation function. The result is then transformed again through a fully connected transformation from the first hidden layer to the second hidden layer, applying a linear transformation with weight matrix W2 (64×128 dimension) and bias vector b2 (128 dimension), followed by the ReLU activation function. The output of the second hidden layer is the extracted node feature vector, with a dimension of 128. Gradient descent uses the Adam optimizer with an initial learning rate of 0.001, β1=0.9, β2=0.999, and ε=10. -8 The batch size is 32, and each batch randomly samples 32 state-action pairs. An empirical replay buffer with a size of 10,000 is used to store training samples. When the buffer is full, the new sample replaces the oldest sample.

[0060] The attention weight coefficients are obtained by calculating the dynamic dependencies between robots. For any two robots i and j, their node feature vectors vi and vj (each with a dimension of 128) are extracted. The two feature vectors are concatenated into a 256-dimensional vector [vi,vj]. The concatenated vector is processed by an attention calculation unit, which is a single-layer neural network with a weight matrix Wa of dimension 256×1 and a bias term ba as a scalar. The original attention score eij = tanh(Wa·[vi,vj]+ba) is calculated, where tanh is the hyperbolic tangent activation function, ensuring that the score is in the range [-1,1]. For robot i, its original attention score with all robots j is calculated, and then normalized by the softmax function to obtain the final attention weight coefficient αij. The softmax calculation process is αij = exp(eij) / ∑k=1 to n exp(eik), where n is the total number of robots, which is 5. After the calculation, a 5×5 attention weight matrix is ​​formed, representing the degree of mutual attention between all robot pairs.

[0061] A message passing mechanism is constructed to update the node feature vector. For robot i, all other robots j are traversed. Weighted feature information from robot j is obtained according to the attention weight coefficient αij, and the weighted feature vector is calculated as αij×vj. All weighted feature vectors are summed to obtain the aggregated information mi=∑j≠i(αij×vj). The aggregated information is combined with the original feature vector vi of robot i itself, and a gating mechanism is used to control the information fusion. The gating parameter gi is calculated through a single-layer neural network, with vi as the input and a scalar between 0 and 1 as the output. The updated feature vector v'i=(1-gi)×vi+gi×mi. The update process is performed simultaneously on all 5 robots to obtain 5 updated node feature vectors, each with a dimension of 128.

[0062] To calculate the local reward value, for robot i, the updated node feature vector v'i is input into the third hidden layer of the target value network; through a fully connected transformation from the third hidden layer to the output layer, the weight matrix W3 (dimension 128×64) and the bias vector b3 (dimension 64) are applied, and then the ReLU activation function is applied; the result is then transformed through the output layer to obtain the Q values ​​of the 8 actions; the task completion state rt is calculated as the difference between the current task completion percentage and the task completion percentage at the previous time step, multiplied by the task priority, with a value range of [-1,1]; The distance state rd is calculated as 1 / (1+d), where d is the Euclidean distance between the robot's current position and the target position, with a value range of (0,1]. The collaborative state rc is calculated based on the dispersion and task coverage of the robot team. The team dispersion is the degree of closeness between the average distance between each robot and the ideal dispersion distance. The task coverage is the ratio of the covered task points to the total task points, with a value range of [-0.5,0.5]. The local reward value r = 0.5×rt + 0.3×rd + 0.2×rc, with a value range of [-0.8,1].

[0063] Update the state-action values. For the current state s and the selected action a, sample a batch of state-action pairs (s, a, r, s') from the experience replay buffer, where s' is the next state. Calculate the target Q-value y = r + γ × maxa'Q'(s', a') using the target value network, where γ is the discount factor 0.95, Q' is the output of the target value network, and maxa' represents the action with the largest Q-value among all actions a'. Calculate the current Q-value Q(s, a) using the online evaluation network. Calculate the temporal difference error δ = yQ(s, a). Calculate the loss function L = δ 2 +λ×∑θ 2 , where λ is the L2 regularization coefficient 0.0001, ∑θ 2The sum of squares of the network parameters is given; the loss function is minimized using gradient descent, and the network parameters are updated online; the gradient is calculated as ∇θL=∇θQ(s,a)×(-2δ)+2λθ; the parameters are updated as θ=θ-α×∇θL, where α is the learning rate of 0.001; the above steps are repeated, and the learning rate is reduced by 0.9 times after processing 1000 batches of data.

[0064] The strategy parameters are dynamically adjusted to generate multi-robot collaborative decision-making results. An ε-greedy strategy is used to select actions, with an initial value of ε of 0.3. During training, the ε value decays linearly according to the formula ε = max(0.05, 0.3 - number of steps / 100000), with a minimum value of 0.05. For robot i, the action with the largest Q value, amax = argmaxaQ(s,a), is selected with a probability of 1-ε, and an action is randomly selected uniformly with a probability of ε. Considering the collaboration between robots, the action selection is adjusted according to the attention weight. Specifically, the adjusted Q value Q'(s,a) = Q(s,a) + β × ∑j≠i(αij × Qj(s,a)) is calculated, where β is the collaboration influence factor of 0.3, and Qj(s,a) is robot j's evaluation of action a. The optimal action is reselected based on the adjusted Q value. To balance exploration and utilization, action noise is introduced, and the final selected action a' = amax + N(0,σ) 2 ), where N(0,σ 2 The noise is Gaussian noise with a mean of 0 and a standard deviation of σ. The initial value of σ is 0.5, which decays exponentially during training with a decay rate of 0.995. The generated decision results include the action instructions of each robot, such as robot R1 selecting "forward and turn right" and robot R2 selecting "decelerate and turn left". The decision results are sent to each robot execution unit via wireless network.

[0065] In practical testing, this method was applied to a collaborative inspection task involving five robots. Compared to traditional independent decision-making methods, the task completion time decreased from an average of 135 minutes to 103.3 minutes, a reduction of 23.5%; energy consumption decreased from an average of 75% of the battery per robot to 61%, a reduction of 18.7%; and task coverage increased from 82.5% to 95.1%, an improvement of 15.2%. In robot expansion testing, when the number of robots increased from 5 to 20, the calculation time for a single decision increased from 25 milliseconds to 85 milliseconds, while still meeting real-time requirements. The tests demonstrate that this method exhibits excellent collaborative decision-making capabilities and robustness in complex task scenarios.

[0066] In one optional implementation, the priority values ​​of each task are updated by inputting the robot's remaining battery power and execution capability as influencing factors into the inspection task priority matrix, and the tasks are reallocated based on the updated priority values ​​to generate an optimized inspection task scheme, including: The remaining battery power and execution capability of the robot performing the inspection task are obtained. The remaining battery power and execution capability are used to construct robot state influencing factors. The robot state influencing factors are weighted and calculated to obtain a comprehensive robot state index. The robot's overall state index is input into the task priority adjustment function, and the priority values ​​in the inspection task priority matrix are updated through the task priority adjustment function. Multi-objective constraints are constructed based on the updated inspection task priority matrix. The multi-objective constraints are input into the task priority adjustment function, and the constraint violation degree is calculated through the task priority adjustment function. The constraint violation degree and the updated inspection task priority matrix are combined to construct a task allocation matrix. Under the condition of satisfying the single task unique allocation constraint, the task allocation matrix is ​​solved to obtain the optimal task allocation scheme. Based on the task-robot matching relationship corresponding to the optimal task allocation scheme, and combined with the priority values ​​in the updated inspection task priority matrix, the inspection tasks are reassigned to generate an optimized inspection task scheme that takes into account robot state constraints.

[0067] Multiple inspection robots are deployed in photovoltaic power plants to perform inspection tasks in various areas within the plant. This method first acquires the status information of the robots performing the inspection tasks, including remaining battery power and execution capability. The remaining battery power is obtained in real time through the battery management system, which collects battery voltage and current data and combines them with battery discharge curves to convert them into a precise percentage of remaining battery power. The execution capability is obtained through a robot performance evaluation module. This module comprehensively considers multiple factors such as the robot's moving speed (m / s), navigation accuracy (m), detection accuracy (pixel error), operational stability (failure interval time), and environmental adaptability. After standardizing each factor, a weighted average method is used to calculate the value, which ranges from 0 to 1; a higher value indicates stronger execution capability.

[0068] After constructing the remaining battery power and execution capability values ​​as robot state influencing factors, the system calculates the robot's comprehensive state index using a linear weighted method. Specifically, the calculation involves multiplying the remaining battery power by a weighting coefficient of 0.6 and the execution capability value by a weighting coefficient of 0.4, respectively. These weighting coefficients are derived through regression analysis based on extensive historical task execution data and experimental results, reflecting that the battery power factor has a slightly greater impact on the continuous execution capability than the execution capability factor. The system performs this calculation on all robots participating in the inspection, obtaining a comprehensive state index value for each robot, providing foundational data for subsequent task allocation. The comprehensive state index ranges from 0 to 1; a higher value indicates a better overall robot state and greater suitability for performing important or complex inspection tasks.

[0069] The robot's overall status index is input into the task priority adjustment function to update the priority values ​​in the inspection task priority matrix. The initial task priority values ​​are preset during the task generation phase based on factors such as equipment importance, fault risk level, and inspection timeliness. The task complexity adjustment factor is determined based on the inspection content, operational difficulty, and environmental complexity, using a three-level division mechanism: for tasks requiring fine inspection, complex operations, or execution in special environments, the adjustment factor is set to 1.2; for moderately complex tasks involving routine scanning and standard operations, the adjustment factor is set to 1.0; and for simple inspections and basic checks, the adjustment factor is set to 0.8.

[0070] For the original priority value of a task, if it is assigned to a specific robot for execution, the updated priority value is calculated as the product of the original priority value, the robot's overall state index, and the task complexity adjustment factor. The system uses this method to calculate the updated priority values ​​for all robot-task combinations, forming a complete priority update matrix. The number of rows in this matrix equals the number of robots, and the number of columns equals the number of tasks.

[0071] Based on the updated inspection task priority matrix, the system constructs multi-objective constraints, including task completion time constraints, robot power consumption constraints, and task priority satisfaction constraints. The task completion time constraint considers task location, inspection content, and robot mobility. The calculation method is to divide the path length by the robot's speed and add the inspection operation time. The path length is pre-calculated using the in-station map and the A* navigation algorithm, while the inspection operation time is obtained from the task database based on the task type. The completion time is set to not exceed a predetermined time window, typically 120 minutes.

[0072] Power consumption constraints consider path length, terrain complexity, and inspection content. The calculation method is a base consumption rate multiplied by the path length, plus a task type consumption coefficient multiplied by the inspection operation time. The base consumption rate is set at 0.03% / meter, taking into account motor power and walking resistance. The task type consumption coefficient is determined based on task complexity: 0.15% / minute for complex tasks, 0.10% / minute for medium tasks, and 0.05% / minute for simple tasks, taking into account sensor power consumption and computational load. Power consumption is set to not exceed 80% of the robot's remaining battery power to ensure the robot can safely return to the charging station. Priority is met by assigning high-priority tasks to robots with higher overall state indicators. This is achieved by defining a task-robot matching degree, calculated as task priority multiplied by the robot's overall state indicator and then multiplied by a complexity adjustment factor (specifically, 1 plus 0.5 multiplied by the complexity adjustment factor). Combinations with high matching degrees have priority in task allocation.

[0073] The system inputs the aforementioned multi-objective constraints into the task priority adjustment function to calculate the constraint violation degree for each robot-task combination. The constraint violation degree calculation process consists of three parts: time constraint violation degree, power constraint violation degree, and priority matching violation degree. The calculation method for time constraint violation degree is as follows: if the expected completion time exceeds the time limit, the violation degree is equal to the ratio of the excess time to the time limit; if the time limit is not exceeded, the violation degree is 0.

[0074] If the expected power consumption exceeds the limit (80% of remaining power), the violation score is equal to the ratio of the excess to the limit; if the limit is not exceeded, the violation score is 0. Priority matching violation score is calculated based on the matching scores of all tasks and robots. First, the maximum matching score is found. Then, the matching scores of the current combination are normalized, and the violation score is calculated as the maximum matching score minus the current matching score, then divided by the maximum matching score. The three violation scores are weighted and summed to obtain the total constraint violation score. The weights are set as follows: time constraint 0.4, power constraint 0.4, and priority matching constraint 0.2. These weight values ​​are determined based on actual operational data analysis and reflect the relative importance of each constraint.

[0075] The system constructs a task allocation matrix by combining the constraint violation rate and the updated inspection task priority matrix. Each element in the task allocation matrix represents a comprehensive evaluation value for a specific robot performing a specific task, calculated by subtracting the constraint violation rate from the updated priority value and multiplying it by an adjustment factor of 0.5. This adjustment factor is determined through sensitivity analysis based on a large amount of experimental data and is used to balance the relative weights of task priority and constraint violation effects. Under the condition of a unique assignment for each task (i.e., each task can only be assigned to one robot), the system solves the task allocation matrix to obtain the optimal task allocation scheme.

[0076] The solution process employs a greedy algorithm, with the following steps: Initialize an empty allocation scheme set; find the robot-task pair with the highest comprehensive evaluation value from the task allocation matrix; check if the pair violates the unique task allocation constraint and robot load constraint; if it does not violate the constraints, add the pair to the allocation scheme set and remove the task from the set of tasks to be assigned; update the remaining load capacity and available power of the relevant robots; repeat the above steps until all tasks are assigned or no pair satisfying the constraints can be found. The robot load capacity is determined based on its maximum working time, typically set to 6 hours of continuous operation, and the task workload is determined based on the expected completion time.

[0077] Based on the task-robot matching relationship corresponding to the optimal task allocation scheme, and combined with the priority values ​​in the updated inspection task priority matrix, the system generates a detailed optimized inspection task scheme. This scheme includes four key components: task execution order, specific path planning, task execution parameters, and emergency handling strategies. The task execution order is determined based on two factors: task priority and proximity. The system first sorts the tasks assigned to the same robot according to priority, executing higher-priority tasks first; then, considering the location factor, for tasks with similar priorities (difference less than 0.1), tasks in close proximity are selected for consecutive execution to reduce robot movement costs.

[0078] The specific path planning utilizes the A* algorithm combined with the power plant map to generate the optimal inspection path. This algorithm considers path length, terrain complexity, and obstacle distribution. Task execution parameters are selected from a parameter library based on the task type, including inspection speed, sensor sampling frequency, image acquisition resolution, and detection threshold. The emergency handling strategy defines the procedures for handling abnormal situations encountered by the robot during execution, including mechanisms for communication interruption, obstacle detour, abnormal data processing, and emergency return.

[0079] The system implements a dynamic adjustment mechanism during task execution. The robot reports its status information to the central control system every 30 seconds, including its current position coordinates, remaining battery percentage, task progress percentage, and exception status code. When one of four triggering conditions is met, the system initiates a task re-optimization process: the robot's remaining battery rate decreases by more than 20% as expected; the robot's execution capability changes significantly, such as a decrease in movement speed of more than 30%; a new high-priority task is discovered during task execution and needs to be inserted; or the execution time of a task exceeds the expected time by more than 50%. The re-optimization process uses the same algorithm framework as the initial optimization process, but uses the latest robot status and task information as input to generate a new optimization plan.

[0080] The new plan is transmitted in real time to each robot execution unit via the station's wireless network. Upon receiving the new plan, the robot completes its current operation and then switches to the new plan to continue execution. Through this dynamic adjustment mechanism, the system can adapt to various changes and anomalies during execution, ensuring the efficient completion of inspection tasks. In practical applications, this method significantly improves the intelligence level and execution efficiency of photovoltaic power station inspection work, reduces robot energy consumption, and enhances inspection quality and reliability.

[0081] Figure 3This diagram illustrates the comparison of task completion rates using different task priority allocation methods according to embodiments of the present invention. The performance comparison chart shows the results of three different task scheduling schemes (fixed priority method, dynamic priority method, and the present invention) on five key performance indicators. Specifically, the fixed priority method (square markers) represents the traditional static task allocation strategy, the dynamic priority method (circle markers) embodies the basic adaptive scheduling mechanism, and the present invention (diamond markers) employs an innovative intelligent scheduling algorithm. The curve trends show that this technical solution demonstrates superiority across all indicators: It achieves an 80% completion rate in handling urgent tasks, a significant improvement over the 60% of the fixed-priority method and the 70% of the dynamic-priority method; it achieves 70% execution efficiency in periodic task management, better than the 50% of the fixed method and the 60% of the dynamic method; and it achieves an even higher efficiency of 90% in handling daily tasks, far exceeding the 70% of the fixed method and the 80% of the dynamic method. The average completion rate remains consistently at 80%, significantly higher than the 60% of the fixed method and the 70% of the dynamic method. Under energy constraints, it maintains 80% performance, compared to only 40% for the fixed method and 60% for the dynamic method. Overall, the data indicates that this technical solution, through its intelligent task scheduling strategy, effectively balances system resource utilization while ensuring efficient execution, especially maintaining stable high performance under energy-constrained conditions. This fully demonstrates the significant advantages and good adaptability of this solution in practical application scenarios.

[0082] In one optional implementation, the task execution status data is fed back as a new influencing factor to the edge computing node to update the inspection task priority matrix, for continuous optimization of subsequent task allocation, including: The task execution status data is transmitted to the edge computing node as a new influencing factor. The priority value in the inspection task priority matrix is ​​updated based on the task execution status data in the edge computing node to obtain the updated inspection task priority matrix that reflects the current task execution status. Based on the priority values ​​in the updated inspection task priority matrix, subsequent inspection tasks to be executed are reallocated to achieve continuous optimization of the inspection task scheme.

[0083] Edge computing nodes, deployed near the production site in industrial inspection systems, can be servers or dedicated devices with computing capabilities, used to process real-time data from inspection robots. When the inspection robot performs an inspection task, it collects task execution status data in real time, including but not limited to: task completion time, task execution difficulty coefficient, anomaly detection, equipment status changes, and other key indicators.

[0084] Task execution status data is transmitted to edge computing nodes via wireless communication networks (such as industrial Wi-Fi or 5G networks). Data transmission uses encrypted channels to ensure data security, while data compression technology is used to reduce transmission bandwidth consumption. For example, after an inspection robot completes an inspection task on a pressure vessel, it will package and transmit status data such as the task execution time of 15 minutes (expected time of 12 minutes), the discovery of a minor oil leak (abnormal level 2), and the ambient temperature of 38℃ (outside the normal range) to the edge computing node.

[0085] After receiving task execution status data, the edge computing node first preprocesses the data, including data cleaning, outlier detection, and data standardization. Data cleaning removes noise and redundant information; outlier detection identifies and processes data points with abnormal values; and data standardization unifies data with different dimensions to the same scale. For example, the task execution time is converted into an execution efficiency index, calculated as the ratio of expected time to actual time, which is 0.8 (12 / 15) in this example.

[0086] The edge computing nodes maintain an inspection task priority matrix, a two-dimensional data structure where rows represent different inspection areas or equipment, and columns represent different influencing factors. Each element in the matrix represents the priority score of a specific area or equipment under a specific influencing factor. For example, in a 10×8 priority matrix, the rows represent 10 key pieces of equipment in the factory, and the columns represent 8 factors affecting priority (such as equipment importance, fault history, last inspection time, etc.).

[0087] When updating the inspection task priority matrix, edge computing nodes process newly received task execution status data using an impact factor adjustment algorithm. Specifically, for equipment that has already undergone inspection, the weight of the "anomaly risk" impact factor is adjusted based on the anomaly detection in the execution status data. For example, for a pressure vessel with a level 2 anomaly detected, the "anomaly risk" factor is increased from 0.4 to 0.7, indicating a need for more frequent inspections. Simultaneously, based on the ambient temperature exceeding the standard, the "environmental impact" factor is increased from 0.3 to 0.6, indicating that environmental conditions accelerate equipment degradation. Furthermore, based on task execution efficiency (0.8), the "inspection difficulty" factor is appropriately adjusted from 0.5 to 0.6, reflecting the need for more resources for inspections in this area.

[0088] For related equipment, the system will also make corresponding adjustments. For example, the "abnormal risk" factor of another piece of equipment sharing a piping system with the pressure vessel will be appropriately increased from 0.3 to 0.4 because of the potential associated risks. This correlation adjustment is based on an equipment relationship map, which records the physical connections and logical dependencies between equipment.

[0089] In the updated inspection task priority matrix, the overall priority of each device is obtained by calculating the weighted sum of all influencing factors. For example, the overall priority of a certain device is calculated as follows: 0.7 × 0.3 (device importance weight) + 0.7 × 0.25 (abnormal risk weight) + 0.6 × 0.15 (environmental impact weight) + 0.6 × 0.1 (inspection difficulty weight) + weighted value of other factors = 0.64.

[0090] The system uses an updated inspection task priority matrix to redistribute subsequent inspection tasks. The task allocation process considers factors such as the availability, battery status, professional capabilities, and current location of multiple inspection robots. For example, for equipment with a priority increased to 0.64, the system will assign the nearest inspection robot with the corresponding detection capabilities to perform the inspection task ahead of schedule, moving the original plan of inspection in 5 days to 2 days later.

[0091] To avoid wasting resources and duplication of work, the system will also optimize the handling of repeated inspections of the same equipment within a short period of time. For example, if a piece of equipment has been inspected within 24 hours, even if it has a high priority, the system will appropriately postpone its next inspection, unless the level of anomaly found reaches the emergency handling standard (such as level 4 or level 5 anomaly).

[0092] Through this dynamic feedback and priority update mechanism, the inspection system can continuously adjust and optimize itself based on actual performance, improving its ability to detect potential problems early. Real-world application data shows that after adopting this method, the anomaly detection rate of critical equipment increased by 23%, the average response time was shortened by 35%, and the efficiency of inspection resource utilization improved by 18%.

[0093] Edge computing nodes also periodically (usually every 24 hours) analyze the historical changes in the priority matrix to identify priority change patterns and trends, providing a basis for optimizing long-term inspection strategies. For example, if the system detects that the frequency of anomalies in a certain type of equipment increases significantly during hot seasons, it can increase the inspection frequency of the relevant equipment in advance to achieve preventative maintenance.

[0094] A second aspect of the present invention provides an intelligent inspection task planning and scheduling system for photovoltaic power plants, comprising: The first unit is used to acquire real-time monitoring data collected by multiple edge computing nodes in the photovoltaic power station, evaluate the equipment status of each area in the photovoltaic power station based on the real-time monitoring data, and generate equipment status evaluation results. The second unit is used to construct an inspection task priority matrix based on the equipment status assessment results, and to construct a regional collaborative decision-making network through distributed edge computing nodes. Each edge computing node generates a local task allocation strategy according to the inspection task priority matrix of its own region. Edge computing nodes in adjacent regions interact with their respective local task allocation strategies and perform iterative optimization to generate a global inspection task allocation result. The third unit is used to generate an initial inspection task plan based on the global inspection task allocation result, obtain the status information of the robot executing the inspection task in real time, dynamically optimize the initial inspection task plan based on the status information, update the priority value of each task by inputting the robot's remaining power and execution capability into the inspection task priority matrix, and redistribute the tasks based on the updated priority values ​​to generate an optimized inspection task plan. The fourth unit is used to distribute the optimized inspection task plan to the robot execution unit. During the execution process, it collects task execution status data including task execution efficiency and completion quality, and feeds the task execution status data as a new influencing factor to the edge computing node to update the inspection task priority matrix for continuous optimization of subsequent task allocation.

[0095] A third aspect of the present invention provides an electronic device, comprising: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.

[0096] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0097] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.

[0098] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for intelligent inspection task planning and scheduling of photovoltaic power plants, characterized in that, include: The system acquires real-time monitoring data collected by multiple edge computing nodes in a photovoltaic power station, evaluates the equipment status of each area within the photovoltaic power station based on the real-time monitoring data, and generates equipment status evaluation results. Based on the equipment status assessment results, an inspection task priority matrix is ​​constructed. A regional collaborative decision-making network is built through distributed edge computing nodes. Each edge computing node generates a local task allocation strategy according to the inspection task priority matrix of its own region. Edge computing nodes in adjacent regions interact with their respective local task allocation strategies and perform iterative optimization to generate a global inspection task allocation result. An initial inspection task plan is generated based on the global inspection task allocation result. The status information of the robot executing the inspection task is obtained in real time. The initial inspection task plan is dynamically optimized based on the status information. The priority value of each task is updated by inputting the robot's remaining power and execution capability into the inspection task priority matrix. The tasks are then redistributed based on the updated priority values ​​to generate an optimized inspection task plan. The optimized inspection task scheme is sent to the robot execution unit. During the execution process, task execution status data, including task execution efficiency and completion quality, is collected. The task execution status data is used as a new influencing factor to feed back to the edge computing node to update the inspection task priority matrix for continuous optimization of subsequent task allocation.

2. The method according to claim 1, characterized in that, Based on the equipment status assessment results, an inspection task priority matrix is ​​constructed. A regional collaborative decision-making network is built through distributed edge computing nodes. Each edge computing node generates a local task allocation strategy based on the inspection task priority matrix of its region, including: The equipment status assessment results are constructed into a status assessment result set. The weights of equipment status features are determined based on equipment runtime, fault frequency, and maintenance records. Temporal correlation analysis and multi-dimensional feature fusion are performed on the equipment status features in the status assessment result set. A deep learning network is used to dynamically predict and identify abnormal patterns of the fused equipment status features. An inspection task priority matrix is ​​generated based on the prediction results and abnormal identification results. Based on the equipment status information in the inspection task priority matrix, a task execution time window and a task priority level are generated. The task execution time window and the task priority level are then input into the deep learning network for multi-objective constraint evaluation to obtain the optimal state action value. Based on the optimal state action value, calculate the spatial clustering coefficient and task correlation degree between devices. According to the spatial clustering coefficient and the task correlation degree, divide the photovoltaic power station into multiple regions with task collaboration characteristics, and deploy edge computing nodes in the divided regions. Based on the edge computing node, the inspection task priority matrix and historical task execution feedback data are input into the deep learning network. A task execution strategy is generated by combining online strategy optimization and offline experience learning. The task execution strategy is updated by gradient, and the updated gradient is used to guide the optimization direction of the strategy, and a local task allocation strategy is output.

3. The method according to claim 2, characterized in that, Based on the optimal state action values, the spatial clustering coefficient and task correlation degree between devices are calculated. According to the spatial clustering coefficient and the task correlation degree, the photovoltaic power station is divided into multiple regions with task collaboration characteristics. Edge computing nodes are deployed in the divided regions, including: Based on the optimal state action value, the spatial clustering coefficient between devices is calculated, and the equipment tasks in the photovoltaic power station are divided into emergency tasks, periodic tasks, and daily tasks. Long Short-Term Memory Network is used to perform load prediction on the historical operation data to obtain the task load prediction values ​​of the emergency task, the periodic task and the daily task. Based on the task load prediction values, corresponding priority weights are set for different types of tasks. The task correlation degree between devices is calculated based on the priority weight. The spatial clustering coefficient and the task correlation degree are weighted and combined to obtain the task collaboration comprehensive evaluation index. A task collaboration execution queue is constructed based on the task collaboration comprehensive evaluation index. Tasks with similar comprehensive evaluation indicators for task collaboration in the task collaboration execution queue are clustered into the same region to obtain multiple regions with task collaboration characteristics. The spatial distribution centroid of the devices in each region is calculated, and the spatial distribution centroid is determined as the deployment location of the edge computing node. The task processing status of the edge computing nodes is monitored in real time. When a change is detected in the task collaborative execution queue, the region division results are dynamically updated based on the changed task collaborative execution queue, and the deployment location of the edge computing nodes is adjusted accordingly.

4. The method according to claim 1, characterized in that, An initial inspection task plan is generated based on the global inspection task allocation result. The status information of the robot executing the inspection task is acquired in real time. The initial inspection task plan is dynamically optimized based on the status information, including: A task dependency graph is constructed based on the global inspection task allocation results. The task nodes in the task dependency graph are connected based on spatiotemporal constraints. The task nodes in the task dependency graph are used to extract features using a graph neural network. The neighborhood information of the task nodes is aggregated through the graph neural network to obtain task relevance features. An initial inspection task plan is generated based on the task relevance features. The state information of the robot performing the inspection task is collected and constructed into a robot state vector. The robot state vector is processed by a multi-agent collaborative decision-making framework, wherein the multi-agent collaborative decision-making framework calculates a state decision value based on a local reward value and realizes collaborative decision-making of multiple robots through the state decision value, generating a decision result that considers the collaboration of multiple robots. A task adjustment function is constructed based on the robot's state vector, and the initial inspection task scheme is dynamically optimized using the calculation results of the task adjustment function and the decision results of the multi-agent collaborative decision-making framework.

5. The method according to claim 4, characterized in that, The robot's state vector is processed through a multi-agent cooperative decision-making framework, wherein the framework calculates state decision values ​​based on local reward values, and uses these state decision values ​​to achieve collaborative decision-making among multiple robots, generating decision results that consider multi-robot collaboration, including: A dual-Q learning network is constructed, which simultaneously maintains an online evaluation network and a target value network through a parallel structure. Based on the online evaluation network, gradient descent is used to extract features from the robot's state vector to obtain node feature vectors. The dynamic dependencies between robots are calculated based on the node feature vectors to obtain attention weight coefficients. Based on the dual Q learning network, message passing is constructed according to the attention weight coefficients, and the node feature vector is updated to obtain the updated node feature vector. The target value network is used to calculate the combination of task completion state, distance state and collaborative state according to the updated node feature vector to obtain the local reward value. Based on the dual Q learning network, the state action value is updated by minimizing the output difference between the online evaluation network and the target value network according to the local reward value. The policy parameters are then dynamically adjusted according to the state action value to generate a collaborative decision-making result for multiple robots.

6. The method according to claim 1, characterized in that, By inputting the robot's remaining battery power and execution capability as influencing factors into the inspection task priority matrix to update the priority values ​​of each task, and then reallocating tasks based on the updated priority values, an optimized inspection task scheme is generated, including: The remaining battery power and execution capability of the robot performing the inspection task are obtained. The remaining battery power and execution capability are used to construct robot state influencing factors. The robot state influencing factors are weighted and calculated to obtain a comprehensive robot state index. The robot's overall state index is input into the task priority adjustment function, and the priority values ​​in the inspection task priority matrix are updated through the task priority adjustment function. Multi-objective constraints are constructed based on the updated inspection task priority matrix. The multi-objective constraints are input into the task priority adjustment function, and the constraint violation degree is calculated through the task priority adjustment function. The constraint violation degree and the updated inspection task priority matrix are combined to construct a task allocation matrix. Under the condition of satisfying the single task unique allocation constraint, the task allocation matrix is ​​solved to obtain the optimal task allocation scheme. Based on the task-robot matching relationship corresponding to the optimal task allocation scheme, and combined with the priority values ​​in the updated inspection task priority matrix, the inspection tasks are reassigned to generate an optimized inspection task scheme that takes into account robot state constraints.

7. The method according to claim 1, characterized in that, The task execution status data is used as a new influencing factor to feed back to the edge computing nodes to update the inspection task priority matrix, which is used for continuous optimization of subsequent task allocation, including: The task execution status data is transmitted to the edge computing node as a new influencing factor. The priority value in the inspection task priority matrix is ​​updated based on the task execution status data in the edge computing node to obtain the updated inspection task priority matrix that reflects the current task execution status. Based on the priority values ​​in the updated inspection task priority matrix, subsequent inspection tasks to be executed are reallocated to achieve continuous optimization of the inspection task scheme.

8. A photovoltaic power station intelligent inspection task planning and scheduling system, used to implement the method of any one of claims 1-7, characterized in that, include: The first unit is used to acquire real-time monitoring data collected by multiple edge computing nodes in the photovoltaic power station, evaluate the equipment status of each area in the photovoltaic power station based on the real-time monitoring data, and generate equipment status evaluation results. The second unit is used to construct an inspection task priority matrix based on the equipment status assessment results, and to construct a regional collaborative decision-making network through distributed edge computing nodes. Each edge computing node generates a local task allocation strategy according to the inspection task priority matrix of its own region. Edge computing nodes in adjacent regions interact with their respective local task allocation strategies and perform iterative optimization to generate a global inspection task allocation result. The third unit is used to generate an initial inspection task plan based on the global inspection task allocation result, obtain the status information of the robot executing the inspection task in real time, dynamically optimize the initial inspection task plan based on the status information, update the priority value of each task by inputting the robot's remaining power and execution capability into the inspection task priority matrix, and redistribute the tasks based on the updated priority values ​​to generate an optimized inspection task plan. The fourth unit is used to distribute the optimized inspection task plan to the robot execution unit. During the execution process, it collects task execution status data including task execution efficiency and completion quality, and feeds the task execution status data as a new influencing factor to the edge computing node to update the inspection task priority matrix for continuous optimization of subsequent task allocation.

9. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 7.

Citation Information

Cited By

  • Intelligent inspection vehicle dynamic cooperation method based on multi-mode perception fusion and related equipment

    CN121349104A

  • Multi-robot collaborative inspection path planning method and platform for energy scene

    CN121578800A

  • Autonomous inspection method and system for substation inspection robot

    CN121791440A

  • A substation inspection robot autonomous inspection method and system

    CN121791440B

  • Multi-agent inspection task allocation method for photovoltaic power station inspection

    CN121903308A