AI Vision-Assisted Adaptive Power Distribution UAV Route Generation System

By introducing AI vision assistance and adaptive control technology into the drone power patrol system, the adaptive generation and optimization of drone routes are achieved, and the problems of flight stability and patrol accuracy of existing systems in complex environments are solved, and patrol efficiency and reliability are improved.

CN119597010BActive Publication Date: 2025-06-03STATE GRID HUBEI ELECTRIC POWER CO LTD +1
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510127391.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-04
Publication Date
2025-06-03
Estimated Expiration
2045-02-04

AI Technical Summary

Technical Problem

The existing drone power patrol system lacks adaptability and cannot adjust the flight route according to real-time conditions, resulting in unsatisfactory patrol results, and flight stability and patrol accuracy are difficult to ensure in complex environments.

Method used

Adaptive generation system for distribution network drone routes based on AI vision assistance is adopted to receive and analyze the drone's status data through the server, generate and optimize control strategies, and use the actor-critician algorithm and policy network to perform real-time data integration and strategy optimization to achieve adaptive generation and optimization of drone routes.

Benefits of technology

It improves the flexibility of drone flight and the accuracy of mission completion, ensures the safe operation of power equipment, enhances the reliability and stability of the system, and reduces the cost of system construction and maintenance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119597010B_ABST
    Figure CN119597010B_ABST
Patent Text Reader

Abstract

The present invention relates to a system for adaptively generating a distribution network drone flight path based on AI vision assistance. The system includes a server, a main drone, and a secondary drone. The server consists of an operation unit and a control unit. The operation unit is used to execute an adaptive algorithm, receive and analyze the status data of the drones to generate a control strategy and send control instructions; the control unit is responsible for transmitting these instructions to the drones. The main drone performs power inspection tasks according to a preset flight path, collects image data of power equipment and transmits it to the server; the secondary drone flies alongside the main drone, collects its attitude and motion data, and transmits the data to the server. The main drone and the secondary drone respectively generate a main record Sa and a secondary record Sb, the server integrates them into a total record S, and the operation unit analyzes and optimizes it using the actor-critic algorithm to generate an adaptive strategy to optimize the drone flight path.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of aircraft, and specifically relates to an adaptive power distribution UAV flight path generation system based on AI vision assistance. Background Art

[0002] With the continuous expansion of the scale and increasing complexity of the power grid, traditional manual inspection methods have significant deficiencies in terms of efficiency and safety. The application of UAV technology provides an efficient and safe solution for power inspection. However, existing UAV power inspection systems mainly rely on preset flight paths for flight, lacking adaptability and unable to adjust the flight route according to real-time situations, resulting in unsatisfactory inspection effects. In addition, UAVs may encounter complex environmental conditions during flight, such as wind changes and obstacles, which will affect the flight stability and inspection accuracy of UAVs. To improve the efficiency and reliability of UAV power inspection, there is an urgent need for an intelligent system that can adjust the flight route in real time and monitor and optimize the flight state. By introducing AI vision assistance and adaptive control technology, precise flight and efficient inspection of UAVs in complex environments can be achieved, ensuring the safe operation of power equipment.

[0003] Referring to relevant publicly available technologies, the technical solution with the publication number US20210224512A1 proposes a UAV inspection system using a dual attention network, which performs image recognition on floating objects on the coastline to determine whether the floating objects are hazardous waste or other threatening objects; the technical solution with the publication number CN106774221A proposes a patrol system combining UAVs and unmanned vehicles, which realizes three-dimensional recognition of various objects on the ground through their cooperation, and the unmanned vehicle can support the endurance working time of the UAV; the technical solution with the publication number KR102440819B1 proposes a security system for UAVs to identify personnel in a region, which verifies whether the personnel in a specified region are security personnel or other personnel through regular inspections by the UAV.

[0004] The above technical solutions all propose UAV application technical solutions using image recognition to judge objects or people, but there is little mention in the current technical solutions regarding how to continuously optimize the strategies of UAVs in recognition and inspection.

[0005] The foregoing discussion of the background art is only intended to facilitate understanding of the present invention. This discussion does not recognize or admit any part of the materials mentioned as common general knowledge. Summary of the Invention

[0006] The object of the present invention is to provide a system for adaptively generating a distribution network UAV flight path based on AI vision assistance; the generation system includes a server, a main UAV, and a secondary UAV; the server consists of an operation unit and a control unit. The operation unit is used to execute an adaptive algorithm, receive and analyze the status data of the UAVs, generate a control strategy and send control instructions; the control unit is responsible for transmitting these instructions to the UAVs. The main UAV performs power inspection tasks according to a preset flight path, collects image data of power equipment and transmits it to the server; the secondary UAV flies alongside the main UAV, collects its attitude and action data, and transmits the data to the server. The main UAV and the secondary UAV respectively generate a main record Sa and a secondary record Sb, and the server integrates them into a total record S, and the operation unit uses the actor-critic algorithm to analyze and optimize it, generating an adaptive strategy to optimize the UAV flight path.

[0007] The present invention adopts the following technical solutions: A system for adaptively generating a distribution network UAV flight path based on AI vision assistance, the generation system includes:

[0008] A server, at least one main UAV, and at least one secondary UAV;

[0009] The server includes:

[0010] An operation unit configured to receive and analyze the status data of the UAVs, generate, save, and optimize a control strategy for controlling the flight of the UAVs, and generate control instructions for controlling the flight of one or more UAVs according to the control strategy; and

[0011] A control unit configured to communicate with one or more UAVs and send the control instructions to the UAVs;

[0012] The main UAV is configured to perform a power inspection task on a preset flight path, including flying and taking pictures of power equipment on the flight path, and transmitting the captured image data to the server;

[0013] The secondary UAV is configured to fly alongside the main UAV, collect the attitude and action data of the main UAV, and transmit the captured image data to the server;

[0014] Among them, the data collected by the main UAV is stored as the main record Sa, and the main record Sa includes the self-status data of the main UAV during flight and the image data of the power equipment captured by the main UAV; the data collected by the secondary UAV is stored as the secondary record Sb, and the secondary record Sb includes the status data and action data of the main UAV captured by the secondary UAV during flight;

[0015] And, the operation unit executes the following steps:

[0016] Integrate the main record Sa and the secondary record Sb to form a total record S;

[0017] Adopt the actor-critic algorithm to evaluate the current control strategy based on the total record S;

[0018] Continuously optimize the current control strategy according to the evaluation results, and finally generate an optimized adaptive strategy for the current flight path;

[0019] Preferably, a policy network and an evaluation network are running in the operation unit, which are used to continuously execute the generation and optimization of the adaptive strategy based on the actor-critic algorithm; among them,

[0020] The policy network is configured to, based on the current control strategy and the execution effect provided by the evaluation network, take the current state and / or action of the unmanned aerial vehicle as input, and correspondingly output the action of the unmanned aerial vehicle at the next time node;

[0021] The evaluation network is configured to evaluate the execution effect of the unmanned aerial vehicle executing the current control strategy, and feedback the execution effect to the policy network;

[0022] Preferably, the control strategy includes multiple strategies, and the control strategy describes the specific actions that the unmanned aerial vehicle should execute based on the current state and / or action of the unmanned aerial vehicle;

[0023] Preferably, the operation unit further includes verifying the stability of the evaluation network by using the invariant set principle;

[0024] Preferably, the evaluation network evaluates the execution effect of the unmanned aerial vehicle executing the current control strategy, and the content of its evaluation includes:

[0025] Whether the attitude of the unmanned aerial vehicle is stable;

[0026] Whether the distance between the unmanned aerial vehicle and the target power equipment to be inspected and its own positioning during inspection are accurate;

[0027] The image quality of the images collected by the unmanned aerial vehicle;

[0028] Preferably, the image quality of the images collected by the unmanned aerial vehicle includes one or more of the following evaluation elements:

[0029] Whether the target part of the target power equipment is captured in the picture;

[0030] The proportion of the target part of the target power equipment in the picture;

[0031] Whether the target part of the target power equipment is clear in the picture;

[0032] Preferably, the flight control of the secondary UAV is achieved by one of the following methods:

[0033] Manual operation by at least one operator through the control unit,

[0034] Automatically executed by the control unit;

[0035] Automatically executed by at least one on-board control unit deployed on the corresponding UAV;

[0036] Preferably, the main UAV and the secondary UAV are each configured with at least one image acquisition device, and the image acquisition device is composed of one or more of the following sensors or image devices: cameras, video cameras, thermal imaging cameras, infrared cameras, night vision sensors, depth cameras, ranging sensors, laser ranging sensors, radio ranging sensors.

[0037] The beneficial effects achieved by the present invention are as follows:

[0038] 1. The generation system of the present technical solution realizes the adaptive generation and optimization of the UAV flight path by adopting the actor-critic algorithm and AI vision assistance; the server can continuously evaluate and adjust the flight control strategy according to the real-time data collected by the main UAV and the secondary UAV, enabling the UAV to dynamically adapt to complex inspection environments, improving the flight flexibility and the accuracy of task completion;

[0039] 2. The generation system of the present technical solution is configured with a cooperation mechanism of a main UAV that autonomously flies along a predetermined flight path and a secondary UAV that can fly automatically or manually; during the process of generating the adaptive strategy, the main UAV focuses on executing the preset power inspection task, while the secondary UAV is responsible for monitoring the attitude and actions of the main UAV and providing data support from a third-person perspective; this multi-angle and multi-view data acquisition method not only improves the UAV's perception ability of its own state but also effectively avoids the collision risk during the process of generating the adaptive strategy, ensuring the safety and reliability of the inspection task; after the adaptive strategy generation is completed, it can significantly optimize the adaptive flight of the current flight path and can be completed by one UAV for the inspection work;

[0040] 3. In the generation system of the present technical solution, the server forms a total record by integrating the main record and the secondary record and uses the actor-critic algorithm for real-time analysis and optimization of the control strategy; the evaluation network's evaluation and feedback on the execution effect of the UAV ensure the continuous improvement and optimization of the control strategy; this mechanism not only improves the intelligent level of UAV inspection but also ensures the efficient processing and utilization of inspection data, improving the response speed and working efficiency of the overall system;

[0041] 4. The software and hardware parts of the generation system of this technical solution adopt a modular design. Each working module, component of the hardware part in the system, as well as the instructions, parameters, and algorithms of the software part can be conveniently replaced and / or upgraded later, thereby reducing the construction cost and maintenance cost of this system. Description of the Drawings

[0042] The present invention can be further understood from the following description in conjunction with the drawings. The components in the drawings are not necessarily drawn to scale, but the emphasis is on showing the principles of the embodiments. In different views, the same reference numerals designate corresponding parts.

[0043] Explanation of the Reference Numerals in the Drawings: 10 - Server; 20 - Main UAV; 30 - Sub UAV; 40 - Network; 50 - Target Power Equipment; 301 - Image Input Layer; 302 - Sensor Input Layer; 311 - Convolution Layer A; 312 - Pooling Layer A; 321 - Convolution Layer B; 322 - Pooling Layer B; 330 - Flattening Layer; 341 - Fully Connected Layer A; 350 - Merging Layer; 351 - Fully Connected Layer B; 352 - Fully Connected Layer C; 360 - Fully Connected Layer D; 500 - Computer System; 502 - Bus; 504 - Processor; 506 - Main Memory; 508 - Read - Only Memory; 510 - Storage Device; 512 - Display; 514 - Input Device; 516 - Cursor Control Device; 518 - Network Device;

[0044] Figure 1 Schematic diagram of the framework of the growth system according to the embodiment of the present invention;

[0045] Figure 2 Schematic diagram of the data acquisition part of the growth system according to the embodiment of the present invention;

[0046] Figure 3 Schematic diagram of the evaluation network according to the embodiment of the present invention;

[0047] Figure 4 Schematic diagram of the performance comparison of the evaluation network verified by the invariant set principle according to the embodiment of the present invention;

[0048] Figure 5 Schematic diagram of the framework of the computer system adopted by the arithmetic unit according to the embodiment of the present invention. Detailed Embodiments

[0049] In order to make the object, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below in conjunction with its embodiments; it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. For those skilled in the art, after referring to the following detailed description, other systems, methods, and / or features of this embodiment will become obvious. It is intended that all such additional systems, methods, features, and advantages be included within this specification, be included within the scope of the present invention, and be protected by the appended claims. Additional features of the disclosed embodiments are described in the following detailed description and will be obvious from the following detailed description.

[0050] In the drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", etc. indicating the orientation or positional relationship, they are based on the orientation or positional relationship shown in the drawings. This is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or component referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the drawings are only for illustrative purposes and cannot be understood as a limitation of this patent. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.

[0051] Embodiment 1: Exemplarily, a system for adaptively generating a distribution network UAV flight path based on AI vision assistance is proposed. The generation system includes:

[0052] A server, at least one main UAV, and at least one secondary UAV;

[0053] The server includes:

[0054] An operation unit, configured to receive and analyze the state data of the UAV, generate, save, and optimize a control strategy for controlling the flight of the UAV, and generate a control instruction for controlling the flight of one or more UAVs according to the control strategy; and

[0055] A control unit, configured to communicate with one or more UAVs and send the control instruction to the UAV;

[0056] The main UAV is configured to perform a power inspection task on a preset flight path, including flying and taking pictures of power equipment on the flight path, and transmitting the captured image data to the server;

[0057] The secondary UAV is configured to fly in company with the main UAV, collect the attitude and action data of the main UAV, and transmit the captured image data to the server;

[0058] Among them, the data collected by the main drone is stored as the main record Sa, and the main record Sa includes the self-state data of the main drone during flight and the image data of the power equipment captured by the main drone; the data collected by the sub-drone is stored as the sub-record Sb, and the sub-record Sb includes the state data and action data of the main drone captured by the sub-drone during flight.

[0059] And, the operation unit performs the following steps:

[0060] Integrate the main record Sa and the sub-record Sb to form the total record S;

[0061] Adopt the actor-critic algorithm to evaluate the current control strategy based on the total record S;

[0062] Continuously optimize the current control strategy according to the evaluation result, and finally generate an adaptive strategy optimized for the current flight path;

[0063] Preferably, a policy network and an evaluation network are running in the operation unit, which are used to continuously execute the generation and optimization of the adaptive strategy based on the actor-critic algorithm; among them,

[0064] The policy network is configured to, based on the current control strategy and the execution effect provided by the evaluation network, take the current state and / or action of the drone as input, and correspondingly output the action of the drone at the next time node;

[0065] The evaluation network is configured to evaluate the execution effect of the drone executing the current control strategy and feedback the execution effect to the policy network;

[0066] Preferably, the control strategy includes multiple strategies, and the control strategy describes the specific actions that the drone should execute based on the current state and / or action of the drone;

[0067] Preferably, the operation unit further includes using the invariant set principle to verify the stability of the adaptive strategy;

[0068] Preferably, the evaluation network evaluates the execution effect of the drone executing the current control strategy, and the content of its evaluation includes:

[0069] Whether the attitude of the drone is stable;

[0070] Whether the distance between the drone and the target power equipment to be inspected and its own positioning during inspection are accurate;

[0071] The image quality of the images collected by the drone;

[0072] Preferably, the image quality of the images collected by the UAV includes one or more of the following evaluation elements:

[0073] Whether the target part of the target power equipment is captured in the picture;

[0074] The proportion of the target part of the target power equipment in the picture;

[0075] Whether the target part of the target power equipment is clear in the picture;

[0076] Preferably, the aircraft control of the secondary UAV is achieved by one of the following methods:

[0077] Manual operation by at least one operator through the control unit,

[0078] Automatically executed by the control unit;

[0079] Automatically executed by at least one on-board control unit deployed on the corresponding UAV;

[0080] Preferably, the main UAV and the secondary UAV are each configured with at least one image acquisition device, and the image acquisition device is composed of one or more of the following sensors or image devices: camera, video camera, thermal imaging camera, infrared camera, night vision sensor, depth camera, range sensor, laser range sensor, radio range sensor;

[0081] As shown in the appendix Figure 1 is a schematic diagram of an exemplary embodiment of the generation system; wherein, the server 10 can be various types of computer devices, including but not limited to: traditional physical servers, virtual servers, cloud computing servers, edge computing devices or other electronic devices with data processing capabilities; the server 10 is communicatively connected to the UAV through the network 40, receives the data sent by the UAV, and sends control instructions to the UAV; the type of the network 40 may be WiFi, radio, cellular network (such as 4G, 5G), satellite communication, Bluetooth, ZigBee or other wireless communication technologies suitable for data transmission; the communication bandwidth and communication distance of the network 40 can ensure real-time data exchange between the server and the UAV, and support remote control, status monitoring and image data transmission of the UAV;

[0082] Exemplarily, both the main UAV 20 and the secondary UAV 30 are communicatively connected to the server 10; the main UAV 20 and the secondary UAV 30 can be of the same type of UAV, or based on the requirements of the power inspection task, since the main UAV 20 needs to be responsible for relatively complex tasks, it can be a different type and different specification of UAV from the secondary UAV 30;

[0083] Further, the operator can operate the UAV by using a control unit on the server; preferably, the main UAV 20 and the sub-UAV 30 can be flight-controlled through a preset automated flight program pre-stored in the server 10, so as to implement the functions of the present generation system; alternatively, the operator can also intervene through the control unit to manually operate the main UAV 20 and / or the sub-UAV 30 to actively take over the UAV.

[0084] In a more specific embodiment, the state s of the UAV represents various state information of the current environment and the UAV itself. The state s includes but is not limited to:

[0085] The position and speed of the UAV, where the position may include the deviation degree of the UAV flying along the route, and when it is necessary to hover at a fixed position to collect images of the target power equipment, the hovering position of the UAV;

[0086] The attitude of the UAV, such as pitch angle, roll angle, and yaw angle;

[0087] The current environmental perception information, such as sensor data, camera images, etc.; preferably, the environmental perception information includes the clarity of the collected images; when the UAV collects images of the target power equipment, the collection points calibrated by the preset route may not be the optimal positions due to errors, for example, they may be too far or too close, resulting in an inability to obtain detection images within a suitable image range, or the selected observation angle may be incorrect, resulting in a large number of irrelevant objects in the viewfinder range, thus causing difficulty in focusing;

[0088] The action a refers to the specific control instructions that the UAV can execute in a specific state, such as: adjusting the speed of the UAV, changing the flight direction of the UAV, adjusting the height of the UAV; for the collection device on the UAV, it includes controlling the angle, focal length, aperture, shutter, sensitivity, anti-shake, etc. of the camera;

[0089] In the case where only one drone is used to perform the task of photographing power equipment, there are multiple limitations and potential risks in its perceptibility of the state s; possible situations include that the field of view of the drone is limited, the camera and sensors are mainly facing forward and the target object, making it difficult to cover all directions around itself, resulting in blind spots and the inability to monitor the situations on the sides and back, presenting a collision risk; in addition, the drone cannot monitor the front and rear environments simultaneously during complex flight operations, ignoring sudden situations of obstacles behind; at the same time, if problems such as sensor failures or communication interruptions occur to the drone, it will cause the interruption of the entire system task because there is no other drone to monitor the state and environment of the main drone; furthermore, the drone cannot comprehensively monitor the surrounding environment in real time in the complex environment of power lines, increasing the difficulty of dynamic obstacle avoidance; the drone also needs to perform flight control, equipment photographing and environmental monitoring simultaneously, increasing the task complexity and system load and insufficient resource allocation, affecting the task execution efficiency;

[0090] Therefore, in a preferred exemplary embodiment, introducing a secondary drone can significantly solve the above problems; firstly, the secondary drone can fly around the main drone, photograph and monitor the main drone and its surrounding environment from different angles, provide an all-round field of view, eliminate the blind spots of the main drone, and provide real-time environmental feedback to ensure that the main drone flies in the correct predetermined route and attitude, and collect the state data of the main drone during flight from a third-person perspective; secondly, as an auxiliary system, the secondary drone can take over part of the monitoring and navigation tasks when the main drone fails or the sensor fails, provide redundant perception information, and ensure the continuity of the task; the secondary drone can also enhance the dynamic obstacle avoidance ability, monitor the dynamic changes around the main drone in real time, discover and feedback potential obstacles in time, improve the reaction speed, and ensure the safe flight of the main drone in a complex environment; finally, the secondary drone shares part of the environmental monitoring and data processing tasks, reduces the burden on the main drone, enables it to focus more on the photographing and inspection tasks, and improves the overall task efficiency through division of labor and cooperation and optimized resource utilization; in addition, the secondary drone can photograph the environmental position where the main drone is located from multiple angles, and by photographing the attitude and actions of the main drone, confirm that the attitude detected by the main drone through its own sensors is consistent, so as to reconfirm that the main drone has no mistakes in its own state monitoring; this dual verification mechanism can ensure the accuracy of the flight attitude and position data of the main drone, and improve the stability and safety of the entire system;

[0091] Based on the above settings, in a preferred exemplary embodiment, the main record Sa formed by the main drone includes:

[0092] Self-state data: including flight parameters of the main UAV, such as current position, speed, altitude, attitude angle, etc.; these data are obtained through the sensors of the main UAV itself (such as GPS, inertial measurement unit) and are used for real-time flight control and navigation;

[0093] Image data of the power equipment collected: Image data of the power equipment taken by the main UAV during flight, which is used for visual inspection and status evaluation of the power equipment; these images can contain detailed visual information of equipment such as wires and transformers;

[0094] In a preferred exemplary embodiment, the secondary record Sb formed by the secondary UAV includes:

[0095] Main UAV state data: State data of the main UAV taken by the secondary UAV, including the position, speed, attitude, etc. of the main UAV, for backing up and verifying the self-state data in the main record Sa;

[0096] Main UAV-environment interaction data: including the mutual distance between the main UAV and objects in the environment, whether there are dangerous objects in the flight route of the main UAV, etc.;

[0097] Main UAV action data: Action data of the main UAV recorded by the secondary UAV, such as control instructions during flight, heading changes, etc.; these data are used to verify the flight dynamics and behavior of the main UAV;

[0098] Furthermore, the main record Sa and the secondary record Sb are integrated into a data set of state s; both the main record Sa and the secondary record Sb are time-series data sequences; among them,

[0099] Main record Sa = {(Sa at , Ia at )}, at = 1, 2, 3... Ta; where Sa at is the state vector of the main UAV at time t, and Ia at is the image data taken by the main UAV at time t; Ta is the total length of the data sequence Sa;

[0100] Secondary record Sb = {(Sb bt , Ib bt )}, bt = 1, 2, 3... Tb; where Sb bt is the state vector of the main UAV at time bt, and Ib bt is the image data taken by the secondary UAV at time bt;

[0101] Both of the above state vectors Sa at and Sb bt include multiple components, and each component represents a numerical value of a state of the main UAV;

[0102] Further, the main record Sa and the secondary record Sb are integrated, including merging and deduplicating the two records, so that the final integrated total record S is obtained; the data of the total record S is used as the sample data of the state s.

[0103] The integrated state s data not only provides the real-time state and environmental information of the main UAV, but also enhances the guarantee of flight safety and data accuracy for the entire system through the data verification mechanism of the secondary UAV; such a design not only improves the reliability and stability of the system, but also optimizes the efficiency and accuracy of power equipment inspection.

[0104] Further, the policy function C W (s) means that in policy C, the parameter set W is used to determine the probability of selecting action a in state s; preferably, the policy function C is implemented by a neural network, where the parameter set W is the set of weights and biases of each node in the hidden layer of the neural network; the output of the policy function C W (s) can be a probability distribution (in a discrete action space) or a specific action (in a continuous action space); by training a large number of parameters in the set of parameter sets W, the policy network is made to learn to select appropriate control actions in different states under the policy model constituted by the policy function C W (s).

[0105] Further, in a preferred embodiment, the evaluation network can collect simulation data as training data by running a simulation environment, or use the historical operation data of the UAV in previous power inspections as training data for neural network training; alternatively, randomly generated data within a limited range can also be used as training data for neural network training; the evaluation network aims to estimate the expected value of future rewards for a given state and action, and its goal is to learn the reward structure and state transition rules of the environment, and this information can be learned through interaction or historical data before the training of the policy network;

[0106] Preferably, in the simulation environment, the initial control policy can be a random policy, a manually designed policy, or a predefined policy; thereafter, multiple trials can be carried out on the initial control policy, and each state-action pair and its corresponding reward and next state are recorded; in this way, a large amount of data for training can be generated;

[0107] Preferably, the historical data experienced by the UAV under the initial control policy can be used to extract state-action pairs and their corresponding rewards from these data; the historical data can come from previous operation records, experimental data, or simulation data;

[0108] Meanwhile, during the reinforcement learning process, preferably, an experience replay buffer (Replay Buffer) is used to store the experiences of the agent interacting with the environment; these experiences include states, actions, rewards, and next states; these data can be reused for training the value evaluation network;

[0109] More specifically, the evaluation network is based on the state-action value function Q of a policy C C (s,a;ψ) as the output model for evaluation;

[0110] Among them, the state-action value function Q() can be expressed as:

[0111] ;

[0112] In the above formula, it represents the expected cumulative return when starting from state s, taking action a based on policy C, and then following policy C; ψ is the parameter set of the state-action value function Q;

[0113] Among them, η is the discount factor, and its value range is [0, 1], which is set by the technical personnel; η t represents the change with the time step t, and η t decreases, which is used to represent the change in the weight value of future rewards, thereby reducing the impact of future returns; this way makes the algorithm pay more attention to recent returns rather than overly focusing on the distant future; E is the expected value

[0114] r t is the sequence of expected reward values obtained when executing each subsequent [state-action] pair according to policy C; it is also formulated by the technical personnel when formulating the initial policy;

[0115] Taking the data in Table 1 below as an example, the meaning of the state-action value function Q is illustrated:

[0116] Table 1

[0117]

[0118] Finally, it can be calculated that:

[0119] Q = r 0 + 0.9 * r 1 + 0.81 * r 2 + 0.729 * r 3 = 4.707;

[0120] When the state s changes, for example, the takeoff is not stable or the pointing is incorrect, the corresponding reward value is reduced; and the reward value r can be negative;

[0121] Based on the above settings, the training process of the evaluation network is as follows:

[0122] (1) Sample data, randomly sample a batch of samples from the training set samples, where each sample includes: a pair (s, a) of the state s of the drone and the action a to be performed according to the established policy C, the reward r corresponding to (s, a), and the expected state s' after performing the action a corresponding to the current state s;

[0123] (2) Preferably, define the target value target for each sample. For a sample (s, a, r, s'), there is:

[0124] target = r + ηMax a´ Q(s', a');

[0125] In the above formula, r is the reward value obtained by taking the action a in the current state s, and this reward value is preset by the technician in the policy C; η is the discount factor, and its value is in the interval [0, 1]; Max a´ Q(s', a') represents the maximum Q value of the future optimal action, which means the maximum Q value that can be obtained among all possible actions a' in the next state s'; in the next state s', the drone can choose different actions a' in order to maximize the Q value and calculate this maximum Q value;

[0126] The significance of the above formula is that it combines the current immediate reward and the possible maximum future return to evaluate the long-term return of the current state-action pair; by continuously updating the Q value, the algorithm can gradually optimize the policy C so that the actions selected in each state can maximize the long-term return;

[0127] (3) In a preferred embodiment, use the mean squared error as the loss function L(), and calculate the value function

[0128] ;

[0129] In the above formula, L(ψ) is the evaluation network model using the parameter set ψ, and E represents the expected value or approximate expected value; (s, a, r, s') ~ D means randomly sampling a quadruple (s, a, r, s') from the experience replay buffer D; by training the parameter set ψ of the evaluation network model, the loss function is minimized; in the above formula, it means taking the expected value of the square of the difference between the total return of the input data (s - a) pairs in all samples and the Q and the target value target;

[0130] In an exemplary embodiment, after obtaining the trained evaluation network, implement the actor-critic algorithm through the following steps to generate the adaptive policy:

[0131] (1) Initialize network parameters: Initialize the policy function C in the policy network W(s) and the evaluation function Q in the evaluation network C The parameter groups W and ψ in Q(s,a;ψ);

[0132] (2) Sampling: Execute the policy C in the environment W π(s), and collect multiple state-action-reward triples (s,a,r);

[0133] (3) Training the evaluation network: Train the evaluation network Q with the collected data C Q(s,a;ψ);

[0134] (4) Calculate the gradient: For each (s,a), calculate ∇ W logπ W (a|s)Q(s,a;ψ);

[0135] (5) Update the policy network: Update the policy network parameter W with the calculated gradient, that is:

[0136] W←W + α∇ W logπ W (a|s)Q(s,a;ψ), where α is the learning rate set by relevant technical personnel;

[0137] (6) Repeat the processes of sampling, training the evaluation network, and updating the policy network until the policy network converges;

[0138] Through the above steps, the training of the policy network is completed, and the policy network can be applied during the flight of the main UAV to generate an adaptive policy, and the main UAV can be controlled according to the adaptive policy to perform the inspection task of power equipment alone;

[0139] In an exemplary embodiment, the structure of the evaluation network is as shown in the appendix Figure 3 ; Preferably, the evaluation network includes two input layers, which are respectively used to process the image and the data input by the sensor; among them, the input layer includes an image input layer 301 for receiving the input RGB image data, and the image size is (H,W,3), indicating that the height is H pixels, the width is W pixels, and each pixel has three color channels (red, green, blue), and this layer provides the image data input interface required by the model; the input layer also includes a sensor input layer 302 for receiving the data from other sensors of the UAV, and the input size is (n), and the specific sensor data may include physical parameters such as speed and angle; this layer provides the sensor data input interface required by the model;

[0140] Further, the evaluation network further includes a feature extraction layer,

[0141] Among them, the part for image processing includes: a first convolutional layer 311: This layer uses 32 convolutional kernels of 8x8, and performs a convolution operation on the input image with a stride of 4; through the ReLU activation function, low-level features in the image are extracted to generate a feature map;

[0142] a first pooling layer 312: This layer uses a pooling window of 2x2 and performs max pooling with a stride of 2, thereby reducing the size of the feature map, while retaining important features and reducing computational complexity;

[0143] a second convolutional layer 321: This layer uses 64 convolutional kernels of 4x4 and performs a convolution operation with a stride of 2 to further extract intermediate features in the image, and processes the feature map through the ReLU activation function.

[0144] a second pooling layer 322: This layer again uses a pooling window of 2x2 and performs max pooling with a stride of 2 to further reduce the size of the feature map while retaining important features;

[0145] a flattening layer 330: This layer flattens the multi-dimensional feature map generated by the convolution operation into a one-dimensional vector for subsequent processing by the fully connected layer;

[0146] The sensor processing part includes a first fully connected layer 341, which contains 128 neurons and uses the ReLU activation function. This layer is used to process and extract high-dimensional features of sensor data and provide input for subsequent merged features;

[0147] Furthermore, the evaluation network further includes a data merging part, including a merging layer 350: used to splice the flattened output of the image processing part and the output of the sensor data processing part; by merging the features of the image and sensor data, comprehensive information is provided for further feature extraction and decision-making;

[0148] a second fully connected layer 351: The second fully connected layer contains 512 neurons and uses the ReLU activation function. This layer performs further non-linear transformation on the merged features to extract higher-level features;

[0149] a third fully connected layer 352: The third fully connected layer contains 256 neurons and uses the ReLU activation function. This layer further processes the features and provides input for the final value evaluation;

[0150] Furthermore, the evaluation network further includes a fourth fully connected layer 360 in the output part. The fourth fully connected layer 360 contains 1 neuron and uses a linear activation function; this layer is used to output the state-action value Q(s,a), evaluate the value of the current state and action combination, and provide feedback for the optimization of the policy network.

[0151] Example 2: This example should be understood as including at least all the features of any of the foregoing examples, and further improvements are made on this basis;

[0152] Furthermore, in a preferred embodiment, the invariant set principle is used to analyze and verify the stability of the evaluation network in this technical solution;

[0153] The steps of verification include:

[0154] (1) Select an appropriate Lyapunov function V(s) that satisfies V(s)>0 when x≠0 and V(0)=0 when x = 0;

[0155] More specifically, set V(s)=Q(s,a;ψ)-Q(s,a;ψ) * , where Q(s,a;ψ) * represents the numerical value of Q(s,a;ψ) under the optimal policy;

[0156] (2) Set s´=f(s,W), where the function f() represents the change law of the state s under the current state s and the control policy parameter W;

[0157] (3) Using the chain rule, calculate the derivative V´(s) of V(s) with respect to time, then we have:

[0158] V´(s)=(∂V(s) / ∂s)·s´=(∂V(s) / ∂s)·f(s,W);

[0159] (4) Verify that for all multi-dimensional variables s in all states, V´(s)≤0 can be ensured, that is, it can be determined that Q(s,a;ψ) is in a stable state, that is, the evaluation network has sufficient stability;

[0160] After comparison, as shown in the appendix Figure 4 shown, the return values of the evaluation network after 100 training sessions are compared; among them, the return value curve of the technical solution adopted is curve Rew.A, and the return value curve of the evaluation network without using the invariant set principle for verification is curve Rew.B; it can be seen that curve Rew.A can reach a steady state faster, and compared with curve Rew.B, under the same number of training sessions, it can obtain a higher return value, that is, better performance.

[0161] Example 3: This example should be understood as including at least all the features of any of the foregoing examples, and further improvements are made on this basis;

[0162] Exemplarily, as shown in the appendix Figure 5As shown, it illustrates the implementation of the computer system 500 adopted by the arithmetic unit in the generation system; the computer system 500 can be applied to the data storage, arithmetic operation, and result output processes of each working module in the recognition and judgment system;

[0163] Exemplarily, the computer system 500 includes a bus 502 or other communication mechanisms for transmitting information, and one or more processors 504 coupled to the bus 502 for processing information; the processor 504 can be, for example, one or more general-purpose microprocessors;

[0164] The computer system 500 further includes a main memory 506, such as a random access memory (RAM), cache, and / or other dynamic storage devices, which are coupled to the bus 502 for storing information and instructions to be executed by the processor 504; the main memory 506 can also be used to store temporary variables or other intermediate information during the execution of instructions executed by the processor 504; when these instructions are stored in a storage medium accessible by the processor 504, the computer system 500 is presented as a special-purpose machine customized to execute the operations specified in the instructions;

[0165] The computer system 500 may also include a read-only memory (ROM) 508 or other static storage devices coupled to the bus 502 for storing static information and instructions of the processor 504; a storage device 510 such as a magnetic disk, optical disk, or USB drive (flash drive) will be coupled to the bus 502 for storing information and instructions;

[0166] Furthermore, coupled to the bus 502 may also include a display 122 for displaying various information, data, media, etc., and an input device 514 for allowing the user of the computer system 500 to control, manipulate, and / or interact with the computer system 500;

[0167] Preferably, a way to interact with the management system can be through a cursor control device 516, such as a computer mouse or similar control / navigation mechanism;

[0168] Furthermore, the computer system 500 may also include a network device 518 coupled to the bus 502; the network device 518 may include components such as a wired network card, wireless network card, switching chip, router, switch, etc.;

[0169] Generally, as used herein, terms such as "engine", "component", "system", "database", etc. may refer to logic embodied in hardware or firmware, or to a collection of software instructions, which may have entry and exit points and be written in a programming language such as Java, C, or C++; software components may be compiled and linked into an executable program, installed in a dynamic link library, or may be written in an interpreted programming language (such as BASIC, Perl, or Python); it should be understood that software components may be called from other components or from themselves, and / or may be called in response to detected events or interrupts;

[0170] Software components configured to execute on a computing device may be provided on a computer-readable medium, such as a compact disc, digital video disc, flash drive, magnetic disk, or any other tangible medium, or as a digital download (and may initially be stored) in a compressed or installable format that requires installation, decompression, or decryption before execution); such software code may be stored, in whole or in part, on the memory device of the executing computing device for execution by the computing device; software instructions may be embedded in firmware, such as an EPROM; it should also be understood that hardware components may be composed of connected logic units (such as gates and flip-flops), and / or may be composed of programmable units (such as programmable gate arrays or processors);

[0171] Computer system 500 includes custom hardwired logic, one or more ASICs or FPGAs, firmware, and / or program logic that may be used to implement the techniques described herein, and the program logic, in combination with the computer system, causes computer system 500 to be a dedicated computing device;

[0172] In accordance with one or more embodiments, the techniques herein are performed by computer system 500 in response to one or more sequences of one or more instructions contained in main memory 506 being executed by processor 504; such instructions may be read into main memory 506 from another storage medium, such as storage device 510; the execution of the instruction sequence contained in main memory 506 causes processor 504 to perform the processing steps described herein; in alternative embodiments, hardwired circuitry may be used in place of or in combination with software instructions;

[0173] As used herein, the term "non-transitory medium" and like terms refer to any medium that stores data and / or instructions that cause a machine to operate in a particular manner; such non-transitory media may include non-volatile media and / or volatile media; non-volatile media includes, for example, optical discs or magnetic disks, such as storage device 510; volatile media includes dynamic memory, such as main memory 506;

[0174] Among them, common forms of non-transitory media include, for example, floppy disks, hard disks, solid state drives, magnetic tapes, or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical media with a hole pattern, RAM, PROM, and EPROM, FLASH-EPROM, NVRAM, any other storage chip or cartridge, and network versions thereof;

[0175] Non-transient media are different from transmission media, but can be used in combination with transmission media; transmission media participate in the transmission of information between non-transient media; for example, transmission media include coaxial cables, copper wires, and optical fibers, including the wires that make up bus 502; transmission media can also take the form of acoustic waves or light waves, such as radio waves and infrared data communication.

[0176] Although the present invention has been described above with reference to various embodiments, it should be understood that many changes and modifications can be made without departing from the scope of the present invention. That is, the methods, systems, and devices discussed above are examples. Various configurations can be appropriately omitted, replaced, or various processes or components added. For example, in alternative configurations, the methods can be executed in a different order than described, and / or various components can be added, omitted, and / or combined. Moreover, the features described with respect to certain configurations can be combined in various other configurations, such as different aspects and elements of the configurations can be combined in a similar manner. In addition, as technology develops, the elements therein can be updated, that is, many elements are examples and do not limit the scope of the present disclosure or the claims.

[0177] Specific details are given in the specification to provide a thorough understanding of the exemplary configurations including the implementation. However, the configurations can be practiced without these specific details. For example, well-known circuits, processes, algorithms, structures, and technologies have been shown without unnecessary details to avoid obscuring the configurations. The description only provides example configurations and does not limit the scope, applicability, or configuration of the claims. Instead, the foregoing description of the configurations will provide those skilled in the art with an enabling description for implementing the described technology. Various changes can be made to the functions and arrangements of the elements without departing from the spirit or scope of the present disclosure.

[0178] In summary, it is intended that the above detailed description be considered illustrative rather than restrictive, and it should be understood that the above embodiments should be construed as only illustrative of the present invention and not as limiting the scope of the present invention. After reading the content described in the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent changes and modifications also fall within the scope defined by the claims of the present invention.

Claims

1. The adaptive generation system of network distribution drone routes based on AI vision assistance is characterized by: include: server; The server comprises: A computing unit configured to receive and analyze the status data of the drone, generate, save, and optimize a control strategy for controlling the flight of the drone, and generate a control instruction for controlling the flight of one or more drones according to the control strategy; and A control unit, configured to communicate with one or more drones and send the control instructions to the drones; A master drone is configured to perform a power inspection task on a preset route, including flying and photographing the power equipment on the route, and transmitting the photographed image data to the server; The secondary drone is configured to fly around the primary drone, collect the posture and motion data of the primary drone and its surrounding environment, timely discover and feedback potential obstacles, and transmit the captured image data to the server; The data collected by the master drone is stored as a master record Sa, which includes the state data of the master drone during flight and the image data of the power equipment taken by the master drone; the data collected by the slave drone is stored as a slave record Sb, which includes the state data, action data and interaction data between the master drone and the environment taken by the slave drone during flight; Furthermore, the operation unit performs the following steps: integrating the main record Sa and the secondary record Sb to form a total record S; using the actor-critic algorithm to evaluate the current control strategy based on the total record S; continuously optimizing the current control strategy according to the evaluation result, and finally generating an adaptive strategy optimized for the current route; The operation unit runs a strategy network and an evaluation network for continuously executing the generation and optimization of adaptive strategies based on the actor-critic algorithm; the training of the evaluation network includes: defining the target value target of each sample, for a sample (s, a, r, s′), target = r + ηMax a′ Q(s′,a′); r is the reward value obtained by taking action a in the current state s, which is pre-set by the technician in strategy C; η is the discount factor, which is in the interval [0, 1]; Max a′ Q(s′,a′) represents the maximum Q value of the optimal action in the future, which means the maximum Q value that can be obtained from all possible actions a′ in the next state s′; in the next state s′, the drone can choose different actions a′ in the hope of maximizing the Q value and calculate the maximum Q value.

2. The generation system according to claim 1, characterized in that: The strategy network is configured to output the action of the drone at the next time node based on the current control strategy and the execution effect provided by the evaluation network, taking the current state and / or action of the drone as input; The evaluation network is configured to evaluate the execution effect of the drone in executing the current control strategy and feed back the execution effect to the strategy network.

3. The generation system according to claim 2, characterized in that: The control strategy includes multiple strategies, and the control strategy describes the specific actions that the drone should perform based on the current state and / or action of the drone.

4. The generation system according to claim 3, characterized in that: The operation unit also includes using an invariant set principle to verify the stability of the evaluation network.

5. The generation system according to claim 4, characterized in that: The evaluation network evaluates the execution effect of the UAV's current control strategy, and its evaluation content includes: Is the drone's attitude stable? The distance between the drone and the target power equipment to be inspected and whether its own positioning during inspection is accurate; Image quality of images collected by drones.

6. The generation system according to claim 5, characterized in that: The image quality of the images collected by the drone includes one or more of the following evaluation factors: Whether the target part of the target power equipment is captured in the picture; The proportion of the target part of the target power equipment in the picture; Whether the target part of the target power equipment is clear in the picture.

7. The generation system according to claim 6, characterized in that: The flight control of the secondary drone is achieved by one of the following methods: manually operated by at least one operator via the control unit, being automatically executed by the control unit; The process is automatically performed by at least one onboard control unit deployed on the corresponding drone.

8. The generation system according to claim 7, characterized in that: The main drone and the slave drone are equipped with at least one image acquisition device, and the image acquisition device is composed of one or more of the following sensors or image devices: camera, video camera, thermal imaging camera, infrared camera, night vision sensor, depth camera, ranging sensor, laser ranging sensor, radio ranging sensor.

9. An electronic device, characterized in that: include: A processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor and the memory communicate via the bus, and when the machine-readable instructions are executed by the processor, the functions of the generation system as described in any one of claims 1 to 8 are performed.

Citation Information

Patent Citations

  • Collaborative patrol of drone and driverless car system and method

    CN106774221A

  • Area patrol system with patrol drone

    KR102440819B1

  • Danet-based drone patrol and inspection system for coastline floating garbage

    US20210224512A1

  • Method, device and system for displaying position of slave drone from master drone vision

    CN107966136A

  • Unmanned aerial vehicle automatic inspection system and method for power line

    CN112904890A