Vehicle and unmanned aerial vehicle combined dispatching method for wide-range low-cost inspection

By optimizing the joint scheduling of UAVs and vehicles using mixed integer nonconvex optimization and Markov decision process optimization, the problem of cost-effective scheduling in large-scale inspections is solved, realizing efficient joint inspections of UAVs and vehicles, expanding the inspection range and reducing costs.

CN120998050APending Publication Date: 2025-11-21GUANGDONG UNIV OF TECH

Patent Information

Application Number
CN202510949806.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-09
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

In large-scale inspections, how to jointly schedule the vehicles and drones involved in the inspection based on the overall driving cost of vehicles and the time and energy consumption cost of all drones in order to complete the task quickly and at low cost is an extremely challenging research task.

Method used

By acquiring prior information, the problem is modeled as a mixed-integer nonconvex optimization problem. The nonconvex bilinear terms are linearized and subjected to piecewise linear approximation discretization. By combining Markov decision processes and deep reinforcement learning methods, the scheduling strategy for UAVs and vehicles is optimized.

Benefits of technology

It enables efficient joint scheduling of drones and vehicles, expands the inspection range, reduces costs, and improves the efficiency of data collection and processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120998050A_ABST
    Figure CN120998050A_ABST
Patent Text Reader

Abstract

The invention relates to a vehicle and unmanned aerial vehicle combined scheduling method for wide-range low-cost inspection. The method comprises the following steps: acquiring prior information; modeling the unmanned aerial vehicle inspection problem of each target area according to the prior information to obtain a mixed integer non-convex optimization problem with the goal of minimizing the weighted sum of the total execution time and the energy consumption of all the inspection unmanned aerial vehicles; performing linearization on a non-convex bilinear term in the mixed integer non-convex optimization problem, and performing discretization processing on a nonlinear function by adopting piecewise linear approximation; an approximate mixed integer linear programming problem is obtained and solved, and an unmanned aerial vehicle scheduling strategy is obtained; modeling according to the unmanned aerial vehicle scheduling strategy and the prior information to obtain an inspection vehicle path planning problem taking the comprehensive driving cost as a target; the routing inspection vehicle path planning problem is converted and modeled into a Markov decision process, a routing inspection vehicle is used as an intelligent agent, a state, an action and a reward function are defined, and a routing inspection vehicle scheduling strategy is obtained. Therefore, combined inspection of the inspection vehicle and the unmanned aerial vehicle is realized, and the inspection range is expanded.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of smart cities, drones and mobile edge computing, and in particular to a method for joint scheduling of vehicles and drones for large-scale, low-cost inspections. Background Technology

[0002] In the construction of smart cities, city-level information sensing and processing will play a crucial role. However, traditional fixed-deployment sensor systems are no longer sufficient to meet the demands for multi-point, dynamic, and low-latency data collection in areas such as urban traffic, environment, and security. Unmanned Aerial Vehicles (UAVs), due to their high mobility, high-altitude perspective, and flexible deployment, have become a common tool for target area inspection and are widely used in inspection scenarios such as target tracking, pedestrian recognition, and traffic flow prediction. During urban inspections, the environment is complex and changeable. Relying solely on a single UAV often faces problems such as long task times, insufficient inspection coverage, and lack of fault tolerance. Therefore, it is necessary to schedule multiple UAVs to perform inspection tasks simultaneously. In addition, combined vehicle and UAV inspections, and joint processing of collected data by UAVs and local mobile edge computing (MEC) servers within the target area, provide feasible technical solutions for large-scale inspections of multiple target areas.

[0003] On the one hand, by fully utilizing the mobility of vehicles and the maneuverability of drones, vehicles can quickly reach parking points near the target area, deploy multiple drones to perform inspection tasks in the target area, and then the vehicles wait for the drones to complete their inspection tasks and retrieve all drones. The vehicles can also recharge the drones in a timely manner, overcoming the limitation of insufficient drone battery life that prevents long-distance flight. This combined vehicle-drone inspection expands the inspectable range. On the other hand, when different drones are performing data collection and processing in the target area, to support drones in completing computationally intensive inspection tasks (such as traffic prediction, target monitoring and tracking, and behavior analysis), 5G-era mobile edge computing technology is combined with the deployment of MEC servers on the local base station side. This provides additional storage space and computing resources for drones with limited storage capacity and computing power, accelerating the collection and on-site processing of surrounding data and improving the responsiveness of inspection services.

[0004] Existing technologies include combining drones and vehicles to perform tasks (such as logistics delivery), drones taking off and landing on vehicle-mounted platforms, drones safely and efficiently reaching designated locations through path planning, and drones jointly processing data with MEC servers. These technologies lay the foundation for the comprehensive solution proposed in this patent, which combines vehicles and drones and uses MEC to assist drone inspections to accelerate data collection and processing. However, in large-scale inspection solutions, how to jointly schedule participating vehicles and drones based on the overall vehicle operating cost and the time and energy consumption cost of all drones to complete large-scale inspections quickly and at low cost is a highly challenging research task. Summary of the Invention

[0005] Therefore, it is necessary to provide a more efficient method for joint scheduling of vehicles and drones for large-scale, low-cost inspections, addressing the aforementioned technical problems. This method includes:

[0006] S1: Obtain prior information;

[0007] S2: Based on prior information, the UAV inspection problem of each target area will be modeled to obtain a mixed integer non-convex optimization problem with the objective of minimizing the weighted sum of the total execution time and energy consumption of all inspection UAVs;

[0008] S3: Linearize the non-convex bilinear terms in the mixed-integer non-convex optimization problem, and discretize the nonlinear function using a piecewise linear approximation; obtain an approximate mixed-integer linear programming problem and solve it to obtain the UAV scheduling strategy;

[0009] S4: Based on the UAV scheduling strategy and prior information, a model is built to obtain the inspection vehicle path planning problem with the overall travel cost as the objective.

[0010] S5: Transform the inspection vehicle path planning problem into a Markov decision process model. Using the inspection vehicle as an agent, define its state, action, and reward function to obtain the inspection vehicle scheduling strategy.

[0011] Furthermore, in step S1, the prior information includes the number of target areas, the observation points that the UAV needs to reach in each target area and the type of inspection task to be performed, and all state parameters of the UAV and inspection vehicle during the inspection process. All state parameters of the UAV and inspection vehicle during the inspection process include: the transmit / receive power of data transmission between the UAV and the base station, channel state parameters, available computing, storage and communication resources of the UAV / MEC server, the unit travel distance of the inspection vehicle, travel speed, power consumption, and maintenance cost.

[0012] Further, step S2 specifically involves modeling based on the communication model, latency model, and energy consumption model, and setting optimization objectives and resource capacity constraints to obtain a mixed integer nonconvex optimization problem with the objective of minimizing the weighted sum of the total execution time and energy consumption of all inspection drones.

[0013] Furthermore, the communication model adopts a quasi-static decision-making model. After the inspection vehicle reaches the parking point of a target area, multiple drones are dispatched to different observation points to perform computationally intensive inspection tasks. When each drone reaches its designated observation point, it remains hovering to ensure its aerial position remains fixed, and it communicates wirelessly with the local MEC server through a line-of-sight channel. The wireless channel is a quasi-static channel, meaning that the channel state remains unchanged during data transmission. The target area is j, and one base station j is deployed within the area, using I... j This represents the set of all drones performing tasks within this area, and a given drone is labeled as i∈I. j The wireless communication between the UAV and the base stations in the area adopts orthogonal frequency division multiple access technology. The channel power gain adopts the free space path loss model, and the downlink data rate from base station j to UAV i is defined as follows:

[0014]

[0015] Among them, b i,j For the communication bandwidth allocated to UAV i, l i,j Let h be the communication distance between UAV i at the observation point and base station j. i,j Here, Q represents the channel gain at a reference distance of 1 meter, α is the channel fading exponent, and Q... i,j Let Ni be the received power of UAVi, and N0 be the noise power; the logarithmic function in equation (1) contains constants, and constants are used. Replace the logarithmic part; define the data uplink rate from UAV i to base station j as

[0016]

[0017] Among them, P i,j The transmit power of UAV i is represented by a constant. Replace the logarithmic part;

[0018] Delay model: Considering the binary computation offloading method, define a 0-1 variable a. i,j a i,j =0 indicates that UAV i performs the computation task locally, a i,j =1 indicates that UAV i will offload the computing task to MEC server j, and has

[0019]

[0020] If the UAV performs the computation task locally, then the total time for UAV i is

[0021]

[0022] Among them, the first term on the right side of the equation is 2τ i,j The first item is the time taken for UAV i to travel to and return to the associated observation point; the second item is the amount of data D collected by UAV i. i,j The time taken, s i,j The first is the data acquisition rate; the third is the time taken for UAV i to download the specified application. i,j m is the amount of data requested by the application. i,j =1 indicates application A i,j Already cached on MEC servers j, m i,j =0 indicates that the current MEC server j does not cache A. i,j A needs to be obtained from the cloud. i,j , This is the data transfer rate from the cloud to the MEC server j; the last item is the CPU computation time of UAV i. For UAV i's local computing resources, W i,j To process data D i,j The required number of CPU cycles, after offloading the UAV i computation task to MEC server j, results in the total time spent on UAV i being...

[0023]

[0024] in, This refers to the computing resources allocated by MEC server j to UAV i task. The total time cost of UAV i is expressed as...

[0025]

[0026] Energy consumption model: Considering the communication, computing, and hovering energy consumption of the UAV in the system, and assuming UAV i performs computing tasks locally, the relevant energy consumption is as follows:

[0027]

[0028] In this equation, the first term on the right-hand side represents the energy consumption of the UAV i when downloading applications, and the second term represents the CPU energy consumption of the UAV i. i Let be the energy consumption factor of UAV i, and represent the effective switching capacitor of the CPU. If the computational tasks of UAV i are offloaded to MEC server j, then the relevant energy consumption is:

[0029]

[0030] The first term on the right side of the equals sign represents the energy consumption of the data collected by the UAV i during upload, and the second term contains the parameter ò. j Let be the CPU power consumption coefficient of MEC server j, and let be the flight power of UAV i.

[0031]

[0032] The first term on the right side of the equals sign represents the blade profile power of UAV i during flight. Let ξ be the blade profile power of UAV i in hovering state. i η is the tip velocity of the UAV i rotor blades. i The first term is the flight speed of UAV i, and the second term is the induced power of UAV i during flight. ηi represents the induced power of UAV i in hovering, η0 represents the average rotor induced velocity of UAV i in hovering, and the third term represents the air drag power of UAV i in flight. These are the fixed parameters related to the characteristics of UAV i; the hovering power of UAV i is...

[0033]

[0034] Then the flight hovering energy consumption of UAV i is

[0035]

[0036] Finally, the system energy consumption resulting from processing UAV i computational tasks is expressed as:

[0037]

[0038] Optimization Objective: In the MEC-assisted UAV inspection problem in target region j, the optimization objective is to minimize the weighted sum of the total energy consumption of all UAVs and MEC servers in this region, as well as the execution time of all UAVs. This minimization problem is denoted as P1, and takes the following form:

[0039]

[0040] in, The maximum time for UAV i to perform an inspection task at parking point j in the target area is, i.e. δ1 and δ2 are the weights for controlling energy consumption and time;

[0041] Resource capacity constraint: Let the storage capacity of MEC server j be C. s,j There are MEC server data storage constraints.

[0042]

[0043] Let B be the total communication bandwidth, with a bandwidth constraint.

[0044]

[0045]

[0046] Let F be the total computing resources of the MEC server. For all unloaded tasks, there is a constraint on the total computing resources.

[0047]

[0048]

[0049] set up To determine the maximum time for UAV i to perform its task at parking point j in the target area, the constraint is:

[0050]

[0051] Furthermore, step S3 specifically includes:

[0052] Based on the modeling transformation method of linearization and discretization, by introducing auxiliary variables and constraints, the non-convex bilinear terms of the mixed integer non-convex optimization problem are linearized. At the same time, the nonlinear function is discretized by piecewise linear approximation. Furthermore, a discrete point generation algorithm is designed to automatically control the discretization accuracy and linearization error, thereby transforming the mixed integer non-convex optimization problem into an approximate mixed integer linear programming problem for solution.

[0053] Furthermore, step S3 specifically includes:

[0054] The convex envoy method is used to linearize the bilinear terms, where equations (6) and (12) contain non-convex bilinear terms. Both are products of a 0-1 variable and a continuous variable, i.e., linearization is performed using the convex hull method; with bilinear terms... For example, we can linearize it equivalently as follows:

[0055]

[0056]

[0057] in, They are The upper and lower bounds; note that at this point, we are using... Treating it as an independent variable; linear inequalities (20) and (21) yield: when a i,j When = 0, we have when a i,j When = 1, we have Using variables that satisfy equations (20) and (21) To equivalently replace the original bilinear term Linearize the other bilinear terms using the same method;

[0058] Next, we introduce new variables. and To replace the original ones respectively and Right now and Therefore, equations (4) and (5) can be rewritten as follows:

[0059]

[0060] Equations (22) and (23) are both linear equations. Equations (7) and (8) can be rewritten as follows:

[0061]

[0062] Among them, equations (24) and (25) are both linear equations, g i,j As an intermediate variable; further relax equation (26) to equivalently

[0063]

[0064] Equation (27) specifies g i,j The lower bound is Optimal g i,j The value will equal Equation (27) is equivalent to replacing equation (26);

[0065] The total bandwidth constraint is rewritten as follows

[0066]

[0067] set up For a sufficiently large value, the total resource constraint is rewritten as follows:

[0068]

[0069] set up For a sufficiently large value; by introducing a new variable The transformed optimization problem still contains nonlinear functions: in equation (27) In equation (28) In the formula (30)

[0070] By employing discretization techniques, a linear function is used to approximate a nonlinear function: For example, let The range of discretized values ​​is Define the set of discrete points Where K1 is the number of discrete points, and Choose any two adjacent discrete points A straight line is defined through these two points.

[0071]

[0072] Then, equations (33)-(34) are used to approximate the nonlinear constraint (30).

[0073]

[0074] Where, θ i,j It is an intermediate variable, and equation (33) contains a bilinear term. right Define another set of discrete points Definition process straight line Then, equation (35) is used to approximate the nonlinear constraint (27).

[0075]

[0076] against Define the set of discrete points Definition process straight line Then, equations (36)-(37) are used to approximate the nonlinear constraint (28);

[0077]

[0078] Where, ω i,j It is an intermediate variable;

[0079] Discrete point generation algorithm:

[0080] function Functions on the positive real number line are convex functions. Discrete points are generated using the properties of convex functions. Let f(r) denote a convex function, where r satisfies 0. <r min ≤r≤r max Define the set of discrete points. There is r min =r1<... <r K =r max After two points r k ,r k+1 Define a straight line

[0081]

[0082] In r k ≤r≤r k+1 Within the range, there is always lk,k+1 (r)≥f(r); therefore, using l k,k+1 The maximum perpendicular error of f(r) approximated by f(r) is

[0083]

[0084] By analyzing its KKT conditions, the optimal solution is obtained as follows:

[0085]

[0086] That is, at point r * At this point, the maximum value of the approximate error Δ can be obtained. * ;

[0087] Generate a set of discrete points Specific steps:

[0088] The algorithm's input parameter δ is the preset maximum permissible error. δ0 = 0.1δ is set to ensure that the actual maximum error Δ... * The point is near δ and less than or equal to δ; initially, there is only one discrete point r1 in the set of discrete points R. Starting from the first point r1, find a second point r2 such that the distance between the two points is Δ. * Near δ, find a third point r3 such that the Δ between r2 and r3 is... * This process is repeated near δ until the last point r is reached. max Each time a discrete point is generated, the maximum vertical error Δ can be guaranteed. * ≤δ, therefore, in the worst case, the approximation error of equations (33) and (36) is Iδ, and the approximation error of equation (35) is δ; based on simple algebraic calculations and using the bisection method to find new points, a suitable set of discrete points R can be quickly obtained; before solving P1, for Run the algorithm once each, and then substitute the generated discrete points into constraints (33), (35), and (36).

[0089] Furthermore, step S4 specifically includes:

[0090] A vehicle travel time cost model, a vehicle electricity consumption cost model, and a vehicle maintenance cost model are established, and a path planning optimization objective and constraints are set to construct an inspection vehicle path planning problem with comprehensive travel cost as the objective.

[0091] Furthermore, step S4 specifically includes:

[0092] Vehicle travel time cost model: The time cost incurred during the inspection vehicle's journey is determined by the distance and speed it travels.

[0093]

[0094] Where d nm The Euclidean distance between parking spots in the target area;

[0095] Vehicle power consumption cost model: The power generated by the inspection vehicle during operation due to rolling resistance and air resistance is...

[0096]

[0097]

[0098] Among them, f r It is the rolling resistance coefficient, C d Here, m is the air drag coefficient, g is the vehicle mass, ρ is the air density, and A is the air drag coefficient. f Where W is the frontal area of ​​the vehicle and W is the wind speed, the electricity cost of the inspection vehicle during operation is...

[0099]

[0100] Vehicle maintenance cost model: Inspection vehicles may experience wear and tear due to prolonged driving, requiring maintenance. The maintenance cost is...

[0101] M nm =μ·d nm (100)

[0102] Where μ is the maintenance factor;

[0103] Path planning optimization objective: To minimize the overall travel cost of inspection vehicles, consisting of travel time, power consumption, and maintenance costs; where x nm =1 indicates that the inspection vehicle passes through the section of road from parking point n to parking point m in the target area, x nm =0 indicates that the inspection vehicle did not pass through the section from target area parking point n to target area parking point m; the set of target area parking points is S = {0, 1, 2, ..., N}, where N is the number of target area parking points, and 0 is the starting point of the inspection vehicle. The problem of minimizing this set is denoted as P2, and takes the following form.

[0104]

[0105] in, Let λ1, λ2, and λ3 be the power consumption required for UAV i to return to the inspection vehicle in target area j and charge it; λ1, λ2, and λ3 are the weights of control time, power consumption, and maintenance cost.

[0106] Constraints:

[0107] Each parking spot in the target area can only be visited once, as expressed by the constraint condition.

[0108]

[0109] Some road sections in the city are under traffic control and cannot be passed. If the traffic-controlled road section (n,m)∈σ is impassable, then the constraint is:

[0110]

[0111] To prevent inspection vehicles from forming sub-loops when visiting parking points in the target area, an auxiliary variable u is introduced. n This indicates the order of visits to parking point n in the target area, with the following constraints:

[0112]

[0113] Where u n It is a continuous variable, representing the order in which task point n is visited;

[0114] Assuming the inspection vehicle must start from and return to starting point 0, the constraint is:

[0115]

[0116] Furthermore, step S5 specifically includes:

[0117] Based on the Double DQN algorithm of deep reinforcement learning, the path planning problem of inspection vehicles is transformed into a Markov decision process. The inspection vehicle is used as the agent, and the state, action and reward functions are defined. The state-action value function is approximated by a neural network. The stability and accuracy of policy learning are improved by combining experience storage and a dual network structure. Finally, the optimal path planning of the inspection vehicle among all parking points in the target area is achieved.

[0118] Furthermore, step S5 specifically includes:

[0119] State: In the inspection vehicle path planning problem, the state is a binary tuple s. t =(s c ,s v ), where s c Indicates the current target area parking point, s v This indicates the target area parking points that have already been visited; the above state design reflects the relevance of the last inspection vehicle's travel path, helping to find the path with the lowest overall travel cost.

[0120] Action: The action is "select the next target area parking spot to visit"; the action serves as input to the environment, which simulates the behavior of the inspection vehicle.

[0121] Rewards: The reward function is used to guide the agent to learn the optimal policy; the goal is to minimize the total cost, and the immediate reward is the inverse of the cost of the current action. This design encourages the agent to avoid high-cost paths and gradually learn a feasible path with the lowest cost.

[0122] DDQN algorithm:

[0123] In each round of training, the agent interacts with the environment to generate training samples; first, the agent observes the current main network state s. t In state s t Below, the agent selects action a. t The ε-Greedy strategy is used for action selection, that is:

[0124]

[0125] Where Q(s) t ,a t ;θ), representing the current state-action value function of the main network, the parameter θ is continuously updated during training, while ε gradually decreases with each training round;

[0126] Perform action a t Afterwards, the agent receives an immediate reward r from the environment. t and transition to the next state s t+1 This generates a complete state transition sample (s). t ,a t ,r t ,s t+1 ); Introducing an experience replay mechanism: state transition samples (s) generated during training. t ,a t ,r t ,s t+1 The data is stored in the experience storage pool R. When the number of samples accumulates to a certain extent, a small batch of data is randomly extracted from it for training, which effectively breaks the temporal correlation between samples and improves training efficiency.

[0127] The training objective is to minimize the mean squared error between the current state-action value and the target value. To mitigate the overestimation problem in Q-value estimation, a Double DQN structure is adopted, and its target Q-value is expressed by the following formula:

[0128]

[0129] Where γ∈[0,1) is the discount factor, and θ is the parameter of the current main network. - For fixed target network parameters; the loss function is defined as:

[0130]

[0131] All parameters θ of the Q-network are updated through gradient backpropagation in the neural network; after every few training rounds, the parameters θ of the main network are synchronously copied to the target network to update θ. - .

[0132] This invention first linearizes the non-convex bilinear terms in the mixed-integer non-convex optimization problem, and then discretizes the nonlinear function using a piecewise linear approximation. This yields an approximate mixed-integer linear programming problem, which is then solved to obtain a drone scheduling strategy. Furthermore, the inspection vehicle path planning problem is transformed into a Markov decision process, with the inspection vehicle as the agent, defining its state, actions, and reward functions to obtain the inspection vehicle scheduling strategy. This enables combined inspection by inspection vehicles and drones, thereby expanding the inspectable area. Attached Figure Description

[0133] Figure 1 This is a flowchart illustrating a vehicle and drone joint scheduling method for large-scale, low-cost inspection in one embodiment.

[0134] Figure 2 This is a schematic diagram of the architecture of a vehicle and drone joint scheduling method for large-scale, low-cost inspection in one embodiment.

[0135] Figure 3 This is a schematic diagram of the Double DQN algorithm architecture in one embodiment; Detailed Implementation

[0136] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0137] Example 1:

[0138] This embodiment provides, as follows: Figure 1 The following is a method for joint scheduling of vehicles and drones for large-scale, low-cost inspections:

[0139] S1: Obtain prior information;

[0140] S2: Based on prior information, the UAV inspection problem of each target area will be modeled to obtain a mixed integer non-convex optimization problem with the objective of minimizing the weighted sum of the total execution time and energy consumption of all inspection UAVs;

[0141] S3: Linearize the non-convex bilinear terms in the mixed-integer non-convex optimization problem, and discretize the nonlinear function using a piecewise linear approximation; obtain an approximate mixed-integer linear programming problem and solve it to obtain the UAV scheduling strategy;

[0142] S4: Based on the UAV scheduling strategy and prior information, a model is built to obtain the inspection vehicle path planning problem with the overall travel cost as the objective.

[0143] S5: Transform the inspection vehicle path planning problem into a Markov decision process model. Using the inspection vehicle as an agent, define its state, action, and reward function to obtain the inspection vehicle scheduling strategy.

[0144] This embodiment first linearizes the non-convex bilinear terms in the mixed-integer non-convex optimization problem, and then discretizes the nonlinear function using a piecewise linear approximation. This yields an approximate mixed-integer linear programming problem, which is then solved to obtain the UAV scheduling strategy. Furthermore, the inspection vehicle path planning problem is transformed into a Markov decision process, with the inspection vehicle as the agent, defining its state, actions, and reward functions to obtain the inspection vehicle scheduling strategy. This enables combined inspection by inspection vehicles and UAVs, thereby expanding the inspectable area.

[0145] Example 2:

[0146] This embodiment is based on Embodiment 1, and further details the following: Figure 2 The method for joint scheduling of vehicles and drones for large-scale, low-cost inspections is further disclosed below:

[0147] Furthermore, in step S1, the prior information includes the number of target areas, the observation points that the UAV needs to reach in each target area and the type of inspection task to be performed, and all state parameters of the UAV and inspection vehicle during the inspection process. All state parameters of the UAV and inspection vehicle during the inspection process include: the transmit / receive power of data transmission between the UAV and the base station, channel state parameters, available computing, storage and communication resources of the UAV / MEC server, the unit travel distance of the inspection vehicle, travel speed, power consumption, and maintenance cost.

[0148] Further, step S2 specifically involves modeling based on the communication model, latency model, and energy consumption model, and setting optimization objectives and resource capacity constraints to obtain a mixed integer nonconvex optimization problem with the objective of minimizing the weighted sum of the total execution time and energy consumption of all inspection drones.

[0149] Furthermore, the communication model adopts a quasi-static decision-making model. After the inspection vehicle reaches the parking point of a target area, multiple drones are dispatched to different observation points to perform computationally intensive inspection tasks. When each drone reaches its designated observation point, it remains hovering to ensure its aerial position remains fixed, and it communicates wirelessly with the local MEC server through a line-of-sight channel. The wireless channel is a quasi-static channel, meaning that the channel state remains unchanged during data transmission. The target area is j, and one base station j is deployed within the area, using I... jThis represents the set of all drones performing tasks within this area, and a given drone is labeled as i∈I. j The wireless communication between the UAV and the base stations in the area adopts orthogonal frequency division multiple access technology. The channel power gain adopts the free space path loss model, and the downlink data rate from base station j to UAV i is defined as follows:

[0150]

[0151] Among them, b i,j For the communication bandwidth allocated to UAV i, l i,j Let h be the communication distance between UAV i at the observation point and base station j. i,j Here, Q represents the channel gain at a reference distance of 1 meter, α is the channel fading exponent, and Q... i,j Let Ni be the received power of UAVi, and N0 be the noise power; the logarithmic function in equation (1) contains constants, and constants are used. Replace the logarithmic part; define the data uplink rate from UAV i to base station j as

[0152]

[0153] Among them, P i,j The transmit power of UAV i is represented by a constant. Replace the logarithmic part;

[0154] Delay model: Considering the binary computation offloading method, define a 0-1 variable a. i,j a i,j =0 indicates that UAV i performs the computation task locally, a i,j =1 indicates that UAV i will offload the computing task to MEC server j, and has

[0155]

[0156] If the UAV performs the computation task locally, then the total time for UAV i is

[0157]

[0158] Among them, the first term on the right side of the equation is 2τ i,j The first item is the time taken for UAV i to travel to and return to the associated observation point; the second item is the amount of data D collected by UAV i. i,j The time taken, s i,j The first is the data acquisition rate; the third is the time taken for UAV i to download the specified application. i,j m is the amount of data requested by the application. i,j =1 indicates application A i,j Already cached on MEC servers j, m i,j=0 indicates that the current MEC server j does not cache A. i,j A needs to be obtained from the cloud. i,j , This is the data transfer rate from the cloud to the MEC server j; the last item is the CPU computation time of UAV i. For UAV i's local computing resources, W i,j To process data D i,j The required number of CPU cycles, after offloading the UAV i computation task to MEC server j, results in the total time spent on UAV i being...

[0159]

[0160] in, This refers to the computing resources allocated by MEC server j to UAV i task. The total time cost of UAV i is expressed as...

[0161]

[0162] Energy consumption model: Considering the communication, computing, and hovering energy consumption of the UAV in the system, and assuming UAV i performs computing tasks locally, the relevant energy consumption is as follows:

[0163]

[0164] In this equation, the first term on the right-hand side represents the energy consumption of the UAV i when downloading applications, and the second term represents the CPU energy consumption of the UAV i. i Let be the energy consumption factor of UAV i, and represent the effective switching capacitor of the CPU. If the computational tasks of UAV i are offloaded to MEC server j, then the relevant energy consumption is:

[0165]

[0166] The first term on the right side of the equals sign represents the energy consumption of the data collected by the UAV i during upload, and the second term contains the parameter ò. j Let be the CPU power consumption coefficient of MEC server j, and let be the flight power of UAV i.

[0167]

[0168] The first term on the right side of the equation is the blade profile power of UAV i during flight, K. i 0 Let ξ be the blade profile power of UAV i in hovering state. i η is the tip velocity of the UAV i rotor blades. i The first term is the flight speed of UAV i, and the second term is the induced power of UAV i during flight. ηi represents the induced power of UAV i in hovering, η0 represents the average rotor induced velocity of UAV i in hovering, and the third term represents the air drag power of UAV i in flight. These are the fixed parameters related to the characteristics of UAV i; the hovering power of UAV i is...

[0169]

[0170] Then the flight hovering energy consumption of UAV i is

[0171]

[0172] Finally, the system energy consumption resulting from processing UAV i computational tasks is expressed as:

[0173]

[0174] Optimization Objective: In the MEC-assisted UAV inspection problem in target region j, the optimization objective is to minimize the weighted sum of the total energy consumption of all UAVs and MEC servers in this region, as well as the execution time of all UAVs. This minimization problem is denoted as P1, and takes the following form:

[0175]

[0176] in, The maximum time for UAV i to perform an inspection task at parking point j in the target area is, i.e. δ1 and δ2 are the weights for controlling energy consumption and time;

[0177] Resource capacity constraint: Let the storage capacity of MEC server j be C. s,j There are MEC server data storage constraints.

[0178]

[0179] Let B be the total communication bandwidth, with a bandwidth constraint.

[0180]

[0181] Let F be the total computing resources of the MEC server. For all unloaded tasks, there is a constraint on the total computing resources.

[0182]

[0183] set up To determine the maximum time for UAV i to perform its task at parking point j in the target area, the constraint is:

[0184]

[0185] Furthermore, step S3 specifically includes:

[0186] Based on the modeling transformation method of linearization and discretization, by introducing auxiliary variables and constraints, the non-convex bilinear terms of the mixed integer non-convex optimization problem are linearized. At the same time, the nonlinear function is discretized by piecewise linear approximation. Furthermore, a discrete point generation algorithm is designed to automatically control the discretization accuracy and linearization error, thereby transforming the mixed integer non-convex optimization problem into an approximate mixed integer linear programming problem for solution.

[0187] Furthermore, step S3 specifically includes:

[0188] The convex envoy method is used to linearize the bilinear terms, where equations (6) and (12) contain non-convex bilinear terms. Both are products of a 0-1 variable and a continuous variable, i.e., linearization is performed using the convex hull method; with bilinear terms... For example, we can linearize it equivalently as follows:

[0189]

[0190] in, They are The upper and lower bounds; note that at this point, we are using... Treating it as an independent variable; linear inequalities (20) and (21) yield: when a i,j When = 0, we have when a i,j When = 1, we have Using variables that satisfy equations (20) and (21) To equivalently replace the original bilinear term Linearize the other bilinear terms using the same method;

[0191] Next, we introduce new variables. and To replace the original ones respectively and Right now and Therefore, equations (4) and (5) can be rewritten as follows:

[0192]

[0193] Equations (22) and (23) are both linear equations. Equations (7) and (8) can be rewritten as follows:

[0194]

[0195] Among them, equations (24) and (25) are both linear equations, g i,jAs an intermediate variable; further relax equation (26) to equivalently

[0196]

[0197] Equation (27) specifies g i,j The lower bound is Optimal g i,j The value will equal Equation (27) is equivalent to replacing equation (26);

[0198] The total bandwidth constraint is rewritten as follows

[0199]

[0200]

[0201] set up For a sufficiently large value, the total resource constraint is rewritten as follows:

[0202]

[0203] set up For a sufficiently large value; by introducing a new variable The transformed optimization problem still contains nonlinear functions: in equation (27) In equation (28) In the formula (30)

[0204] By employing discretization techniques, a linear function is used to approximate a nonlinear function: For example, let The range of discretized values ​​is Define the set of discrete points Where K1 is the number of discrete points, and Choose any two adjacent discrete points A straight line is defined through these two points.

[0205]

[0206] Then, equations (33)-(34) are used to approximate the nonlinear constraint (30).

[0207]

[0208] Where, θ i,j It is an intermediate variable, and equation (33) contains a bilinear term. right Define another set of discrete points Definition process straight line Then, equation (35) is used to approximate the nonlinear constraint (27).

[0209]

[0210] against Define the set of discrete points Definition process straight line Then, equations (36)-(37) are used to approximate the nonlinear constraint (28);

[0211]

[0212] Where, ω i,j It is an intermediate variable;

[0213] Discrete point generation algorithm:

[0214] function Functions on the positive real number line are convex functions. Discrete points are generated using the properties of convex functions. Let f(r) denote a convex function, where r satisfies 0. <r min ≤r≤r max Define the set of discrete points. There is r m i n =r1<... <r K =r max After two points r k ,r k+1 Define a straight line

[0215]

[0216] In r k ≤r≤r k+1 Within the range, there is always l k,k+1 (r)≥f(r); therefore, using l k,k+1 The maximum perpendicular error of f(r) approximated by f(r) is

[0217]

[0218] By analyzing its KKT conditions, the optimal solution is obtained as follows:

[0219]

[0220] That is, at point r * At this point, the maximum value of the approximate error Δ can be obtained. * ;

[0221] Generate a set of discrete points Specific steps:

[0222] The algorithm's input parameter δ is the preset maximum permissible error. δ0 = 0.1δ is set to ensure that the actual maximum error Δ... * The point is near δ and less than or equal to δ; initially, there is only one discrete point r1 in the set of discrete points R. Starting from the first point r1, find a second point r2 such that the distance between the two points is Δ. * Near δ, find a third point r3 such that the Δ between r2 and r3 is... * This process is repeated near δ until the last point r is reached. max Each time a discrete point is generated, the maximum vertical error Δ can be guaranteed. * ≤δ, therefore, in the worst case, the approximation error of equations (33) and (36) is Iδ, and the approximation error of equation (35) is δ; based on simple algebraic calculations and using the bisection method to find new points, a suitable set of discrete points R can be quickly obtained; before solving P1, for Run the algorithm once each, and then substitute the generated discrete points into constraints (33), (35), and (36).

[0223] This embodiment first linearizes the non-convex bilinear terms in the mixed-integer non-convex optimization problem, and then discretizes the nonlinear function using a piecewise linear approximation. This yields an approximate mixed-integer linear programming problem, which is then solved to obtain the UAV scheduling strategy. Furthermore, the inspection vehicle path planning problem is transformed into a Markov decision process, with the inspection vehicle as the agent, defining its state, actions, and reward functions to obtain the inspection vehicle scheduling strategy. This enables combined inspection by inspection vehicles and UAVs, thereby expanding the inspectable area.

[0224] Example 3:

[0225] This embodiment further discloses information based on Embodiment 2:

[0226] Furthermore, step S4 specifically includes:

[0227] A vehicle travel time cost model, a vehicle electricity consumption cost model, and a vehicle maintenance cost model are established, and a path planning optimization objective and constraints are set to construct an inspection vehicle path planning problem with comprehensive travel cost as the objective.

[0228] Furthermore, step S4 specifically includes:

[0229] Vehicle travel time cost model: The time cost incurred during the inspection vehicle's journey is determined by the distance and speed it travels.

[0230]

[0231] Where d nm The Euclidean distance between parking spots in the target area;

[0232] Vehicle power consumption cost model: The power generated by the inspection vehicle during operation due to rolling resistance and air resistance is...

[0233]

[0234]

[0235] Among them, f r It is the rolling resistance coefficient, C d Here, m is the air drag coefficient, g is the vehicle mass, ρ is the air density, and A is the air drag coefficient. f Where W is the frontal area of ​​the vehicle and W is the wind speed, the electricity cost of the inspection vehicle during operation is...

[0236]

[0237] Vehicle maintenance cost model: Inspection vehicles may experience wear and tear due to prolonged driving, requiring maintenance. The maintenance cost is...

[0238] M nm =μ·d nm (155)

[0239] Where μ is the maintenance factor;

[0240] Path planning optimization objective: To minimize the overall travel cost of inspection vehicles, consisting of travel time, power consumption, and maintenance costs; where x nm =1 indicates that the inspection vehicle passes through the section of road from parking point n to parking point m in the target area, x nm =0 indicates that the inspection vehicle did not pass through the section from target area parking point n to target area parking point m; the set of target area parking points is S = {0, 1, 2, ..., N}, where N is the number of target area parking points, and 0 is the starting point of the inspection vehicle. The problem of minimizing this set is denoted as P2, and takes the following form.

[0241]

[0242] in, For the target region j, UAV i The power consumption incurred when returning to the inspection vehicle and charging it; λ1, λ2, and λ3 are the weights of control time, power consumption, and maintenance cost;

[0243] Constraints:

[0244] Each parking spot in the target area can only be visited once, as expressed by the constraint condition.

[0245]

[0246] Some road sections in the city are under traffic control and cannot be passed. If the traffic-controlled road section (n,m)∈σ is impassable, then the constraint is:

[0247]

[0248] To prevent inspection vehicles from forming sub-loops when visiting parking points in the target area, an auxiliary variable u is introduced. n This indicates the order of visits to parking point n in the target area, with the following constraints:

[0249]

[0250] Where u n It is a continuous variable, representing the order in which task point n is visited;

[0251] Assuming the inspection vehicle must start from and return to starting point 0, the constraint is:

[0252]

[0253] Furthermore, step S5 specifically includes:

[0254] Based on the Double DQN algorithm of deep reinforcement learning, the path planning problem of inspection vehicles is transformed into a Markov decision process. The inspection vehicle is used as the agent, and the state, action and reward functions are defined. The state-action value function is approximated by a neural network. The stability and accuracy of policy learning are improved by combining experience storage and a dual network structure. Finally, the optimal path planning of the inspection vehicle among all parking points in the target area is achieved.

[0255] Furthermore, step S5 specifically includes:

[0256] State: In the inspection vehicle path planning problem, the state is a binary tuple s. t =(s c ,s v ), where s c Indicates the current target area parking point, s v This indicates the target area parking points that have already been visited; the above state design reflects the relevance of the last inspection vehicle's travel path, helping to find the path with the lowest overall travel cost.

[0257] Action: The action is "select the next target area parking spot to visit"; the action serves as input to the environment, which simulates the behavior of the inspection vehicle.

[0258] Rewards: The reward function is used to guide the agent to learn the optimal policy; the goal is to minimize the total cost, and the immediate reward is the inverse of the cost of the current action. This design encourages the agent to avoid high-cost paths and gradually learn a feasible path with the lowest cost.

[0259] DDQN algorithm:

[0260] In each round of training, the agent interacts with the environment to generate training samples; first, the agent observes the current main network state s. t In state s t Below, the agent selects action a. t The ε-Greedy strategy is used for action selection, that is:

[0261]

[0262] Where Q(s) t ,a t ;θ), representing the current state-action value function of the main network, the parameter θ is continuously updated during training, while ε gradually decreases with each training round;

[0263] Perform action a t Afterwards, the agent receives an immediate reward r from the environment. t and transition to the next state s t+1 This generates a complete state transition sample (s). t ,a t ,r t ,s t+1 ); Introducing an experience replay mechanism: state transition samples (s) generated during training. t ,a t ,r t ,s t+1 The data is stored in the experience storage pool R. When the number of samples accumulates to a certain extent, a small batch of data is randomly extracted from it for training, which effectively breaks the temporal correlation between samples and improves training efficiency.

[0264] The training objective is to minimize the mean squared error between the current state-action value and the target value. To mitigate the overestimation problem in Q-value estimation, a Double DQN structure is adopted, and its target Q-value is expressed by the following formula:

[0265]

[0266] Where γ∈[0,1) is the discount factor, and θ is the parameter of the current main network. - For fixed target network parameters; the loss function is defined as:

[0267]

[0268] All parameters θ of the Q-network are updated through gradient backpropagation in the neural network; after every few training rounds, the parameters θ of the main network are synchronously copied to the target network to update θ. - .

[0269] Detailed process of DDQN algorithm

[0270] First, the parameters of the main network and the target network are initialized to be identical initially, and an experience storage pool R is constructed to store training samples. Relevant hyperparameters are set, including the learning rate, discount factor, exploration rate, and their decay policy. In each training round, the agent first observes the current state s. t And select action a based on the ε-Greedy policy. t This involves randomly exploring with a certain probability and selecting the action with the highest current Q value with a relatively high probability. After the action is executed, the environment returns an immediate reward r. t With the next state s t+1 Forming state transition samples (s) t ,a t ,r t ,s t+1 The data is stored in the experience storage pool R. Once the accumulated experience samples reach a set threshold, a small batch of data is randomly sampled from this pool for training. In the DDQN structure, action a... t+1 The target Q-value is calculated by selecting the target Q-value from the primary network, but its corresponding Q-value is estimated by the target network. Then, the primary network parameters θ are updated using gradient descent by minimizing the mean squared error L(θ) between the target Q-value and the Q-value output by the primary network. To maintain training stability, the primary network parameters are copied to the target network at regular intervals to update θ. - During training, the exploration rate will be gradually reduced to encourage the strategy to converge. Training will terminate after reaching the maximum number of rounds, ultimately yielding the path planning result with the lowest overall vehicle operating cost.

[0271] This embodiment considers multiple drones deployed by an inspection vehicle to different observation points within a target area. When the drones and the local MEC server jointly process the data collected on-site, the drone computation offloading problem needs to be addressed. Specifically, with the optimization objective of minimizing the time and energy consumption cost of all drones, decisions are made on how each drone should offload its computational tasks, how the local base station should allocate communication bandwidth, and how the MEC server should allocate computing resources to process the offloaded computational tasks. Based on this, after determining the parking time of the inspection vehicle at different target area parking points, the next step is to solve the path planning problem of how the inspection vehicle, starting from a designated starting point, should sequentially travel to different target area parking points. Considering that the inspection vehicle is a new energy vehicle, the overall travel cost, including travel time, energy consumption, and maintenance costs, is the objective of vehicle path planning. Through the above design and optimization, this invention proposes a method for large-scale, low-cost inspection in smart cities. The inspection range is expanded by combining inspection vehicles and drones. With the assistance of MEC servers, drones improve their inspection service response capabilities and reduce their energy consumption. Inspection vehicles reduce operating costs by optimizing their travel routes. Ultimately, the efficient scheduling of vehicles and drones in collaboration enables timely completion of large-scale inspection tasks and achieves cost savings.

Claims

1. A method for joint scheduling of vehicles and drones for large-scale, low-cost inspections, characterized in that: include: S1: Obtain prior information; S2: Based on prior information, the UAV inspection problem of each target area will be modeled to obtain a mixed integer non-convex optimization problem with the objective of minimizing the weighted sum of the total execution time and energy consumption of all inspection UAVs; S3: Linearize the non-convex bilinear terms in the mixed integer non-convex optimization problem, and discretize the nonlinear function using a piecewise linear approximation; An approximate mixed-integer linear programming problem is obtained and solved to derive a drone scheduling strategy; S4: Based on the UAV scheduling strategy and prior information, a model is built to obtain the inspection vehicle path planning problem with the overall travel cost as the objective. S5: Transform the inspection vehicle path planning problem into a Markov decision process model. Using the inspection vehicle as an agent, define its state, action, and reward function to obtain the inspection vehicle scheduling strategy.

2. The vehicle and UAV joint scheduling method for large-scale, low-cost inspection according to claim 1, characterized in that, In step S1, the prior information includes the number of target areas, the observation points that the UAV needs to reach in each target area and the type of inspection task to be performed, and all state parameters of the UAV and inspection vehicle during the inspection process. All state parameters of the UAV and inspection vehicle during the inspection process include: the transmit / receive power of the UAV and base station data transmission, channel state parameters, available computing, storage and communication resources of the UAV / MEC server, the unit travel distance of the inspection vehicle, travel speed, power consumption and maintenance cost.

3. The vehicle and UAV joint scheduling method for large-scale, low-cost inspection according to claim 1, characterized in that, Step S2 specifically involves modeling based on the communication model, latency model, and energy consumption model, and setting optimization objectives and resource capacity constraints to obtain a mixed integer nonconvex optimization problem with the objective of minimizing the weighted sum of the total execution time and energy consumption of all inspection drones.

4. The vehicle and UAV joint scheduling method for large-scale, low-cost inspection according to claim 3, characterized in that, The communication model employs a quasi-static decision-making model. After the inspection vehicle reaches the parking point in a target area, multiple drones are dispatched to different observation points to perform computationally intensive inspection tasks. Once each drone reaches its designated observation point, it remains hovering to ensure its aerial position remains fixed, and it communicates wirelessly with the local MEC server via a line-of-sight channel. The wireless channel is a quasi-static channel, meaning that the channel state remains unchanged during data transmission. The target area is j, and one base station j is deployed within the area, using I... j This represents the set of all drones performing tasks within this area, and a given drone is labeled as i∈I. j The wireless communication between the UAV and the base stations in the area adopts orthogonal frequency division multiple access technology, and the channel power gain adopts the free space path loss model, defining the path loss from base station j to the UAV. i The data downlink rate is Among them, b i,j For the communication bandwidth allocated to UAV i, l i,j Let h be the communication distance between UAV i at the observation point and base station j. i,j Here, Q represents the channel gain at a reference distance of 1 meter, α is the channel fading exponent, and Q... i,j Let Ni be the received power of UAVi, and N0 be the noise power; the logarithmic function in equation (1) contains constants, and constants are used. Replace the logarithmic part; define from UAV i The uplink data rate to base station j is Among them, P i,j The transmit power of UAV i is represented by a constant. Replace the logarithmic part; Delay model: Considering the binary computation offloading method, define a 0-1 variable a. i,j a i,j =0 indicates that UAV i performs the computation task locally, a i,j =1 indicates that UAV i will offload the computing task to MEC server j, and has If the UAV performs the computing task locally, then the UAV i The total time is Among them, the first term on the right side of the equation is 2τ i,j The first term is the time taken for UAV i to travel to and return to the associated observation point; the second term is the amount of data D collected by UAV i. i,j The time taken, s i,j The first is the data acquisition rate; the third is the time taken for UAV i to download the specified application. i,j m is the amount of data requested by the application. i,j =1 indicates application A i,j Already cached on MEC servers j, m i,j =0 indicates that the current MEC server j does not cache A. i,j A needs to be obtained from the cloud. i,j , This is the data transfer rate from the cloud to the MEC server j; the last item is the CPU computation time of UAV i. For UAV i's local computing resources, W i,j To process data D i,j The required number of CPU cycles, after offloading the UAV i computation task to MEC server j, results in the total time spent on UAV i being... in, It is the MEC server j that is assigned to the UAV i The task's computational resources, UAV i The total time spent is expressed as Energy consumption model: Considering the communication, computing, and hovering energy consumption of the UAV in the system, UAV i When performing computing tasks locally, the relevant energy consumption is... The first term on the right side of the equals sign represents the power consumption of UAV i when downloading applications, and the second term represents the CPU power consumption of UAV i. For UAV i The energy efficiency factor represents the effective switching capacitor of the CPU, which controls the UAV. i If the computational task is offloaded to MEC server j, the relevant energy consumption is... The first term on the right side of the equals sign is UAV. i The energy consumption for uploading the collected data, the parameter in the second item. UAV is the CPU power consumption coefficient of MEC server j. i Flight power is The first term on the right side of the equals sign is UAV. i Blade profile power during flight Let ξ be the blade profile power of UAV i in hovering state. i η is the tip velocity of the UAV i rotor blade. i For UAV i Flight speed, the second item is UAV i Induced power during flight, For UAV i The induced power in hovering state, η0 is the average rotor induced velocity of UAV i during hovering, and the third term is the induced power of UAV i during hovering. i Air resistance power during flight For UAV i The characteristics are related to fixed parameters; the hovering power of UAV i is UAV i The energy consumption for hovering is Finally, processing UAVs i The system energy consumption caused by the computational task is expressed as: Optimization Objective: In the MEC-assisted UAV inspection problem in target region j, the optimization objective is to minimize the weighted sum of the total energy consumption of all UAVs and MEC servers in this region, as well as the execution time of all UAVs. This minimization problem is denoted as P1, and takes the following form: in, To park at point j in the target area, UAV i The maximum time for performing inspection tasks, i.e. δ1 and δ2 are the weights for controlling energy consumption and time; Resource capacity constraint: Let the storage capacity of MEC server j be C. s,j There are MEC server data storage constraints. Let B be the total communication bandwidth, with the following bandwidth constraint: Let F be the total computing resources of the MEC server. For all unloaded tasks, there is a constraint on the total computing resources. set up To determine the maximum time for UAV i to perform its task at parking point j in the target area, the constraint is:

5. The vehicle and UAV joint scheduling method for large-scale, low-cost inspection according to claim 4, characterized in that, Step S3 is as follows: Based on the modeling transformation method of linearization and discretization, by introducing auxiliary variables and constraints, the non-convex bilinear terms of the mixed integer non-convex optimization problem are linearized. At the same time, the nonlinear function is discretized by piecewise linear approximation. Furthermore, a discrete point generation algorithm is designed to automatically control the discretization accuracy and linearization error, thereby transforming the mixed integer non-convex optimization problem into an approximate mixed integer linear programming problem for solution.

6. The vehicle and UAV joint scheduling method for large-scale, low-cost inspection according to claim 5, characterized in that, Step S3 is as follows: The convex envoy method is used to linearize the bilinear terms, where equations (6) and (12) contain non-convex bilinear terms. Both are products of a 0-1 variable and a continuous variable, i.e., linearization is performed using the convex hull method; with bilinear terms... For example, we can linearize it equivalently as follows: in, They are The upper and lower bounds; note that at this point, we are using... Treating it as an independent variable; linear inequalities (20) and (21) yield: when a i,j When = 0, there is when a i,j When = 1, we have Using variables that satisfy equations (20) and (21) To equivalently replace the original bilinear term Linearize the other bilinear terms using the same method; Next, we introduce new variables. and To replace the original ones respectively and Right now and Therefore, equations (4) and (5) can be rewritten as follows: Equations (22) and (23) are both linear equations. Equations (7) and (8) can be rewritten as follows: Among them, equations (24) and (25) are both linear equations, g i,j As an intermediate variable; further relax equation (26) to equivalently Equation (27) specifies g i,j The lower bound is Optimal g i,j The value will equal Equation (27) is equivalent to replacing equation (26); The total bandwidth constraint is rewritten as follows set up For a sufficiently large value, the total resource constraint is rewritten as follows: set up For a sufficiently large value; by introducing a new variable The transformed optimization problem still contains nonlinear functions: in equation (27) In equation (28) In the formula (30) By employing discretization techniques, a linear function is used to approximate a nonlinear function: For example, let The range of discretized values ​​is Define the set of discrete points Where K1 is the number of discrete points, and Choose any two adjacent discrete points A straight line is defined through these two points. Then, equations (33)-(34) are used to approximate the nonlinear constraint (30). Where, θ i,j It is an intermediate variable, and equation (33) contains a bilinear term. right Define another set of discrete points Definition process straight line Then, equation (35) is used to approximate the nonlinear constraint (27). against Define the set of discrete points Definition process straight line Then, equations (36)-(37) are used to approximate the nonlinear constraint (28); Where, ω i,j It is an intermediate variable; Discrete point generation algorithm: function Functions on the positive real number line are convex functions. Discrete points are generated using the properties of convex functions. Let f(r) denote a convex function, where r satisfies 0. <r min ≤r≤r max Define the set of discrete points. There is r min =r1<... <r K =r max After two points r k ,r k+1 Define a straight line In r k ≤r≤r k+1 Within the range, there is always l k,k+1 (r)≥f(r); therefore, using l k,k+1 The maximum perpendicular error of f(r) approximated by f(r) is By analyzing its KKT conditions, the optimal solution is obtained as follows: That is, at point r * At this point, the maximum value of the approximate error Δ can be obtained. * ; Generate a set of discrete points Specific steps: The algorithm's input parameter δ is the preset maximum permissible error. δ0 = 0.1δ is set to ensure that the actual maximum error Δ... * The point is near δ and less than or equal to δ; initially, there is only one discrete point r1 in the set of discrete points R. Starting from the first point r1, find a second point r2 such that the distance between the two points is Δ. * Near δ, find a third point r3 such that the Δ between r2 and r3 is... * This process is repeated near δ until the last point r is reached. max Each time a discrete point is generated, the maximum vertical error Δ can be guaranteed. * ≤δ, therefore, in the worst case, the approximation error of equations (33) and (36) is Iδ, and the approximation error of equation (35) is δ; based on simple algebraic calculations and using the bisection method to find new points, a suitable set of discrete points R can be quickly obtained; before solving P1, for Run the algorithm once each, and then substitute the generated discrete points into constraints (33), (35), and (36).

7. The vehicle and UAV joint scheduling method for large-scale, low-cost inspection according to claim 1, characterized in that, Step S4 is as follows: A vehicle travel time cost model, a vehicle electricity consumption cost model, and a vehicle maintenance cost model are established, and a path planning optimization objective and constraints are set to construct an inspection vehicle path planning problem with comprehensive travel cost as the objective.

8. The vehicle and UAV joint scheduling method for large-scale, low-cost inspection according to claim 7, characterized in that, Step S4 is as follows: Vehicle travel time cost model: The time cost incurred during the inspection vehicle's journey is determined by the distance and speed it travels. Where d nm The Euclidean distance between parking spots in the target area; Vehicle power consumption cost model: The power generated by the inspection vehicle during operation due to rolling resistance and air resistance is... Among them, f r It is the rolling resistance coefficient, C d Here, m is the air drag coefficient, g is the vehicle mass, ρ is the air density, and A is the air drag coefficient. f Where W is the frontal area of ​​the vehicle and W is the wind speed, the electricity cost of the inspection vehicle during operation is... Vehicle maintenance cost model: Inspection vehicles may experience wear and tear due to prolonged driving, requiring maintenance. The maintenance cost is... M nm =μ·d nm (45) Where μ is the maintenance factor; Path planning optimization objective: To minimize the overall travel cost of inspection vehicles, consisting of travel time, power consumption, and maintenance costs; where x nm =1 indicates that the inspection vehicle passes through the section of road from parking point n to parking point m in the target area, x nm =0 indicates that the inspection vehicle did not pass through the section from target area parking point n to target area parking point m; the set of target area parking points is S = {0, 1, 2, ..., N}, where N is the number of target area parking points, and 0 is the starting point of the inspection vehicle. The problem of minimizing this set is denoted as P2, and takes the following form. in, For the target region j, UAV i The power consumption incurred when returning to the inspection vehicle and charging it; λ1, λ2, and λ3 are the weights of control time, power consumption, and maintenance cost; Constraints: Each parking spot in the target area can only be visited once, as expressed by the constraint condition. Some road sections in the city are under traffic control and cannot be passed. If the traffic-controlled road section (n,m)∈σ is impassable, then the constraint is: To prevent inspection vehicles from forming sub-loops when visiting parking points in the target area, an auxiliary variable u is introduced. n This indicates the order of visits to parking point n in the target area, with the following constraints: Where u n It is a continuous variable, representing the order in which task point n is visited; Assuming the inspection vehicle must start from and return to starting point 0, the constraint is:

9. The vehicle and UAV joint scheduling method for large-scale, low-cost inspection according to claim 1, characterized in that, Step S5 is as follows: Based on the Double DQN algorithm of deep reinforcement learning, the path planning problem of inspection vehicles is transformed into a Markov decision process. The inspection vehicle is used as the agent, and the state, action and reward functions are defined. The state-action value function is approximated by a neural network. The stability and accuracy of policy learning are improved by combining experience storage and a dual network structure. Finally, the optimal path planning of the inspection vehicle among all parking points in the target area is achieved.

10. The vehicle and UAV joint scheduling method for large-scale, low-cost inspection according to claim 9, characterized in that, Step S5 is as follows: State: In the inspection vehicle path planning problem, the state is a binary tuple s. t =(s c ,s v ), where s c Indicates the current target area parking point, s v This indicates the target area parking points that have already been visited; the above state design reflects the relevance of the last inspection vehicle's travel path, helping to find the path with the lowest overall travel cost. Action: The action is "Select the next target area parking spot to visit"; Actions serve as input to the environment, which in turn simulates the behavior of the inspection vehicle. Reward: The reward function is used to guide the agent to learn the optimal policy; The goal is to minimize the total cost, with the immediate reward being the inverse of the cost of the current action. This design encourages the agent to avoid high-cost paths and gradually learn a feasible path with the lowest cost. DDQN algorithm: In each round of training, the agent interacts with the environment to generate training samples; first, the agent observes the current main network state s. t In state s t Below, the agent selects action a. t The ε-Greedy strategy is used for action selection, that is: Where Q(s) t ,a t ;θ), representing the current state-action value function of the main network, the parameter θ is continuously updated during training, while ε gradually decreases with each training round; Perform action a t Afterwards, the agent receives an immediate reward r from the environment. t and transition to the next state s t+1 This generates a complete state transition sample (s). t ,a t ,r t ,s t+1 ); Introducing an experience replay mechanism: state transition samples (s) generated during training. t ,a t ,r t ,s t+1 The data is stored in the experience storage pool R. When the number of samples accumulates to a certain extent, a small batch of data is randomly extracted from it for training, which effectively breaks the temporal correlation between samples and improves training efficiency. The training objective is to minimize the mean squared error between the current state-action value and the target value; To mitigate the overestimation problem in Q-value estimation, a Double DQN structure is adopted, and its target Q-value is expressed by the following formula: Where γ∈[0,1) is the discount factor, and θ is the parameter of the current main network. - For fixed target network parameters; the loss function is defined as: All parameters θ of the Q-network are updated through gradient backpropagation in the neural network; after every few training rounds, the parameters θ of the main network are synchronously copied to the target network to update θ. - .

Citation Information

Patent Citations

  • Computing task unloading method based on deep reinforcement learning in electric power internet of things

    CN114065963A

  • Joint cache decision and trajectory optimization method under unmanned aerial vehicle assisted Internet of Vehicles

    CN116847293A

  • Electric vehicle charging planning method and device, computer equipment and storage medium

    CN117610763A

  • Power distribution network inspection planning method based on double-layer optimization algorithm

    CN119354219A

  • Unmanned aerial vehicle mobile edge computing resource allocation method based on deep reinforcement learning

    CN119922617A

Cited By

  • Vehicle-aircraft cooperative unmanned aerial vehicle power inspection path and energy configuration optimization method

    CN121457749A

  • Vehicle-aircraft cooperative scheduling method for onshore wind power plant inspection

    CN121599433A

  • Autonomous inspection method and system for substation inspection robot

    CN121791440A

  • A substation inspection robot autonomous inspection method and system

    CN121791440B