Method and system for dynamically scheduling disinfection robot based on multi-objective optimization
Through the combination of multi-objective optimization and deep reinforcement learning, the optimal scheduling solution for disinfection robots is generated, which solves the resource conflicts and dynamic environmental adaptability problems of disinfection robots in complex scenarios, and achieves efficient and energy-saving disinfection effects.
Patent Information
- Application Number
- CN202510701198.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-08-29
AI Technical Summary
The existing disinfection robot scheduling scheme cannot simultaneously optimize disinfection coverage, energy consumption, time cost and disinfectant dosage, and it is difficult to cope with dynamic environmental changes.
A multi-objective optimization method is adopted, combined with the NSGA-III algorithm and deep reinforcement learning, an optimization model with maximum coverage, minimized energy consumption, minimized time and disinfectant consumption constraints is built, and the optimal scheduling scheme is generated through the non-dominant sorting genetic algorithm, and the deep reinforcement learning model is used to adjust the path and disinfection strategy in real time.
It has achieved the improvement of disinfection efficiency in complex scenarios, significant energy-saving effects, the amount of disinfectant is within the safety threshold, high resource utilization rate and strong adaptability.
Smart Images

Figure CN120560269A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of dynamic scheduling of pest control robots, and in particular to a method and system for dynamic scheduling of pest control robots based on multi-objective optimization. Background Art
[0002] With the continuous improvement of public health and safety requirements, disinfection robots are increasingly used in medical, logistics, public places and other scenarios. Traditional disinfection robots mostly adopt single-objective scheduling strategies, such as static planning methods based on the shortest path or maximum coverage. However, actual disinfection tasks often face multi-dimensional constraints and dynamic environmental challenges: on the one hand, there is a significant conflict between disinfection efficiency, energy consumption cost, time window and disinfectant dosage, and single-objective optimization can easily lead to waste of other resources or task failure; on the other hand, problems such as the dynamic appearance of obstacles in complex scenarios and changes in the level of contaminated areas require robots to have real-time response and adaptive adjustment capabilities.
[0003] The inventors discovered that the technical problems existing in the prior art are that it is impossible to generate a multi-objective scheduling plan for the disinfection robot, and it is impossible to realize the path planning of the disinfection robot. Summary of the Invention
[0004] In order to address the shortcomings of the existing technology, the present invention provides a dynamic scheduling method and system for disinfection robots based on multi-objective optimization; the present invention is based on a multi-objective dynamic scheduling framework based on Pareto optimization, and incorporates maximizing disinfection coverage, minimizing energy consumption and time costs, and constraining disinfectant consumption as hard constraints into the optimization model. Combined with the improved NSGA-III algorithm and reinforcement learning, it overcomes the difficulties of target conflict resolution under multi-resource constraints, dynamic environment adaptive decision-making, and other difficult problems, providing an efficient, safe, and energy-saving integrated solution for disinfection tasks in complex scenarios.
[0005] On the one hand, a dynamic scheduling method for disinfection robots based on multi-objective optimization is provided, including:
[0006] Acquire environmental perception information of the disinfection robot's operating environment, construct a two-dimensional grid map based on the environmental perception information, divide the two-dimensional grid map into a number of sub-areas, and mark the area of the current area and the pollution level of the current area for each sub-area;
[0007] Based on the two-dimensional grid map, a multi-objective optimization model for the disinfection robot is constructed; the multi-objective optimization model includes: a coverage maximization objective function, an energy consumption minimization objective function, an operation time minimization objective function, and a disinfectant consumption constraint condition;
[0008] A non-dominated sorting genetic algorithm is used to solve the multi-objective optimization model and obtain the optimal scheduling plan for the disinfection robot. Based on the optimal scheduling plan for the disinfection robot, the optimal planning path for the disinfection robot is generated.
[0009] While the disinfection robot is walking along the optimal planned path, the real-time position information, real-time environmental perception information, real-time battery status detection information and the remaining amount of disinfectant of the disinfection robot are input into the trained deep reinforcement learning model. The model outputs a new optimization target and modifies the pre-set optimization target based on the new optimization target; the optimization target includes: the real-time obstacle avoidance path of the disinfection robot, the disinfectant spraying density and the walking speed of the disinfection robot.
[0010] On the other hand, a dynamic dispatching system for disinfection robots based on multi-objective optimization is provided, including:
[0011] an acquisition module configured to: acquire environmental perception information of the disinfection robot's operating environment, construct a two-dimensional grid map based on the environmental perception information, divide the two-dimensional grid map into a plurality of sub-regions, and mark the area of the current region and the pollution level of the current region for each sub-region;
[0012] A model construction module is configured to: construct a multi-objective optimization model for the disinfection robot based on the two-dimensional grid map; the multi-objective optimization model includes: a coverage maximization objective function, an energy consumption minimization objective function, an operation time minimization objective function, and a disinfectant consumption constraint condition;
[0013] The solution module is configured to: solve the multi-objective optimization model using a non-dominated sorting genetic algorithm to obtain an optimal scheduling solution for the disinfection robot; and generate an optimal planning path for the disinfection robot based on the optimal scheduling solution for the disinfection robot;
[0014] The output module is configured to: input the real-time position information, real-time environmental perception information, real-time battery status detection information and the remaining amount of disinfectant of the disinfection robot into the trained deep reinforcement learning model during the process of the disinfection robot walking according to the optimal planned path. The model outputs a new optimization target and modifies the pre-set optimization target based on the new optimization target; the optimization target includes: the real-time obstacle avoidance path of the disinfection robot, the disinfectant spraying density and the walking speed of the disinfection robot.
[0015] In another aspect, an electronic device is provided, comprising:
[0016] a memory for non-transitory storage of computer-readable instructions; and
[0017] a processor for executing said computer-readable instructions,
[0018] When the computer-readable instructions are executed by the processor, the method described in the first aspect is executed.
[0019] On the other hand, a storage medium is provided, which non-temporarily stores computer-readable instructions, wherein when the non-temporary computer-readable instructions are executed by a computer, the method described in the first aspect is executed.
[0020] On the other hand, a computer program product is provided, comprising a computer program, wherein the computer program is configured to implement the method described in the first aspect when running on one or more processors.
[0021] The above technical solution has the following advantages or beneficial effects:
[0022] Through a multi-objective optimization model, the pest control coverage rate is maximized, energy consumption and time costs are minimized, and the consumption of pesticides is constrained. The method includes: constructing an environmental map and path network for the pest control task, designing a multi-objective optimization function that integrates coverage, energy consumption, time, and pesticide consumption, generating a non-dominated solution set based on the Pareto front, selecting the optimal scheduling scheme in combination with a dynamic weight adjustment strategy, and adapting to environmental changes in real time through reinforcement learning. The present invention solves the problems of single target, low resource utilization, and poor dynamic adaptability in the scheduling of traditional pest control robots, significantly improving the pest control efficiency and energy-saving effect in complex scenarios, while ensuring that the amount of pesticide used is within the safety threshold. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0024] Figure 1 This is a flow chart of the method of embodiment 1. DETAILED DESCRIPTION
[0025] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.
[0026] Example 1
[0027] This embodiment provides a dynamic scheduling method for disinfection robots based on multi-objective optimization;
[0028] like Figure 1 As shown in FIG, a dynamic scheduling method for disinfection robots based on multi-objective optimization includes:
[0029] S101: Acquire environmental perception information of the disinfection robot's operating environment, construct a two-dimensional grid map based on the environmental perception information, divide the two-dimensional grid map into a plurality of sub-regions, and mark each sub-region with the area of the current region and the pollution level of the current region;
[0030] S102: Constructing a multi-objective optimization model for the disinfection robot based on the two-dimensional grid map; the multi-objective optimization model includes: a coverage maximization objective function, an energy consumption minimization objective function, an operation time minimization objective function, and a disinfectant consumption constraint condition;
[0031] S103: Using a non-dominated sorting genetic algorithm, solve the multi-objective optimization model to obtain an optimal scheduling plan for the disinfection robot; based on the optimal scheduling plan for the disinfection robot, generate an optimal planning path for the disinfection robot;
[0032] S104: While the disinfection robot is walking along the optimal planned path, the real-time position information, real-time environmental perception information, real-time battery status detection information and the remaining amount of disinfectant of the disinfection robot are input into the trained deep reinforcement learning model. The model outputs a new optimization target and modifies the pre-set optimization target based on the new optimization target; the optimization target includes: the real-time obstacle avoidance path of the disinfection robot, the disinfectant spraying density and the walking speed of the disinfection robot.
[0033] Furthermore, the method further includes:
[0034] S105: Control the disinfection robot to execute a real-time obstacle avoidance path, control the robot to execute the disinfectant spraying density, and control the robot to walk according to the output walking speed;
[0035] S106: Update obstacle information, update environmental pollution level, update battery status and update disinfectant remaining amount, and return to S102.
[0036] Furthermore, the S101: obtains environmental perception information of the disinfection robot's operating environment, and the environmental perception information includes: obstacle locations, boundaries of the area to be disinfected, preset key point locations, and sub-area pollution levels.
[0037] It should be understood that the environmental perception information is collected through lidar, infrared sensors and visual sensors.
[0038] The two-dimensional grid map includes the location of obstacles, the boundaries of the area to be eliminated, and the locations of preset key points. The two-dimensional grid map is constructed using the SLAM method.
[0039] The pollution level is divided into level 1, level 2, level 3 and level 4. The pollution level is determined by collecting gas from the polluted area through a gas sensor and then classifying the degree of gas pollution.
[0040] Dividing the two-dimensional grid map into several sub-areas is to rasterize the map into N sub-areas, each of which has an associated area A. i and PollutionLevel i , obtain structured environmental data; the structured environmental data includes: grid map matrix δ i ∈{0,1} (1 means to be disinfected) and pollution level vector: PollutionLevel∈R N .
[0041] The beneficial effect of this technical solution is that through the multi-source fusion of lidar, infrared sensors, and visual sensors, a high-precision two-dimensional map is constructed in real time to identify obstacle locations and detect regional pollution levels (such as virus concentration distribution). The sensor data is processed by the SLAM algorithm, and the output is a dynamic environmental model that includes topology, pollution hotspots, and safety boundaries, providing input for subsequent optimization.
[0042] Furthermore, the coverage maximization objective function is specifically expressed as follows:
[0043] Cover at least 90% of the area to be disinfected through path planning. Define the coverage rate f1 as the proportion of the target area covered by the disinfection path:
[0044]
[0045] Where N represents the total number of areas to be disinfected; A i represents the area of the i-th sub-region; δ i represents a binary indicator function, which is 1 if the grid is covered and 0 otherwise; Indicates the total area to be disinfected.
[0046] Furthermore, the energy consumption minimization objective function is specifically expressed as follows:
[0047] The total energy consumption consisting of movement energy consumption and disinfection energy consumption is:
[0048]
[0049] Among them, P disinfect Indicates the power of the disinfection equipment; Δt t represents the duration of time period t.
[0050] Robot moving power P move The quadratic function relationship with velocity v:
[0051] P move =k1v 2 +k2;
[0052] Where k1 represents the air resistance coefficient (typical value: 0.05kW·s 2 / m 2 ); k2 represents the basic power consumption (typical value: 0.1kW); v represents the moving speed (unit: m / s).
[0053] Furthermore, the operation time minimization objective function includes:
[0054]
[0055] Where M represents the number of path segments; L j represents the length of the jth path; t stop,k Indicates the time of the kth pause (such as turning, obstacle avoidance, etc., unit: second); v avg Indicates the average moving speed.
[0056] Furthermore, the pesticide consumption constraints include:
[0057] Total amount limit of disinfectant:
[0058] DU≤D max ;
[0059] Among them, D max Indicates the maximum capacity of the device (e.g. 500mL).
[0060] Safety threshold of usage per unit area:
[0061]
[0062] Among them, d min and d max Respectively represent the minimum and maximum usage per unit area.
[0063] Movement speed:
[0064] v min ≤v≤v max .
[0065] Among them, the amount of disinfectant per unit area is expressed as:
[0066] g=d i =d base +α·PollutionLevel i
[0067] Among them, d irepresents the amount of disinfectant used per unit area of the ith sub-region (dynamically adjusted according to the pollution level); d base Indicates the basic dosage density (such as 0.1mL / m2); α indicates the pollution level coefficient (such as 0.05mL / m 2 ·level); PollutionLevel i Indicates the pollution level of the i-th grid (level 1-4).
[0068] Total pesticide consumption, expressed as:
[0069]
[0070] The beneficial effects of the above technical solution are: based on the environmental model, a four-dimensional objective function is established to maximize coverage, minimize energy consumption, optimize time and constrain pesticides.
[0071] Furthermore, S103: a non-dominated sorting genetic algorithm is used to solve the multi-objective optimization model to obtain the optimal scheduling plan for the disinfection robot, which specifically includes:
[0072] Generate an initial solution that satisfies the constraints and generate an initialized population;
[0073] Use the parent population to perform crossover operations;
[0074] Use the parent population to perform mutation operations;
[0075] Merge the parent generation, the offspring generated by crossover, and the offspring generated by mutation to obtain a new population. The new population is a set of corresponding solutions.
[0076] Sort and select the new population to get the optimal solution.
[0077] A non-dominated sorting genetic algorithm (NSGA-III) was used to generate a Pareto-optimal solution set. The algorithm eliminated excessive solutions for the pesticide using an adaptive penalty function and employed a reference point method to maintain solution diversity, ensuring a balance between global optimization and local refinement.
[0078] Furthermore, S103: a non-dominated sorting genetic algorithm is used to solve the multi-objective optimization model to obtain the optimal scheduling plan for the disinfection robot, which specifically includes:
[0079] S103-1: Initialize the population and calculate the target value:
[0080] Generate an initial population P0 with a size of N. Chromosome encoding: Real number encoding is used. Each individual represents a scheduling scheme. The gene includes: path node sequence (grid coordinates), moving speed v (value range [v min ,v max]), disinfectant distribution density d i (Constraint d min ≤d i ≤d max );
[0081] Population generation: randomly generate N individuals to ensure that the initial solution satisfies the speed constraint and d i range, but allows DU to temporarily exceed D max (Subsequently processed by penalty function.) Calculate f1, f2, f3 of the initial population and check constraint g;
[0082] S103-2: Non-dominated Sorting:
[0083] Domination relationship: Individual x dominates y if and only if: f i (x)≤f i (y) and
[0084] Frontier partitioning: Divide the population into multiple non-dominated levels F1, F2, ... F N , where F1 is the optimal frontier.
[0085] S103-3: Generate reference point Z r :In the 4-dimensional target space (CR, EC, TT, DU), the boundary intersection construction method is used to generate uniformly distributed reference points, and each target dimension is divided into p equal points (p = 4), generating reference points. And after normalizing the reference point coordinates, the following relationship is satisfied:
[0086]
[0087] Among them, CR represents coverage, EC represents total energy consumption, TT represents the total execution time of the task, and DU represents the total dose of the disinfectant.
[0088] S103-4: Normalized objective function
[0089] Ideal point update: find the ideal point z of the current population min Population:
[0090] Extreme point calculation: find the extreme point of each target direction through the scalarization function
[0091] Normalization formula:
[0092] S103-5: Associated reference points: Vertical distance calculation: For each individual x, calculate its vertical distance to all reference points ωk ∈Z r Vertical distance:
[0093] Association operation: associate the individual to the nearest reference point and record the association relationship ρ k (The number of associated individuals of reference point k).
[0094] S103-6: Niche Preservation: Fill the new population according to the frontier level until it approaches the capacity N; for the unfilled new population, give priority to the associated reference point ρ k The smallest individual; if multiple individuals are associated with the same reference point, the closest individual is selected;
[0095] S103-7: Crossover / Mutation Operation:
[0096] Crossover: Sequential crossover (OX) is used for path sequence, and arithmetic crossover is used for speed and disinfectant density:
[0097] v child =αv parent1 +(1-α)v parent2 ,α~U[0,1].
[0098] Mutation: Randomly swap positions in the path sequence, add Gaussian perturbations to the speed and disinfectant density:
[0099]
[0100] σ=0.1×(v max -v min );
[0101] in, Represents a random number.
[0102] S103-8: Merge populations, merge the parent generation, offspring generated by crossover, and offspring generated by mutation into a new population. Merge parent generation and offspring: R t =P t ∪Q t , of size 2N.
[0103] S103-9: Constraint processing: Dynamically adjust the penalty function, increasing linearly with the number of iterations t:
[0104]
[0105] Where, λ0=0.1,λ=1.0,T max Indicates the maximum number of iterations.
[0106] Penalty design: impose penalties on solutions that violate constraints:
[0107] f penalty (x)=f(x)-λ(t)·max(0,DU(x)-D max ).
[0108] Furthermore, the optimal planning path of the disinfection robot is generated based on the optimal dispatching plan of the disinfection robot, specifically including:
[0109] The optimal scheduling plan for the disinfection robot includes: the location of the target disinfection area corresponding to each time node;
[0110] The initial position of the disinfection robot and the target disinfection area position of the disinfection robot at each time node are input into the A* algorithm to obtain the optimal planning path of the disinfection robot.
[0111] Furthermore, the cost function of the A* algorithm is:
[0112] Cost(p)=α·f1+β·f2+γ·f3
[0113] Among them, α, β, and γ represent weight coefficients, which are set according to user needs.
[0114] Furthermore, the heuristic function of the A* algorithm is:
[0115] f(n)=g(n)+h(n)+λ·Cost(p)
[0116] Among them, g(n) is the actual cost from the starting point to node n, h(n) is the traditional straight-line distance heuristic value, and λ is the weight coefficient.
[0117] Path planning engine: combines the Pareto optimal solution with the A algorithm to generate a feasible path that takes into account multiple objective constraints; motion control: adjusts the movement speed ([0.2, 1.5] m / s}) and obstacle avoidance strategy; disinfection equipment control: dynamically allocates the disinfectant density according to the pollution level to achieve precise spraying.
[0118] Furthermore, in step S104, while the disinfection robot is walking along the optimally planned path, the real-time position information, real-time environmental perception information, real-time battery status detection information, and the remaining amount of disinfectant of the disinfection robot are input into the trained deep reinforcement learning model, and the model outputs a new optimization target. The deep reinforcement learning model training process includes:
[0119] Construct a training set, wherein the training set is the real-time position of the robot, the real-time power level, the remaining amount of disinfectant, and the obstacle position, given the real-time obstacle avoidance path of the disinfection robot, the disinfectant spraying density, and the walking speed of the disinfection robot;
[0120] Constructing a deep reinforcement learning model, the deep reinforcement learning model comprising: an input layer, a first hidden layer, a second hidden layer, a third hidden layer, and an output layer connected in sequence; wherein the first, second, and third hidden layers are all implemented by a fully connected network;
[0121] Input the training set into the deep reinforcement learning model and set the reward function to calculate the reward value after each action performed by the disinfection robot;
[0122] The robot's real-time position, real-time power, disinfectant remaining and obstacle position are used as the input values of the model, and the disinfection robot's real-time obstacle avoidance path, disinfectant spraying density and disinfection robot's walking speed are used as the output values of the model. When the loss function value of the model no longer decreases, the training is stopped to obtain the trained deep reinforcement learning model.
[0123] It should be understood that through the real-time decision-making capabilities of deep reinforcement learning models, the long-term reward optimization of the Bellman equation, and the dynamic data-driven experience replay, the disinfection robot can effectively cope with complex scenarios such as sudden obstacles and resource shortages. The core advantages of this solution are: Immediate response: millisecond-level action updates ensure mission continuity. Long-term optimization: Discount factors and target network design avoid short-sighted strategies. Collaborative learning: Combined with the multi-objective planning layer of NSGA-III, it achieves the integration of global resource allocation and local dynamic adjustment.
[0124] Division of labor and collaboration: Deep reinforcement learning model (real-time layer): handles dynamic decision-making (obstacle avoidance, speed adjustment) within seconds. NSGA-III (planning layer): optimizes global path and resource allocation within minutes (based on energy consumption and time data fed back by DQN).
[0125] Data Interoperability: The deep reinforcement learning model transmits real-time status (such as remaining disinfection dose and average speed) to NSGA-III for updating multi-objective weights. NSGA-III outputs a global Pareto solution set, which serves as the long-term reference goal of the deep reinforcement learning model (such as maximizing coverage and minimizing total time).
[0126] It should be understood that the input layer of the deep reinforcement learning model is used to input dynamic environment features: robot coordinates (2 dimensions), real-time power (1 dimension), disinfectant remaining (1 dimension), obstacle position (n dimensions, dynamic changes)
[0127] Real-time guarantee: The status update frequency is synchronized with the sensor (for example, 10 samples per second) to ensure immediate response to environmental changes.
[0128] It should be understood that the hidden layers of deep reinforcement learning models have nonlinear modeling capabilities: they capture the complex relationship between state and action (such as reducing speed to extend battery life when the battery is low) through a multi-layer neural network. The activation function uses ReLU, balancing computational efficiency and nonlinear expression capabilities.
[0129] It should be understood that the output layer of the deep reinforcement learning model outputs the action space Q(s,a): Multi-dimensional action decision-making: Path adjustment: Avoiding new obstacles or congested areas, generating a real-time obstacle avoidance path. Pesticide density: Dynamically adjusting the spray density based on the remaining pesticide dosage (for example, reducing the amount per unit area when the remaining amount is insufficient). Speed adjustment: Adjusting the movement speed based on the battery level and path complexity (for example, reducing the speed in highly complex environments to improve safety).
[0130] Bellman equation and loss function: Target Q value calculation:
[0131] Dynamic reward design: reward function R(s t ,a t ) contains real-time objectives (such as coverage area gain, energy consumption penalty, and mission time penalty).
[0132] Loss function:
[0133] L(θ)=E (s,a,r,s′) ~ReplayBuffer[Q target -Q(s, a; θ) 2 ].
[0134] Where γ = 0.99 is a discount factor that emphasizes long-term returns and avoids short-sighted decisions (such as excessive power consumption to quickly cover a certain area); θ - Represents the target network parameters, which are updated synchronously every 1000 steps.
[0135] Experience replay mechanism: storage transfer samples (s t ,a t ,r t ,s t+1 ) to the buffer, randomly sampling small batches (batch size = 64) to train the network. Sampling weights are increased for high-importance samples (such as battery drops and sudden failures) to accelerate learning efficiency in key scenarios.
[0136] During the execution process, the battery status, disinfectant remaining amount, new obstacles and pollution level changes are monitored in real time. The data is fed back to the perception module, triggering the update of the environmental model and strategy re-optimization, forming a closed-loop adaptive mechanism.
[0137] This architecture significantly improves resource utilization and scenario adaptability of disinfection operations through multi-objective collaborative optimization and dynamic closed-loop control, providing a systematic solution for autonomous robot operations in complex environments.
[0138] Closed-loop feedback: Real-time data drives continuous system iteration to cope with dynamic environmental challenges; multi-objective balance: Pareto frontier analyzes goal conflicts, and reinforcement learning enhances real-time decision-making; precise control: Dynamic distribution of disinfectant density and path planning are coordinated to ensure both safety and efficiency.
[0139] Furthermore, after the step of constructing a multi-objective optimization model of the disinfection robot according to the two-dimensional grid map, before the step of solving the multi-objective optimization model using a non-dominated sorting genetic algorithm, the method further includes:
[0140] S102-31: Determine whether an emergency has occurred, where the emergency includes: the remaining battery power is less than a first set threshold, the remaining disinfectant amount is less than a second set threshold, or a new obstacle is discovered;
[0141] S102-32: If an emergency occurs, the dynamic weight is calculated; the dynamic weight is used to update the multi-objective optimization model to obtain an updated multi-objective optimization model; the updated multi-objective optimization model is solved using a non-dominated sorting genetic algorithm to obtain a non-dominated solution set; all solutions in the non-dominated solution set are sorted to obtain the optimal solution.
[0142] Furthermore, the calculation of dynamic weights; using the dynamic weights to update the multi-objective optimization model to obtain an updated multi-objective optimization model includes:
[0143] (1) Calculate the state factor:
[0144] Remaining coverage requirements:
[0145]
[0146] State of Charge Factor:
[0147]
[0148] Among them, Battery remaining (t) indicates the current remaining battery capacity. total Indicates the battery capacity.
[0149] Environmental complexities:
[0150]
[0151] Wherein, FreePathLength(t) represents the total length of the remaining paths, and TotalPathLength represents the total length of all paths.
[0152] Pesticide residual factor:
[0153]
[0154] Among them, Disinfectant remaining (t) represents the remaining capacity of the disinfectant, D max Indicates the maximum capacity of the device to add disinfectant. The choice of parameters is related to the degree of obstruction of the path. The higher the degree of obstruction, the closer S3(t) is to 1.
[0155] (2) Dynamic weight update formula:
[0156]
[0157] Among them, the initial weight ω 1,base =0.4,ω 2,base =0.3,ω 3,base =0.2,ω 4,base =0.1;
[0158] Each target is independently multiplied by the weight to modify the objective function:
[0159] f′ j =ω j (t)·f j
[0160] f1′=ω1(t)·f1, f2′=ω2(t)·f2, f3′=ω3(t)·f3, g1′=ω3(t)·g,
[0161] The revised targets {f1′, f2′, f3′, g1′} are non-dominated sorted and associated with reference points.
[0162] Dynamic weight adjustment process description:
[0163] Dynamic weight adjustment is the core module of the disinfection robot's multi-target scheduling system. By monitoring key status parameters in real time and dynamically adjusting and optimizing target priorities, it ensures the robot's efficiency and safety in complex scenarios. The process is mainly divided into the following three stages:
[0164] 1. Real-time monitoring and conditional triggering
[0165] The system continuously collects information on battery power, disinfectant levels, and environmental obstacles. When the remaining battery power drops below 20%, the disinfectant level drops below 30%, or an unexpected obstacle is detected, a weight adjustment mechanism is triggered. For example, battery threshold warnings prioritize battery life, while low disinfectant levels prioritize resource conservation. And sudden obstacles prioritize route timeliness.
[0166] 2. Dynamic update of weights
[0167] Based on the triggering conditions, the normalized formula is used to adjust the objective function weight:
[0168]
[0169] Among them S i (t) is the state factor (e.g., the remaining battery percentage). In emergency scenarios, weighted mandatory coverage rules take effect: for example, when the battery is low, the energy consumption weight increases to 0.5; when the disinfectant is insufficient, the coverage weight decreases to 0.3; and in the event of an unexpected obstacle, the time weight is set to 0.5, ensuring that key objectives are prioritized.
[0170] 3. Plan reselection and implementation
[0171] After updating the weights, the system reorders the Pareto-optimal solution set and, using the A* algorithm, generates a new path that balances energy consumption, time, and pesticide constraints. For example, it selects a short, low-power path in low-battery scenarios, or employs a high-utilization spiral coverage strategy when pesticides are scarce. After the new solution is implemented, real-time data is fed back to the monitoring module, forming a closed loop of "perception → decision → execution → feedback" to continuously adapt to dynamic environmental changes.
[0172] This process effectively solves the problem of multi-objective conflicts through dynamic priority switching and closed-loop optimization. While ensuring disinfection coverage, it significantly improves resource utilization efficiency and system robustness.
[0173] Innovations: Multi-objective collaborative optimization: For the first time, coverage, energy consumption, time and pesticide consumption are incorporated into a unified Pareto framework to solve the problem of target conflicts under multiple resource constraints; dynamic pesticide management: The pesticide allocation strategy is adjusted in real time according to the pollution level to balance the disinfection effect and dosage restrictions; hybrid solution strategy: Combining the global search of NSGA-III, the real-time decision-making of reinforcement learning and the adaptive penalty function to ensure that the solution set meets complex constraints.
[0174] Example 2
[0175] This embodiment provides a dynamic dispatching system for disinfection robots based on multi-objective optimization, including:
[0176] an acquisition module configured to: acquire environmental perception information of the disinfection robot's operating environment, construct a two-dimensional grid map based on the environmental perception information, divide the two-dimensional grid map into a plurality of sub-regions, and mark the area of the current region and the pollution level of the current region for each sub-region;
[0177] A model construction module is configured to: construct a multi-objective optimization model for the disinfection robot based on the two-dimensional grid map; the multi-objective optimization model includes: a coverage maximization objective function, an energy consumption minimization objective function, an operation time minimization objective function, and a disinfectant consumption constraint condition;
[0178] The solution module is configured to: solve the multi-objective optimization model using a non-dominated sorting genetic algorithm to obtain an optimal scheduling solution for the disinfection robot; and generate an optimal planning path for the disinfection robot based on the optimal scheduling solution for the disinfection robot;
[0179] The output module is configured to: input the real-time position information, real-time environmental perception information, real-time battery status detection information and the remaining amount of disinfectant of the disinfection robot into the trained deep reinforcement learning model during the process of the disinfection robot walking according to the optimal planned path. The model outputs a new optimization target and modifies the pre-set optimization target based on the new optimization target; the optimization target includes: the real-time obstacle avoidance path of the disinfection robot, the disinfectant spraying density and the walking speed of the disinfection robot.
[0180] It should be noted that the acquisition module, model building module, solution module, and output module described above correspond to steps S101 to S104 in Example 1. The examples and application scenarios implemented by these modules and the corresponding steps are the same, but are not limited to the contents disclosed in Example 1. It should be noted that the above modules, as part of the system, can be executed in a computer system, such as a set of computer-executable instructions.
[0181] The descriptions of the various embodiments in the above embodiments have different focuses. For parts not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0182] The proposed system can be implemented in other ways. For example, the system embodiment described above is merely illustrative. For example, the above module division is only a logical function division. In actual implementation, other division methods may be used. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not implemented.
[0183] Example 3
[0184] This embodiment also provides an electronic device, comprising: one or more processors, one or more memories, and one or more computer programs; wherein the processor is connected to the memory, and the one or more computer programs are stored in the memory. When the electronic device is running, the processor executes the one or more computer programs stored in the memory, so that the electronic device executes the method described in the above embodiment one.
[0185] It should be understood that in this embodiment, the processor may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), off-the-shelf field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0186] The memory may include a read-only memory and a random access memory, and provides instructions and data to the processor. A portion of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.
[0187] During implementation, each step of the above method may be completed by an integrated logic circuit of hardware in a processor or by instructions in the form of software.
[0188] The method in Example 1 can be directly implemented as being executed by a hardware processor, or by a combination of hardware and software modules within the processor. The software module can be located in a storage medium well-established in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. The storage medium is located in the memory, and the processor reads the information in the memory and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, a detailed description is not given here.
[0189] Those skilled in the art will appreciate that the units and algorithm steps of the various examples described in conjunction with this embodiment can be implemented using electronic hardware or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.
[0190] Example 4
[0191] This embodiment further provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the method described in the first embodiment is performed.
[0192] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A dynamic scheduling method for disinfection robots based on multi-objective optimization is characterized by: include: Acquire environmental perception information of the disinfection robot's operating environment, construct a two-dimensional grid map based on the environmental perception information, divide the two-dimensional grid map into a number of sub-areas, and mark the area of the current area and the pollution level of the current area for each sub-area; Based on the two-dimensional grid map, a multi-objective optimization model for the disinfection robot is constructed; The multi-objective optimization model includes: a coverage maximization objective function, an energy consumption minimization objective function, an operation time minimization objective function, and a pesticide consumption constraint condition; A non-dominated sorting genetic algorithm is used to solve the multi-objective optimization model and obtain the optimal scheduling plan for the disinfection robot. Based on the optimal scheduling plan for the disinfection robot, the optimal planning path for the disinfection robot is generated. While the disinfection robot is walking along the optimal planned path, the real-time position information, real-time environmental perception information, real-time battery status detection information and the remaining amount of disinfectant of the disinfection robot are input into the trained deep reinforcement learning model. The model outputs a new optimization target and modifies the pre-set optimization target based on the new optimization target; the optimization target includes: the real-time obstacle avoidance path of the disinfection robot, the disinfectant spraying density and the walking speed of the disinfection robot.
2. The dynamic scheduling method for disinfection robots based on multi-objective optimization according to claim 1 is characterized in that: The coverage maximization objective function is specifically expressed as follows: Cover at least 90% of the area to be disinfected through path planning. Define the coverage rate f1 as the proportion of the target area covered by the disinfection path: Where N represents the total number of areas to be disinfected; A i represents the area of the i-th sub-region; δ i represents a binary indicator function, which is 1 if the grid is covered and 0 otherwise; Indicates the total area to be disinfected; The energy consumption minimization objective function is specifically expressed as follows: The total energy consumption consisting of movement energy consumption and disinfection energy consumption is: Among them, P disinfect Indicates the power of the disinfection equipment; Δt t represents the duration of time period t; Robot moving power P move The quadratic function relationship with velocity v: P move =k1v 2 +k2; Among them, k1 represents the air resistance coefficient; k2 represents the basic power consumption; and v represents the moving speed.
3. The dynamic scheduling method for disinfection robots based on multi-objective optimization according to claim 1 is characterized in that: The operation time minimization objective function includes: Where M represents the number of path segments; L j represents the length of the jth path; t stop,k represents the kth pause time; v avg represents the average moving speed; The pesticide consumption constraints include: Total amount limit of disinfectant: DU≤D max ; Among them, D max Indicates the maximum capacity of the device; Safety threshold of usage per unit area: d min ≤d i ≤d max ; Among them, d min and d max Respectively represent the minimum and maximum usage per unit area; Movement speed: in min ≤v≤v max ; Among them, the amount of disinfectant per unit area is expressed as: g=d i =d base +α·PollutionLevel i ; Among them, d i represents the amount of pesticide used per unit area of the i-th sub-region; d base Indicates basic usage density; α indicates pollution level coefficient; PollutionLevel i represents the pollution level of the i-th grid; Total pesticide consumption, expressed as:
4. The dynamic scheduling method for disinfection robots based on multi-objective optimization according to claim 1 is characterized in that: As the disinfection robot walks along the optimal planned path, its real-time location information, real-time environmental perception information, real-time battery status detection information, and the remaining amount of disinfectant are input into the trained deep reinforcement learning model. The model then outputs a new optimization target. The deep reinforcement learning model training process includes: Construct a training set, wherein the training set is the real-time position of the robot, the real-time power level, the remaining amount of disinfectant, and the obstacle position, given the real-time obstacle avoidance path of the disinfection robot, the disinfectant spraying density, and the walking speed of the disinfection robot; Constructing a deep reinforcement learning model, the deep reinforcement learning model comprising: an input layer, a first hidden layer, a second hidden layer, a third hidden layer, and an output layer connected in sequence; wherein the first, second, and third hidden layers are all implemented by a fully connected network; Input the training set into the deep reinforcement learning model and set the reward function to calculate the reward value after each action performed by the disinfection robot; The robot's real-time position, real-time power, disinfectant remaining and obstacle position are used as the input values of the model, and the disinfection robot's real-time obstacle avoidance path, disinfectant spraying density and disinfection robot's walking speed are used as the output values of the model. When the loss function value of the model no longer decreases, the training is stopped to obtain the trained deep reinforcement learning model.
5. The dynamic scheduling method for disinfection robots based on multi-objective optimization according to claim 1 is characterized in that: After the step of constructing a multi-objective optimization model for the disinfection robot according to the two-dimensional grid map, and before the step of solving the multi-objective optimization model using a non-dominated sorting genetic algorithm, the following steps are further included: Determining whether an emergency event has occurred, the emergency event including: the remaining power is less than a first set threshold, the remaining disinfectant is less than a second set threshold, or a new obstacle is discovered; If an emergency occurs, the dynamic weight is calculated; the dynamic weight is used to update the multi-objective optimization model to obtain an updated multi-objective optimization model; the non-dominated sorting genetic algorithm is used to solve the updated multi-objective optimization model to obtain a non-dominated solution set; all solutions in the non-dominated solution set are sorted to obtain the optimal solution.
6. The dynamic scheduling method for disinfection robots based on multi-objective optimization according to claim 5 is characterized in that: The calculation of dynamic weights; The multi-objective optimization model is updated using dynamic weights to obtain an updated multi-objective optimization model, including: (1) Calculate the state factor: Remaining coverage requirements: State of Charge Factor: Among them, Battery remaining (t) indicates the current remaining battery capacity; Battery total Indicates battery capacity; Environmental complexities: Where FreePathLength(t) represents the total length of the remaining paths; TotalPathLength represents the total length of all paths; Pesticide residual factor: Among them, Disinfectant remaining (t) represents the remaining capacity of the disinfectant, D max Indicates the maximum capacity of disinfectant that can be added to the equipment; (2) Dynamic weight update formula: Among them, the initial weight ω 1,base =0.4,ω 2,base =0.3,ω 3,base =0.2,ω 4,base =0.1; Each target is independently multiplied by the weight to modify the objective function: f′ j =ω j (t)·f j f1′=ω1(t)·f1, f2′=ω2(t)·f2, f3′=ω3(t)·f3, g1′=ω3(t)·g, The revised targets {f1′, f2′, f3′, g1′} are non-dominated sorted and associated with reference points.
7. The dynamic scheduling method for disinfection robots based on multi-objective optimization according to claim 1 is characterized in that: The non-dominated sorting genetic algorithm is used to solve the multi-objective optimization model and obtain the optimal scheduling plan for the disinfection robot, which includes: Generate an initial solution that satisfies the constraints and generate an initialized population; Use the parent population to perform crossover operations; Use the parent population to perform mutation operations; Merge the parent generation, the offspring generated by crossover, and the offspring generated by mutation to obtain a new population. The new population is a set of corresponding solutions. Sort and select the new population to get the optimal solution; The optimal planning path of the disinfection robot is generated based on the optimal dispatching plan of the disinfection robot, specifically including: The optimal scheduling plan for the disinfection robot includes: the location of the target disinfection area corresponding to each time node; The initial position of the disinfection robot and the target disinfection area position of the disinfection robot at each time node are input into the A* algorithm to obtain the optimal planning path of the disinfection robot.
8. The dynamic scheduling system of disinfection robots based on multi-objective optimization is characterized by: include: an acquisition module configured to: acquire environmental perception information of the disinfection robot's operating environment, construct a two-dimensional grid map based on the environmental perception information, divide the two-dimensional grid map into a plurality of sub-regions, and mark the area of the current region and the pollution level of the current region for each sub-region; A model building module is configured to: build a multi-objective optimization model of the disinfection robot based on the two-dimensional grid map; The multi-objective optimization model includes: a coverage maximization objective function, an energy consumption minimization objective function, an operation time minimization objective function, and a pesticide consumption constraint condition; The solution module is configured to: solve the multi-objective optimization model using a non-dominated sorting genetic algorithm to obtain an optimal scheduling solution for the disinfection robot; and generate an optimal planning path for the disinfection robot based on the optimal scheduling solution for the disinfection robot; The output module is configured to: input the real-time position information, real-time environmental perception information, real-time battery status detection information and the remaining amount of disinfectant of the disinfection robot into the trained deep reinforcement learning model during the process of the disinfection robot walking according to the optimal planned path. The model outputs a new optimization target and modifies the pre-set optimization target based on the new optimization target; the optimization target includes: the real-time obstacle avoidance path of the disinfection robot, the disinfectant spraying density and the walking speed of the disinfection robot.
9. An electronic device, comprising: a memory for non-transitory storage of computer-readable instructions; as well as a processor for executing said computer-readable instructions, When the computer-readable instructions are executed by the processor, the method according to any one of claims 1 to 7 is executed.
10. A storage medium, characterized in that: Non-transitory storage of computer-readable instructions, wherein when the non-transitory computer-readable instructions are executed by a computer, the method according to any one of claims 1 to 7 is performed.
Citation Information
Cited By
Dynamic adjustment method, system and equipment of disinfection path and storage medium
CN121983347A