Return charging scheduling method for mobile charging robot
By calculating the behavioral urgency index and energy competition index and combining them with the learning model to generate a scheduling parameter group, the problem that the scheduling method in the existing technology fails to consider task urgency and resource congestion is solved, and efficient charging path decision-making of mobile robots in complex environments is achieved, which improves the intelligence and adaptability of the system.
Patent Information
- Application Number
- CN202511158556.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-19
- Publication Date
- 2025-10-03
AI Technical Summary
Existing mobile robot scheduling methods fail to effectively consider task urgency and charging resource congestion, resulting in scheduling failure in high-load or queuing scenarios, affecting continuous and stable operation.
By collecting multi-dimensional state information to calculate the behavior urgency index and energy competition index, the learning model is used to generate a scheduling parameter group, and the charging path decision is dynamically adjusted, including the scheduling trigger mark, path sorting factor and path trade-off factor, to achieve reasonable charging path selection.
It improves the accuracy and efficiency of scheduling decisions, ensures that robots can be charged in a timely manner in complex environments, avoids mission interruptions, and enhances the intelligence level and adaptability of the system.
Smart Images

Figure CN120746199A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of mobile robot scheduling management, and more particularly, to a method for scheduling the return charging of a mobile charging robot. Background Art
[0002] With the widespread deployment of intelligent robots in diverse mission-based scenarios, such as factory logistics, indoor delivery, and warehouse inspections, achieving intelligent energy management decisions during operation has become a key issue limiting their sustained and stable operation. This is especially true in environments where mobile robots must perform long missions and share limited charging infrastructure. Improper scheduling strategies can easily lead to charging resource conflicts, mission interruptions, and even operational anomalies.
[0003] Existing mobile robot scheduling methods often make return charging decisions based on task priority, remaining battery power, or fixed-path strategies. These methods often suffer from the following drawbacks: Most methods determine whether to trigger charging based solely on remaining battery power, without considering the urgency of task completion or the congestion of charging resources. Current scheduling models often rely on static judgments, failing to adapt scheduling behavior to dynamic resource changes in real time. This can easily lead to scheduling failures in high-load or queuing scenarios. Therefore, this paper proposes a method for scheduling mobile charging robots to return charging, aiming to address these issues. Summary of the Invention
[0004] To achieve the above object, the present invention provides the following technical solutions: The mobile charging robot returns to the charging scheduling method, including the following steps: The first step is to collect status information of multiple mobile charging robots during operation. The status information includes current location, remaining battery power, remaining mission path length, mission urgency, battery performance indicators, historical charging records, current usage status of candidate charging locations, and occupancy status within the predicted time period. In the second step, an initial evaluation is performed when the status information meets the scheduling evaluation trigger conditions. In the initial evaluation, the behavior urgency index and energy competition index are calculated based on the status information to quantify the urgency of the robot's current task and the degree of competition for charging resources. In the third step, the behavioral urgency index and energy competition index are used as inputs and substituted into a learning model trained based on historical operating data. The model outputs a set of decision parameters, including a scheduling trigger flag, a path ranking factor, and a path trade-off factor. These factors indicate whether to execute a return schedule, the priority order of candidate paths, and the dynamic weight allocation between behavioral urgency and energy competition, respectively. In the fourth step, the trade-off between behavioral urgency and energy competition is calculated based on the decision parameter group, and a target path is selected from the candidate charging locations. The target path is the charging path that is determined to be the comprehensive optimal result after the trade-off.
[0005] In a preferred embodiment, during the process of collecting status information of multiple mobile charging robots during operation, the task urgency is calculated by calculating the ratio of the remaining time of the task to the estimated time of the shortest path from the current position to the task target position, using the formula: task urgency = remaining time of the task / estimated time of the path; Battery performance indicators are calculated by combining the normalized mean of the historical charge and discharge cycles and the current voltage recovery rate. The normalization method is to divide each parameter by the historical optimal reference value. The combined formula is: ; The historical charging record generates a weighted value for charging behavior by counting the total number of charging times and the duration of each charging session within a set operating cycle. The calculation method is to multiply the duration of each charging session by the weight of the time period in which it occurs and then take the average. The formula is: , where tn is the duration of the nth charging, wn is the corresponding weight preset in the period of the nth charging, and N is the total number of charging times; The occupancy of candidate charging locations during the prediction period is estimated by introducing an exponentially weighted moving average model by setting a five-minute sliding window. The number of queues, task type density, and continuous occupancy time of the charging pile within the most recent fixed time window are used for prediction. The exponential weighting function is: , where X_t is the current observation value, which is obtained by calculating the Euclidean distance between the observation vector composed of the number of queues, task type density and continuous occupancy time of the charging pile and the preset standard vector, E_t and are respectively the tense value of the occupancy situation in the current corresponding prediction time period and the tense value of the occupancy situation in the previous corresponding prediction time period, and α is the preset weight attenuation coefficient.
[0006] In a preferred embodiment, the task type density is the ratio of the number of high-priority tasks assigned to the target area to the total number of tasks per unit time, using the formula: task type density = number of high-priority tasks / total number of tasks; The continuous occupancy time of the charging pile is the cumulative value of the period of time when the continuous connection state exceeds the set connection threshold, and the total duration of all connection segments is counted; After calculation, the battery performance index is corrected by the temperature factor and the overload flag. The temperature factor is the proportional difference between the current ambient temperature and the standard ambient temperature. The correction method is: ; ; T is the current temperature, T0 is the standard temperature, and the temperature attenuation coefficient is the empirical setting value; The overload flag is a logic value. If the current peak detected in the last three charge and discharge cycles exceeds 120% of the rated value, it is marked as 1, otherwise it is 0, and then the multiplication factor is used: , and lower the revised values of battery performance indicators.
[0007] In a preferred embodiment, the scheduling evaluation trigger condition is that the remaining power is insufficient to support the completion of the remaining path length of the task, and there is at least one backup charging location in the candidate path that is reachable within the predicted time period, or the total expected queuing time of the candidate charging locations in the predicted time period exceeds the maximum waiting time limit. If there is no backup charging location in the candidate path that is reachable within the predicted time period, an active power supply operation is performed.
[0008] In a preferred embodiment, the behavior urgency index is calculated in the initial assessment based on the estimated path time between the current location and the task target location, the remaining power, the battery performance index, and the remaining time of the task. The nonlinear safety margin model is used to construct the index value. The formula is: , where R is the ratio of the remaining task time to the path estimation time, C is the normalized value of the battery performance index after the current correction and down-adjustment operation, and E is the ratio of the difference between the remaining power and the path estimation power consumption to the path estimation power consumption.
[0009] In a preferred embodiment, the energy competition index is modeled in the initial evaluation based on the current usage status of the candidate charging locations, occupancy during the forecast period, task type density, historical charging records, and path navigation costs. The local congestion coefficient is used to construct the model. The energy competition index is calculated as follows: ,in: D is the normalized result of the navigation cost, calculated as the ratio of the actual path cost to the benchmark shortest path cost. The benchmark shortest path cost is the path with the smallest cost among all candidate charging locations from the current location. Q is the number of high-priority tasks whose predicted target is the current candidate charging location per unit time multiplied by the mapping weight corresponding to their task type density. The number of high-priority tasks is the number of tasks within the prediction time period whose target is the charging location and whose task urgency exceeds the preset threshold. Z is the weighted value of the charging behavior. A is the expected availability value of the charging location within the prediction time period, defined as the normalized tension value E_t.
[0010] In a preferred embodiment, the behavior urgency index and the energy competition index are substituted as inputs into a learning model trained based on historical operating data. The learning model is constructed based on labeled sample data, and a two-dimensional input vector consisting of the behavior urgency index and the energy competition index is mapped to the scheduling execution result space. A multi-layer neural network structure is used to perform prediction calculations. The decision parameter group output by the model includes a scheduling trigger flag, a path ranking factor, and a path trade-off factor, where: The scheduling trigger flag is output through the sigmoid activation function, indicating whether return scheduling is executed in the current state. When the output result is greater than 0.5, it is determined to be immediately scheduled; The path ranking factor is the normalized value obtained by ranking and scoring all candidate paths based on the feature set consisting of historical task completion rate, path success rate, and failure readjustment rate. It is used to determine the priority order of candidate paths. The path trade-off factor is an adjustment coefficient used to dynamically allocate between behavioral urgency and energy competition, and is calculated as the interval [-1, 1] corresponding to the tanh activation output.
[0011] In a preferred embodiment, in the process of calculating the trade-off between behavior urgency and energy competition based on the decision parameter group, the path trade-off factor uses the ratio function of the behavior urgency index and the energy competition index as the basic input, and is dynamically adjusted in combination with the historical scheduling offset mode. The calculation formula is: , where: B represents the behavior urgency index; E represents the energy competition index; H represents the historical scheduling offset factor, which is defined as the product of the path length deviation between the actual execution path and the optimal recommended path in the past three consecutive schedulings and the change in the scheduling success rate; σ is the preset adjustment sensitivity parameter; The path trade-off factor participates in the comprehensive optimization scoring of each path in the candidate charging location set. The comprehensive scoring result is used to screen the target path, which is the charging path with the highest score and the current remaining power to support the execution of its complete navigation path. Comprehensive optimization score = path tradeoff factor × remaining power redundancy ratio × equivalent waiting time adjustment factor, where the remaining power redundancy ratio is the ratio of the current remaining power to the estimated power required for the target path. If this ratio is less than 1, the comprehensive score of the path is forced to zero. The equivalent waiting time adjustment factor is the ratio of the predicted waiting time to the standard maximum acceptable waiting time, calculated as: , where DD is the predicted waiting time at the current path's target location, YZ is the set maximum waiting tolerance threshold, and the predicted waiting time is set to the preset soft lock limit of the target charging location. The soft lock limit is defined as the average service processing time of the candidate charging location over the most recent consecutive scheduling cycles, where the average service processing time is the arithmetic mean of the time it takes for the charging location to complete a complete charging task. The ratio of path costs: the benchmark shortest path cost is the path with the lowest cost from the current location to all candidate charging locations. Q is the number of high-priority tasks predicted to target the current candidate charging location per unit time, multiplied by the mapping weight corresponding to its task type density. The number of high-priority tasks is the number of tasks within the prediction time period whose target is the charging location and whose task urgency exceeds the preset threshold. Z is the charging behavior weight, and A is the expected availability value of the charging location within the prediction time period, defined as the normalized stress value E_t.
[0012] In a preferred embodiment, the behavior urgency index and the energy competition index are substituted as inputs into a learning model trained based on historical operating data. The learning model is constructed based on labeled sample data, and a two-dimensional input vector consisting of the behavior urgency index and the energy competition index is mapped to the scheduling execution result space. A multi-layer neural network structure is used to perform prediction calculations. The decision parameter group output by the model includes a scheduling trigger flag, a path ranking factor, and a path trade-off factor, where: The scheduling trigger flag is output through the sigmoid activation function, indicating whether return scheduling is executed in the current state. When the output result is greater than 0.5, it is determined to be immediately scheduled; The path ranking factor is the normalized value obtained by ranking and scoring all candidate paths based on the feature set consisting of historical task completion rate, path success rate, and failure readjustment rate. It is used to determine the priority order of candidate paths. The path trade-off factor is an adjustment coefficient used to dynamically allocate between behavioral urgency and energy competition, and is calculated as the interval [-1, 1] corresponding to the tanh activation output.
[0013] In a preferred embodiment, in the process of calculating the trade-off between behavior urgency and energy competition based on the decision parameter group, the path trade-off factor uses the ratio function of the behavior urgency index and the energy competition index as the basic input, and is dynamically adjusted in combination with the historical scheduling offset mode. The calculation formula is: , where: B represents the behavior urgency index; E represents the energy competition index; H represents the historical scheduling offset factor, which is defined as the product of the path length deviation between the actual execution path and the optimal recommended path in the past three consecutive schedulings and the change in the scheduling success rate; σ is the preset adjustment sensitivity parameter; The path trade-off factor participates in the comprehensive optimization scoring of each path in the candidate charging location set. The comprehensive scoring result is used to screen the target path, which is the charging path with the highest score and the current remaining power to support the execution of its complete navigation path. Comprehensive optimization score = path tradeoff factor × remaining power redundancy ratio × equivalent waiting time adjustment factor, where the remaining power redundancy ratio is the ratio of the current remaining power to the estimated power required for the target path. If this ratio is less than 1, the comprehensive score of the path is forced to zero. The equivalent waiting time adjustment factor is the ratio of the predicted waiting time to the standard maximum acceptable waiting time, calculated as: , where DD is the predicted waiting time of the target location on the current path, YZ is the set maximum waiting tolerance threshold, and the predicted waiting time is set as the preset soft lock time limit of the target charging location. The soft lock time limit is defined as the average service processing time of the candidate charging location in the last several consecutive scheduling cycles, where the average service processing time is the arithmetic mean of the time it takes for the charging location to complete a complete charging task.
[0014] Technical effects and advantages of the present invention: The present invention comprehensively collects multi-dimensional status information of the mobile charging robot during operation, making the judgment basis before scheduling more complete and the data source more authentic and reliable. Before scheduling is executed, the mobile charging robot body may be in a complex and frequently changing operating environment. If there is a lack of detailed perception of its status, the scheduling accuracy will be seriously restricted. The information collected in the first step of the present invention covers the robot's spatial position, energy status, task progress, battery health and historical behavior, and introduces the usage of candidate charging positions within the predicted time period. The combination of the above information constitutes the necessary data set for scheduling evaluation, which not only realizes the accurate capture of the current task execution status, but also provides comprehensive and synchronous support for the subsequent scheduling mechanism, ensuring that subsequent logical judgments do not fail due to missing data.
[0015] The present invention introduces a quantitative evaluation mechanism to convert the multi-dimensional factors of task status and resource environment that need to be considered during scheduling into index values with clear physical meanings, so as to achieve reasonable judgment and standard unification before scheduling is triggered. By calculating the behavioral urgency index and the energy competition index, the two core dimensions of whether the task is urgent and whether charging resources are scarce are expressed in a clear numerical way, so that the scheduling system can quickly judge whether there is a need to return to scheduling under complex conditions. The index value is calculated from the status information collected in the first step, and is dynamic and real-time, which solves the problem that traditional methods rely on fixed thresholds and are difficult to adapt to different operating scenarios. This evaluation method also provides a structured input basis for subsequent machine learning models, effectively improving the overall intelligence level and adaptive scheduling capabilities of the system.
[0016] The present invention establishes a scheduling parameter generation mechanism based on the index evaluation results, and uses this mechanism to guide the screening and selection process of the target path, thereby realizing an integrated linkage from judgment to action and improving the efficiency of the scheduling decision-making closed loop. The present invention introduces the behavior urgency index and the energy competition index as inputs into the learning model, and forms an internal basis for scheduling behavior decisions through a decision parameter group composed of a scheduling trigger mark, a path ranking factor, and a path trade-off factor output by the model. In the fourth step, the priority between behavior urgency and energy competition is weighed accordingly, and finally the target path is selected from the candidate path set to complete the establishment of the candidate path for the scheduling behavior. This process not only ensures the continuity and explainability of the scheduling behavior, but also builds a complete data-driven mechanism within the scheduling system, improving the response speed and selection quality from judgment to execution. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to facilitate understanding by those skilled in the art, the present invention will be further described below with reference to the accompanying drawings; Figure 1 This is a schematic diagram of the return charging scheduling method for the mobile charging robot in the present invention. DETAILED DESCRIPTION
[0018] The following will provide a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0019] Reference Figure 1 The following examples were obtained: Example 1: A method for scheduling a mobile charging robot to return to charging, comprising the following steps: The first step involves collecting status information from multiple mobile charging robots during operation. This information includes current location, remaining battery life, remaining task path length, task urgency, battery performance indicators, historical charging records, the current usage status of candidate charging locations, and occupancy during the predicted time period. This step builds the fundamental data support for scheduling decisions. By continuously collecting the operating status of multiple mobile charging robots, a multi-dimensional view of the system's current operational status is established. The current location reflects the robot's current spatial coordinates, providing a starting point for path planning. The remaining battery life is a core indicator for determining whether charging is required during scheduling. The remaining task path length is used to assess the accessibility of task completion in conjunction with the remaining battery life. Task urgency reflects the strength of the task completion time constraint and helps distinguish between high and low priority. Battery performance indicators are used to determine the current battery health and energy release capacity, preventing system scheduling from being based on incorrect assumptions. Historical charging records reflect the robot's long-term dependence on charging resources. The current usage status of candidate charging locations provides static occupancy information, while occupancy during the predicted time period is used to estimate potential future resource constraints using a dynamic model. This step provides the data foundation and real-time feedback support for subsequent behavior judgment and scheduling strategy development.
[0020] The second step is an initial assessment, performed when the status information meets the scheduling evaluation trigger conditions. This initial assessment calculates the behavioral urgency index and energy competition index based on the status information to quantify the urgency of the robot's current task and the degree of competition for charging resources. This step is a logical judgment process initiated when the robot's operating state meets specific scheduling evaluation trigger conditions. Through in-depth analysis of the status information, the urgency of the robot's task execution and the pressure on external resource usage are quantified. The core of the behavioral urgency index calculation lies in characterizing the matching relationship between the current task time constraint and available energy, reflecting whether the robot is in a state of urgent execution or potential failure. The energy competition index reflects the load status of the overall scheduling environment by constructing judgment logic for the future competition for candidate charging resources. These two indices are independent of each other, yet together they form the dual-dimensional foundation of scheduling decisions, providing a reliable basis for subsequent decision-making on whether to return to scheduling and how to return.
[0021] In the third step, the behavioral urgency index and energy competition index are fed into a learning model trained on historical operational data. The model outputs a set of decision parameters, including a scheduling trigger flag, a path ranking factor, and a path tradeoff factor. These factors, respectively, indicate whether to execute a return schedule, the priority order of candidate paths, and the dynamic weighting between behavioral urgency and energy competition. This step utilizes a data-driven learning mechanism, incorporating the behavioral urgency index and energy competition index as structured inputs of the scheduling context into the learning model trained on historical scheduling data, yielding a structured scheduling decision output. By deeply learning the relationship between historical scheduling behaviors and outcomes, the model achieves reasoning and prediction capabilities under current input conditions. The output scheduling parameter set includes a scheduling trigger flag as a logical decision metric, determining whether to immediately execute a scheduling action under the current state; a path ranking factor, which prioritizes multiple candidate paths to ensure a reasonable path bias for the scheduling objective; and a path tradeoff factor, which dynamically weights the two evaluation objectives of behavioral urgency and resource competition to ensure that the scheduling result achieves the optimal compromise under multiple objective conditions. This step is the core component of the system's intelligent decision-making process, establishing a bridge between data perception and policy output.
[0022] The fourth step is to calculate the trade-off between behavioral urgency and energy competitiveness based on the decision parameter set. A target path is then selected from the candidate charging locations. This target path is the charging path that is determined to be the optimal overall path after these trade-offs and is recorded as the candidate path for scheduling behavior. This step, based on the path ranking factor and path trade-off factor output by the learning model in the previous step, quantitatively calculates a comprehensive score between behavioral urgency and energy competitiveness, which is used to compare and judge all candidate charging paths. Among all candidate charging locations, the target path that is optimal overall in terms of the scoring dimension is selected, taking into account the current power accessibility and the resource pressure of the path. This path is marked as a candidate path for scheduling behavior in the scheduling logic and serves as an important intermediate input before the final behavioral decision is executed, ensuring that the execution of the return path considers both individual urgency and global resource balance. This step completes the transformation from model output parameters to specific behavioral implementation paths, ensuring the controllability and explainability of the scheduling process.
[0023] When collecting status information from multiple mobile charging robots during operation, task urgency is calculated by the ratio of the remaining time to the estimated time of the shortest path from the current location to the task's target location. The formula is: Task Urgency = Remaining Time / Estimated Path Time. The remaining time is the time remaining until the robot completes the task it is currently performing. It can be calculated by inferring the difference between the task's dispatch time and the current system time, taking the task's deadline window as the upper limit. The estimated path time is the expected time required for the robot to move from its current location to the task's target location along the optimal path (usually the shortest path). This path time can be calculated by dividing the length of the navigation path output by conventional path planning algorithms (such as Dijkstra or A*) by the robot's current speed. For example, if the remaining time for a task is 120 seconds and the estimated path time from the current location to the task's target location is 80 seconds, the task urgency is 1.5, indicating that there is still some margin of error. If the ratio is less than 1, the task may not be completed on schedule and has a scheduling priority.
[0024] Battery performance indicators are calculated by combining the normalized mean of the historical charge and discharge cycles and the current voltage recovery rate. The normalization method is to divide each parameter by its historical optimal reference value (such as the factory indicator or the best performance record of the robot model) to eliminate the dimensional and benchmark deviations between different samples. The combined formula is: ; Among them, the historical number of charge and discharge cycles refers to the number of complete charge and discharge cycles completed by the robot since it was put into operation, reflecting the degree of battery aging; the current voltage recovery rate refers to the speed at which the battery voltage recovers within a certain period of time after stopping discharge, indicating its internal resistance and activity level. The normalization calculation method is as follows: if the current number of cycles is 600 times and the historical reference best value is 500 times, then the normalized value is 600 / 500=1.2; if the voltage recovery rate is 0.1V / min and the reference value is 0.2V / min, then the normalized value is 0.5. The combined weighted average is the battery performance index. This indicator can be used to determine whether the robot has the ability to complete the scheduling path. If the battery performance deteriorates significantly, charging scheduling will be triggered first.
[0025] The historical charging record is generated by counting the total number of charging times and the duration of each charging within a set operating cycle (such as 72 consecutive hours) to generate a weighted value for charging behavior. The calculation method is to multiply the duration of each charging time by the weight of the time period and then take the average. The formula is: Where tn is the duration of the nth charge, i.e., the time from connecting to the charging station to disconnecting. wn is the preset weight for the time period in which the nth charge occurs. Weights can be set to 1.2 during peak charging periods (such as nighttime and overlapping missions), 1.0 during normal periods, and 0.8 during idle periods, reflecting the pressure that charging places on resources. N is the total number of charges, used to normalize the weighted average. For example, if a robot charges four times during a certain operation cycle, each for 60, 45, 30, and 20 minutes, respectively, falling into time periods with weights of 1.2, 1.0, 0.8, and 0.8, respectively, the weighted values are: (60 × 1.2 + 45 × 1.0 + 30 × 0.8 + 20 × 0.8) / 4 = (72 + 45 + 24 + 16) / 4 = 157 / 4 = 39.25 minutes. A higher weight indicates a greater reliance on system charging resources and should be prioritized during scheduling.
[0026] The occupancy of candidate charging locations during the prediction period is estimated by introducing an exponentially weighted moving average model by setting a five-minute sliding window. The number of queues, task type density, and continuous occupancy time of the charging pile within the most recent fixed time window are used for prediction. The exponential weighting function is: ; Among them, α is the weight attenuation coefficient, with a recommended value of 0.3, which is used to control the weight of the influence of the current observation value on the prediction result. X_t is the current observation value, and the observation vector composed of three dimensions is the current number of people in the queue, the density of high-priority tasks per unit time, and the continuous occupancy time of the charging pile (referring to the cumulative time of maintaining the connection state exceeding the preset threshold). The three are uniformly calculated through the Euclidean distance and the standard reference vector (such as the system idle benchmark state), indicating the degree of deviation of the current system charging resource tension. E_t and Represents the tension value of the current and previous time points in the predicted time period. The model smoothes the sudden change trend by weighted superposition to predict the future load of the system. For example: the current value of X_t is calculated to be 3.2 (unit: Euclidean distance) away from the baseline state. If it is 2.5, then E_t=0.3×3.2+0.7×2.5=0.96+1.75=2.71, which means that the charging location may enter a peak period of resource competition in the future.
[0027] Task type density is the ratio of the number of high-priority tasks assigned to the target area to the total number of tasks per unit time. The formula is: Task type density = number of high-priority tasks / total number of tasks, where unit time refers to the scheduling evaluation window period set in the scheduling system (such as 5 minutes or 10 minutes). The target area is the physical service area where the current candidate charging location is located. Its range can be defined based on the charging location radiation radius or the navigation map topology level: When the charging location radius is used as the delineation basis, the target area is defined as a two-dimensional circular region with the candidate charging location as the geometric center and a preset distance threshold (e.g., 30 meters or 50 meters) as the radius. All scheduled tasks within this region whose target points or execution paths overlap or intersect with this circular region are considered to belong to the charging location's target area. This approach is suitable for scenarios with relatively regular map structures and point-to-point robot movement, facilitating direct spatial mapping and distance determination. When the navigation map topology is used as the delineation basis, the target area is defined as the set of navigation nodes directly adjacent to the candidate charging location. The topology map consists of a graph structure consisting of all intersections of navigable paths, with connections between nodes representing traversable paths. The target area is defined as all path segments covered by n layers (e.g., two or three layers) of adjacent nodes extending outward from the charging location node as the center. This approach is suitable for large, multi-block sites, complex map structures, or environments with closed navigation paths to ensure that the delineated area aligns with actual drivable routes. For example, if charging location A is the central node and two layers of adjacent nodes are expanded outward, the path connected in the topology diagram has 9 sections and covers 5 task target points. Then all tasks whose path targets fall within these 5 nodes are considered to belong to the target area of the charging location.
[0028] High-priority tasks are those whose task type rating exceeds a preset urgency threshold (e.g., response time below a set value, task objectives uninterruptible, etc.). For example, within a scheduling cycle, the system assigns 20 tasks to a target area, 8 of which are high-priority tasks. The task type density is 8 / 20 = 0.4. This density value reflects the proportion of high-priority tasks in the local task backlog. A higher value indicates greater scheduling pressure in that area, requiring more cautious resource scheduling.
[0029] The continuous occupancy time of a charging pile is the cumulative value of the period of time in which the continuous connection state remains above the set connection threshold, and the total duration of all connection segments is counted. Among them, the connection state is the period when the mobile charging robot is in the charging interface connected state (i.e., contact charging or wireless coupling current is closed) and maintains a stable charging process; the connection threshold is set to the minimum effective charging duration preset by the system (for example, set to 15 minutes). If it is less than this duration, it is considered invalid occupation. When counting, if a charging position has been occupied by three different robots in the past hour, with connection durations of 10, 20, and 35 minutes respectively, and only the latter two are greater than the set threshold, the continuous occupancy time is 20 + 35 = 55 minutes. This value is used to assess the saturation of recent resource occupancy at the location and reflect its exclusive impact on the scheduling system.
[0030] After calculation, the battery performance index is corrected by the temperature factor and the overload flag. The temperature factor is the proportional difference between the current ambient temperature and the standard ambient temperature, which is used to reflect the thermal impact of the operating environment on battery performance. The correction method is: ; Where T is the current ambient temperature (detected by the robot's sensors), and T0 is the standard ambient temperature (typically set at 25°C). The temperature attenuation coefficient is an empirically set value (for example, 0.05) that adjusts the impact of temperature differences on performance. For example, if T = 35°C and T0 = 25°C, the temperature difference ratio is 0.4. If the original battery performance index is 0.85, the corrected value is 0.85 × (1 − 0.4 × 0.05) = 0.85 × 0.98 = 0.833. This mechanism ensures accurate judgment of battery capacity during dispatch in high or low temperature environments.
[0031] The overload flag is a logical value used to determine whether the battery has been in an overload state recently. The specific judgment is: if the current peak detected in any cycle within the last three charge and discharge cycles exceeds 120% of the rated current value, the flag is 1, indicating an overload risk; otherwise, it is 0, indicating that the battery is operating within the normal load range. The impact of the overload flag on battery performance indicators is multiplied by the following factors: , the revised values of battery performance indicators are adjusted downwards, and the correction method is as follows: For example, if the battery performance index after temperature correction in the previous stage is 0.833 and the overload flag is 1, the final correction value = 0.833 × (1 − 0.1) = 0.75. If the overload flag is 0, the final correction value remains at 0.833. This mechanism is used to prevent potential battery aging risks and internal heating anomalies before they occur, ensuring that the robot does not mistakenly enter a high-intensity return scheduling path if the battery load capacity is insufficient.
[0032] The scheduling evaluation trigger conditions are: the remaining power is insufficient to complete the remaining path length of the task, and there is at least one backup charging location on the candidate path that is reachable within the predicted time period, or the estimated waiting time of all candidate charging locations exceeds the maximum waiting time within the predicted time period. If there is no backup charging location on the candidate path that is reachable within the predicted time period, active power supply is performed. The remaining power refers to the available energy remaining in the mobile charging robot's current battery, which can be expressed as a percentage or watt-hour. This value is obtained in real time by the robot's power system detection module.
[0033] The remaining path length of a task refers to the portion of the robot's path that has not yet completed the task. It is an estimate of the shortest feasible path from the current location to the task's target location, expressed in meters or seconds. "Insufficient remaining battery life to complete the remaining path length" indicates that, based on the ratio of the current remaining battery life to the energy consumption per unit distance / unit time, the robot will be unable to complete the current task path, posing a risk of power outage. This is the primary prerequisite for the scheduling system to initiate an assessment. A candidate path refers to the set of all charging paths that the robot can navigate to from its current location. These paths are typically generated by existing path planning algorithms based on map structure and charging locations. Alternative charging locations are identified as charging resource nodes within the candidate path and are alternative locations to the current default target path. Reachability within the predicted time period means that the robot's current remaining battery life is sufficient to reach the alternative charging location within the predicted time window, and that there are no resource exclusivity conflicts or navigation obstacles at that location. Reachability requires two conditions: the current remaining battery life ≥ the estimated power consumption of the predicted path; and the charging location is expected to be unoccupied or accessible in the queue buffer during the predicted time period. For example: Robot A has 8% remaining battery power, and the power required to navigate from its current location to candidate charging location P1 is 6%. The predicted arrival time is 10 minutes, and the number of people queuing at P1 during this time period is expected to be 0, indicating that it is reachable. If the power required for P2 is 9%, it is unreachable.
[0034] The maximum waiting time is the service response threshold set by the scheduling system, for example, it is set to 20 minutes. Exceeding this value will cause task delays or resource imbalance. Therefore, "the total expected queuing time of candidate charging locations within the predicted time period exceeds the maximum waiting time limit" means that all possible charging nodes in the current path are in a state of severe resource congestion in the future. Even if the robot's power can still support navigation, there is no effective charging supply point, which triggers the scheduling evaluation mechanism. Active power supply operation is an emergency scheduling strategy: when there is no backup charging location that is reachable within the predicted time period in the candidate path, that is, the charging resources cannot meet the conditions in space or time, the scheduling system will execute the task interruption or task transfer mechanism, and instruct other resources or platforms (such as backup robots, power supply interfaces or scheduling stations) to perform on-site power supply or transfer processing on the current robot to avoid complete power outage at the path interruption point. This strategy is a built-in fault tolerance protection mechanism of the system.
[0035] In the initial assessment, the behavioral urgency index is calculated based on the estimated path time between the current location and the task target location, the remaining power, the battery performance index, and the remaining time of the task. The nonlinear safety margin model is used to construct the index value. The formula is: , where R is the ratio of the remaining time of the task to the time consumed by the path estimation, C is the normalized value of the battery performance index after the current correction and downward adjustment operations, and E is the ratio of the difference between the remaining power and the power consumption of the path estimation to the power consumption of the path estimation. The closer the behavior urgency index value is to 1, the smaller the time and power redundancy of the robot task completion is, and the higher the priority required for behavior scheduling.
[0036] The estimated path duration between the current location and the mission target is the estimated time required for the mobile charging robot to travel the shortest path from the current location to the mission target, based on the current map and navigation status. This duration is calculated by the path planning algorithm, typically in seconds, and is used to measure the time cost of reaching the mission target. The remaining battery capacity is the real-time value of the current available battery capacity, expressed as a percentage or watt-hours. This value is output in real time by the robot's power management module and is used to measure whether there is sufficient power to continue the mission. In this model, the battery performance index is a normalized value that reflects the historical number of cycles, voltage recovery rate, and temperature and overload corrections. It represents the battery's overall health and continuous discharge capacity. The normalized value is generally controlled within the range (0, 1). The remaining time of a task is the amount of time remaining between the current time and the scheduled completion time of the task in the scheduling system. It is calculated based on the execution time window set when the task is issued and is expressed in seconds. The variables in the exponential formula are defined as follows: R represents the ratio of the remaining time of the task to the estimated path duration, calculated as: R = Remaining time of task / Estimated path duration. This value indicates the degree of time redundancy. A smaller R indicates that the remaining time is closer to the minimum duration, and the scheduling pressure is greater. C represents the normalized value of the current battery performance index, which is defined in the battery performance section and includes the results after temperature and overload corrections. The value range is limited to 0 to 1, with a smaller value indicating poorer battery performance. E represents the remaining charge redundancy ratio, calculated as: This value reflects the safety margin of remaining power relative to the task's power consumption. A positive value for E indicates power redundancy; a negative value indicates insufficient power. This formula uses a negative exponential function to ensure that the behavior urgency index approaches 0 when both the task and energy redundancy are sufficient (large R, large E). When both power and time are limited (small R, small or even negative E), the index approaches 1, reflecting the execution urgency of the task in terms of spatial and temporal resources.
[0037] In the initial evaluation, the energy competition index is modeled based on the current usage status of the candidate charging locations, occupancy during the forecast period, task type density, historical charging records, and path navigation costs. The local congestion coefficient is used to construct the model. The calculation formula for the energy competition index is: ,in: D is the normalized result of the navigation cost, calculated as the ratio of the actual path cost to the benchmark shortest path cost. The benchmark shortest path cost is the path with the smallest cost among all candidate charging locations from the current location. Q is the number of high-priority tasks whose predicted target is the current candidate charging location per unit time multiplied by the mapping weight corresponding to their task type density. The number of high-priority tasks is the number of tasks within the prediction time period whose target is the charging location and whose task urgency exceeds the preset threshold. Z is the weighted value of the charging behavior. A is the expected availability value of the charging location within the prediction time period, defined as the normalized tension value E_t.
[0038] The current usage status of a candidate charging location indicates whether it is currently occupied by another robot and whether it is in a "charging" or "available" state. This information is provided in real time by the dispatch system's charging pile status feedback module. The occupancy status during the forecast period represents the resource pressure trend at that location within the future forecast window. This prediction output is the stress value E_t, which is defined here as A after normalization. The stress value is normalized to a standard interval [0, 1], for example, with the maximum stress value for the same time period in the last 24 hours set to 1 and the minimum to 0, to maintain consistency in the model input dimensions. Historical charging records are obtained by calculating a weighted value for charging behavior. This value comprehensively considers the duration of each charge and the weight of the time period, representing a robot's historical preference and reliance on a specific charging location. Here, Z is the weighted value of the current robot's charging behavior at the target candidate charging location. A larger value indicates a higher historical reliance on that location and a greater sensitivity to resource competition.
[0039] D = Actual Path Cost / Baseline Shortest Path Cost, where the actual path cost is the cost of navigating from the robot's current location to a candidate charging location. This cost not only includes the path distance but also factors such as obstacle complexity and congestion penalty. The base shortest path cost is the cost of the lowest-cost path from the robot's current location to all candidate charging locations, serving as a normalization benchmark. Through normalization, D reflects the relative path disadvantage of the robot from its current location to the candidate charging location. A D closer to 1 indicates a near-optimal path; a larger D indicates a higher navigation cost and a higher scheduling cost.
[0040] It should be noted that in most path planning and robot scheduling systems, path cost (PathCost) is typically used to evaluate the overall efficiency of a navigation path. Its calculation method often includes the following components: Path distance: The length of the straight or curved path from the robot's starting point to the destination, the most basic cost element. Path obstacle complexity: The density or distribution of static or dynamic obstacles (such as equipment or personnel) in the path, resulting in travel difficulty. For example, the more obstacles and the narrower the passage, the higher the cost. Congested area penalty coefficient: If a path passes through areas with high traffic or high task density in the map, a certain cost penalty is added to simulate the possibility of localized travel delays or scheduling conflicts. This type of comprehensive path cost model can be found in, for example, the weighted version of the A* algorithm, the Dijkstra enhancement model, and the reward correction strategy in reinforcement learning. It is a general path cost evaluation mechanism and is commonly used by those skilled in the art. In the "Mobile Charging Robot Return Charging Scheduling Method" of the present invention, the calculation of the normalized path navigation cost D is based on the above-mentioned existing modeling method. Its purpose is to normalize the actual path cost with the shortest path cost for subsequent energy competition index modeling. Therefore, the present invention does not advocate "path cost = path distance + obstacle + congestion penalty" as an innovation point and will not elaborate on it; instead, it innovates in that the path cost normalization result D is combined with Q, A, and Z in the energy competition index formula to construct a nonlinear scheduling index.
[0041] The number of high-priority tasks targeting the location per unit time multiplied by the task type density (Q). The number of high-priority tasks is the number of tasks targeting the charging location within the forecast period whose urgency exceeds a preset threshold. The urgency is defined in the behavioral urgency index, with a typical threshold of 0.7. The task type density is the ratio of the number of high-priority tasks assigned to the target area to the total number of tasks per unit time. Density weight mappings, such as low density = 0.8, medium density = 1.0, and high density = 1.2, emphasize the amplified degree of competition in areas with high task density. For example, within a 5-minute period, a charging location is expected to receive six high-priority tasks, with a medium density level and a corresponding weight of 1.0, where Q = 6 × 1.0 = 6. The energy competition index essentially constructs a local resource competition assessment model. It uses a logarithmic function to nonlinearly compress the ratio to ensure numerical stability and highlight the steep exponential rise in high-competition areas. This nonlinear compression makes the energy competition index highly distinguishable across different scheduling scenarios.
[0042] The behavioral urgency index and energy competition index are fed into a learning model trained on historical operational data. The model is constructed based on labeled sample data. The two-dimensional input vector consisting of the behavioral urgency index and the energy competition index is mapped to the scheduling execution result space. Predictive calculations are performed using a multi-layer neural network structure. The model outputs a set of decision parameters, including a scheduling trigger flag, a path ranking factor, and a path trade-off factor. The scheduling trigger flag is output via a sigmoid activation function, indicating whether return scheduling is executed in the current state. An output greater than 0.5 indicates immediate scheduling. The path ranking factor is a normalized value obtained by ranking and scoring all candidate paths based on a feature set consisting of historical task completion rates, path success rates, and failure rescheduling rates. This value is used to determine the priority of candidate paths. The path trade-off factor is an adjustment coefficient used to dynamically balance behavioral urgency and energy competition. Input data structure: The input is a two-dimensional real number vector [B, E], where B is the behavioral urgency index, ranging from [0, 1], and E is the energy competition index, ranging from 0 to 1. This input is collected from all historical scheduling samples and serves as the input to the model training set. Label Structure (Supervised Training): Label data consists of three dimensions of scheduling execution results: whether the schedule was executed: a Boolean label (1 indicates immediate scheduling, 0 indicates no scheduling); path ranking target order: candidate path ranking number or normalized weight; and historical weights between behavioral urgency and energy competition, used to derive path trade-off factors. This supervision data comes from real scheduling records or simulated labels generated by empirical strategies. The model adopts a typical multi-layer perceptron (MLP) architecture, comprising: an input layer with two nodes corresponding to [B, E]; a hidden layer with at least two layers, each with 32 or 64 neurons, and a ReLU or LeakyReLU activation function; an output layer with output 1 (scheduling trigger flag): one node, with a sigmoid activation function; an output 2 (path ranking factor set): N nodes (N is the number of candidate paths), with a softmax activation function; and an output 3 (path trade-off factor): one node, with a tanh activation function. The scheduling trigger flag is output using a sigmoid activation function, with an output range of (0,1). The judgment criteria are: if the output result is > 0.5, it is determined that the return schedule should be executed immediately under the current state. Scheduling trigger mark: used to determine whether the scheduling behavior should be initiated; the output value is approximately the scheduling probability and is used in conjunction with the policy gating mechanism; path ranking factor: all candidate paths are scored and ranked by multiple features, and the output is a normalized weight. In the process of calculating the trade-off between behavioral urgency and energy competition based on the decision parameter group, the path trade-off factor uses the ratio function composed of the behavioral urgency index and the energy competition index as the basic input, and is dynamically adjusted in combination with the historical scheduling offset pattern. The calculation formula is: , where B represents the behavior urgency index; E represents the energy competition index; H represents the historical scheduling deviation factor, defined as the product of the path length deviation between the actual execution path and the optimal recommended path in the past three consecutive schedulings and the change in the scheduling success rate; σ represents the preset adjustment sensitivity parameter, an empirical value determined during system training based on the scheduling behavior stability assessment; path length deviation refers to the average sum of the differences between the actual execution path length and the model-recommended path length in each scheduling, reflecting the degree of execution path deviation; and the scheduling success rate change value is the range of success rate changes (maximum minus minimum) over the past three schedulings, used to assess scheduling stability fluctuations. A larger H indicates a larger historical scheduling deviation, and the current model output path trade-off factor should be adjusted to suppress system oscillations. The adjustment sensitivity parameter σ is a real constant set empirically during model training or scheduling system evaluation. It controls the impact of historical deviation on the path trade-off factor and is typically between 0.1 and 2.0. A larger σ indicates a greater impact of historical deviation on adjustment and a more cautious system; a smaller σ indicates a greater reliance on the current evaluation index itself.
[0043] If B≫E, it means that the task urgency is much higher than the resource shortage, the path trade-off factor tends to the positive range, and the system will give more priority to behavior; if E≫B, the system tends to adopt a resource-conservative strategy, and the path selection tends to be low-competition path; the larger H is, the impact of the difference on the final output will be compressed (divided by a larger denominator), that is, the model tends to have a stable and conservative output to prevent drastic path jumps. For example, the following parameters are set: current behavior urgency index B = 0.82; current energy competition index E = 0.47; the actual path lengths for the last three dispatches were 90m, 110m, and 100m, respectively; the recommended path lengths were 80m, 100m, and 90m, respectively; → average deviation: [(90-80)+(110-100)+(100-90)] / 3 = 10; the dispatch success rate changes were: 100%, 60%, and 80% for the last three dispatches, with the maximum minus the minimum = 0.4; → H = 10 × 0.4 = 4.0; σ = 0.5; the path trade-off factor is: A positive output indicates that the current path selection is more inclined toward behavioral urgency, but the existence of the historical offset factor makes the output cautious.
[0044] The path tradeoff factor contributes to the comprehensive optimization score of each path in the candidate charging location set. The comprehensive score is used to select the target path, which is the charging path with the highest score and the current remaining battery capacity to support its complete navigation path. This scoring mechanism quantifies the overall scheduling value of multiple candidate paths. By dynamically integrating factors such as behavioral urgency, resource competition, battery redundancy, and waiting cost, a single scoring metric is formed to ensure that the selected path achieves the global optimal balance between scheduling responsiveness and resource coordination. The comprehensive optimization score = path tradeoff factor × remaining battery redundancy ratio × equivalent waiting time adjustment factor. The remaining battery redundancy ratio is the ratio of the current remaining battery capacity to the estimated battery capacity required for the target path, indicating whether the robot can complete the navigation task to the end of the path without interruption. The calculation formula is: Remaining power redundancy ratio = current remaining power / estimated power consumption of the target path. If this ratio is less than 1, it means that the current energy cannot support complete path navigation, and the comprehensive score of the path is forcibly set to zero, making it considered an unselectable path. This rule ensures that the scheduling logic does not mistakenly select a path that exceeds the energy capacity due to other high-scoring factors, ensuring scheduling accessibility and practical feasibility.
[0045] The equivalent waiting time adjustment factor is used to suppress the expected queuing delay risk of different candidate paths at the target charging location. It reflects the relative ratio of the waiting time to the acceptable threshold and implements a nonlinear penalty for the waiting cost. Its calculation form is: , where DD is the predicted waiting time at the target location on the current path; YZ is the maximum waiting tolerance threshold, representing the upper limit of delays the scheduling system accepts in this charging scenario. The specific value is set by the scheduling policy and is typically derived from simulation experience or operational parameters. Values closer to 1 indicate acceptable waiting times; values closer to 0 indicate extremely high waiting costs for the current path, significantly reducing its overall preference score. The predicted waiting time is set as the preset soft lockout period for the target charging location, representing the service window the robot reserves for this candidate charging location in the scheduling policy. The soft lockout period is calculated as follows: Soft lockout period = average service processing time for the charging location over the most recent consecutive scheduling cycles. Average service processing time is defined as the arithmetic mean of the time it takes for the charging location to complete a complete charging task. The statistical window is typically set over the past 5 to 10 scheduling cycles to ensure a representative sample size without sacrificing real-time accuracy. A complete charging task refers to the entire process from robot insertion, charging initiation, and recovery of power above the scheduling lower limit before disconnection. This average value reflects the processing capacity of the charging point under typical workloads and can be used to estimate the expected response time in future time periods, ultimately playing a key penalty role in the comprehensive optimization score.
[0046] For example, a path has a remaining power redundancy ratio of 1.2, a path tradeoff factor of +0.65, a predicted wait time of 12 minutes, and a maximum tolerance threshold of 10 minutes. The equivalent adjustment factor is 1 / (1+12 / 10)≈0.45. This path's overall optimal score is 0.65×1.2×0.45≈0.351. Compared to another path with a wait time of only 4 minutes and slightly lower urgency (e.g., a tradeoff factor of 0.5), this path's overall score may be superior, thus achieving a tradeoff between feasibility and latency for scheduling paths.
[0047] Example 2: Based on Example 1, a fifth step is added. In the fifth step, a fuzzy logic device constructed based on the behavior urgency index, energy competition index and delay factor is used to determine whether the scheduling behavior is triggered immediately and its behavior type. The scheduling behavior type is selected based on the attribution result. The scheduling behavior types include immediate insertion scheduling, delayed queuing execution, scheduling after task transfer or target position replacement scheduling. Before the scheduling behavior takes effect, the path queuing simulation verification and scheduling stability confirmation are completed to trigger the robot to perform the return charging operation based on the candidate path.
[0048] During the fuzzy logic modeling process, the behavior urgency index and energy competition index were modeled separately in the previous steps. These are used to quantify resource conflict pressure and task delay risk in task scheduling. The delay factor is the ratio of the predicted wait time of the current candidate path to the maximum acceptable wait time threshold, defined as: DD / YZ, where DD is the predicted wait time and YZ is the maximum tolerable wait time. This factor reflects whether the target path is resource saturated. The three input variables are divided into three levels: "low," "medium," and "high" using a fuzzy mapping function, forming an input membership space. The fuzzy rule base is used to perform behavior type inference. An example rule is as follows: "If the behavior urgency index is high, the energy competition index is high, and the delay factor is high, then recommend alternative scheduling at the target location." All rule sets are developed through a combination of empirical data training and policy definition. The fuzzy inference results in a behavior label. Based on this label classification, the system executes scheduling path construction, navigation planning, and priority insertion strategies for the corresponding behavior type. To ensure system stability and scheduling prediction accuracy, after the behavior attribution results are generated, the behavior is applied to a well-established technology: path queuing simulation. The path queuing simulation module simulates the path conflict, service waiting and resource response effects of the behavior in the current navigation topology. If the simulation verification passes, the behavior officially enters the execution process; if there is a conflict, delay or blocking risk, the reasoning result of this time is rolled back and the next level of behavior priority is entered to re-evaluate the path.
[0049] In the fifth step, the fuzzy logic outputs the following four types of scheduling behaviors. The system assigns these behaviors based on the fuzzy inference results and completes the selection and triggering of the scheduling response path: Immediately insert scheduling interrupts the existing task execution process while the current task is incomplete, immediately triggering the navigation planning and scheduling instructions for the mobile charging robot to return to the charging path. This behavior is applicable to the following situations: a high behavior urgency index indicates that the task execution process is approaching the lower limit of available power and there is a risk of interruption; a low energy competition index indicates that the current candidate charging location has good resource utilization within the predicted time period; and a low delay factor indicates that the waiting time in the queue for the target path is within an acceptable range. Once this behavior is triggered, the robot will jump out of the task execution queue and execute charging path navigation based on the scheduling behavior candidate path generated in the previous stage. The scheduling system automatically suspends the task and locks the charging resources.
[0050] Delayed queue execution means that when the scheduling behavior is determined to be valid but does not meet the immediate triggering conditions, the scheduling request will be suspended and placed in the delayed queue, and the scheduling will be triggered in the queue order after the execution conditions are met.
[0051] This behavior applies to the following situations: a medium or low Behavior Urgency Index indicates that the robot still has some execution buffer; a medium or high Energy Competition Index indicates that candidate locations are experiencing queuing pressure during the forecast period; and a high Delay Factor indicates that the current path's wait time is approaching or exceeding the maximum acceptable threshold. The system records the dispatch request and path information in a delayed dispatch cache queue and reassesses its execution based on the following events: completion of the current task; an update to the candidate path's status (e.g., a decrease in the number of queued passengers); or a change in energy status that causes the Behavior Urgency Index to increase.
[0052] Post-task-handover scheduling means that the current robot must exit the original task flow prematurely due to scheduling requirements and hand over the task to another robot in the task pool, releasing the robot to immediately return to the charging path. This behavior applies to the following situations: the behavior urgency index is high; the energy competition index is high; and the delay factor is medium, meaning that the candidate paths have a certain degree of resource tension but not to the point of complete exclusion. The system transfers the original task to the task reallocation module, matching it with other idle or lightly loaded robots for task inheritance, ensuring that the original task is not interrupted by the charging operation, while freeing up the original robot's resources for charging. The task inheritance process considers factors such as the target area, remaining range, and scheduling queue to maintain overall system stability.
[0053] Target location replacement scheduling means that the original candidate path is cancelled as the target path due to reasons such as unacceptable queuing or loss of reachability. The system reselects the path with the smallest delay and the second highest comprehensive score in the candidate set as the new target path based on the comprehensive preference score, triggering navigation planning and scheduling operations. This behavior applies to the following situations: the behavior urgency index is high; the energy competition index is high; the delay factor is high, indicating that the currently selected path is no longer worth executing. The system will eliminate the current path and re-mark the path with the second highest comprehensive preference score in the candidate set and that meets the remaining power redundancy and reachability conditions as the target path, and generate the corresponding navigation sequence and queuing evaluation process, and enter the simulation verification module. This behavior ensures local optimal response in unavoidable high-load scenarios to avoid scheduling failures or path blockages.
[0054] The above formulas are all dimensionless and numerical calculations. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain the most recent real situation. The preset parameters in the formulas are set by technicians in this field according to actual conditions.
[0055] It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0056] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0057] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0058] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A method for scheduling the return charging of a mobile charging robot, characterized in that: The following steps are involved: The first step is to collect status information of multiple mobile charging robots during operation. The status information includes current location, remaining battery power, remaining mission path length, mission urgency, battery performance indicators, historical charging records, current usage status of candidate charging locations, and occupancy status within the predicted time period. In the second step, an initial evaluation is performed when the status information meets the scheduling evaluation trigger conditions. In the initial evaluation, the behavior urgency index and energy competition index are calculated based on the status information to quantify the urgency of the robot's current task and the degree of competition for charging resources. In the third step, the behavioral urgency index and energy competition index are used as inputs and substituted into a learning model trained based on historical operating data. The model outputs a set of decision parameters, including a scheduling trigger flag, a path ranking factor, and a path trade-off factor. These factors indicate whether to execute a return schedule, the priority order of candidate paths, and the dynamic weight allocation between behavioral urgency and energy competition, respectively. In the fourth step, the trade-off between behavioral urgency and energy competition is calculated based on the decision parameter group, and a target path is selected from the candidate charging locations. The target path is the charging path that is determined to be the comprehensive optimal result after the trade-off.
2. The method for scheduling the return charging of a mobile charging robot according to claim 1, characterized in that: In the process of collecting the status information of multiple mobile charging robots during operation, the task urgency is calculated by the ratio of the remaining time of the task to the estimated time of the shortest path from the current position to the task target position, using the formula: task urgency = remaining time of the task / estimated time of the path; Battery performance indicators are calculated by combining the normalized mean of the historical charge and discharge cycles and the current voltage recovery rate. The normalization method is to divide each parameter by the historical optimal reference value. The combined formula is: ; The historical charging record generates a weighted value for charging behavior by counting the total number of charging times and the duration of each charging session within a set operating cycle. The calculation method is to multiply the duration of each charging session by the weight of the time period in which it occurs and then take the average. The formula is: , where tn is the duration of the nth charging, wn is the corresponding weight preset in the period of the nth charging, and N is the total number of charging times; The occupancy of candidate charging locations during the prediction period is estimated by introducing an exponentially weighted moving average model by setting a five-minute sliding window. The number of queues, task type density, and continuous occupancy time of the charging pile within the most recent fixed time window are used for prediction. The exponential weighting function is: , where X_t is the current observation value, which is obtained by calculating the Euclidean distance between the observation vector composed of the number of queues, task type density and continuous occupancy time of the charging pile and the preset standard vector, E_t and are respectively the tense value of the occupancy situation in the current corresponding prediction time period and the tense value of the occupancy situation in the previous corresponding prediction time period, and α is the preset weight attenuation coefficient.
3. The method for scheduling the return charging of a mobile charging robot according to claim 2, characterized in that: Task type density is the ratio of the number of high-priority tasks assigned to the target area to the total number of tasks per unit time, using the formula: Task type density = number of high-priority tasks / total number of tasks; The continuous occupancy time of the charging pile is the cumulative value of the period of time when the continuous connection state exceeds the set connection threshold, and the total duration of all connection segments is counted; After calculation, the battery performance index is corrected by the temperature factor and the overload flag. The temperature factor is the proportional difference between the current ambient temperature and the standard ambient temperature. The correction method is: ; ; T is the current temperature, T0 is the standard temperature, and the temperature attenuation coefficient is the empirical setting value; The overload flag is a logic value. If the current peak detected in the last three charge and discharge cycles exceeds 120% of the rated value, it is marked as 1, otherwise it is 0, and then the multiplication factor is used: , and lower the revised values of battery performance indicators.
4. The method for scheduling the return charging of a mobile charging robot according to claim 3, characterized in that: The scheduling evaluation trigger condition is that the remaining power is insufficient to support the completion of the remaining path length of the task, and there is at least one backup charging location in the candidate path that is reachable within the predicted time period, or the estimated queuing time of all candidate charging locations in the predicted time period exceeds the maximum waiting time limit. If there is no backup charging location in the candidate path that is reachable within the predicted time period, the active power supply operation is performed.
5. The method for scheduling the return charging of a mobile charging robot according to claim 4, characterized in that: In the initial assessment, the behavioral urgency index is calculated based on the estimated path time between the current location and the task target location, the remaining power, the battery performance index, and the remaining time of the task. The nonlinear safety margin model is used to construct the index value. The formula is: , where R is the ratio of the remaining task time to the path estimation time, C is the normalized value of the battery performance index after the current correction and down-adjustment operation, and E is the ratio of the difference between the remaining power and the path estimation power consumption to the path estimation power consumption.
6. The method for scheduling the return charging of a mobile charging robot according to claim 5, characterized in that: In the initial evaluation, the energy competition index is modeled based on the current usage status of the candidate charging locations, occupancy during the forecast period, task type density, historical charging records, and path navigation costs. The local congestion coefficient is used to construct the model. The calculation formula for the energy competition index is: ,in: D is the normalized result of the navigation cost, calculated as the ratio of the actual path cost to the benchmark shortest path cost. The benchmark shortest path cost is the path with the smallest cost among all candidate charging locations from the current location. Q is the number of high-priority tasks whose predicted target is the current candidate charging location per unit time multiplied by the mapping weight corresponding to their task type density. The number of high-priority tasks is the number of tasks within the prediction time period whose target is the charging location and whose task urgency exceeds the preset threshold. Z is the weighted value of the charging behavior. A is the expected availability value of the charging location within the prediction time period, defined as the normalized tension value E_t.
7. The method for scheduling the return charging of a mobile charging robot according to claim 6, characterized in that: The behavioral urgency index and energy competition index are substituted as input into a learning model trained based on historical operating data. The learning model is constructed based on labeled sample data. The two-dimensional input vector consisting of the behavioral urgency index and the energy competition index is mapped to the scheduling execution result space. The prediction calculation is completed through a multi-layer neural network structure. The decision parameter group output by the model includes a scheduling trigger flag, a path ranking factor, and a path trade-off factor, where: The scheduling trigger flag is output through the sigmoid activation function, indicating whether return scheduling is executed in the current state. When the output result is greater than 0.5, it is determined to be immediately scheduled; The path ranking factor is the normalized value obtained by ranking and scoring all candidate paths based on the feature set consisting of historical task completion rate, path success rate, and failure readjustment rate. It is used to determine the priority order of candidate paths. The path trade-off factor is an adjustment coefficient used to dynamically allocate between behavioral urgency and energy competition, and is calculated as the interval [-1, 1] corresponding to the tanh activation output.
8. The method for scheduling the return charging of a mobile charging robot according to claim 7, characterized in that: In the process of calculating the trade-off between behavioral urgency and energy competition based on the decision parameter group, the path trade-off factor uses the ratio function of the behavioral urgency index and the energy competition index as the basic input and is dynamically adjusted in combination with the historical scheduling offset pattern. The calculation formula is: , where: B represents the behavior urgency index; E represents the energy competition index; H represents the historical scheduling offset factor, which is defined as the product of the path length deviation between the actual execution path and the optimal recommended path in the past three consecutive schedulings and the change in the scheduling success rate; σ is the preset adjustment sensitivity parameter; The path trade-off factor participates in the comprehensive optimization scoring of each path in the candidate charging location set. The comprehensive scoring result is used to screen the target path, which is the charging path with the highest score and the current remaining power to support the execution of its complete navigation path. Comprehensive optimization score = path tradeoff factor × remaining power redundancy ratio × equivalent waiting time adjustment factor, where the remaining power redundancy ratio is the ratio of the current remaining power to the estimated power required for the target path. If this ratio is less than 1, the comprehensive score of the path is forced to zero. The equivalent waiting time adjustment factor is the ratio of the predicted waiting time to the standard maximum acceptable waiting time, calculated as: , where DD is the predicted waiting time of the target location on the current path, YZ is the set maximum waiting tolerance threshold, and the predicted waiting time is set as the preset soft lock time limit of the target charging location. The soft lock time limit is defined as the average service processing time of the candidate charging location in the last several consecutive scheduling cycles, where the average service processing time is the arithmetic mean of the time it takes for the charging location to complete a complete charging task.
Citation Information
Patent Citations
Charging control method and device for transfer robot and electronic equipment
CN116455033A
Automatic charging management system for four-differential wheel type inspection robot
CN117748688A
Automatic charging method, device and equipment of auxiliary robot and storage medium
CN119539435A
Multi-robot collaborative scheduling system in automatic warehousing system
CN120255517A
Unmanned aerial vehicle charging network dynamic optimization system and method based on multi-agent cooperation
CN120450184A
Cited By
Battery replacement scheduling method and system for unmanned mine car
CN121032141A
A battery replacement scheduling method and system for unmanned mine cars
CN121032141B