An Intelligent Navigation Method and System for Mobile Charging Robots Based on a Distillation Strategy
By adopting an intelligent navigation method based on distillation strategy in the mobile charging robot navigation system, and using real-time monitoring and deep learning technology to build a dynamic node road network, the problem of difficulty in updating navigation strategies in dynamic and complex environments is solved, and more efficient and intelligent navigation path planning is achieved.
Patent Information
- Application Number
- CN202510403340.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-04-01
AI Technical Summary
The existing intelligent navigation methods of mobile charging robots cannot update their navigation strategies in a timely manner in dynamic and complex environments, resulting in inefficient navigation.
The intelligent navigation method of mobile charging robot based on distillation strategy is adopted. By monitoring the state of charging piles and robots in real time, nonlinear road network information is obtained, environmental features are extracted using convolutional neural networks, and dynamic node road network is built, and the teacher model is trained to generate the optimal path strategy, and the student model is trained to optimize the navigation strategy based on knowledge distillation.
It improves the navigation efficiency of mobile charging robots in dynamic and complex environments, enhances the adaptability and intelligence of path planning, optimizes the generalization ability of decision-making strategies, and improves the real-time response ability of the navigation system.
Smart Images

Figure CN119915298B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of robot navigation, and particularly relates to an intelligent navigation method and system for a mobile charging robot based on a distillation strategy. Background Art
[0002] A mobile charging robot is a device that combines robot technology and charging technology and can automatically provide charging services for devices such as electric vehicles. It usually consists of a chassis, a battery pack, a charging pile, a charging gun, machinery, sensors, a control system, a charging port sensing system, a voice interaction system, etc., and mostly appears in the form of a "vehicle" in appearance, similar to a movable large "portable power bank".
[0003] A mobile charging robot usually needs to have an autonomous navigation function, be able to move freely in an indoor or outdoor environment, and find the location of the device to be charged. Once the target device is found, it can automatically connect to the charging interface and charge the device without manual intervention.
[0004] Existing intelligent navigation algorithms can implement path recommendations in linear or non-linear scenarios, but the dynamic scenario adaptability of these path recommendations is mostly based on information updates on the Internet. For dynamic complex environments (such as weak network signals in multi-layer underground parking lots, partial vehicle blockages on roads, inability to reach the destination, movement of the location of the device to be charged, etc.), the navigation strategy cannot be updated in a timely manner. At present, the intelligent navigation method for mobile charging robots in large-scale weak network signal non-linear dynamic scenarios is still in a blank stage. Summary of the Invention
[0005] In view of the problems such as update delay caused by weak network signals, complication of the graph topology structure caused by non-linear scenarios, and difficulty in real-time updating of the navigation strategy caused by dynamic scenarios, the present invention proposes an intelligent navigation method and system for a mobile charging robot based on a distillation strategy to improve the navigation efficiency of the mobile charging robot in dynamic complex environments.
[0006] The present invention realizes the above object through the following technical solutions:
[0007] An intelligent navigation method for a mobile charging robot based on a distillation strategy, the method comprising:
[0008] Real-time monitoring of the usage status of charging piles and the usage status of the mobile charging robot in the target area, and obtaining non-linear road network information data in the target area, the non-linear road network information data at least including geographical coordinates, altitude, and network signal strength, and the network signal strength being described by Wi-Fi fingerprints;
[0009] Process the non - linear road network information data, convert it into a linear structure, and use a convolutional neural network to extract environmental features from the processed linear structured data. Construct a dynamic node road network by combining the usage status of charging piles and the usage status of mobile charging robots in the target area;
[0010] Use the dynamic node road network as the input, and use the DQN algorithm to train the teacher model for path planning to generate the optimal path strategy in a dynamic environment;
[0011] Judge whether a knowledge distillation process is required. If so, train the student model based on the distillation strategy, output the training result of the teacher model as a soft label, optimize the decision probability of the soft label according to the usage status of charging piles and mobile charging robots, and re - rank it to obtain the optimal navigation strategy in the current state;
[0012] Determine the real - time positions of the mobile charging robot and the target charging device through Wi - Fi fingerprint technology, and visually display the navigation path planned based on the teacher model or the student model.
[0013] A further improvement of the present invention is that: the non - linear road network information data is a non - linear data set containing static environment information and dynamic environment information. The static environment information is directly obtained through an open - source database or collected on - site by a handheld geographic information collection device, and the dynamic environment information is collected in real - time through visual sensors and lidar; present the obtained static and dynamic environment information in the form of ordered geographic information points;
[0014] The usage status of the charging pile at least includes four types: in use, not in use, damaged, and the parking position is occupied although not in use;
[0015] The usage status of the mobile charging robot at least includes six types: charging, on the way, in use, not in use, damaged, and insufficient power to wait for charging.
[0016] A further improvement of the present invention is that: the method for processing the non - linear road network information data and converting it into a linear structure includes:
[0017] Path linear decomposition: Decompose the positions where the non - linear road is disconnected due to altitude into linear path segments;
[0018] Parameter optimization: For the problem of parameter inflation caused by environmental complexity, adjust the path weights and node attributes to suppress the computational complexity;
[0019] Data conversion: Concatenate and convert the original non - linear data into linear structured data to adapt to the training requirements of the reinforcement learning model.
[0020] A further improvement of the present invention is that: the method for constructing a dynamic node road network includes:
[0021] Static framework: Based on geographical coordinates and altitude, a basic node road network is constructed, and the network signal strength is used as an additional attribute of the nodes.
[0022] Dynamic update: Combining the multi-modal data collected by sensors in real time and the usage status of charging piles and mobile charging robots monitored in real time, the attributes of nodes and edges are dynamically adjusted.
[0023] A further improvement of the present invention lies in that when training the teacher model using the DQN algorithm for path planning, the following environment modeling steps are included:
[0024] Define the state space S to represent the position, moving speed of the mobile charging robot in the environment, and the distance from obstacles, and the expression is:
[0025] ;
[0026] In the formula, is t the state of the mobile charging robot at time is t the position coordinates of the mobile charging robot at time is t the moving speed of the mobile charging robot at time is t the distance between the mobile charging robot and the obstacle at time
[0027] Define the action space A to represent the set of actions that the mobile charging robot can execute, and the expression is:
[0028] ;
[0029] In the formula, is t the action executed by the mobile charging robot at time is the set of adjacent nodes of the current node ;
[0030] Define the state transition probability , indicating that after taking the action t at time , the probability that the system transfers from the state to ; The state transition probability is determined by the road network topology and the power consumption model, and the power consumption model is expressed as:
[0031] ;
[0032] In the formula, is tBattery power at a moment is t the battery power at the +1 moment; represents the power consumption per unit distance, represents the current node selected action and the length to be traveled after that;
[0033] Define the reward function , and the expression is:
[0034] ;
[0035] In the formula, is t the reward obtained at the moment; represents t the shortest distance for the mobile charging robot to move to the target charging device position at the moment, represents t the shortest distance from the new position reached by the mobile charging robot after performing the action at the +1 moment to the target charging device position; represents the safe power threshold.
[0036] A further improvement of the present invention lies in that: the training steps of the teacher model include:
[0037] Define the Q value to represent the expected future cumulative reward after performing the action in the state s, and evaluate the value of each state-action pair through the Q value function. The Q value update formula is:
[0038] ;
[0039] In the formula, represents t the Q value of performing the action in the state at the moment; is the learning rate; is the reward given according to the path selection; is the discount factor;
[0040] Adopt the neural network parameter to approximate the optimal Q value function, and the formula is:
[0041] ;
[0042] In the formula, represents the Q value output by the neural network; is the optimal Q value;
[0043] The neural network structure is represented as:
[0044] ;
[0045] In the formula, , , are the weight matrices of the neural network; , , are the bias terms of the neural network; is the activation function;
[0046] The model training adopts the experience replay method, and the sample is stored and transferred to the replay buffer where is the navigation end flag. If it is 1, it means it is a termination state; if it is 0, it means it is not a termination state;
[0047] The mean square error is used to describe the loss function in the model training :
[0048] ;
[0049] The calculated value of the target value is:
[0050] ;
[0051] In the formula, represents the expected value of the loss function obtained when the relevant parameters of the state space are put back into the replay buffer ;
[0052] The policy network is updated using the gradient descent method:
[0053] ;
[0054] In the formula, is the learning rate of the neural network, is the loss function with respect to the neural network parameter gradient;
[0055] The ε-greedy strategy is used for policy exploration and probability decay:
[0056] ;
[0057] After the model training is completed, the greedy strategy is used for model testing. The initialization of the state space is used to reset and initialize the definition of the test environment, and then the cyclic decision-making of the navigation behavior is carried out:
[0058] ;
[0059] In the formula, represents the initial state; represents the position of the mobile charging robot at the start of the navigation action; represents the randomly given initial power in the test environment; represents the shortest moving distance before formal navigation; represents the optimal Q-value function, is the optimal neural network parameter value;
[0060] During each decision-making, DQN calculates the Q-value of each action according to the current state, selects the action with the maximum Q-value using the greedy strategy, and completes the path optimization by continuously updating the Q-value function;
[0061] The trained teacher model can select the optimal path for navigation given the starting point and the ending point.
[0062] A further improvement of the present invention lies in that: the evaluation indexes of the teacher model include:
[0063] Path length , and the formula is:
[0064] ;
[0065] Travel time , and the formula is:
[0066] ;
[0067] Smoothness , and the formula is:
[0068] ;
[0069] Node error rate , and the formula is:
[0070] ;
[0071] In the formula, represents the distance between nodes, represents the th node coordinate or position on the path, is the node number, m is the number of nodes; AverageSpeed represents the average moving speed of the mobile charging robot on the path; represents the supplementary angle of the steering angle of node ; std represents the standard deviation.
[0072] A further improvement of the present invention lies in that: the method for judging whether the knowledge distillation process needs to be carried out includes:
[0073] Obtain the current position of the mobile charging robot as the starting point and the position of the target charging device as the ending point. If both the starting point and the ending point are vertices in the dynamic node road network, directly call the teacher model for path recommendation, solve the model through the ε-greedy strategy, and output the path navigation result;
[0074] If at least one of the starting point and the ending point is not a vertex in the dynamic node road network, or the dynamic node road network changes during navigation, then perform knowledge distillation and use the distilled student model for navigation path planning.
[0075] A further improvement of the present invention lies in: training the student model based on the distillation strategy, and the method includes:
[0076] Transfer the path strategy of the teacher model to the student model using the distillation strategy. If at least one of the starting point and the ending point is not a vertex in the dynamic node road network, the graph topology structure is updated during the training of the student model, and the update method is: add new nodes, and update the attributes and weights of the corresponding edges, and use the ε-greedy strategy to solve the shortest path;
[0077] If the dynamic node road network changes during navigation, the student model performs real-time update on the graph topology structure, specifically including:
[0078] When a node disappears during navigation, if the disappearing node is not the ending point, delete the node and the edges adjacent to the node on the graph topology, re-plan the navigation, and solve the shortest path; if the disappearing node is the ending point, re-set the ending point and plan the navigation after updating the graph topology according to the charging requirements of the target charging device, and solve the shortest path;
[0079] When a road is impassable during navigation, the student model updates the graph topology structure, and the update method is: delete the impassable road on the graph topology. If the deletion causes the ending point to be unreachable, re-set the ending point and plan the navigation, and solve the shortest path;
[0080] During the training of the student model, it is necessary to align the state encodings in the model, and introduce a graph encoder to achieve the soft label and feature vector learning of the student model and the teacher model:
[0081] Teacher encoder It is expressed as: ;
[0082] Student encoder It is expressed as: ;
[0083] In the formula, It is the topological graph in the teacher model, and it is the time-varying topological graph of the student model; It represents the neural network parameters after the teacher model is processed by the encoder, and these parameters need to be retrained in the student model;
[0084] During model training, the soft labels are classified into and , where represents the abnormal action distribution, and
[0085] represents the normal action distribution; The formula for the KL divergence loss is:
[0086] ;
[0087] ;
[0088] The formula for the cross-entropy loss is:
[0089] ;
[0090] In the formula, E represents the expectation; , are the action probability distributions of the teacher model and the student model in state s respectively; , are the policy functions of the teacher model and the student model respectively, , are the distribution descriptions of the teacher model and the student model for the action respectively; represents the probability that the teacher model selects the action in state s, represents the probability that the student model selects the action in state s, and the parameter is .
[0091] An intelligent navigation system for a mobile charging robot based on a distillation strategy is applied to an intelligent navigation method for a mobile charging robot based on a distillation strategy as described above. The system includes:
[0092] A data acquisition module, which is used to monitor the usage status of charging piles and the usage status of mobile charging robots in the target area in real time, and obtain the non-linear road network information data in the target area;
[0093] The node road network construction module is used to process the non-linear road network information data, convert it into a linear structure, extract environmental features from the processed linear structured data using a convolutional neural network, and construct a dynamic node road network in combination with the usage status of charging piles and mobile charging robots in the target area;
[0094] The teacher model training module is used to take the dynamic node road network as input, train the teacher model for path planning using the DQN algorithm, and generate the optimal path strategy in a dynamic environment;
[0095] The student model training module is used to train the student model based on the distillation strategy, output the soft label of the training result of the teacher model, optimize the decision probability of the soft label according to the usage status of charging piles and mobile charging robots and reorder it, and obtain the optimal navigation strategy in the current state;
[0096] The positioning and navigation module is used to determine the real-time positions of the mobile charging robot and the target charging device through Wi-Fi fingerprint technology, and visually display the navigation path planned based on the teacher model or the student model.
[0097] By improving the intelligent navigation method for traditional linear static scenarios, the navigation problem of mobile charging robots in weak network signals and non-linear dynamic scenarios is solved. Combining reinforcement learning with the knowledge distillation strategy and integrating Wi-Fi fingerprint positioning technology can improve the adaptability and intelligence of the navigation path, enhance the path planning efficiency and accuracy, optimize the generalization ability of the decision-making strategy, and improve the real-time response ability of the navigation system. BRIEF DESCRIPTION OF THE DRAWINGS
[0098] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. Among them:
[0099] Figure 1 is the method flow chart of the present invention;
[0100] Figure 2 is the modular structure diagram of the system of the present invention;
[0101] Figure 3 is the distribution map of parking spaces and charging piles in the embodiment of the present invention;
[0102] Figure 4 is the linearized point-line charging pile distribution in the parking lot in the embodiment of the present invention Figure 1 ;
[0103] Figure 5 The point-line charging pile distribution for the linearization of the parking lot in the embodiment of the present invention Figure 2 ;
[0104] Figure 6 The point-line charging pile distribution for the linearization of the parking lot in the embodiment of the present invention Figure 3 ;
[0105] Figure 7 Schematic diagram of the comparison of the decision results of the teacher model with the A* algorithm and the Dijkstra algorithm in the static scenario in the embodiment of the present invention;
[0106] Figure 8 Schematic diagram of the decision result of the teacher model in the dynamic scenario in the embodiment of the present invention;
[0107] Figure 9 Schematic diagram of the decision result of the student model in the dynamic scenario in the embodiment of the present invention;
[0108] Figure 10 Schematic diagram of the probability distribution of the user's charging waiting time under the teacher model in the embodiment of the present invention;
[0109] Figure 11 Schematic diagram of the probability distribution of the user's charging waiting time under the student model in the embodiment of the present invention. Detailed implementation manners
[0110] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present invention fall within the scope of protection of the present invention.
[0111] As Figure 1 shown, it is an embodiment of the present invention. This embodiment provides an intelligent navigation method for a mobile charging robot based on a distillation strategy, and the main application scenarios include weak signal areas (scenarios where GPS signals are missing or network signals are unstable, such as underground parking lots, tunnels, etc.), dynamic obstacles (temporarily parked vehicles, moving pedestrians, construction facilities, etc.), multi-layer structures (three-dimensional navigation requirements that need to consider floor height differences), etc.
[0112] The method mainly includes the following steps:
[0113] S1: Real-time monitor the usage status of the charging piles and the usage status of the mobile charging robot in the target area, and obtain the non-linear road network information data in the target area, including at least geographical coordinates (represented by longitude and latitude), altitude, network signal strength, etc., where the network signal strength is described by Wi-Fi fingerprints.
[0114] Step S1 is mainly used to collect data, and the target area is the working environment that the mobile charging robot needs to navigate. The usage status of the charging pile means information such as whether the charging pile in the monitoring area is available, whether it is charging, and the battery level, which at least includes four states: in use, unused, damaged, and unused but the parking space is occupied. The usage status of the mobile charging robot is mainly used to monitor the current charging situation, location, etc. of the robot, which at least includes six states: charging, on the way, in use, unused, damaged, and low battery waiting to be charged.
[0115] The non-linear road network information data is a non-linear data set containing static environment information (road nodes, fixed obstacles, etc.) and dynamic environment information (temporary obstacles, moving targets, road state changes, real-time network signal fluctuations, etc.), and the obtained environment information is presented in the form of ordered geographical information points.
[0116] The static environment information is directly obtained through open source databases (open street map, UJIIndoorLoc, etc.) or collected on-site with handheld geographic information (GIS) collection devices. For example, download basic geographic data including road layout, floor distribution, etc. from the open street map platform, or select the UJIIndoorLoc data set designed for indoor weak network signal positioning problems. In the case of no available data set, a handheld geographic information data collector can also be used for data collection to obtain data through on-site road information collection.
[0117] In this embodiment, the UJIIndoorLoc data set is preferably used. This data set contains the intensity of the network signal and the position coordinates of the user, longitude, latitude, and altitude, and includes multiple test areas where the network signal intensity and the corresponding coordinate positions are measured and marked. These data provide a reliable verification basis for the training of convolutional neural networks and deep reinforcement learning models and path planning algorithms.
[0118] The dynamic environment information is collected in real time through visual sensors, lidar, etc. Further, sensors such as visual sensors and lidar are directly embedded in the body of the mobile charging robot to support the low latency and high precision characteristics of real-time environment perception and path planning.
[0119] Wi-Fi Fingerprint is an indoor positioning technology based on Wi-Fi signal characteristics. In the offline phase, Wi-Fi signal strength information (such as RSSI) at different positions in the environment is pre-collected and stored as "fingerprint" data to establish a fingerprint database; in the online phase, the device real-time collects the current Wi-Fi signal characteristics and matches them with the fingerprints in the fingerprint database to determine its own position. It has the characteristics of low cost and strong applicability, but the fingerprint database needs to be updated regularly to cope with environmental changes. It is difficult to use GPS positioning in weak network signal scenarios. Wi-Fi fingerprint is an important solution for spatial positioning in weak network scenarios. Therefore, the present invention uses Wi-Fi fingerprint to describe the network signal strength in the scenario.
[0120] S2: Process the non-linear road network information data, convert it into a linear structure, and use a convolutional neural network to extract environmental features (such as longitude and latitude, altitude, network signal strength, fixed obstacle positions, and dynamic targets) from the processed linear structured data. Combine the usage status of charging piles and the usage status of mobile charging robots in the target area to construct a dynamic node road network (graph topology).
[0121] The methods for processing the non-linear road network information data and converting it into a linear structure include:
[0122] Path linearization decomposition: Decompose the positions (such as ramps and floor joints) where the non-linear road is disconnected due to altitude into linear path segments to ensure that the topological structure is resolvable;
[0123] Parameter optimization: In response to the problem of parameter inflation caused by environmental complexity (such as an increase in dynamic obstacles), adjust the path weights and node attributes to suppress the computational complexity;
[0124] Data conversion: Stitch and convert the original non-linear data (such as curved roads and three-dimensional structures) into linear structured data to adapt to the training requirements of the reinforcement learning model.
[0125] Use the environmental perception ability of the convolutional neural network to perform convolutional operations on the converted linear structured data, extract spatial features from the high-dimensional perception data of the dataset, and transform the environmental information into graph topology network data that the agent can process. The node attributes include geographical coordinates (longitude and latitude), altitude, signal strength, etc., and the edge attributes include road length, traffic status, etc.
[0126] The convolutional layer of the convolutional neural network is set as follows:
[0127] The first convolutional layer: Contains 32 convolutional kernels, with a size of 3*3, a stride of 1, and padding "same", extracting low-level features (such as road edges and obstacle contours);
[0128] Second Convolutional Layer: It contains 64 convolutional kernels with a size of 3*3, a stride of 1, and padding "same" to extract complex features (such as dynamic target positions and passable areas).
[0129] Third Convolutional Layer: It contains 128 convolutional kernels with a size of 3*3, a stride of 1, and padding "same" to extract advanced features (such as road connectivity and network signal distribution).
[0130] After each convolution operation, the max-pooling layer reduces the dimension of the feature map, thereby further reducing the computational complexity for the agent. Through the operations of the three convolutional layers, the convolutional neural network can effectively identify obstacles, dynamic targets, and other environmental information, providing real-time perception support for path planning.
[0131] The method for constructing a dynamic node road network includes:
[0132] Static Framework: Based on geographical coordinates and altitude, construct a basic node road network (such as road nodes and fixed obstacles), with network signal strength as an additional attribute of the nodes;
[0133] Dynamic Update: Combine multi-modal data collected in real-time by sensors (such as temporary obstacles, moving vehicles, and road availability data), as well as the usage status of charging piles and mobile charging robots monitored in real-time, to dynamically adjust the attributes of nodes and edges (such as deleting blocked roads and adding temporary nodes).
[0134] S3: Use the dynamic node road network as the input, and train a teacher model using the DQN (Deep Q-Network) algorithm for path planning to generate an optimal path strategy in a dynamic environment;
[0135] DQN optimizes the path selection of the mobile charging robot by continuously interacting with the environment, enabling it to find the optimal path in a dynamic environment and using this model as the teacher model.
[0136] When using the DQN algorithm to train a teacher model for path planning, the following environmental modeling steps are included:
[0137] Define the state space S to represent the position, moving speed, and distance from obstacles of the mobile charging robot in the environment. The expression is:
[0138] ;
[0139] In the formula, is t the state of the mobile charging robot at time is t the position coordinates of the mobile charging robot at time is tThe moving speed of the mobile charging robot at a moment, is t the distance between the mobile charging robot and the obstacle at a moment;
[0140] Define the action space A to represent the set of actions that the mobile charging robot can execute. The expression is:
[0141] ;
[0142] In the formula, is t the action executed by the mobile charging robot at a moment; is the set of adjacent nodes of the current node ;
[0143] Define the state transition probability , which represents that after taking the action t at the moment , the probability that the system transfers from the state to ; The state transition probability is determined by the road network topology and the power consumption model. The power consumption model is expressed as:
[0144] ;
[0145] In the formula, is t the battery power at the moment is t the battery power at the t +1 moment. The action to be taken is determined by the battery power at the +1 moment represents the power consumption per unit distance, represents the current node selected action the length to be traveled;
[0146] Define the reward function , and the expression is:
[0147] ;
[0148] In the formula, represents t the shortest distance from the mobile charging robot to the target charging device position at a moment, represents t when the mobile charging robot reaches a new position after executing the action at the +1 moment, the shortest distance to the target charging device position; that is, when the robot shortens the distance to the target charging device through the action ( ), the reward increases, otherwise the reward decreases; It represents the safe power threshold. When the battery power is lower than the safe power threshold, an alarm will occur or the robot will be unable to move. When the battery power in the next moment is lower than the safe power threshold, a penalty of -5 will be imposed to ensure that the robot maintains a safe power level.
[0149] In this embodiment, the training steps of the teacher model include:
[0150] Define the Q value to represent the expected future cumulative reward after executing the action (s, is a general variable representing any one of all possible states of the mobile charging robot and all executable action sets), and evaluate the value of each state-action pair through the Q value function. The Q value update formula is:
[0151] ;
[0152] In the formula, represents t the Q value of executing the action at time in the state ; is the learning rate; is the immediate reward value given according to the path selection; is the discount factor.
[0153] The parameter settings during the model training process are shown in Table 1:
[0154] Table 1 Parameters for training the teacher model by the DQN algorithm
[0155]
[0156] Adopt the neural network parameter to approximate the optimal Q value function, and the formula is:
[0157] ;
[0158] In the formula, represents the Q value output by the neural network; is the optimal Q value;
[0159] The neural network structure is expressed as:
[0160] ;
[0161] In the formula, , , are the weight matrices of the neural network; , , are the bias terms of the neural network; is the activation function;
[0162] The model training adopts the experience replay method, and the samples are stored and transferred to the replay buffer where is the navigation end flag. If it is 1, it means the termination state; if it is 0, it means it is not the termination state;
[0163] The mean squared error is used to describe the loss function in model training :
[0164] ;
[0165] The target value The calculated value is:
[0166] ;
[0167] In the formula, represents the expected value of the loss function obtained by putting the relevant parameters of the state space back into the replay buffer. Describing the loss function with the expected value can reduce the influence of the noise generated by random data during training;
[0168] The policy network is updated using the gradient descent method:
[0169] ;
[0170] In the formula, is the learning rate of the neural network, is the loss function The gradient of the neural network parameters ;
[0171] The ε-greedy strategy is used for policy exploration and probability decay:
[0172] ;
[0173] After the model training is completed, the greedy strategy is used for model testing. The initialization of the state space is used to reset and initialize the definition of the test environment, and then the loop decision of the navigation behavior is carried out:
[0174] ;
[0175] In the formula, represents the initial state; represents the position of the mobile charging robot at the start of the navigation action; represents the randomly given initial power under the test environment; Represents the shortest moving distance before formal navigation; Represents the optimal Q-value function, Is the optimal neural network parameter value;
[0176] At each decision-making moment, DQN calculates the Q-value of each action according to the current state, selects the action with the maximum Q-value using the greedy strategy, and completes path optimization by continuously updating the Q-value function;
[0177] After the training of the teacher model is completed, it can select the optimal path for navigation given the starting point and the ending point.
[0178] To improve the training stability, the DQN algorithm in this embodiment uses the experience replay mechanism. Each time it interacts with the environment, DQN will store each experience (i.e., the current state, the action taken, the reward obtained, the next state, etc.) into the experience pool. DQN randomly extracts a batch of experiences from the experience pool for training and updates the Q-value. This mechanism breaks the correlation between experiences, helps to avoid overfitting, and improves the learning efficiency.
[0179] In this embodiment, the evaluation indicators of the teacher model include:
[0180] Path length , the formula is:
[0181] ;
[0182] Travel time , the formula is:
[0183] ;
[0184] Smoothness , the formula is:
[0185] ;
[0186] Node error rate , the formula is:
[0187] ;
[0188] In the formula, Represents the distance between nodes, Represents the th node's coordinate or position on the path, Is the node number, m is the number of nodes; AverageSpeed represents the average moving speed of the mobile charging robot on the path; Represents the complementary angle of the steering angle of node; std represents the standard deviation.
[0189] S4: Determine whether a knowledge distillation process is required. If so, train the student model based on the distillation strategy, output the training results of the teacher model as soft labels, optimize the decision probability of the soft labels according to the usage status of the charging pile and the mobile charging robot, re - sort them, obtain the optimal navigation strategy in the current state, and perform visual display.
[0190] Specifically, the method for determining whether a knowledge distillation process is required includes:
[0191] Obtain the current position of the mobile charging robot as the starting point and the position of the target charging device as the ending point. If both the starting point and the ending point are vertices in the dynamic node road network (i.e., both the starting point and the ending point are preset nodes), directly call the teacher model for path recommendation, perform model solution through the ε - greedy strategy, and output the path navigation result;
[0192] If at least one of the starting point and the ending point is not a vertex in the dynamic node road network, or the dynamic node road network (such as node disappearance, road congestion) changes during the navigation process, then perform knowledge distillation and use the distilled student model for navigation path planning.
[0193] In one embodiment, the method for training the student model based on the distillation strategy includes:
[0194] Transfer the path strategy of the teacher model to the student model using the distillation strategy. If at least one of the starting point and the ending point is not a vertex in the dynamic node road network, update the graph topology during the training of the student model. The update method is: add new nodes, and update the attributes and weights of the corresponding edges, and perform the shortest path solution using the ε - greedy strategy;
[0195] If the dynamic node road network changes during the navigation process, the student model performs real - time update of the graph topology, specifically including:
[0196] When a node disappears during the navigation process, if the disappeared node is not the ending point, delete the node and the edges adjacent to the node on the graph topology, re - plan the navigation, and perform the shortest path solution; if the disappeared node is the ending point, re - set the ending point according to the charging requirements of the target charging device after updating the graph topology and plan the navigation, and perform the shortest path solution;
[0197] When a road is impassable during the navigation process, the student model updates the graph topology. The update method is: delete the impassable road on the graph topology. If the deletion causes the ending point to be unreachable, re - set the ending point and plan the navigation, and perform the shortest path solution.
[0198] During the training of the student model, it is necessary to align the state encoding in the model. By introducing a graph encoder Implement soft label and feature vector learning for the student model and the teacher model:
[0199] Teacher encoder Denoted as: ;
[0200] Student encoder Denoted as: ;
[0201] Wherein, 、 are the states of the teacher model and the student model respectively; is the topological graph in the teacher model, is the time-varying topological graph of the student model; represents the neural network parameters of the teacher model after being processed by the encoder. These parameters need to be retrained in the student model, but the training process is more concise than that of the teacher model;
[0202] Classify the soft labels in model training into and , where represents the abnormal action distribution, represents the normal action distribution;
[0203] Evaluate the distillation loss through KL divergence and cross-entropy evaluation. The formula for the KL divergence loss is:
[0204] ;
[0205] ;
[0206] The formula for the cross-entropy loss is:
[0207] ;
[0208] Wherein, E represents the expectation; 、 are the action probability distributions of the teacher model and the student model in state s respectively; 、 are the policy functions of the teacher model and the student model respectively, 、 are the distribution descriptions of the teacher model and the student model for the action respectively; represents the probability that the teacher model selects the action in state s, represents the probability that the student model selects the action in state s.
[0209] S5: Determine the real-time positions of the mobile charging robot and the target charging device through Wi-Fi fingerprint technology, and visually display the navigation path planned based on the teacher model or the student model.
[0210] As Figure 2 shown, this is another embodiment of the present invention. This embodiment provides an intelligent navigation system for a mobile charging robot based on a distillation strategy, which is applied to the intelligent navigation method for a mobile charging robot based on a distillation strategy as described above, and includes:
[0211] A data acquisition module, which is used to monitor the usage status of charging piles and the usage status of the mobile charging robot in real time in the target area, and obtain the non-linear road network information data in the target area;
[0212] A node road network construction module, which is used to process the non-linear road network information data, convert it into a linear structure, and use a convolutional neural network to extract environmental features from the processed linearly structured data, and construct a dynamic node road network in combination with the usage status of charging piles and the usage status of the mobile charging robot in the target area;
[0213] A teacher model training module, which is used to take the dynamic node road network as the input, use the DQN algorithm to train the teacher model for path planning, and generate the optimal path strategy in a dynamic environment;
[0214] A student model training module, which is used to train the student model based on the distillation strategy, perform soft label output on the training results of the teacher model, optimize the decision probability of the soft label according to the usage status of the charging pile and the mobile charging robot, and reorder it to obtain the optimal navigation strategy in the current state;
[0215] A positioning and navigation module, which is used to determine the real-time positions of the mobile charging robot and the target charging device through Wi-Fi fingerprint technology, and visually display the navigation path planned based on the teacher model or the student model.
[0216] Next, taking an underground parking lot as an example, the working principle of the system in this embodiment is further described.
[0217] The first step: Collect data - draw a "parking lot map"
[0218] Use devices such as cameras and lidar to scan the underground parking lot, record information such as road positions, obstacles, floor heights, and network signal strengths, and generate a.csv format data file containing longitude and latitude, altitude, network signals, etc. through sensor scanning, so as to let the robot understand where there are roads, where there is traffic congestion, and where the signal is poor.
[0219] For positioning in a weak network, Wi-Fi fingerprints are used instead of GPS to determine the robot's position. The Wi-Fi signal distributions at different locations are pre-recorded, and the robot compares the current signals in real time to match the preset positions, achieving precise positioning.
[0220] As Figure 3 shown, this underground parking lot has a total of 4 floors, with transition platforms between each floor. There are also parking spaces and charging positions on the platforms for electric vehicles to park, that is, there are a total of 8 layers of map data. The 8 layers of map data overlap in longitude and latitude, forming a relatively complex non-linear road network structure.
[0221] Step 2: Process the data - turn the complex roads into a "point-line graph"
[0222] Use a convolutional neural network (CNN) to automatically identify road features (such as intersections and obstacles), generate a simplified map composed of points and lines, and convert the scanned complex roads into a "node road network graph" that the computer can understand, thereby simplifying the environmental information and facilitating the robot to plan paths.
[0223] As Figures 4 - 6 shown, the 8 layers of map data are disassembled according to different altitudes and combined according to the platform layers. The complex non-linear road network data is split into 3 different planar data, that is, the road network structure is converted into a point-line type linear road network data that is easier for the agent to understand (0 indicates idle, 1 indicates damaged, 2 indicates in use or occupied). Figure 4 、 Figure 6 The parking space distribution in Figure 4 is in the shape of a "king" as the main floor, with a relatively large number of parking spaces and a more uniform distribution. However, Figure 5 the floor represented by Figures 4 - 6 has a relatively large number of damaged charging piles.
[0224] Step 3: Train the "teacher model" - teach the robot to find the optimal path
[0225] Train a "teacher model" using deep reinforcement learning (DQN) to let the robot learn how to avoid obstacles and choose the shortest path through trial and error. The robot continuously tries different routes, adjusts its strategy according to the arrival time and path length, and finally learns to navigate efficiently. The teacher model is complex but intelligent and can handle various complex situations, but it requires a large amount of computing resources.
[0226] As Figure 7As shown, when there is sufficient computing resources, the path decisions given by the teacher model trained using CNN + DQN are relatively smooth, that is, it can avoid taking unnecessary detours. In contrast, the path decisions given by traditional models (A* algorithm and Dijkstra algorithm) are less smooth, with situations of spinning in place and turning around.
[0227] Step 4: Train the "student model" - lightweight navigation assistant
[0228] Compress the knowledge of the teacher model into a lighter student model to adapt to the dynamically changing environment. The student model is faster and more flexible, suitable for real-time route adjustment.
[0229] For example: when the starting point or ending point is not at a preset node, the student model can automatically add new nodes on the map and re-plan the route (such as the sudden appearance of a temporary parking space). When a road is blocked or a node disappears, the student model can delete the nodes of the blocked road and find a new route (such as bypassing illegally parked vehicles). When the vehicle to be charged moves, the student model can update the ending position in a timely manner and recalculate the route (following the moving target vehicle).
[0230] As Figure 8 、 Figure 9 shown, in a dynamic scenario where a linear road network is reorganized into a non-linear road network, directly using the teacher model for training will result in parameter inflation on the model side due to the dynamic scenario, leading to online delays on the user side. The dynamic changes of charging piles, mobile charging robots, and road conditions cannot be updated in a timely manner. Although the target position is finally reached successfully, the smoothness of the path is poor and the efficiency is low. While Figure 9 the navigation path provided by the student model in [reference] has better smoothness and higher efficiency.
[0231] As Figure 10 、 Figure 11 shown, in [reference], the waiting time distributions of the teacher model and the student model from when the user issues a charging request to when the charging is successfully completed are compared. The waiting time of the user under the teacher model is significantly longer than that under the student model.
[0232] Compare the reinforcement learning model of the present invention with the A* algorithm and the Dijkstra algorithm respectively, and evaluate the path planning effects of all algorithms for mobile charging robots in the same test area. In a non-linear scenario, the A* algorithm and the Dijkstra algorithm did not find an effective path within the specified time. The path planning in the non-linear scenario was successfully carried out using the knowledge distillation method and the path length was shorter, showing higher path optimization ability. Table 2 shows the evaluation results of the three algorithms.
[0233] Table 2 Comparison experiment results
[0234]
[0235] During the model training process, the distillation strategy of reinforcement learning utilizes the Wi-Fi Fingerprint of the scenario, thus showing good performance in multiple weak signal areas. The test results show that once dynamic changes occur in the scenario, the A* algorithm and Dijkstra algorithm cannot make adaptive adjustments, resulting in the mobile charging robot being stranded in place due to blocked roads, moving or disappearing target nodes. Table 3 shows the test results in multiple weak signal areas.
[0236] Table 3 Test Results in Indoor Weak Network Signal Areas
[0237]
[0238] The experimental results prove that the intelligent navigation model based on the distillation strategy of reinforcement learning proposed in the present invention improves the path planning and obstacle avoidance capabilities of the mobile charging robot to reach the vehicle to be navigated in dynamic and complex environments. Compared with the traditional A* algorithm and Dijkstra algorithm, the reinforcement learning algorithm reduces the path length by 15.9% compared with the traditional algorithm and shows strong robustness in multiple weak signal areas. In addition, the method proposed in the present invention also shows better performance in terms of real-time performance and path smoothness, can effectively cope with dynamic changing environments and sudden obstacles, and effectively reduces the probability of accidents occurring during the movement of the mobile charging robot. The experimental results fully verify the adaptability and accuracy of the present invention in complex environments.
[0239] In summary, the present invention realizes more refined and dynamic navigation data acquisition by real-time monitoring the usage status of charging piles and mobile charging robots in the target area and integrating non-linear road network information (such as geographical coordinates, altitude, and network signal strength), ensuring that the path planning process fully adapts to complex and variable environmental factors.
[0240] By structuring the non-linear road network information and using a convolutional neural network for environmental feature extraction, the accuracy of path modeling is effectively enhanced. On this basis, the constructed dynamic node road network makes the path planning more in line with the actual dynamic scenario and improves the overall intelligence level of the navigation system.
[0241] Introduce a reinforcement learning path training mechanism based on DQN. Generate the optimal path strategy in a dynamic environment through the teacher model, and train the student model in combination with the distillation strategy, which significantly reduces the computational complexity while ensuring the model accuracy, and improves the generalization ability and deployment flexibility of the model in different scenarios.
[0242] Based on Wi-Fi fingerprint technology, high-precision positioning of robots and target devices is achieved, and the optimal path is generated and visually displayed in real time in combination with teacher or student models, effectively improving the response speed and path visibility of the system in practical applications and enhancing the user operation experience.
[0243] Through multi-source heterogeneous data fusion, deep learning path planning, and knowledge distillation strategies, the present invention effectively improves the navigation intelligence and execution efficiency of mobile charging robots in complex environments, and has broad application prospects and industrial promotion value.
[0244] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product, which includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another.
[0245] In addition, in each embodiment of the present application, each functional unit can be integrated in a processing module, or each unit can exist physically alone, or two or more units can be integrated in a module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. When the above integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. The storage medium can be a read-only memory, a magnetic disk, an optical disk, etc.
[0246] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various changes or substitutions, and these should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A mobile charging robot intelligent navigation method based on distillation strategy, characterized in that: The method comprises: Monitor the use status of charging piles and mobile charging robots in the target area in real time, and obtain nonlinear road network information data in the target area, wherein the nonlinear road network information data at least includes geographic coordinates, altitude, and network signal strength, wherein the network signal strength is described by Wi-Fi fingerprint; The nonlinear road network information data is processed and converted into a linear structure, and the environmental features of the processed linear structured data are extracted using a convolutional neural network, and a dynamic node road network is constructed in combination with the usage status of the charging piles and the usage status of the mobile charging robot in the target area; Taking the dynamic node road network as input, using the DQN algorithm to train the teacher model for path planning, and generating the optimal path strategy in the dynamic environment; Determine whether the knowledge distillation process is needed. If necessary, train the student model based on the distillation strategy, output the training results of the teacher model with soft labels, optimize the decision probability of the soft labels according to the usage status of the charging pile and the mobile charging robot, and re-rank them to obtain the optimal navigation strategy under the current state; The real-time location of the mobile charging robot and the target charging device is determined through Wi-Fi fingerprint technology, and the navigation path planned by the teacher model or student model is visualized.
2. The intelligent navigation method of a mobile charging robot based on a distillation strategy as claimed in claim 1, characterized in that: The nonlinear road network information data is a nonlinear data set including static environmental information and dynamic environmental information. The static environmental information is directly acquired through an open source database or collected on-site by a handheld geographic information acquisition device, and the dynamic environmental information is collected in real time by a visual sensor and a laser radar. The acquired static and dynamic environmental information is presented in the form of ordered geographic information points. The use status of the charging pile includes at least four types: in use, unused, damaged, and unused but the parking space is occupied; The use status of the mobile charging robot includes at least six types: charging, on the way, in use, not in use, damaged, and low on power and waiting to be charged.
3. The intelligent navigation method of a mobile charging robot based on a distillation strategy as claimed in claim 1, characterized in that: The method for processing the nonlinear road network information data and converting it into a linear structure includes: Path linearization: decompose the locations in nonlinear roads that are disconnected by altitude into linear path segments; Parameter optimization: To address the parameter expansion problem caused by environmental complexity, we adjust path weights and node attributes to reduce computational complexity. Data conversion: Convert the original nonlinear data into linear structured data to meet the training requirements of reinforcement learning models.
4. The mobile charging robot intelligent navigation method based on distillation strategy as claimed in claim 1, characterized in that: Methods for constructing a dynamic node road network include: Static framework: Based on geographic coordinates and altitude, a basic node network is constructed, and network signal strength is used as an additional attribute of the node; Dynamic update: Combine the multimodal data collected by sensors in real time with the usage status of charging piles and mobile charging robots monitored in real time to dynamically adjust the properties of nodes and edges.
5. The mobile charging robot intelligent navigation method based on distillation strategy as claimed in claim 1, characterized in that: When the DQN algorithm is used to train the teacher model for path planning, the following environment modeling steps are included: The state space S is defined to represent the position, moving speed and distance from obstacles of the mobile charging robot in the environment. The expression is: ; In the formula, for t The status of the mobile charging robot at all times; for t The position coordinates of the mobile charging robot at all times, for t The moving speed of the charging robot at all times, for t Keep the distance between the mobile charging robot and obstacles at all times; The action space A is defined to represent the set of actions that the mobile charging robot can perform. The expression is: ; In the formula, for t The actions performed by the mobile charging robot at all times; For the current node The set of adjacent nodes of Define state transition probability , indicating that t Always take action After that, the system changes from state Transfer to The state transition probability is determined by the road network topology and the power consumption model, and the power consumption model is expressed as: ; In the formula, for t The battery level at the moment, for t +1 moment of battery power; Indicates the power consumption per unit distance. Indicates the current node Select Action The length of the walk you need to take; Defining the reward function , the expression is: ; In the formula, for t Rewards received at all times; express t The shortest distance that the charging robot moves to the target charging device at any time. express t +1: The shortest distance to the target charging device when the mobile charging robot reaches the new position after executing the action at time 1; Indicates the safe power threshold.
6. The mobile charging robot intelligent navigation method based on distillation strategy as claimed in claim 5, characterized in that: The training steps of the teacher model include: Define Q value to represent the action taken in state s The future cumulative reward expectation is calculated by using the Q value function to evaluate the value of each state-action pair. The Q value update formula is: ; In the formula, express t Always in the state Next action Q value; is the learning rate; The reward given according to the path selection; is the discount factor; Using neural network parameters Approximate optimal Q value function, the formula is: ; In the formula, Represents the Q value of the neural network output; is the optimal Q value; Neural Network Architecture It is expressed as: ; In the formula, , , is the weight matrix of the neural network; , , is the bias term of the neural network; is the activation function; Model training uses the experience replay method to Storage transferred to playback buffer Among them Navigation end flag, if it is 1, it means it is in the end state, if it is 0, it means it is not in the end state; Mean square error is used to describe the loss function in model training : ; Target value The calculated value of is: ; In the formula, Indicates that the relevant parameters of the state space Put back into the playback buffer The expected value of the loss function obtained under the condition of Update the policy network using gradient descent: ; In the formula, is the neural network learning rate, is the loss function Neural network parameters The gradient of Use the ε-greedy strategy for strategy exploration and probability decay: ; After the model training is completed, the greedy strategy is used to test the model, and the initialization of the state space is used to reset and initialize the test environment, and then the navigation behavior is cyclically decided: ; In the formula, Indicates the initial state; Indicates the position of the mobile charging robot when the navigation action starts; Indicates the initial power randomly given in the test environment; Indicates the shortest moving distance before formal navigation; represents the optimal Q-value function, Take the value of the optimal neural network parameters; At each decision, DQN calculates the Q value of each action based on the current state, uses a greedy strategy to select the action with the largest Q value, and completes path optimization by continuously updating the Q value function; After training, the teacher model can select the optimal navigation path given the starting point and the end point.
7. The mobile charging robot intelligent navigation method based on distillation strategy as claimed in claim 6, characterized in that: The evaluation indicators of the teacher model include: Path length , the formula is: ; Travel time , the formula is: ; Smoothness , the formula is: ; Node error rate , the formula is: ; In the formula, represents the distance between nodes, Indicates the first The coordinates or positions of the nodes, is the node number, m is the number of nodes; AverageSpeed represents the average moving speed of the mobile charging robot on the path; Representation Node The complementary angle of the steering angle; std means standard deviation.
8. The mobile charging robot intelligent navigation method based on distillation strategy as claimed in claim 1, characterized in that: The method for determining whether a knowledge distillation process is required includes: The current position of the mobile charging robot is obtained as the starting point, and the position of the target charging device is obtained as the end point. If the starting point and the end point are both vertices in the dynamic node road network, the teacher model is directly called for path recommendation, and the model is solved through the ε-greedy strategy to output the path navigation result. If at least one of the starting point and the end point is not a vertex in the dynamic node network, or the dynamic node network changes during navigation, knowledge distillation is performed and the distilled student model is used for navigation path planning.
9. The mobile charging robot intelligent navigation method based on distillation strategy as claimed in claim 8, characterized in that: The method for training the student model based on the distillation strategy includes: The path strategy of the teacher model is transferred to the student model using the distillation strategy. If at least one of the starting point and the end point is not a vertex in the dynamic node network, the graph topology structure is updated during the student model training. The updating method is: adding new nodes, updating the attributes and weights of the corresponding edges, and using the ε-greedy strategy to solve the shortest path. If the dynamic node network changes during navigation, the student model will update the graph topology in real time, including: When a node disappears during navigation, if the disappeared node is not the end point, the node and the edges adjacent to the node are deleted in the graph topology, and the navigation is replanned to solve the shortest path; if the disappeared node is the end point, the end point is reset and the navigation is planned after updating the graph topology according to the charging needs of the target charging device, and the shortest path is solved; When a road becomes impassable during navigation, the student model updates the graph topology by deleting the impassable road on the graph topology. If the deletion causes the destination to be unreachable, the destination is reset and navigation is planned to solve the shortest path. When training the student model, it is necessary to align the state encoding in the model. By introducing the graph encoder Implement soft label and feature vector learning for student model and teacher model: Teacher Coder It is expressed as: ; Student Coder It is expressed as: ; In the formula, , are the states of the teacher model and the student model respectively; is the topological graph in the teacher model, is the time-varying topological graph of the student model; Represents the neural network parameters of the teacher model after being processed by the encoder. This parameter needs to be retrained in the student model; Classify the soft labels in model training into and ,in represents the abnormal action distribution, represents the normal action distribution; Distillation loss evaluation is performed through KL divergence and cross entropy evaluation, KL divergence loss The formula is: ; ; Cross Entropy Loss The formula is: ; In the formula, E represents expectation; , are the action probability distributions of the teacher model and the student model in state s respectively; , are the strategy functions of the teacher model and the student model respectively, , The teacher model and the student model are respectively Description of the distribution of Indicates that the teacher model selects an action in state s The probability of Indicates that the student model selects an action in state s The probability of , the parameter is .
10. A mobile charging robot intelligent navigation system based on a distillation strategy, applied to a mobile charging robot intelligent navigation method based on a distillation strategy as claimed in any one of claims 1 to 9, characterized in that: The system comprises: A data acquisition module is used to monitor the usage status of charging piles and mobile charging robots in the target area in real time, and obtain nonlinear road network information data in the target area; A node road network construction module is used to process the nonlinear road network information data, convert it into a linear structure, and use a convolutional neural network to extract environmental features from the processed linear structured data, and build a dynamic node road network based on the usage status of the charging piles and the usage status of the mobile charging robot in the target area; A teacher model training module is used to take the dynamic node road network as input, use the DQN algorithm to train the teacher model for path planning, and generate an optimal path strategy in a dynamic environment; The student model training module is used to train the student model based on the distillation strategy, output the training results of the teacher model as soft labels, optimize the decision probability of the soft labels according to the usage status of the charging pile and the mobile charging robot, and re-rank them to obtain the optimal navigation strategy under the current state; The positioning and navigation module is used to determine the real-time location of the mobile charging robot and the target charging device through Wi-Fi fingerprint technology, and to visualize the navigation path planned by the teacher model or student model.
Citation Information
Patent Citations
Quadruped robot motion planning method based on privileged knowledge distillation
CN116203945A
Unmanned aerial vehicle autonomous path planning method based on lightweight continuous SAC algorithm
CN116430904A