A method and system for controlling the movement of a micro-robot
By perceiving and predicting environmental changes in real time, combining AI search and reinforcement learning algorithms to optimize paths, emotional computing and Kalman filters are introduced to adjust the motion trajectory, solving the problems of poor flexibility and low intelligence in complex environments, and achieving higher motion control accuracy and efficiency.
Patent Information
- Application Number
- CN202510030896.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-01-08
AI Technical Summary
Existing motion control methods of micro-robots show poor flexibility and low intelligence when facing complex and dynamic environments, making it difficult to effectively deal with environmental changes, resulting in possible disorientation or collisions.
Non-contact sensors are used to collect environmental status information in real time, generate comprehensive perception images, and analyze images through object detection and recognition technology to build an environmental status change prediction model. Use a fast random tree to generate candidate paths, combined with AI search algorithm optimization, a three-dimensional dynamic probability map is constructed and updated through reinforcement learning algorithms to generate an instant decision map. Introduce emotion computing models to simulate human decision-making processes, adjust the motion trajectory, and use Kalman filters to perform position estimation to ensure that the actual motion trajectory meets the ideal trajectory.
The flexibility and adaptability of micro-robots is improved, allowing them to make more reasonable decisions in complex and changing environments, ensuring the accuracy and efficiency of motion trajectories.
Smart Images

Figure CN119635658B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of robot control, and particularly relates to a motion control method and system for a micro-miniature robot. Background Art
[0002] Due to their small size and flexibility, micro-miniature robots have shown broad application prospects in multiple key fields. They achieve precise operations and quality inspections in industrial automation; assist in surgeries, perform drug delivery, and provide patient care in the medical field; carry out crop monitoring and pollution detection in environmental monitoring; and offer daily assistance such as cleaning and security in home services, significantly improving the efficiency and safety of various industries.
[0003] Currently, the motion control of micro-miniature robots mainly relies on traditional path planning algorithms and simple feedback control systems. Using predefined rules to generate paths, which is suitable for structured and static environments but difficult to handle complex and changing scenarios. Basic position tracking is achieved through proportional-integral-derivative (PID) control, but its performance is limited by the accuracy of the model and the stability of the environment. It is used to describe the behavior pattern of the robot but lacks the ability to adapt to environmental changes.
[0004] Although existing methods perform well under certain specific conditions, there are still significant deficiencies: Traditional methods usually assume that the environment is static or slowly changing and cannot effectively handle dynamically changing environments, resulting in the robot possibly getting lost or colliding. Most solutions lack advanced cognitive capabilities and learning mechanisms and are difficult to make optimal decisions according to the actual situation, especially when facing emergencies. Existing path planning methods often focus on finding a feasible path rather than an optimal path, which may lead to low efficiency or resource waste. To ensure real-time performance, many systems sacrifice the accuracy of position estimation; while pursuing high precision may lead to computational delays and affect real-time response. Summary of the Invention
[0005] The present invention aims to provide a motion control method and system for a micro-miniature robot to solve the problems of poor flexibility and low intelligence of micro-miniature robots in the prior art.
[0006] In a first aspect, an embodiment of the present invention provides a motion control method for a micro-miniature robot, including:
[0007] Using a non-contact sensor to collect real-time environmental state information and performing synchronization and calibration processing on the environmental state information to generate a comprehensive perception image;
[0008] Analyze the comprehensive perception image using object detection and recognition technology to determine the objects present in the current environment and the categories of the objects, obtain a classification result, and construct an environmental state change prediction model based on the classification result and the environmental state information. The environmental state change prediction model is used to predict the environmental state changes that affect the movement trajectory of the micro-miniature robot within a future period of time through historical environmental state information and the current environmental state information, and in combination with the classification result;
[0009] According to the comprehensive perception image and the environmental state changes predicted by the environmental state change prediction model, generate candidate paths using the Rapidly-exploring Random Tree (RRT), and optimize the candidate paths using an AI search algorithm to obtain a target path;
[0010] Based on the target path, construct a three-dimensional dynamic probability map using a Bayesian network. The three-dimensional dynamic probability map takes the environmental state at a specific location as nodes and the transition probability between environmental states as edges, and continuously updates the three-dimensional dynamic probability map using a reinforcement learning algorithm to generate an immediate decision-making map;
[0011] When it is detected that the environmental state has changed, based on the immediate decision-making map, and introducing an emotion computing model to simulate the emotional factors in the human decision-making process, adjust the movement trajectory of the micro-miniature robot to generate an adaptive behavior rule set;
[0012] Use a Kalman filter to estimate the position of the micro-miniature robot to obtain the position state of the micro-miniature robot. Based on the position state and the adaptive behavior rule set, adjust the speed and steering angle of the drive motor of the micro-miniature robot to ensure that the actual movement trajectory conforms to the ideal movement trajectory, and generate the trajectory tracking of the micro-miniature robot. The trajectory tracking refers to the process in which the micro-miniature robot moves according to a predetermined movement trajectory during the task execution, and adjusts its own speed and steering angle in real time to minimize the deviation between the actual movement trajectory and the ideal trajectory.
[0013] Optionally, it further includes:
[0014] After the micro-robot completes a path planning and executes the trajectory tracking, compare the deviation between the actual movement trajectory and the ideal movement trajectory to obtain a comparison result;
[0015] Based on the comparison result, determine effective behavior rules from the adaptive behavior rule set. The effective behavior rules refer to the behavior rules that can reduce the deviation between the actual movement trajectory and the ideal trajectory;
[0016] Incorporate the effective behavior rules into the motion control method library of the micro-robot for implementing the motion control of the micro-robot by using the effective behavior rules in the motion control method library in future tasks.
[0017] Optionally, based on the target path, construct a three-dimensional dynamic probability map by using a Bayesian network. The three-dimensional dynamic probability map takes the environmental state at a specific location as nodes and the transition probability between the environmental states as edges, and continuously update the three-dimensional dynamic probability map by using a reinforcement learning algorithm to generate an immediate decision-making map, including:
[0018] Based on the comprehensive perception image and the environmental state change prediction model, plan and optimize the actual motion trajectory of the micro-robot to obtain a target path. Use a rapidly-exploring random tree or a variant algorithm combined with an AI search algorithm to perform path planning and optimization processing, and further refine and optimize the target path to generate an optimal path;
[0019] According to the optimal path, use a Bayesian network to model the environmental state and the transition probability at the specific location, construct a three-dimensional dynamic probability map, and based on the three-dimensional dynamic probability map, perform prediction processing on the environmental change through spatio-temporal correlation analysis;
[0020] Based on the three-dimensional dynamic probability map, use a reinforcement learning algorithm to continuously perform update processing to generate an immediate decision-making map, so as to learn and optimize the behavior strategy of the micro-robot through the immediate decision-making map and improve the decision-making quality.
[0021] Optionally, based on the three-dimensional dynamic probability map, use a reinforcement learning algorithm to continuously perform update processing to generate an immediate decision-making map, including:
[0022] Perform selection processing based on the initial motion mode of the micro-robot, integrate the non-contact sensing data of the micro-robot, and combine the environmental state change prediction model to evaluate the current environmental state, and adjust the weight of the non-contact sensing data in real time to obtain an optimized initial environmental evaluation result;
[0023] Whenever a new comprehensive perception image is obtained, apply differential privacy protection technology to protect the security of the non-contact sensing data and effectively shield sensitive information, realize the security of the non-contact sensing data without affecting the accuracy of the three-dimensional dynamic probability map, and combine the newly obtained comprehensive perception image with the three-dimensional dynamic probability map to update the three-dimensional dynamic probability map to generate an updated three-dimensional dynamic probability map, and the updated three-dimensional dynamic probability map is used to reflect the latest environmental state;
[0024] Based on the updated three-dimensional dynamic probability map, an immediate decision-making map is generated using a decision-making framework that combines deep reinforcement learning and evolutionary algorithms. The framework is used to combine integrated multi-agent reinforcement learning to enable multiple micro-robots to cooperate with each other to complete complex tasks, including group obstacle avoidance and collaborative transportation, and adversarial training elements are used to simulate opponents or unforeseen environmental factors to generate the immediate decision-making map, which is used to guide the micro-robots to make optimal decisions.
[0025] Optionally, the method for generating an immediate decision-making map based on the updated three-dimensional dynamic probability map using a decision-making framework that combines deep reinforcement learning and evolutionary algorithms, where the framework is used to combine integrated multi-agent reinforcement learning to enable multiple micro-robots to cooperate with each other to complete complex tasks, including group obstacle avoidance and collaborative transportation, and adversarial training elements are used to simulate opponents or unforeseen environmental factors to generate the immediate decision-making map, which is used to guide the micro-robots to make optimal decisions, includes:
[0026] Based on the updated three-dimensional dynamic probability map, use a decision-making framework that combines deep reinforcement learning and evolutionary algorithms to automatically adjust the parameters of the environmental state change prediction model to obtain an optimized decision-making framework;
[0027] Based on the optimized decision-making framework, introduce multi-agent reinforcement learning to enable multiple micro-robots to cooperate with each other to complete complex tasks, including group obstacle avoidance and collaborative transportation. Each micro-robot not only considers its own behavior strategy but also shares information with other micro-robots to jointly formulate an optimal action plan to generate a collaborative behavior pattern;
[0028] Based on the collaborative behavior pattern, use adversarial training elements to generate an adversarial network to model and simulate opponents or unforeseen environmental factors, enhance the collaborative ability of the micro-robots, and generate a collaboratively enhanced model;
[0029] Generate the immediate decision-making map according to the optimized decision-making framework and the collaboratively enhanced model.
[0030] Optionally, use the Kalman filter to estimate the position of the micro robot to obtain the position state of the micro robot. Based on the position state and the adaptive behavior rule set, adjust the speed and steering angle of the drive motor of the micro robot to ensure that the actual motion trajectory conforms to the ideal motion trajectory, and generate the trajectory tracking of the micro robot. The trajectory tracking refers to the process in which the micro robot moves along a predetermined motion trajectory during the task execution, and adjusts its own speed and steering angle in real time to minimize the deviation between the actual motion trajectory and the ideal trajectory, including:
[0031] Use the Kalman filter to estimate the position of the micro robot in the environmental state to obtain the accurate position state of the micro robot. Based on the accurate position state, adjust the speed and steering angle of the drive motor to ensure that the actual motion trajectory conforms to the ideal motion trajectory, and obtain the adjustment result of the position state;
[0032] According to the adjustment result of the position state, perform visual servo control processing on the micro robot to obtain the actually adjusted motion trajectory in real time, ensure that the ideal motion trajectory can be achieved even when the environmental state changes, and obtain the trajectory tracking data of the micro robot;
[0033] Based on the trajectory tracking data, compare the difference between the actual motion trajectory and the ideal motion trajectory, generate an optimized motion mode, and use the optimized motion mode to improve the motion control method of the micro robot to obtain the trajectory tracking of the micro robot.
[0034] Optionally, according to the adjustment result of the position state, perform visual servo control processing on the micro robot to obtain the actually adjusted motion trajectory in real time, ensure that the ideal motion trajectory can be achieved even when the environmental state changes, and obtain the trajectory tracking data of the micro robot, including:
[0035] Use the adjustment result of the position state to perform visual servo control processing on the micro robot. By introducing image processing technology and computer vision algorithms, the actual motion trajectory is adjusted in real time to obtain the adjusted actual motion trajectory;
[0036] According to the adjusted actual motion trajectory, combined with multi-sensor fusion technology, further improve the accuracy of position estimation, enable the micro robot to achieve sub-centimeter positioning accuracy in a complex environment, and based on the high-precision position information, fine-tune the speed and steering angle of the drive motor to generate an optimized motion command;
[0037] Based on the optimized motion instructions, applying the adaptive control theory and predictive control algorithm, calculate and compensate in advance for the factors that may affect the actual motion trajectory, enabling the micro-robot to move smoothly along the ideal motion trajectory, and obtain the trajectory tracking data of the micro-robot. Based on the trajectory tracking data, combined with the object tracking technology in deep learning, continuously monitor and correct the traveling route of the micro-robot to make the traveling route meet the task requirements.
[0038] Optionally, based on the updated three-dimensional dynamic probability map, use a decision-making framework that combines deep reinforcement learning and evolutionary algorithms to generate an immediate decision-making map. The framework is used to combine integrated multi-agent reinforcement learning to enable multiple micro-robots to cooperate with each other to complete complex tasks, including group obstacle avoidance and collaborative transportation, and simulate opponents or unforeseen environmental factors through adversarial training elements to generate an immediate decision-making map. The immediate decision-making map is used to guide the micro-robot to make optimal decisions, including:
[0039] Based on the updated three-dimensional dynamic probability map, use a decision-making framework that combines deep reinforcement learning and evolutionary algorithms to automatically adjust the parameters of the environmental state change prediction model to obtain an optimized decision-making framework;
[0040] Among them, the parameters of the environmental state change prediction model include a smoothing coefficient, and the dynamically adjusted smoothing coefficient η t It is calculated by the following formula:
[0041]
[0042] Among them, η t represents the dynamically adjusted smoothing coefficient, k represents the speed of controlling the change of the smoothing coefficient, and Δs t represents the change in the environmental state between the current moment and the previous moment, and b represents a threshold value used to determine whether the environmental change is significant;
[0043] Based on the optimized decision-making framework, introduce multi-agent reinforcement learning to enable multiple micro-robots to cooperate with each other to complete complex tasks, including group obstacle avoidance and collaborative transportation. Each micro-robot not only considers its own behavior strategy but also shares information with other micro-robots to jointly formulate an optimal action plan to generate a collaborative behavior pattern;
[0044] Among them, a parameter local value weight w local is included in the micro-robot's consideration of its own behavior strategy, and a parameter global value weight w global is included in the micro-robot's sharing of information with other robots. The local value weight w localand the global value weight w global is calculated by the following formula:
[0045]
[0046] where w local represents the defined local value weight, w global represents the global value weight, d(s, s0) represents the distance metric between the current state s and the initial state s0, T represents the total task duration, and t represents the current time;
[0047] Based on the collaborative behavior pattern, an adversarial training element is used to generate an adversarial network, modeling and simulating opponents or unforeseen environmental factors. The simulated generated reference data contains an exponential factor to enhance the collaborative ability of the micro-robots and generate an enhanced collaborative model; among them, the exponential factor γ is calculated by the following formula:
[0048]
[0049] where γ represents the introduced exponential factor, α represents a constant that controls the base value of the exponential factor, β represents a constant that controls the growth rate of the exponential factor, C represents the current environmental complexity evaluation value, and C0 represents the benchmark complexity evaluation value;
[0050] According to the optimized decision-making framework and the enhanced collaborative model, the immediate decision-making map is generated; among them, the immediate decision-making map is calculated by the following formula:
[0051] M t+1 (s, a) = (1 - η t )M t (s, a) + η t ·(w local Q local (s, a) + w global Q global (s, a)) γ
[0052] where M t+1 (s, a) represents the updated value of taking action a in state s in the immediate decision-making map at time t + 1, s represents the state, a represents the action, M t (s, a) represents the value of taking action a in state s in the immediate decision-making map at time t, Q local (s, a) represents the local value function estimated by a rapidly-exploring random tree or other local search algorithms, Q global (s, a) represents the global value function estimated by a Bayesian network or other global optimization algorithms.
[0053] Optionally, perform visual servo control processing on the micro-robot according to the adjustment result of the position state to obtain the actual motion trajectory after real-time adjustment, ensuring that the ideal motion trajectory can be achieved even under the condition of environmental state change, and obtain the trajectory tracking data of the micro-robot, including:
[0054] Use the adjustment result of the position state to perform visual servo control processing on the micro-robot. By introducing image processing technology and computer vision algorithms, the actual motion trajectory is adjusted in real time to obtain the adjusted actual motion trajectory; among them, the visual servo control update u t+1 Is calculated by the following formula:
[0055]
[0056] Among them, u t+1 Represents the control input of the visual servo control update at time t + 1, u t Represents the control input at time t, K v Represents the visual servo controller gain matrix, which is used to adjust the control force, x d,t Represents the desired target position or feature point, x t Represents the actual position or feature point at the current moment, α t Represents the adaptive learning rate, which is dynamically adjusted according to the environmental change rate, Represents the gradient of the loss function with respect to the control input, which is used to optimize the control strategy;
[0057] According to the adjusted actual motion trajectory, combined with multi-sensor fusion technology, further improve the accuracy of position estimation, so that the micro-robot can achieve sub-centimeter positioning accuracy in a complex environment. Based on the high-precision position information, fine-tune the speed and steering angle of the drive motor to generate an optimized motion instruction; among them, the fused high-precision position estimation p t And the drive motor fine-tuning v t+1 Is calculated by the following formula:
[0058] Among them, p t Represents the fused high-precision position estimation, N s Represents the number of sensors, w i (t) represents the dynamic weight of the i-th sensor, which reflects the credibility of the sensor data and is dynamically adjusted according to the sensor error and environmental conditions, z i,t Represents the measurement value of the i-th sensor at time t, p t Represents the fused high-precision position estimation, v tDenote the driving motor speed and steering angle of the driving motor at time t, v t+1 Denote the driving motor speed and steering angle of the driving motor at time t+1, K p Denote the proportional control gain matrix for adjusting the fine-tuning force, p d,t Denote the desired target position, p t Denote the high-precision position estimate at the current moment, β t Denote the compensation coefficient, dynamically adjusted according to task requirements, σ act,t And σ des,t Denote the standard deviations of the actual and desired position uncertainties for compensating the position error;
[0059] Based on the optimized motion instruction, apply the adaptive control theory and predictive control algorithm to calculate and compensate in advance for the factors that may affect the actual motion trajectory, enabling the micro-robot to move smoothly along the ideal motion trajectory and obtain the trajectory tracking data of the micro-robot; among them, the adaptive control updates θ t+1 And the predictive compensation c t+1 Are obtained by calculating through the following formula:
[0060]
[0061] Among them, θ t Denote the adaptive controller parameters at time t, θ t+1 Denote the adaptive controller parameters at time t+1, α represents the learning rate, controlling the step size of each update, Denote the gradient of the loss function with respect to the parameters, used to guide the parameter update direction, γ t Denote the evolution factor, dynamically adjusted according to the task difficulty and environmental complexity, Δθ t Denote the change amount of the best hyperparameter combination searched by the evolutionary algorithm, c t Denote the compensation amount at time t, c t+1 Denote the compensation amount at time t+1, β represents the compensation coefficient, controlling the compensation force, f pred,t Denote the predicted disturbance force, f act,t Denote the actual encountered disturbance force, δ t Denote the exponential factor, dynamically adjusted according to the environmental complexity, enhancing the attention to key disturbances;
[0062] Based on the trajectory tracking data, combined with the target tracking technology in deep learning, continuously monitor and correct the traveling route of the micro-robot to make the traveling route meet the task requirements; among them, the target tracking update and incremental learning update are obtained by calculating through the following formula:
[0063]
[0064] Among them, x track,t+1 represents the tracking state vector of the target tracking update at time t + 1, and x track,t represents the tracking state vector at time t, and K t represents the Kalman gain, which is used to adjust the correction strength of the tracking error, and y t represents the observation value, h() represents the observation model, and λ t is an adaptive adjustment coefficient, which is dynamically adjusted according to the tracking error and maps the state vector to the observation space, and ρ t represents the exponential decay factor, which is used to balance the influence of long-term and short-term tracking errors, and e t represents the tracking error at the current moment, and w t+1 represents the behavior rule weight of the incremental learning update at time t + 1, and w t represents the behavior rule weight at time t, and η represents the learning rate, which controls the step size of each update. represents the gradient of the loss function with respect to the weight, which is calculated based on the current data set and μ t represents the incremental learning adjustment coefficient, which is dynamically adjusted according to the task progress, and σ err,t and σ ref represent the standard deviations of the current and reference behavior rule errors, which are used to optimize the learning process.
[0065] In a second aspect, an embodiment of the present invention provides a motion control system for a micro-miniature robot, including:
[0066] A collection and processing module, configured to collect current environmental state information in real time by using a non-contact sensor, and perform synchronization and calibration processing on the environmental state information to generate a comprehensive perception image;
[0067] A classification and prediction module, configured to analyze the comprehensive perception image by using target detection and recognition technologies to determine the objects existing in the current environment and the categories of the objects, obtain a classification result, and construct an environmental state change prediction model based on the classification result and the environmental state information. The environmental state change prediction model is used to predict the environmental state changes that affect the motion trajectory of the micro-miniature robot in a future period of time through historical environmental state information and the current environmental state information, and in combination with the classification result;
[0068] An optimization processing module, configured to generate candidate paths by using a fast random tree according to the comprehensive perception image and the environmental state changes predicted by the environmental state change prediction model, and perform optimization processing on the candidate paths in combination with an AI search algorithm to obtain a target path;
[0069] An update generation module, configured to construct a three-dimensional dynamic probability map based on the target path by using a Bayesian network. The three-dimensional dynamic probability map takes the environmental states at specific positions as nodes and the transition probabilities between the environmental states as edges, and continuously updates the three-dimensional dynamic probability map by using a reinforcement learning algorithm to generate an immediate decision-making map;
[0070] A simulation adjustment module, configured to, when detecting a change in the environmental state, based on the immediate decision-making map and introducing an emotion calculation model to simulate the emotional factors in the human decision-making process, adjust the motion trajectory of the micro-robot to generate an adaptive behavior rule set;
[0071] A position adjustment module, configured to estimate the position of the micro-robot by using a Kalman filter to obtain the position state of the micro-robot. Based on the position state and the adaptive behavior rule set, adjust the speed and steering angle of the drive motor of the micro-robot to ensure that the actual motion trajectory conforms to the ideal motion trajectory, and generate the trajectory tracking of the micro-robot. The trajectory tracking refers to the process in which the micro-robot moves along a predetermined motion trajectory during the task execution and adjusts its own speed and steering angle in real time to minimize the deviation between the actual motion trajectory and the ideal trajectory.
[0072] In the embodiments of the present invention, a non-contact sensor is used to collect real-time information on the current environmental state, and the environmental state information is synchronized and calibrated to generate a comprehensive perception image. The target detection and recognition technology is applied to analyze the comprehensive perception image to determine the objects existing in the current environment and the categories of the objects, obtaining a classification result. And an environmental state change prediction model is constructed based on the classification result and the environmental state information. The environmental state change prediction model is used to predict the environmental state changes affecting the movement trajectory of the micro-miniature robot within a future period of time through the historical environmental state information and the current environmental state information, and in combination with the classification result. According to the comprehensive perception image and the environmental state changes predicted by the environmental state change prediction model, a candidate path is generated using the rapidly-exploring random tree, and the candidate path is optimized using the AI search algorithm to obtain a target path. Based on the target path, a three-dimensional dynamic probability map is constructed using the Bayesian network. The three-dimensional dynamic probability map takes the environmental state at a specific position as nodes and the transition probability between the environmental states as edges, and the three-dimensional dynamic probability map is continuously updated using the reinforcement learning algorithm to generate an immediate decision-making map. When it is detected that the environmental state changes, based on the immediate decision-making map and introducing an emotion calculation model to simulate the emotional factors in the human decision-making process, the movement trajectory of the micro-miniature robot is adjusted to generate an adaptive behavior rule set. The position of the micro-miniature robot is estimated using a Kalman filter to obtain the position state of the micro-miniature robot. Based on the position state and the adaptive behavior rule set, the speed and steering angle of the drive motor of the micro-miniature robot are adjusted to ensure that the actual movement trajectory conforms to the ideal movement trajectory, and the trajectory tracking of the micro-miniature robot is generated. The trajectory tracking refers to the process in which the micro-miniature robot moves along a predetermined movement trajectory during the task execution, and adjusts its own speed and steering angle in real time to minimize the deviation between the actual movement trajectory and the ideal trajectory.
[0073] The technical solution of the present invention has the following beneficial effects:
[0074] The present invention uses a non-contact sensor to collect real-time environmental state information, ensuring the accuracy and timeliness of data. By synchronizing and calibrating the environmental state information, a comprehensive perception image is generated, guaranteeing the consistency and reliability of multi-source data. Applying object detection and recognition technologies to analyze the comprehensive perception image to determine the objects and their categories existing in the environment improves the understanding ability of the environment. Based on the classification results and environmental state information, an environmental state change prediction model is constructed to predict the environmental changes affecting the robot's motion trajectory in a future period of time, enhancing the foresight of the system. The Rapidly-exploring Random Tree (RRT) is used to generate candidate paths, and the candidate paths are optimized by combining with an AI search algorithm to obtain the optimal target path. This method not only improves the speed of path planning but also guarantees the quality of the path. A three-dimensional dynamic probability map is constructed using a Bayesian network and continuously updated through a reinforcement learning algorithm to generate an immediate decision-making map. This dynamic update mechanism enables the robot to flexibly respond to complex and changing environments. When it is detected that the environmental state changes, an emotion computing model is introduced to simulate the emotional factors in the human decision-making process, adjust the robot's motion trajectory, and generate an adaptive behavior rule set. This increases the flexibility and adaptability of the robot, enabling it to make more reasonable decisions in an uncertain environment. A Kalman filter is used to estimate the position of the micro and small robot to obtain an accurate position state. Based on this, the speed and steering angle of the driving motor are adjusted to ensure that the actual motion trajectory conforms to the ideal motion trajectory. By adjusting its own speed and steering angle in real time, the deviation between the actual motion trajectory and the ideal trajectory is reduced, ensuring that the robot moves according to the predetermined motion trajectory and improving the accuracy of task execution.
[0075] Furthermore, by comparing the deviation between the actual motion trajectory and the ideal motion trajectory, effective behavior rules are determined and incorporated into the motion control method library. This enables the robot to utilize these experiences in future tasks to reduce trajectory deviation and improve the accuracy of path execution. The system can continuously optimize the behavior rule set according to the results of each task, enabling the robot to have stronger adaptability when facing different environments and be able to quickly adjust strategies to cope with new situations. Storing the effective behavior rules in the method library for reference in subsequent tasks helps to continuously improve and optimize the long-term performance.
[0076] Furthermore, the Rapidly-exploring Random Tree or its variant algorithm is used in combination with an AI search algorithm for path planning and optimization processing to ensure that the target path is not only feasible but also optimal, improving the quality and efficiency of path planning. Based on the three-dimensional dynamic probability map, prediction processing of environmental changes is carried out through spatio-temporal correlation analysis, enhancing the robot's understanding and pre-judgment ability of future environmental states. The three-dimensional dynamic probability map is continuously updated using a reinforcement learning algorithm to generate an immediate decision-making map, helping the robot to make quick and reasonable decisions in a complex environment and improving the decision-making quality and flexibility.
[0077] These aspects or other aspects of the present invention will be more clearly understood in the following description of the embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0078] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0079] Figure 1 It is a flowchart of a method for controlling the movement of a micro - mini robot provided by an embodiment of the present invention;
[0080] Figure 2 It is a schematic structural diagram of a system for controlling the movement of a micro - mini robot provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0081] In order to enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention.
[0082] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.
[0083] Figure 1 There is provided a flowchart of a method for controlling the movement of a micro - mini robot according to an embodiment of the present invention. As Figure 1 shown, the method includes:
[0084] Based on this, the present invention provides a method for controlling the movement of a micro - mini robot. As Figure 1 , including:
[0085] Step 101: Use a non - contact sensor to collect real - time environmental state information, and synchronize and calibrate the environmental state information to generate a comprehensive perception image;
[0086] In this step, the non - contact sensor includes a lidar (LiDAR), a camera, an infrared sensor, etc. These sensors can obtain environmental data without directly contacting the object. For example, the lidar determines the distance of an object by emitting a laser beam and measuring the reflection time; the camera captures visual images.
[0087] Environmental status information: Refers to the physical properties of objects in the environment, such as position, shape, color, temperature, etc., as well as external factors such as lighting conditions and weather conditions.
[0088] Synchronization and calibration processing: Ensure that data from different sensors is consistent in time and space. This usually involves operations such as timestamp alignment and coordinate system conversion to fuse multi-source data into a unified perception image.
[0089] In the embodiments of the present invention, in industrial automation, micro and small robots can use a combination of lidar and cameras to scan obstacles and other robots in the working area in real time to ensure safe navigation. Home service robots can detect heat sources through infrared sensors and identify family members in combination with cameras to provide personalized services.
[0090] Step 102: Analyze the comprehensive perception image using target detection and recognition technology to determine the objects present in the current environment and the categories of the objects, obtain a classification result, and construct an environmental status change prediction model based on the classification result and the environmental status information. The environmental status change prediction model is used to predict environmental status changes that affect the movement trajectory of the micro and small robot within a future period of time through historical environmental status information and the current environmental status information, and in combination with the classification result;
[0091] In this step, target detection and recognition technology: Use computer vision algorithms (such as convolutional neural network CNN) to identify specific objects and their categories from the comprehensive perception image. For example, distinguish different types of objects such as people, furniture, and equipment.
[0092] Classification result: Refers to the specific category label of each object obtained after target detection.
[0093] Environmental status change prediction model: Based on historical and current environmental status information, combined with the classification result, predict environmental changes that may occur within a future period of time. This helps to plan paths in advance or adjust behavior strategies.
[0094] In the embodiments of the present invention, medical surgical assistance robots can identify surgical tools, patient organs, etc. in the operating room and predict the operating actions of doctors, so as to optimize their own positions to avoid interference. Environmental monitoring robots can identify the health status of crops and predict the occurrence trends of pests and diseases to provide support for precision agriculture.
[0095] Step 103: Generate candidate paths using the fast random tree according to the comprehensive perception image and the environmental status changes predicted by the environmental status change prediction model, and optimize the candidate paths in combination with the AI search algorithm to obtain the target path;
[0096] In this step, the Rapidly-exploring Random Tree (RRT), an algorithm for path planning, can quickly find a feasible path in a complex environment. It is achieved by randomly sampling points in the space and gradually expanding a tree structure.
[0097] AI search algorithms, such as traditional search algorithms like A* and Dijkstra, or more advanced deep learning methods, are used to evaluate and select the best path.
[0098] The target path is the final path after optimization, which takes into account the information provided by the environmental change prediction model to ensure the safety and efficiency of the path.
[0099] In the embodiments of the present invention, a logistics distribution robot can quickly generate multiple candidate paths using RRT in a dynamic pedestrian flow environment and select the shortest and best path that avoids crowded areas through an AI search algorithm. A disaster rescue robot can adjust the path planning in real time according to the changes in the structure of the ruins to ensure that it can reach the destination safely.
[0100] Step 104: Based on the target path, construct a three-dimensional dynamic probability map using a Bayesian network. The three-dimensional dynamic probability map takes the environmental state at a specific location as nodes and the transition probability between the environmental states as edges, and continuously updates the three-dimensional dynamic probability map using a reinforcement learning algorithm to generate an immediate decision-making map.
[0101] In this step, a Bayesian network is a probabilistic graphical model used to represent the dependencies between variables. Here, the nodes represent the environmental states at specific locations, and the edges represent the transition probabilities between the states.
[0102] The three-dimensional dynamic probability map models the environment as a series of interrelated states, each of which may change, forming a dynamic probability distribution map.
[0103] The immediate decision-making map is a map that guides the robot to make optimal decisions based on the current environmental state and predicted changes.
[0104] In the embodiments of the present invention, an autonomous driving vehicle can use the three-dimensional dynamic probability map to predict the behavior of other vehicles, thereby adjusting the driving route in advance to avoid potential collisions. A household cleaning robot can update the immediate decision-making map in real time according to the changes in the room layout to ensure efficient coverage of all areas that need to be cleaned.
[0105] Step 105: When it is detected that the environmental state has changed, based on the immediate decision-making map and introducing an emotion computing model to simulate the emotional factors in the human decision-making process, adjust the motion trajectory of the micro robot to generate an adaptive behavior rule set.
[0106] In this step, the real-time decision-making map: a map that guides the robot to make optimal decisions based on the current environmental state and predicted changes.
[0107] The emotion computing model: mimics the human emotion response mechanism, enabling the robot to exhibit a decision-making mode similar to that of humans when facing emergencies. For example, it preferentially selects a safe path in an emergency.
[0108] The adaptive behavior rule set: an action guide adjusted according to the new environmental state and the real-time decision-making map, ensuring that the robot can flexibly respond to various changes.
[0109] In the embodiments of the present invention, when a medical care robot discovers that a patient is suddenly unwell, it can quickly adjust its behavior rules to give priority to ensuring the safety of the patient. A service robot can adjust its service route during peak restaurant hours, give priority to handling customer needs, and avoid crowded areas at the same time.
[0110] Step 106: Use a Kalman filter to estimate the position of the micro-robot, obtain the position state of the micro-robot, and based on the position state and the adaptive behavior rule set, adjust the speed and steering angle of the drive motor of the micro-robot to ensure that the actual motion trajectory conforms to the ideal motion trajectory, and generate the trajectory tracking of the micro-robot. The trajectory tracking refers to the process in which the micro-robot moves according to a predetermined motion trajectory during the task execution, and adjusts its own speed and steering angle in real time to minimize the deviation between the actual motion trajectory and the ideal trajectory.
[0111] In this step, the Kalman filter: a recursive algorithm used to estimate the state of a system, especially in a noisy environment. It combines data from different sensors to provide a more accurate position estimate.
[0112] The position state: refers to the exact position and attitude of the robot at present.
[0113] The trajectory tracking: ensures that the robot moves according to a predetermined motion trajectory, and adjusts the speed and steering angle in real time to minimize the deviation between the actual motion trajectory and the ideal trajectory.
[0114] In the embodiments of the present invention, an industrial automation robot can accurately complete tasks on an assembly line and maintain stable trajectory tracking even when running at high speed. An indoor delivery robot can use a Kalman filter to accurately locate itself to ensure that it can deliver items to the designated location accurately every time.
[0115] In summary, it can be seen that the method of the present invention not only covers the entire process from environmental perception, target recognition, path planning to behavior adjustment, but also particularly emphasizes the importance of intelligent decision-making and adaptive capabilities, providing strong technical support for the efficient operation of micro-robots in complex and changing environments.
[0116] Optionally, the steps further include:
[0117] After the micro-robot completes a path planning and executes the trajectory tracking, compare the deviation between the actual motion trajectory and the ideal motion trajectory to obtain a comparison result; based on the comparison result, determine an effective behavior rule from the adaptive behavior rule set, where the effective behavior rule refers to a behavior rule that can reduce the deviation between the actual motion trajectory and the ideal trajectory; incorporate the effective behavior rule into the motion control method library of the micro-robot for using the effective behavior rule in the motion control method library to implement the motion control of the micro-robot in future tasks.
[0118] In this step, the key concepts involved in this solution include: actual motion trajectory, ideal motion trajectory, deviation comparison result, adaptive behavior rule set, and motion control method library. The actual motion trajectory refers to the path actually traveled by the micro-robot during the task execution; the ideal motion trajectory is the best path pre-planned and expected for the robot to follow. The deviation comparison result is a quantitative analysis of the difference between the two, used to evaluate the execution accuracy of the robot. The adaptive behavior rule set is a set of dynamically adjusted behavior guidelines that can automatically optimize the robot's action strategy according to environmental changes. The motion control method library is a database storing effective behavior rules, and these rules have been verified to reduce the deviation between the actual motion trajectory and the ideal trajectory, and are called for in future tasks to improve performance.
[0119] In the embodiment of the present invention, after the micro-robot completes a path planning and executes the trajectory tracking, the system will automatically compare the deviation between the actual motion trajectory and the ideal motion trajectory to generate a detailed comparison result. Based on these comparison results, the algorithm will screen out the effective behavior rules from the adaptive behavior rule set that can significantly reduce the deviation. These rules are then incorporated into the motion control method library of the robot for reference in future similar tasks to help the robot execute path planning and trajectory tracking more accurately. By continuously accumulating and applying effective behavior rules, the robot can gradually optimize its motion control strategy and improve the overall operation efficiency and accuracy.
[0120] In an application scenario of an intelligent logistics warehouse, micro and small-sized handling robots are responsible for transporting goods from one location to another. After each transportation task is completed, the system automatically records and analyzes the deviation between the actual path (actual movement trajectory) of the robot and the preset ideal path (ideal movement trajectory). For example, if the robot slightly deviates from the predetermined route when encountering an obstacle, the system calculates the specific value of this deviation and compares it with the previously stored behavior rules. If a certain specific behavior rule has effectively reduced similar deviations in the past, then this rule will be marked as an "effective behavior rule" and added to the robot's motion control method library. When performing the same or similar tasks in the future, the robot can call these verified behavior rules to prevent possible deviations in advance, thereby ensuring a more accurate and smooth path execution. Over time, this method enables the robot to gradually learn and improve its behavior pattern in various complex environments, ultimately achieving a higher task completion rate and lower operation error.
[0121] Optionally, based on the target path described in step 104, a three-dimensional dynamic probability map is constructed using a Bayesian network. The three-dimensional dynamic probability map takes the environmental state of a specific location as nodes and the transition probability between the environmental states as edges, and the three-dimensional dynamic probability map is continuously updated using a reinforcement learning algorithm to generate an immediate decision map, including:
[0122] Based on the comprehensive perception image and the environmental state change prediction model, the actual movement trajectory of the micro and small-sized robot is planned and optimized to obtain a target path. The path planning and optimization are performed using a fast random tree or a variant algorithm combined with an AI search algorithm to further refine and optimize the target path to generate an optimal path; based on the optimal path, the environmental state and the transition probability of the specific location are modeled using a Bayesian network to construct a three-dimensional dynamic probability map. Based on the three-dimensional dynamic probability map, through spatio-temporal correlation analysis, the environmental change is predicted; based on the three-dimensional dynamic probability map, using a reinforcement learning algorithm, continuous update processing is performed to generate an immediate decision map, so as to learn and optimize the behavior strategy of the micro and small-sized robot through the immediate decision map to improve the decision-making quality.
[0123] Optionally, the continuous update processing using a reinforcement learning algorithm based on the three-dimensional dynamic probability map to generate an immediate decision map includes:
[0124] Based on the initial motion pattern of the micro-robot, selection processing is performed, the non-contact sensing data of the micro-robot is integrated, and the current environmental state is evaluated in combination with the environmental state change prediction model. The weight of the non-contact sensing data is adjusted in real time to obtain an optimized initial environmental evaluation result; whenever a new comprehensive sensing image is obtained, differential privacy protection technology is applied to protect the security of the non-contact sensing data and effectively shield sensitive information, realizing the security of the non-contact sensing data without affecting the accuracy of the three-dimensional dynamic probability map. The newly obtained comprehensive sensing image is combined with the three-dimensional dynamic probability map to update the three-dimensional dynamic probability map to generate an updated three-dimensional dynamic probability map, and the updated three-dimensional dynamic probability map is used to reflect the latest environmental state; based on the updated three-dimensional dynamic probability map, a decision-making framework combining deep reinforcement learning and evolutionary algorithm is used to generate an immediate decision-making map. The framework is used to combine integrated multi-agent reinforcement learning to enable multiple micro-robots to cooperate with each other to complete complex tasks, including group obstacle avoidance and collaborative transportation, and adversarial training elements are used to simulate opponents or unforeseen environmental factors to generate an immediate decision-making map, and the immediate decision-making map is used to guide the micro-robot to make optimal decisions.
[0125] Optionally, based on the updated three-dimensional dynamic probability map, a decision-making framework combining deep reinforcement learning and evolutionary algorithm is used to generate an immediate decision-making map. The framework is used to combine integrated multi-agent reinforcement learning to enable multiple micro-robots to cooperate with each other to complete complex tasks, including group obstacle avoidance and collaborative transportation, and adversarial training elements are used to simulate opponents or unforeseen environmental factors to generate an immediate decision-making map, and the immediate decision-making map is used to guide the micro-robot to make optimal decisions, including:
[0126] Based on the updated three-dimensional dynamic probability map, a decision-making framework combining deep reinforcement learning and evolutionary algorithm is used to automatically adjust the parameters of the environmental state change prediction model to obtain an optimized decision-making framework; based on the optimized decision-making framework, multi-agent reinforcement learning is introduced to enable multiple micro-robots to cooperate with each other to complete complex tasks, including group obstacle avoidance and collaborative transportation. Each micro-robot not only considers its own behavior strategy but also shares information with other micro-robots to jointly formulate an optimal action plan to generate a collaborative behavior pattern; based on the collaborative behavior pattern, adversarial training elements are used to generate an adversarial network to model and simulate opponents or unforeseen environmental factors to enhance the collaborative ability of the micro-robots and generate a collaboratively enhanced model; according to the optimized decision-making framework and the collaboratively enhanced model, the immediate decision-making map is generated.
[0127] In this step, several key technical concepts involved in this solution include: three-dimensional dynamic probability map, Bayesian network, reinforcement learning algorithm, Rapidly-exploring Random Tree (RRT), differential privacy protection technology, deep reinforcement learning, evolutionary algorithm, multi-agent reinforcement learning, and adversarial training elements. The three-dimensional dynamic probability map is a model used to represent the environmental states and their transition probabilities at different positions in the environment; the Bayesian network is used to model the conditional dependencies between nodes, and here it is used to model the environmental states and transition probabilities at specific positions. The reinforcement learning algorithm learns the optimal behavior strategy by interacting with the environment; the Rapidly-exploring Random Tree is a path planning algorithm, and combined with the AI search algorithm, the path can be further refined and optimized. The differential privacy protection technology ensures the security and privacy protection of non-contact sensing data without affecting the accuracy of the three-dimensional dynamic probability map. Deep reinforcement learning is a method that combines deep learning and reinforcement learning, while the evolutionary algorithm is an optimization algorithm based on natural selection and genetic mechanisms. Multi-agent reinforcement learning enables multiple robots to cooperate with each other to complete tasks; the adversarial training elements are used to simulate opponents or unforeseen environmental factors to enhance the adaptability of the robots.
[0128] In the embodiment of the present invention, first, according to the comprehensive sensing image and the environmental state change prediction model, the Rapidly-exploring Random Tree or its variant algorithm is used in combination with the AI search algorithm to plan and optimize the actual motion trajectory of the micro-robot, and an optimal path is generated. Then, the Bayesian network is used to model the environmental states and transition probabilities at specific positions, a three-dimensional dynamic probability map is constructed, and the environmental changes are predicted through spatio-temporal correlation analysis. Next, the three-dimensional dynamic probability map is continuously updated through the reinforcement learning algorithm to generate an immediate decision-making map, so as to learn and optimize the behavior strategy of the robot and improve the decision-making quality. To ensure data security, the differential privacy protection technology is used to process the non-contact sensing data. Finally, through the framework combining deep reinforcement learning and evolutionary algorithm, combined with integrated multi-agent reinforcement learning, multiple micro-robots cooperate with each other to complete complex tasks, such as group obstacle avoidance and collaborative transportation, and the adversarial training elements are used to simulate opponents or unforeseen environmental factors to generate an immediate decision-making map to guide the robot to make the optimal decision.
[0129] In an intelligent logistics warehouse, a group of micro and small-sized handling robots equipped with sensors such as lidar and cameras need to efficiently complete the task of transporting goods. These robots first use the Rapidly-Exploring Random Tree (RRT) algorithm combined with the A* search algorithm to plan an optimal path from the starting point to the ending point, taking into account the possible presence of other robots. Subsequently, they use a Bayesian network to construct a three-dimensional dynamic probability map, which not only reflects the current state of each location in the warehouse but also predicts the change trend in the future for a period of time. Whenever a new perception image arrives, this probability map is updated immediately, and the behavior strategy of the robot is continuously optimized through a reinforcement learning algorithm. To protect sensitive information, all data collected by the sensors has undergone differential privacy processing. In addition, through a framework that combines deep reinforcement learning and evolutionary algorithms, the robots can share information and coordinate their work. For example, when encountering obstacles, they jointly decide the best detour route, or when handling large goods, they cooperate to improve efficiency. To better handle unforeseen situations, the system also introduces adversarial training elements to simulate possible problems, enabling the robot team to maintain efficient operation in any situation.
[0130] Optionally, use a Kalman filter to estimate the position of the micro and small-sized robot according to what is described in step 106 to obtain the position state of the micro and small-sized robot. Based on the position state and the adaptive behavior rule set, adjust the speed and steering angle of the drive motor of the micro and small-sized robot to ensure that the actual motion trajectory conforms to the ideal motion trajectory, and generate the trajectory tracking of the micro and small-sized robot. The trajectory tracking refers to the process in which the micro and small-sized robot moves along a predetermined motion trajectory during the task execution, and adjusts its own speed and steering angle in real time to minimize the deviation between the actual motion trajectory and the ideal trajectory, including:
[0131] Use the Kalman filter to estimate the position of the environmental state of the micro and small-sized robot to obtain the accurate position state of the micro and small-sized robot. Based on the accurate position state, adjust the speed and steering angle of the drive motor to ensure that the actual motion trajectory conforms to the ideal motion trajectory, and obtain the adjustment result of the position state; according to the adjustment result of the position state, perform visual servo control processing on the micro and small-sized robot to obtain the actually adjusted motion trajectory in real time, ensure that the ideal motion trajectory can be achieved even under the condition of environmental state change, and obtain the trajectory tracking data of the micro and small-sized robot; based on the trajectory tracking data, compare the difference between the actual motion trajectory and the ideal motion trajectory, generate an optimized motion mode, and use the optimized motion mode to improve the motion control method of the micro and small-sized robot to obtain the trajectory tracking of the micro and small-sized robot.
[0132] Optionally, perform visual servo control processing on the micro robot according to the adjustment result of the position state to obtain the actual motion trajectory after real-time adjustment, ensure that the ideal motion trajectory can be achieved even under the condition of environmental state change, and obtain the trajectory tracking data of the micro robot, including:
[0133] Use the adjustment result of the position state to perform visual servo control processing on the micro robot. By introducing image processing technology and computer vision algorithms, the actual motion trajectory is adjusted in real time to obtain the adjusted actual motion trajectory; according to the adjusted actual motion trajectory, combined with multi-sensor fusion technology, further improve the accuracy of position estimation, enable the micro robot to achieve sub-centimeter positioning accuracy in a complex environment, and based on the high-precision position information, fine-tune the speed and steering angle of the drive motor to generate an optimized motion instruction; based on the optimized motion instruction, apply adaptive control theory and predictive control algorithms to calculate and compensate in advance the factors that may affect the actual motion trajectory, enable the micro robot to move smoothly along the ideal motion trajectory, and obtain the trajectory tracking data of the micro robot. Based on the trajectory tracking data, combined with the object tracking technology in deep learning, continuously monitor and correct the travel route of the micro robot to make the travel route meet the task requirements.
[0134] In this step, the key concepts involved in this solution include: Kalman filter, position estimation, adaptive behavior rule set, visual servo control, multi-sensor fusion technology, adaptive control theory, predictive control algorithm, and object tracking technology. The Kalman filter is a recursive statistical estimation algorithm used to estimate the state of a system from a series of incomplete and noisy measurements; position estimation refers to determining the exact position of a robot in the environment by using sensor data; the adaptive behavior rule set is a set of rules that can be automatically adjusted according to environmental changes to guide the actions of the robot; visual servo control is to use visual feedback to adjust the motion of the robot in real time to achieve more precise task execution; multi-sensor fusion technology is to combine data from different sensors to improve the quality and reliability of information; adaptive control theory allows the control system to automatically adjust its performance according to changes in environmental or system parameters; the predictive control algorithm is a model-based control method that can calculate and compensate in advance the factors affecting the motion trajectory; the object tracking technology is part of the deep learning field, specifically used to continuously monitor and correct the path of a moving object.
[0135] In the embodiment of the present invention, the Kalman filter is first used to accurately estimate the position of the micro robot to ensure that its position state is accurate. Then, based on the obtained position state and the preset adaptive behavior rule set, the speed and steering angle of the drive motor are dynamically adjusted to make the actual motion trajectory as close to the ideal trajectory as possible. Next, in order to ensure that the ideal motion trajectory can be achieved even when the environmental state changes, visual servo control processing is used to further optimize the actual motion trajectory and obtain trajectory tracking data. Finally, by analyzing the trajectory tracking data, the difference between the actual motion trajectory and the ideal trajectory is compared, and an optimized motion pattern is generated to improve the robot's motion control method. In this process, image processing technology and computer vision algorithms, multi-sensor fusion technology, adaptive control theory, and predictive control algorithms are also combined to ensure that the robot can maintain high-precision positioning and smooth movement in a complex environment, and continuously monitor and correct the route through target tracking technology.
[0136] In a smart agriculture scenario, micro robots are deployed in greenhouses to automatically navigate to designated locations for crop monitoring or irrigation tasks. These robots are equipped with a variety of sensors such as lidar, cameras, and inertial measurement units (IMUs). When the robot starts to move, it first applies a Kalman filter to integrate all sensor data to provide an accurate position estimate to ensure that the robot knows its exact position. Then, based on this position information and pre-programmed behavior rules, the robot adjusts the speed and steering angle of the drive motor to ensure that it moves along the predetermined ideal trajectory. If it encounters an obstacle or the ground conditions change, the robot uses visual servo control to instantly adjust its speed and direction with the visual feedback provided by the camera to maintain the correct path. In addition, the robot also uses multi-sensor fusion technology to combine information from different sensors to obtain higher positioning accuracy. Throughout the process, the robot continuously applies adaptive control theory and predictive control algorithms to respond in advance to various factors that may affect its path. For example, it slows down before approaching a turning point, or plans a detour route in advance when an obstacle is detected ahead. Ultimately, the robot uses target tracking technology to ensure that its route always meets the mission requirements, while continuously improving its movement efficiency and accuracy through continuous learning and optimization.
[0137] In the motion control of micro-robots, especially in complex and changing environments, ensuring that robots can efficiently cooperate to complete tasks (such as group obstacle avoidance and collaborative transportation) is a key challenge. To address this challenge, this solution introduces a framework that combines deep reinforcement learning and evolutionary algorithms, and incorporates an immediate decision-making map generation mechanism, enabling multiple robots to cooperate with each other and adapt to environmental changes. The framework optimizes the behavioral strategies of robots by dynamically adjusting parameters such as smoothing coefficients, local and global value weights, and exponential factors to achieve optimal decisions.
[0138] Optionally, based on the updated three-dimensional dynamic probability map, a decision-making framework that combines deep reinforcement learning and evolutionary algorithms is used to generate an immediate decision-making map. The framework is used to enable multiple micro-robots to cooperate with each other to complete complex tasks through integrated multi-agent reinforcement learning. The complex tasks include group obstacle avoidance and collaborative transportation, and adversarial training elements are used to simulate opponents or unforeseen environmental factors to generate an immediate decision-making map. The immediate decision-making map is used to guide the micro-robots to make optimal decisions, including:
[0139] Based on the updated three-dimensional dynamic probability map, using a decision-making framework that combines deep reinforcement learning and evolutionary algorithms, the parameters of the environmental state change prediction model are automatically adjusted to obtain an optimized decision-making framework; where the parameters of the environmental state change prediction model include a smoothing coefficient, and the dynamically adjusted smoothing coefficient η t It is calculated by the following formula:
[0140]
[0141] where η t represents the dynamically adjusted smoothing coefficient, k represents the speed of controlling the change of the smoothing coefficient, Δs t represents the change in the environmental state between the current moment and the previous moment, and b represents a threshold value used to determine whether the environmental change is significant;
[0142] By calculating the smoothing coefficient η as described above t and making adjustments, multi-agent reinforcement learning is introduced into the optimized decision-making framework, enabling multiple micro-robots to cooperate with each other to complete complex tasks, including group obstacle avoidance and collaborative transportation. Each micro-robot not only considers its own behavioral strategy but also shares information with other micro-robots to jointly formulate an optimal action plan to generate a collaborative behavior pattern; where, in the micro-robot's consideration of its own behavioral strategy, there is a parameter local value weight w local , and in the information sharing between the micro-robot and other robots, there is a parameter global value weight w global , the local value weight w localand the global value weight w global is calculated by the following formula:
[0143]
[0144] where, w local represents the defined local value weight, w global represents the global value weight, d(s, s0) represents the distance metric between the current state s and the initial state s0, T represents the total task duration, and t represents the current time;
[0145] In the said collaborative behavior pattern, adversarial training elements are adopted to generate an adversarial network, modeling and simulating opponents or unforeseen environmental factors. An exponential factor is included in the simulated generated reference data to enhance the collaborative ability of the micro-robots, and a model after collaborative enhancement is generated; where, the exponential factor γ is calculated by the following formula:
[0146]
[0147] where, γ represents the introduced exponential factor, α represents a constant controlling the base value of the exponential factor, β represents a constant controlling the growth rate of the exponential factor, C represents the current environmental complexity evaluation value, and C0 represents the benchmark complexity evaluation value;
[0148] According to the optimized decision-making framework and the model after collaborative enhancement, the immediate decision-making map is generated; where, the immediate decision-making map is calculated by the following formula:
[0149] M t+1 (s, a) = (1 - η t )M t (s, a) + η t ·(w local Q local (s, a) + w global Q global (s, a)) γ
[0150] where, M t+1 (s, a) represents the updated value of taking action a in state s in the immediate decision-making map at time t + 1, s represents the state, a represents the action, M t (s, a) represents the value of taking action a in state s in the immediate decision-making map at time t, Q local (s, a) represents the local value function estimated by a rapidly-exploring random tree or other local search algorithms, Q global (s, a) represents the global value function estimated by a Bayesian network or other global optimization algorithms.
[0151] The purpose of the overall formula is to create a flexible and efficient decision-making system that can automatically adjust its behavioral strategies according to changes in the environment. It is used to smoothly adjust the model parameters to ensure that the decision-making framework can quickly respond to environmental changes without excessive fluctuations. Local and global value weights: Balance the relationship between local and global optimal solutions, enabling each robot to not only focus on its own best path but also consider the progress of the entire team's tasks. Enhance sensitivity to environmental complexity to help robots better adapt to unforeseen situations. Incorporating all the above factors, generate the latest decision-making map to guide the robot to make optimal decisions.
[0152] The following briefly explains the design reasons for each item in the formula:
[0153] Smoothing coefficient for dynamic adjustment The change amount of the environmental state Δs_t reflects the speed of environmental change, and the change amount is converted into a smoothing coefficient η through the Sigmoid function t , which can effectively control the adjustment speed of the model parameters. k controls the speed of change of the smoothing coefficient, and b is used as a threshold to determine whether the environmental change is significant.
[0154] Local value weight and global value weight The local value weight emphasizes the action priority within a short distance, while the global value weight focuses on long-term goals. The distance metric d(s, s0) measures the distance between the current state and the initial state, represents the proportion of the remaining time, and these two indicators jointly determine the value of each action.
[0155] Exponential factor The exponential factor enhances the attention to high-complexity environments. When the environment becomes more complex, γ increases, thereby increasing the importance of certain specific behaviors.
[0156] Instantaneous decision-making map M t+1 (s, a) = (1 - η t )M t (s, a) + η t ·(w local Q local (s, a) + w global Q global (s, a)) γ : By fusing the local and global value functions and applying the exponential factor γ for weighting, generate a new instantaneous decision-making map to guide the robot to take optimal actions.
[0157] The following briefly explains the acquisition methods of the parameters of each item in the formula:
[0158] Smoothing coefficient η for dynamic adjustment t: k and b are hyperparameters that can be set through experiments or based on experience. Δs t is calculated from sensor data and represents the change in the environmental state between the current moment and the previous moment.
[0159] The local value weight w local and the global value weight w global : d(s, s0) is calculated by the Euclidean distance or other appropriate distance metrics. T and t are the total task duration and the current time respectively, which are directly obtained from the task planning.
[0160] The exponential factor γ: α and β are hyperparameters that can be set through experiments or based on experience. C and C0 are the current environmental complexity evaluation value and the benchmark complexity evaluation value respectively, which are usually calculated in real time through the environmental perception module.
[0161] The immediate decision-making map M t+1 (s, a): Q local (s, a) and Q global (s, a) are estimated by the rapidly-exploring random tree and other local search algorithms, and global optimization algorithms such as Bayesian networks respectively. M t (s, a) is the immediate decision-making map of the previous time period and is stored in memory.
[0162] Parameter setting and value substitution:
[0163] Suppose there are 5 micro small handling robots in an intelligent logistics warehouse responsible for transporting goods from point A to point B. During this process, the robots need to avoid other moving objects and coordinate their work to improve efficiency. There are the following calculations:
[0164] In the parameter setting, set k = 2.0, b = 0.5 (determined according to historical data analysis), set α = 0.5, β = 0.2 (determined according to experimental results), set the total task duration T = 60 seconds, the current time t = 30 seconds, set the distance d(s, s0) between the current state s and the initial state s0 = 5 meters, set the current environmental complexity evaluation value C = 0.8, and the benchmark complexity evaluation value C0 = 0.5.
[0165] For the smoothing coefficient η t :
[0166]
[0167] Suppose Δs t = 0.7, and the calculation gives:
[0168]
[0169] For the local value weight w local and the global value weight wglobal Calculated as follows:
[0170]
[0171] Calculating the exponential factor γ gives:
[0172]
[0173] Updating the immediate decision-making map M t+1 (s,a):
[0174] Assume M t (s,a) = 0.8, Q local (s,a) = 0.9, Q global (s,a) = 0.7
[0175] Calculated as follows:
[0176] M t+1 (s,a)=(1 - 0.731)×0.8 + 0.731×(0.167×0.9 + 0.5×0.7) 0.684 ≈0.727 Through the above calculations, the new value M in the immediate decision-making map t+1 (s,a)≈0.727 shows an improvement in the value of taking a specific action in the current state, which reflects that the robot has optimized its action strategy after considering environmental changes, local and global value weights, and environmental complexity.
[0177] In the field of robotics, especially for micro and small robots, precise control of their motion trajectories is crucial. To ensure that robots can accurately track a predetermined path in a dynamically changing environment and achieve high-precision position estimation and motion control, technologies such as visual servo control, multi-sensor fusion, adaptive control, and target tracking are widely used. These technologies describe how to adjust the robot's behavior in real time through a series of formulas to adapt to environmental changes and optimize motion performance.
[0178] Optionally, according to the adjustment result of the position state, performing visual servo control processing on the micro and small robot to obtain the actually adjusted actual motion trajectory, ensuring that the ideal motion trajectory can be achieved even under the condition of environmental state change, and obtaining the trajectory tracking data of the micro and small robot, including:
[0179] Using the adjustment result of the position state, performing visual servo control processing on the micro and small robot, and through introducing image processing technology and computer vision algorithms, performing real-time adjustment on the actual motion trajectory to obtain the adjusted actual motion trajectory; where visual servo control at time t + 1 is represented as visual servo control update, and visual servo control update ut+1 Obtained by calculating using the following formula:
[0180]
[0181] where, u t+1 represents the control input of the visual servo control update at time t + 1, u t represents the control input at time t, K v represents the visual servo controller gain matrix for adjusting the control force, x d,t represents the desired target position or feature point, x t represents the actual position or feature point at the current moment, α t represents the adaptive learning rate, which is dynamically adjusted according to the environmental change rate, represents the gradient of the loss function with respect to the control input for optimizing the control strategy;
[0182] According to the adjusted actual motion trajectory, combined with the multi-sensor fusion technology, further improve the accuracy of position estimation, enable the micro and small robot to achieve sub-centimeter positioning accuracy in a complex environment, and based on the high-precision position information, fine-tune the speed and steering angle of the drive motor. The data integration in the fine-tuning process is the fine-tuning of the drive motor to generate an optimized motion instruction; among them, the fused high-precision position estimation p t and the drive motor fine-tuning v t+1 are obtained by calculating using the following formula:
[0183]
[0184] where, p t represents the fused high-precision position estimation, N s represents the number of sensors, w i (t) represents the dynamic weight of the i-th sensor, reflecting the credibility of the data of this sensor, and is dynamically adjusted according to the sensor error and environmental conditions, z i,t represents the measurement value of the i-th sensor at time t, p t represents the fused high-precision position estimation, v t represents the drive motor speed and steering angle of the drive motor at time t, v t+1 represents the drive motor speed and steering angle of the drive motor at time t + 1, K p represents the proportional control gain matrix for adjusting the fine-tuning force, p d,t represents the desired target position, p t represents the high-precision position estimation at the current moment, β t represents the compensation coefficient, which is dynamically adjusted according to the task requirements, σ act,t and σ des,tRepresents the standard deviation of the actual and expected position uncertainties, used to compensate for position errors;
[0185] Based on the optimized motion instruction, applying the adaptive control theory and predictive control algorithm, the adaptive control theory includes adaptive control updates, calculating and compensating in advance for factors that will affect the actual motion trajectory, and the calculated and compensated data is represented by predictive compensation, enabling the micro-robot to move smoothly along the ideal motion trajectory and obtaining the trajectory tracking data of the micro-robot; among them, adaptive control θ t+1 and predictive compensation c t+1 Are calculated by the following formula:
[0186]
[0187] Among them, θ t Represents the adaptive controller parameter at time t, θ t+1 Represents the adaptive controller parameter at time t + 1, α represents the learning rate, controlling the step size of each update, Represents the gradient of the loss function with respect to the parameter, used to guide the parameter update direction, γ t Represents the evolution factor, dynamically adjusted according to the task difficulty and environmental complexity, Δθ t Represents the change in the best hyperparameter combination searched by the evolutionary algorithm, c t Represents the compensation amount at time t, c t+1 Represents the compensation amount at time t + 1, β represents the compensation coefficient, controlling the compensation strength, f pred,t Represents the predicted disturbance force, f act,t Represents the actual encountered disturbance force, δ t Represents the exponential factor, dynamically adjusted according to the environmental complexity, enhancing the attention to key disturbances;
[0188] Based on the trajectory tracking data, combining the object tracking technology in deep learning, the object tracking technology includes object tracking updates, continuously monitoring and correcting the travel route of the micro-robot, making the travel route meet the task requirements, and during the monitoring and correction process, integrating the continuously changing travel route data received into incremental learning updates; among them, object tracking update x track,t+1 and incremental learning update w t+1 Are calculated by the following formula:
[0189]
[0190] Among them, x track,t+1 Represents the tracking state vector of the object tracking update at time t + 1, x track,t Represents the tracking state vector at time t, Kt represents the Kalman gain, which is used to adjust the correction strength of the tracking error, y t represents the observed value, h() represents the observation model, λ t Adaptive adjustment coefficient, dynamically adjusted according to the tracking error, mapping the state vector to the observation space, ρ t represents the exponential decay factor, which is used to balance the influence of long-term and short-term tracking errors, e t represents the tracking error at the current moment, w t+1 represents the behavior rule weight of the incremental learning update at time t+1, w t represents the behavior rule weight at time t, η represents the learning rate, which controls the step size of each update, represents the gradient of the loss function with respect to the weight, calculated based on the current dataset calculated, μ t represents the incremental learning adjustment coefficient, dynamically adjusted according to the task progress, σ err,t and σ ref represent the standard deviations of the current and reference behavior rule errors, which are used to optimize the learning process.
[0191] The purpose of the overall formula is to create an intelligent control system that can adapt to complex environmental changes, enabling the micro and small robots to maintain good motion performance even under uncertain conditions. It is used to adjust the actual motion trajectory of the robot in real time according to image processing techniques and computer vision algorithms, making it as close as possible to the ideal motion trajectory. Improve the accuracy of position estimation, ensure sub-centimeter positioning accuracy, and thus provide precise fine-tuning instructions for the speed and steering angle of the drive motor. Calculate in advance the factors that may affect the motion trajectory and make compensations, so that the robot can move smoothly along the ideal trajectory. Continuously monitor and correct the travel route to ensure that the task requirements are met, and at the same time continuously optimize the behavior rules through incremental learning.
[0192] The following briefly explains the design reasons for each item in the formula:
[0193] Visual servo control update This formula aims to combine visual feedback information to adjust the control input to minimize the difference between the desired position x d,t and the current position x t while optimizing the control strategy through the gradient descent method.
[0194] High-precision position estimation and drive motor fine-tuning By fusing the data of multiple sensors, improve the accuracy of position estimation and accordingly fine-tune the control of the drive motor to achieve more precise actions.
[0195] Adaptive control update and predictive compensation These two formulas enable the robot to automatically adjust its own parameters according to environmental changes and make predictive compensation for potential interference factors, ensuring stability and response speed.
[0196] Target tracking update x track,t+1 = x track,t + K t (y t - h(x track,t )) + λ t · exp(-ρ t ||e t ||) and incremental learning update These two formulas are used to continuously monitor the robot's travel route and continuously improve its behavior rules through incremental learning to better complete tasks.
[0197] The following briefly explains the acquisition methods of each sub-item parameter of the formula:
[0198] Visual servo control update u t+1 : K v is the visual servo controller gain matrix, usually determined based on experiments. α t The adaptive learning rate can be dynamically adjusted according to the environmental change rate. The gradient of the loss function with respect to the control input can be obtained through the backpropagation algorithm.
[0199] High-precision position estimation p t and drive motor fine-tuning v t+1 : w i (t) The dynamic weight reflects the credibility of each sensor data and is dynamically adjusted based on sensor errors and environmental conditions. z i,t is the sensor measurement value, directly read from the sensor. K p is the proportional control gain matrix, used to adjust the fine-tuning strength, also set based on experiments or experience. β t The compensation coefficient is dynamically adjusted according to task requirements, while σ act,t and σ des,t respectively represent the standard deviations of the actual and expected position uncertainties, used to compensate for position errors.
[0200] Adaptive control update θ t+1 and predictive compensation c t+1 : α The learning rate controls the step size of each update, the gradient of the loss function with respect to the parameter, γ t the evolution factor, all need to be set based on experiments. f pred,t and f act,tare the predicted and actual encountered interference forces, δ t The exponential factor enhances the attention to key interferences, and these parameters usually come from the environmental perception module.
[0201] Target tracking updates x track,t+1 and incremental learning updates w t+1 : K t Kalman gain, λ t Adaptive adjustment coefficient, ρ t Exponential decay factor, η learning rate, μ t Incremental learning adjustment coefficient, all of which are hyperparameters and can be set through experiments or based on experience. e t Tracking error at the current moment, y t Observation value, h() observation model, Gradient of the loss function with respect to the weight, based on the current dataset D t Calculation, σ err,t and σ ref Standard deviation of the behavioral rule error, used to optimize the learning process.
[0202] Consider a scenario where a formation of 5 drones performs a reconnaissance mission. During this process, the drones need to maintain specific relative positions and deal with suddenly emerging obstacles or other unforeseen situations.
[0203] Parameter setting and numerical substitution
[0204] Set K in the parameter setting v = [0.2, 0.2], α t = 0.1 (based on the environmental change rate), set N s = 3 (using three different types of sensors), w i (t) = [0.4, 0.3, 0.3] (based on sensor credibility), assume z i,t = [2.5, 3.0, 2.8] meters (measured values of each sensor), set K p = [0.5, 0.5], β t = 0.05, σ act,t = 0.2, σ des,t = 0.1, set α = 0.01, γ t = 0.1, Δθ t = [0.05, -0.03], set f pred,t = 0.5 N (predicted interference force), f act,t = 0.6 N (actual encountered interference force), δ t = 1.5, set K t = [0.7, 0.7], λ t = 0.2, ρt = 0.1, η = 0.001, μ t = 0.1, σ err,t = 0.3, σ ref = 0.2。
[0205] For visual servo control update u t+1 :
[0206] Assume u t = [0.5, 0.5], x d,t = [3.0, 3.0], x t = [2.8, 2.9],
[0207] Calculated:
[0208] u t+1 = [0.5, 0.5] + [0.2, 0.2] × ([3.0, 3.0] - [2.8, 2.9]) + 0.1 × [0.1, 0.1]
[0209] u t+1 ≈ [0.54, 0.53]
[0210] For high-precision position estimation p t and drive motor fine-tuning v t+1 :
[0211] Among them, calculate high-precision position estimation p t = 0.4 × 2.5 + 0.3 × 3.0 + 0.3 × 2.8 = 2.77
[0212] Assume v t = [0.6, 0.6], p d,t = 3.0
[0213] Calculated:
[0214]
[0215] v t+1 ≈ [0.67, 0.67]
[0216] For adaptive control update θ t+1 and predictive compensation c t+1 :
[0217] Assume θ t = [0.4, 0.4],
[0218] Calculated:
[0219] θ t+1= [0.4, 0.4] + 0.01×[0.05, -0.05] + 0.1×[0.05, -0.03]
[0220] θ t+1 ≈[0.405, 0.397]
[0221] Assume c t = 0.1
[0222] Calculated as:
[0223] c t+1 = 0.1 + 0.05×(0.5 - 0.6) 1.5
[0224] c t+1 ≈0.097
[0225] For target tracking, update x track,t+1 and incremental learning update w t+1 :
[0226] Assume x track,t = [2.8, 2.9], y t = [3.0, 3.1], h(x track,t ) = [2.9, 3.0], e t = [0.1, 0.1]
[0227] Calculate:
[0228] x track,t+1 = [2.8, 2.9] + [0.7, 0.7]×([3.0, 3.1] - [2.9, 3.0]) + 0.2·exp(-0.1×
[0229] ∥[[0.1, 0.1]])
[0230] Obtained:
[0231] x track,t+1 ≈[2.87, 2.97]
[0232] Assume
[0233] Calculate:
[0234] Obtained:
[0235] w t+1 ≈[0.514, 0.486]
[0236] Through the above calculations, the navigation ability and mission execution efficiency of the UAV formation in complex environments have been greatly enhanced. It not only ensures the coordination among UAV groups, but also improves the ability of individual UAVs to handle emergencies, thus ensuring that the entire formation can complete the reconnaissance mission safely and efficiently.
[0237] Figure 2 The following is a schematic structural diagram of a motion control system for a micro-miniature robot provided by an embodiment of the present invention. As Figure 2 shown, the system includes:
[0238] A collection and processing module 21, configured to collect real-time environmental state information using a non-contact sensor, and perform synchronization and calibration processing on the environmental state information to generate a comprehensive perception image;
[0239] A classification and prediction module 22, configured to analyze the comprehensive perception image using target detection and recognition technologies to determine the objects present in the current environment and the categories of the objects, obtain a classification result, and construct an environmental state change prediction model based on the classification result and the environmental state information. The environmental state change prediction model is used to predict environmental state changes that affect the motion trajectory of the micro-miniature robot within a future period of time through historical environmental state information and the current environmental state information, and in combination with the classification result;
[0240] An optimization processing module 23, configured to generate candidate paths using a rapid random tree according to the comprehensive perception image and the environmental state changes predicted by the environmental state change prediction model, and perform optimization processing on the candidate paths in combination with an AI search algorithm to obtain a target path;
[0241] An update and generation module 24, configured to construct a three-dimensional dynamic probability map based on the target path using a Bayesian network. The three-dimensional dynamic probability map uses the environmental state at a specific position as a node and the transition probability between environmental states as an edge, and continuously updates the three-dimensional dynamic probability map using a reinforcement learning algorithm to generate an immediate decision-making map;
[0242] A simulation and adjustment module 25, configured to, when it is detected that the environmental state changes, based on the immediate decision-making map, introduce an emotion calculation model to simulate the emotional factors in the human decision-making process, and adjust the motion trajectory of the micro-miniature robot to generate an adaptive behavior rule set;
[0243] A position adjustment module 26 is configured to estimate the position of the micro-robot by using a Kalman filter, obtain the position state of the micro-robot, and based on the position state and the adaptive behavior rule set, adjust the speed and steering angle of the driving motor of the micro-robot to ensure that the actual motion trajectory conforms to the ideal motion trajectory, and generate the trajectory tracking of the micro-robot. The trajectory tracking refers to the process in which the micro-robot moves along a predetermined motion trajectory during the task execution, and adjusts its own speed and steering angle in real time to minimize the deviation between the actual motion trajectory and the ideal trajectory.
[0244] Figure 2 The described micro-robot motion control system can execute Figure 1 The micro-robot motion control method described in the illustrated embodiment, and its implementation principle and technical effects will not be elaborated. For the micro-robot motion control system in the above embodiment, the specific manners in which each module and unit perform operations have been described in detail in the embodiment related to the method, and will not be elaborated here.
[0245] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A micro robot motion control method, characterized in that: include: Using non-contact sensors to collect current environmental status information in real time, and synchronizing and calibrating the environmental status information to generate a comprehensive perception image; Applying target detection and recognition technology to analyze the comprehensive perception image to determine the objects existing in the current environment and the categories of the objects, obtain a classification result, and construct an environment state change prediction model based on the classification result and the environment state information, wherein the environment state change prediction model is used to predict the environment state changes that affect the motion trajectory of the micro robot in the future period of time through historical environment state information and the current environment state information in combination with the classification result; According to the comprehensive perception image and the environmental state change predicted by the environmental state change prediction model, a candidate path is generated using a fast random tree, and the candidate path is optimized in combination with an AI search algorithm to obtain a target path; Based on the target path, a three-dimensional dynamic probability graph is constructed using a Bayesian network, wherein the three-dimensional dynamic probability graph uses the environmental state at a specific location as a node and the transition probability between the environmental states as an edge, and the three-dimensional dynamic probability graph is continuously updated using a reinforcement learning algorithm to generate an instant decision map; When a change in the environmental state is detected, based on the instant decision map, an emotional computing model is introduced to simulate the emotional factors in the human decision-making process, and the motion trajectory of the micro-robot is adjusted to generate an adaptive behavior rule set; A Kalman filter is used to estimate the position of the microrobot to obtain the position state of the microrobot. Based on the position state and the adaptive behavior rule set, the speed and steering angle of the drive motor of the microrobot are adjusted to ensure that the actual motion trajectory conforms to the ideal motion trajectory, and the trajectory tracking of the microrobot is generated. The trajectory tracking refers to the process in which the microrobot moves according to a predetermined motion trajectory during the execution of a task, and adjusts its own speed and steering angle in real time to minimize the deviation between the actual motion trajectory and the ideal motion trajectory.
2. The method according to claim 1, characterized in that: Also includes: After the micro-robot completes a path planning and performs the trajectory tracking, comparing the deviation between the actual motion trajectory and the ideal motion trajectory to obtain a comparison result; Based on the comparison result, determining an effective behavior rule from the adaptive behavior rule set, wherein the effective behavior rule refers to a behavior rule that can reduce the deviation between the actual motion trajectory and the ideal trajectory; The effective behavior rules are incorporated into the motion control method library of the micro robot, so that the effective behavior rules in the motion control method library can be used to realize the motion control of the micro robot in future tasks.
3. The method according to claim 1, characterized in that Based on the target path, a three-dimensional dynamic probability graph is constructed using a Bayesian network, wherein the three-dimensional dynamic probability graph uses the environmental state at a specific location as a node and the transition probability between the environmental states as an edge, and the three-dimensional dynamic probability graph is continuously updated using a reinforcement learning algorithm to generate an instant decision map, including: Based on the comprehensive perception image and the environmental state change prediction model, the actual motion trajectory of the micro robot is planned and optimized, and a target path is generated by using a fast random tree combined with an AI search algorithm; According to the target path, the environmental state of the specific location and the transition probability are modeled using a Bayesian network to construct a three-dimensional dynamic probability map, and based on the three-dimensional dynamic probability map, the change of the environmental state is predicted through spatiotemporal correlation analysis; Based on the three-dimensional dynamic probability map, a reinforcement learning algorithm is used to continuously update and generate a real-time decision map, so as to learn and optimize the behavior strategy of the micro robot through the real-time decision map to improve the decision quality.
4. The method according to claim 3, characterized in that Based on the three-dimensional dynamic probability map, the reinforcement learning algorithm is used to continuously update and generate an instant decision map, including: Selecting and processing based on the initial motion mode of the micro-robot, integrating the non-contact sensing data of the micro-robot, evaluating the current environmental state in combination with the environmental state change prediction model, adjusting the weight of the non-contact sensing data in real time, and obtaining an optimized initial environmental assessment result; Whenever a new comprehensive perception image is acquired, differential privacy protection technology is applied to protect the security of the non-contact perception data and effectively shield sensitive information, thereby achieving the security of the non-contact perception data without affecting the accuracy of the three-dimensional dynamic probability map, and combining the newly acquired comprehensive perception image with the three-dimensional dynamic probability map to update the three-dimensional dynamic probability map to generate an updated three-dimensional dynamic probability map, wherein the updated three-dimensional dynamic probability map is used to reflect the latest environmental status; Based on the updated three-dimensional dynamic probability map, a decision-making framework combining deep reinforcement learning and evolutionary algorithm is used to generate a real-time decision map. The framework is used to combine multi-agent reinforcement learning to enable multiple micro-robots to collaborate with each other to complete complex tasks. The complex tasks include group obstacle avoidance and collaborative transportation, and adversarial training elements are used to simulate opponents or unforeseen environmental factors to generate a real-time decision map. The real-time decision map is used to guide the micro-robots to make optimal decisions.
5. The method according to claim 4, characterized in that The method of generating an instant decision map based on the updated three-dimensional dynamic probability map and utilizing a decision framework combining deep reinforcement learning and an evolutionary algorithm includes: Based on the updated three-dimensional dynamic probability map, the decision framework combining deep reinforcement learning and evolutionary algorithm is used to automatically adjust the parameters of the environmental state change prediction model to obtain an optimized decision framework; Based on the optimized decision-making framework, multi-agent reinforcement learning is introduced to enable multiple micro-robots to collaborate with each other to complete complex tasks, including group obstacle avoidance and collaborative transportation. Each micro-robot not only considers its own behavior strategy, but also shares information with other micro-robots to jointly develop the best action plan to generate a collaborative behavior model. Based on the collaborative behavior pattern, an adversarial network is generated using adversarial training elements to model and simulate opponents or unforeseen environmental factors, enhance the collaborative ability of the micro-robot, and generate a collaboratively enhanced model; The instant decision map is generated according to the optimized decision framework and the collaboratively enhanced model.
6. The method according to claim 1, characterized in that The method uses a Kalman filter to estimate the position of the micro robot, obtains the position state of the micro robot, adjusts the speed and steering angle of the drive motor of the micro robot based on the position state and the adaptive behavior rule set, ensures that the actual motion trajectory conforms to the ideal motion trajectory, and generates the trajectory tracking of the micro robot. The trajectory tracking refers to the process in which the micro robot moves according to the predetermined motion trajectory during the execution of the task, and adjusts its own speed and steering angle in real time to minimize the deviation between the actual motion trajectory and the ideal motion trajectory, including: Using the Kalman filter to estimate the position of the environmental state of the micro robot, obtain the precise position state of the micro robot, and based on the precise position state, adjust the speed and steering angle of the drive motor to ensure that the actual motion trajectory conforms to the ideal motion trajectory, and obtain the adjustment result of the precise position state; According to the adjustment result of the precise position state, visual servo control processing is performed on the micro robot to obtain the actual motion trajectory after real-time adjustment, to ensure that the ideal motion trajectory can be achieved even when the environmental state changes, and to obtain trajectory tracking data of the micro robot; Based on the trajectory tracking data, the difference between the actual motion trajectory and the ideal motion trajectory is compared to generate an optimized motion pattern. The optimized motion pattern is used to improve the motion control method of the micro robot to obtain the trajectory tracking of the micro robot.
7. The method according to claim 6, characterized in that The method of performing visual servo control processing on the micro robot according to the adjustment result of the precise position state to obtain the actual motion trajectory after real-time adjustment, ensuring that the ideal motion trajectory can be achieved even when the environmental state changes, and obtaining the trajectory tracking data of the micro robot, includes: Using the adjustment result of the precise position state, visual servo control processing is performed on the micro robot, and the actual motion trajectory is adjusted in real time by introducing image processing technology and computer vision algorithm to obtain the adjusted actual motion trajectory; According to the adjusted actual motion trajectory, combined with multi-sensor fusion technology, the accuracy of position estimation is further improved, so that the micro robot can achieve sub-centimeter positioning accuracy in a complex environment, and based on high-precision position information, the speed and steering angle of the drive motor are fine-tuned to generate optimized motion instructions; Based on the optimized motion instructions, adaptive control theory and predictive control algorithm are applied to calculate and compensate in advance the factors affecting the actual motion trajectory, so that the micro robot can move smoothly along the ideal motion trajectory, and the trajectory tracking data of the micro robot is obtained. Based on the trajectory tracking data, combined with the target tracking technology in deep learning, the travel route of the micro robot is continuously monitored and corrected to make the travel route meet the task requirements.
8. The method according to claim 4, characterized in that Based on the updated three-dimensional dynamic probability map, a decision framework combining deep reinforcement learning and evolutionary algorithm is used to generate an instant decision map. The framework is used to combine multi-agent reinforcement learning to enable multiple micro-robots to cooperate with each other to complete complex tasks. The complex tasks include group obstacle avoidance and collaborative transportation, and adversarial training elements are used to simulate opponents or unforeseen environmental factors to generate an instant decision map. The instant decision map is used to guide the micro-robot to make the best decision, including: Based on the updated three-dimensional dynamic probability map, the decision framework combining deep reinforcement learning and evolutionary algorithm is used to automatically adjust the parameters of the environmental state change prediction model to obtain an optimized decision framework; The parameters of the environmental state change prediction model include a smoothing coefficient, a dynamically adjusted smoothing coefficient Calculated by the following formula: , in, Indicates the speed at which the smoothing coefficient changes. Indicates the change in the environmental state between the current moment and the previous moment. represents the threshold value used to determine whether the environmental change is significant; Based on the optimized decision-making framework, multi-agent reinforcement learning is introduced to enable multiple micro-robots to collaborate with each other to complete complex tasks, including group obstacle avoidance and collaborative transportation. Each micro-robot not only considers its own behavior strategy, but also shares information with other micro-robots to jointly develop the best action plan to generate a collaborative behavior model. Among them, the behavior strategy of the micro robot considering itself includes a parameter local value weight , and the information shared by the micro robot and other robots includes a parameter global value weight , local value weight and global value weight , calculated by the following formula: , , in, Indicates the current status With the initial state The distance measure between Indicates the total duration of the task. Indicates the current time; Based on the collaborative behavior pattern, an adversarial network is generated using adversarial training elements to model and simulate opponents or unforeseen environmental factors. An exponential factor is included in the simulated reference data to enhance the collaborative ability of the micro-robot and generate a collaboratively enhanced model; wherein the exponential factor Calculated by the following formula: , in, represents a constant that controls the base value of the exponential factor. represents a constant that controls the growth rate of the exponential factor. Indicates the current environment complexity evaluation value, represents the benchmark complexity evaluation value; The instant decision map is generated according to the optimized decision framework and the collaboratively enhanced model; wherein the instant decision map is calculated by the following formula: , in, Indicates that at time In the real-time decision map, the state Take action The value of Indicates the status, Indicates action, Indicates at time In the real-time decision map, the state Take action The value of represents the local value function estimated by fast random trees, represents the global value function estimated by the Bayesian network.
9. The method according to claim 6, characterized in that The method of performing visual servo control processing on the micro robot according to the adjustment result of the precise position state to obtain the actual motion trajectory after real-time adjustment, ensuring that the ideal motion trajectory can be achieved even when the environmental state changes, and obtaining the trajectory tracking data of the micro robot, includes: The adjustment result of the precise position state is used to perform visual servo control processing on the micro robot. By introducing image processing technology and computer vision algorithm, the actual motion trajectory is adjusted in real time to obtain the adjusted actual motion trajectory, which is calculated by the following formula: , in, represents the visual servo control update at time The control input, Indicates at time The control input, represents the visual servo controller gain matrix, which is used to adjust the control strength. represents the desired target position, Indicates the actual position at the current moment. Represents the adaptive learning rate, which is dynamically adjusted according to the rate of change of the environment. Represents the gradient of the loss function with respect to the control input, which is used to optimize the control strategy; According to the adjusted actual motion trajectory, combined with multi-sensor fusion technology, the accuracy of position estimation is further improved, so that the micro robot can achieve sub-centimeter positioning accuracy in complex environments. Based on high-precision position information, the speed and steering angle of the drive motor are fine-tuned to generate optimized motion instructions, which are calculated by the following formula: , , in, represents the high-precision position estimate after fusion, represents the number of sensors, Indicates The dynamic weight of each sensor reflects the credibility of the sensor data and is dynamically adjusted according to sensor errors and environmental conditions. Indicates Sensors at time The measured value of Indicates the time at which the drive motor The drive motor speed, Indicates the time at which the drive motor The drive motor speed, Represents the proportional control gain matrix, which is used to adjust the fine-tuning intensity. represents the desired target position, Represents the compensation coefficient, which is dynamically adjusted according to task requirements. represents the actual position uncertainty standard deviation, represents the expected position uncertainty standard deviation, which is used to compensate for position error; Based on the optimized motion instructions, the adaptive control theory and predictive control algorithm are applied to calculate and compensate the factors affecting the actual motion trajectory in advance, so that the micro robot can move smoothly along the ideal motion trajectory, and obtain the trajectory tracking data of the micro robot, which is calculated by the following formula: , , in, Indicates at time The adaptive controller parameters are Indicates at time The adaptive controller parameters are Represents the learning rate, which controls the step size of each update. Represents the gradient of the loss function with respect to the parameters, which is used to guide the direction of parameter update. Represents the evolution factor, which is dynamically adjusted according to the task difficulty and environment complexity. represents the change in the best hyperparameter combination searched by the evolutionary algorithm, Indicates at time The amount of compensation, Indicates at time The amount of compensation, Indicates the compensation coefficient, controls the compensation strength, represents the predicted interference force, represents the actual interference force encountered, represents an exponential factor, which is dynamically adjusted according to the complexity of the environment to enhance the focus on key interference; Based on the trajectory tracking data, combined with the target tracking technology in deep learning, the route of the micro robot is continuously monitored and corrected to make the route meet the task requirements; wherein, the target tracking update and the incremental learning update are calculated by the following formula: , , in, Indicates the target tracking update at time The tracking state vector, Indicates at time The tracking state vector, represents the Kalman gain, which is used to adjust the correction strength of the tracking error. represents the observed value, represents the observation model, Represents the adaptive adjustment coefficient, which is dynamically adjusted according to the tracking error to map the state vector to the observation space. represents an exponential decay factor, which is used to balance the effects of long-term and short-term tracking errors. represents the tracking error at the current moment, Incremental learning updates at time The behavioral rule weight, Indicates at time The behavioral rule weight, Represents the learning rate, which controls the step size of each update. Represents the gradient of the loss function with respect to the weights, based on the current dataset calculate, represents the incremental learning adjustment coefficient, which is dynamically adjusted according to the progress of the task. represents the standard deviation of the current behavior rule error, Represents the standard deviation of the reference behavior rule error, which is used to optimize the learning process.
10. A micro robot motion control system, characterized in that: include: A collection and processing module, used to collect current environmental status information in real time using a non-contact sensor, and synchronize and calibrate the environmental status information to generate a comprehensive perception image; A classification prediction module, used to apply target detection and recognition technology to analyze the comprehensive perception image to determine the objects existing in the current environment and the categories of the objects, obtain classification results, and build an environment state change prediction model based on the classification results and the environment state information. The environment state change prediction model is used to predict the environment state changes that affect the motion trajectory of the micro robot in the future period of time through historical environment state information and the current environment state information, combined with the classification results; An optimization processing module, used to generate candidate paths using a fast random tree according to the comprehensive perception image and the environmental state changes predicted by the environmental state change prediction model, and optimize the candidate paths in combination with an AI search algorithm to obtain a target path; An update generation module is used to construct a three-dimensional dynamic probability map based on the target path using a Bayesian network, wherein the three-dimensional dynamic probability map uses the environmental state at a specific location as a node and the transition probability between the environmental states as an edge, and continuously updates the three-dimensional dynamic probability map using a reinforcement learning algorithm to generate an instant decision map; A simulation adjustment module is used to adjust the motion trajectory of the micro robot when a change in the environmental state is detected, based on the instant decision map and introducing an emotional computing model to simulate emotional factors in the human decision-making process, so as to generate an adaptive behavior rule set; The position adjustment module is used to use a Kalman filter to estimate the position of the micro robot, obtain the position state of the micro robot, and adjust the speed and steering angle of the drive motor of the micro robot based on the position state and the adaptive behavior rule set to ensure that the actual motion trajectory conforms to the ideal motion trajectory, and generate the trajectory tracking of the micro robot. The trajectory tracking refers to the process in which the micro robot moves according to a predetermined motion trajectory during the execution of a task, and adjusts its own speed and steering angle in real time to minimize the deviation between the actual motion trajectory and the ideal motion trajectory.
Citation Information
Patent Citations
Automatic driving system based on 5G vehicle-road cooperation
CN116811916A
Real-time motion planning system based on dynamic environment perception
CN119146965A