Control system and control method of programming educational robot based on artificial intelligence
By adopting multi-sensor components and deep learning technology in programming education robots, combining multi-modal data fusion and deep reinforcement learning, the robots are able to make accurate decisions and real-time adaptation in complex environments, solving the problems of insufficient flexibility, lack of adaptability and insufficient interactivity in the existing technology, and significantly improving task execution efficiency and students' learning effect.
Patent Information
- Application Number
- CN202510181391.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-02-19
AI Technical Summary
The performance of existing programming education robots in dealing with complex tasks and dynamic environments is limited, lacking real-time learning and adaptability, insufficient interactivity, and lack of personalized design, which affects students' learning effects and motivation.
The programming education robot control system based on artificial intelligence is adopted, including multi-sensor components, deep learning and feature extraction module, multi-modal data fusion module, robot control and behavior decision-making module, path planning and obstacle avoidance module, and adaptive optimization module, to achieve accurate decision-making and real-time adaptation of robots through deep reinforcement learning and multi-modal data fusion.
It significantly improves the efficiency and accuracy of the task execution of robots in complex environments, enhances the adaptability and robustness of robots, improves the interactivity and fun of programming learning, and improves students' learning effects through personalized learning paths.
Smart Images

Figure CN119635687B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of programming educational robots, and in particular to a programming educational robot control system and a control method based on artificial intelligence. Background Art
[0002] At present, programming educational robots are widely used in the field of STEM education, especially in helping students learn programming and robot control. Existing programming educational robot systems usually enable students to directly control robots to perform simple tasks through graphical programming or text programming. Graphical programming interfaces (such as Blockly or Scratch) enable beginners to intuitively understand programming logic, while text programming provides a more flexible programming experience for students with a certain programming foundation. Most of these systems rely on fixed tasks and preset control logic. Students design tasks through programming and let the robot execute the corresponding instructions.
[0003] However, existing technologies still have obvious limitations in handling complex tasks and dynamic environments. First, most traditional systems lack real-time learning and adaptation capabilities. Robots only follow predetermined paths and instructions and cannot make intelligent decisions or adjustments based on environmental changes, which limits the flexibility and autonomy of robots. Secondly, path planning and obstacle avoidance functions are usually based on simple rules and fixed path designs. When the environment changes, the robot cannot adjust the path in real time, resulting in low efficiency in task execution. In addition, although existing programming education platforms can provide certain programming learning tools, they are not interactive enough, students cannot get timely feedback and guidance, and it is difficult to quickly identify and correct programming errors. Finally, programming tasks and teaching content usually lack personalized design, and the difficulty of tasks cannot be flexibly adjusted according to students' learning progress and abilities, which affects students' learning effects and motivation. Summary of the invention
[0004] In view of the deficiencies in the prior art, the present invention provides an artificial intelligence-based programming education robot control system and control method, which solves the problems in the prior art of insufficient flexibility, lack of real-time adaptability, insufficient feedback and interactivity, and lack of personalized learning paths of programming education robots in complex environments.
[0005] To achieve the above objectives, the present invention is implemented through the following technical solutions: a control system and control method of a programming educational robot based on artificial intelligence, comprising:
[0006] The hardware module includes a plurality of sensor components for collecting environmental data, wherein the sensors include:
[0007] Camera: used to collect RGB images and depth images of the environment;
[0008] LiDAR: used to generate 3D point cloud data of the environment;
[0009] Ultrasonic sensor: used for close-range obstacle detection;
[0010] Infrared sensor: used for obstacle detection in low-light or dark environments;
[0011] The sensor data processing module is used to pre-process the environmental data collected by the sensor, remove noise and perform feature extraction;
[0012] Deep learning and feature extraction module, which is used to further process the data features from different sensors using deep neural networks to extract important environmental perception features;
[0013] Multimodal data fusion module, which is used to intelligently fuse feature data from different sensors to generate more accurate and comprehensive environmental perception information;
[0014] The robot control and behavior decision module uses a deep reinforcement learning algorithm to make decisions based on the integrated environmental perception information and control the robot to perform tasks;
[0015] The path planning and obstacle avoidance module uses a path planning algorithm to generate the optimal path based on environmental perception information and target tasks, and avoids obstacles in real time during the path execution process;
[0016] An adaptive optimization module that continuously adjusts decision-making strategies and control algorithms based on feedback from the robot’s task execution;
[0017] A programming education platform that provides students with a programming interface, allowing them to control robots to perform tasks through graphical programming or text programming, and provides real-time feedback and scoring based on the results of task execution.
[0018] Preferably, the deep learning and feature extraction module includes:
[0019] Extracting environmental features from the image data using a convolutional neural network, wherein the convolutional neural network is trained to extract deep features related to obstacles, targets, and terrain;
[0020] The PointNet++ point cloud processing network is used to extract features from LiDAR point cloud data to obtain more accurate spatial perception information.
[0021] Preferably, the multimodal data fusion module adopts a self-attention mechanism or a Transformer network to perform weighted fusion of features from different sensors, combining time series information and spatial information.
[0022] Preferably, the robot control and behavior decision module uses a deep reinforcement learning algorithm to select and execute corresponding control actions based on the fused environmental perception information and task requirements, and continuously optimizes the decision-making strategy through a reward feedback mechanism.
[0023] Preferably, the deep reinforcement learning algorithm adopts deep Q learning or proximal strategy optimization algorithm, and rewards and punishes according to the actual effect of the robot in performing tasks, thereby optimizing the robot's behavioral decision-making ability.
[0024] Preferably, the path planning and obstacle avoidance module adopts the A* algorithm or the D*Lite algorithm, and dynamically adjusts the path planning in combination with real-time environmental information, so that the robot can efficiently avoid obstacles and complete the task objectives when performing tasks.
[0025] A control method for a programming educational robot based on artificial intelligence, comprising:
[0026] S1. Collect environmental data through sensors, including RGB images, depth images, LiDAR point cloud data, ultrasonic sensor data, and infrared sensor data;
[0027] S2. Preprocessing the environmental data and performing feature extraction on the preprocessed data, wherein the feature extraction includes extracting features from the image data using a convolutional neural network and extracting features from the LiDAR data using a point cloud processing network;
[0028] S3, using multimodal data fusion technology to intelligently fuse feature information from different sensors to generate fused environmental perception information;
[0029] S4. Based on the environmental perception information, use a deep reinforcement learning algorithm to make robot behavior decisions, select appropriate actions and control the robot to execute;
[0030] S5. Based on the current state and target position of the robot, the path planning algorithm is used to generate the optimal path, and the obstacle avoidance algorithm is used to ensure the safety of the robot during the path execution;
[0031] S6. Based on the feedback from task execution, use adaptive optimization algorithms to adjust decision strategies and control models.
[0032] Preferably, the deep reinforcement learning algorithm is optimized based on the value functions of states and actions, evaluates the pros and cons of different actions through Q-value functions or policy functions, and updates the policy based on real-time feedback.
[0033] Preferably, the path planning algorithm includes an A* algorithm or a D*Lite algorithm, which generates a shortest path by evaluating the cost from the current position to the target position and adjusts the path in real time to avoid obstacles.
[0034] Preferably, the adaptive optimization module adjusts the robot control strategy, reward function and action strategy based on the task execution results through online learning and transfer learning technology to adapt to new tasks and environments.
[0035] The present invention provides a control system and a control method for a programming educational robot based on artificial intelligence, which has the following beneficial effects:
[0036] 1. The present invention adopts deep reinforcement learning and multimodal data fusion technology to achieve accurate decision-making and real-time obstacle avoidance of robots in complex environments. Through deep learning and real-time optimization, the robot can dynamically adjust its strategy to cope with different situations. Compared with the fixed strategy and preset path method in the prior art, the present invention enables the robot to self-learn and optimize according to real-time feedback, thereby significantly improving the efficiency and accuracy of task execution.
[0037] 2. The present invention uses an adaptive optimization module to enable the robot to automatically adjust its behavior strategy according to different environments and task requirements. This flexible optimization mechanism avoids the limitations of static control in traditional programming systems, allowing the robot to quickly adapt to changing environments and task requirements. Compared with the prior art, traditional systems are often unable to cope with dynamically changing complex environments. The present invention solves this shortcoming and improves the adaptability and robustness of the robot.
[0038] 3. The present invention combines graphical programming and text programming to help students of different age groups learn programming in a way that suits them. Students can easily get started through graphical programming, and can gradually deepen their programming knowledge through text programming. Compared with existing traditional programming education methods, the present invention greatly reduces the difficulty of getting started with programming, while improving the interactivity and fun of the learning process.
[0039] 4. The present invention uses a task and feedback module to provide real-time feedback of the robot's task execution results to students, and evaluates students' programming tasks through a scoring system. This real-time feedback and dynamic evaluation mechanism helps students to understand problems in task execution in a timely manner and optimize programming ideas. Compared with the single teaching model in the prior art, the present invention ensures students' continuous progress and goal achievement in the programming learning process through precise feedback and scoring mechanisms. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 It is a system architecture diagram of the present invention. DETAILED DESCRIPTION
[0041] The following will be combined with the drawings in the specification of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0042] Please see attached Figure 1 The embodiment of the present invention provides a control system and a control method of a programming education robot based on artificial intelligence, including:
[0043] The hardware module includes multiple sensor components for collecting environmental data, including: a camera for collecting RGB images and depth images of the environment; a lidar for generating three-dimensional point cloud data of the environment; an ultrasonic sensor for close-range obstacle detection; an infrared sensor for obstacle detection in low-light or dark environments;
[0044] Cameras are mainly used to acquire image data, LiDAR provides high-precision spatial data, and ultrasonic and infrared sensors are used to detect close-range obstacles, especially in low-light environments.
[0045] Computing Platform:
[0046] In this embodiment, the computing platform uses the embedded computing module NVIDIA Jetson AGX Xavier, which has powerful processing capabilities and can support real-time reasoning of deep learning algorithms. The NVIDIA Jetson AGX Xavier platform supports multi-tasking and can simultaneously perform data acquisition, preprocessing, and reasoning of deep learning models from cameras, LiDAR, ultrasonic and infrared sensors. Its built-in GPU accelerated computing unit performs particularly well in processing computationally intensive tasks and can support the operation of real-time control systems. The computing platform also needs to communicate effectively with other hardware modules to transmit sensor data and control instructions.
[0047] Actuator:
[0048] The actuator of the robot body includes the drive system and the servo system. The design goal of the actuator is to ensure the robot's flexible movement and precise execution of tasks in complex environments.
[0049] The drive system uses omni-wheels to support the robot's free movement in two-dimensional space. Omni-wheels allow the robot to move in any direction without being restricted by the direction of traditional wheel systems. The motor drive module adjusts the speed and direction of the omni-wheels, allowing the robot to be precisely controlled in complex environments.
[0050] The actuators work closely with the computing platform, which drives the motors and servos by outputting control instructions to ensure that the robot's behavior is consistent with task requirements.
[0051] Communication module:
[0052] In order to achieve efficient collaboration between various hardware modules, the communication module plays a vital role in the hardware system. The role of the communication module is to transmit the environmental data collected by the sensor to the computing platform and send the control instructions generated by the computing platform to the actuator.
[0053] The communication system usually includes a wireless communication module, such as a Wi-Fi or Bluetooth module, to ensure that the robot can exchange data with external devices (such as programming education platforms, servers, etc.). Through wireless communication, the robot can upload perception data in real time and receive external control instructions and task updates. In the education platform, students set tasks through the programming interface, and the robot interacts with the platform through the communication module to perform the corresponding programming tasks.
[0054] Sensor data processing and interaction:
[0055] The image data collected by the camera will be processed by image enhancement, denoising, color correction, etc., and then input into the convolutional neural network (CNN) for deep feature extraction. The point cloud data generated by LiDAR is cleaned by a spatial filter and then enters a point cloud processing network such as PointNet++ to extract spatial features such as object shape and position. The data from the ultrasonic and infrared sensors will be passed to the control module after distance calculation for real-time obstacle avoidance.
[0056] These processed data become the robot's environmental perception information, which directly affects the output of the robot's control and behavior decision modules. Through the high integration and processing of sensor data, the robot can adapt to complex environmental changes in real time and perform precise tasks.
[0057] The sensor data processing module is used to pre-process the environmental data collected by the sensor, remove noise and perform feature extraction to facilitate subsequent deep learning processing;
[0058] In this embodiment, the sensor data processing module mainly includes the following steps: data reception and preprocessing, feature extraction, data standardization and format conversion. Each step is optimized for different types of data so that the subsequent deep learning module can effectively process and extract environmental features. The following content will describe these technical steps in detail and illustrate them through relevant formulas and algorithms.
[0059] Data reception and preprocessing:
[0060] Data reception is the first step of the sensor data processing module, which refers to receiving the data from each sensor to the processing unit. In general, the data from the sensor may contain noise or incomplete information and needs to be cleaned and supplemented. Data preprocessing is a work to prepare for subsequent feature extraction and deep learning processing to ensure the quality of input data.
[0061] In this embodiment, after the image data is acquired from the camera, it is first subjected to denoising. Denoising uses traditional image processing algorithms, such as Gaussian filtering or median filtering, to remove noise caused by environmental factors (such as unstable lighting or sensor noise). Subsequently, the image data is color corrected to ensure that the image acquired by the camera conforms to the actual environmental color and can be correctly recognized by the deep learning model.
[0062] For LiDAR data, its point cloud data usually contains many irrelevant background points or erroneous data caused by noise. In some embodiments, spatial filtering algorithms (such as statistical filtering or Voxel filtering) are used to remove these invalid points. This process helps to reduce the amount of calculation and improve the accuracy of point cloud data.
[0063] Ultrasonic sensor data and infrared sensor data are usually used for close-range obstacle detection. The distance information they provide is affected by the environment, such as multipath effects or interference signals. Therefore, the data processing module first applies multiple sampling technology to smooth the data and improves the accuracy through the distance correction algorithm.
[0064] Feature extraction:
[0065] After data preprocessing is completed, the feature extraction module begins to conduct in-depth analysis of different types of sensor data. In some embodiments, after preprocessing, the image data will enter the convolutional neural network (CNN) for feature extraction. CNN extracts high-level features in the image layer by layer through multi-layer convolution operations, such as edges, textures, shapes, etc., and uses these features for subsequent obstacle detection and target recognition tasks.
[0066] For LiDAR data, feature extraction mainly relies on point cloud processing networks such as PointNet++. The point cloud data generated by LiDAR sensors is highly sparse, so it needs to be processed by an appropriate network structure. The data from ultrasonic and infrared sensors, despite their lower resolution, are critical for detecting close-range obstacles. Generally, after smoothing, the data from ultrasonic sensors can be used to extract obstacle information using a threshold-based method. Infrared sensors infer the presence of objects by measuring light of different reflected intensities, and feature extraction methods usually determine whether an object is within the detection range by setting a threshold.
[0067] Data standardization and format conversion:
[0068] For data from different sensors, their formats and units are usually inconsistent, so data standardization is required. For example, the data of camera images are usually in pixels, while LiDAR data is represented in three-dimensional space coordinates. Therefore, data standardization includes not only unit conversion, but also mapping data from different sources into a unified feature space.
[0069] Specifically, the data standardization step in this embodiment includes:
[0070] Normalization of image data: First, resize the image data to meet the input requirements of the deep learning model. Generally, the image size is scaled to a fixed size (such as 224×224 pixels), and the pixel values are normalized to limit their range to between 0 and 1.
[0071] Standardization of LiDAR data: Point cloud data usually needs to be scaled and transformed to map the raw point cloud data into a unified coordinate system so that it can be effectively connected with data from other sensors.
[0072] Normalization of ultrasonic and infrared sensor data: This type of data is usually based on distance values, and data normalization is required to convert the distance values into standardized feature values so that they can be fused with other sensor data.
[0073] Data fusion preparation:
[0074] After the sensor data processing module completes data preprocessing and feature extraction, the various sensor data obtained will be passed to the multimodal data fusion module. At this point, the processed data already has a unified format and corresponding features, which can provide effective input for the subsequent deep learning model.
[0075] Generally, data from different sensors (such as images, LiDAR, ultrasound, infrared, etc.) will be intelligently weighted and fused in the multimodal data fusion module to generate more accurate environmental perception information. The data fusion process uses a self-attention mechanism, which can help the robot adjust the weight of each sensor data according to the current task requirements. For example, in a low-light environment, the weight of the infrared sensor may be enhanced, while in an indoor environment, the weight of the camera may be larger.
[0076] Mathematical expression of feature extraction and fusion:
[0077] In some embodiments, the mathematical expression of the feature extraction and fusion process can be as follows:
[0078] Formula 1: Feature extraction of image data
[0079] ;
[0080] in, is the input image data, is the height of the image, that is, the number of pixels in the vertical direction of the image, is the width of the image, that is, the number of pixels in the horizontal direction of the image, is the number of channels, usually 3 for color images (RGB) and 1 for grayscale images. is the extracted image feature, is the feature dimension.
[0081] Formula 2: Feature extraction from LiDAR data
[0082] ;
[0083] in, It is the point cloud data collected by the LiDAR sensor. is the number of points in the point cloud, and 3 is the spatial coordinate. Indicates the use Model for LiDAR point cloud data Processing and feature extraction. is the extracted LiDAR feature, To extract the dimension of the feature, it is usually represented as a high-dimensional feature vector.
[0084] Formula 3: Multimodal data fusion
[0085] ;
[0086] in, For the The characteristics of the sensor, To extract the dimension of features, For sensor The corresponding weight represents the weight coefficient of the sensor, weight The contribution of each sensor in multimodal fusion can be adjusted dynamically through self-attention mechanism. is the total number of sensors, indicating how many sensors provide data for fusion. It is the fused feature vector, which combines the features of all sensors.
[0087] In this embodiment, the sensor data processing module ensures that the subsequent modules can use high-quality input data for decision-making and control by accurately preprocessing, feature extracting, standardizing and formatting the data from each sensor. The multi-step processing of data preprocessing, feature extraction and standardization ensures that sensor data from different sources can be seamlessly integrated, providing a solid foundation for deep learning and robot control. Through the refined processing of data, the robot can achieve more accurate environmental perception and task execution, further improving the efficiency and reliability of the programming education robot.
[0088] Deep learning and feature extraction module, which is used to further process the data features from different sensors using deep neural networks to extract important environmental perception features;
[0089] In this embodiment, the deep learning and feature extraction module mainly extracts features from sensor data through convolutional neural networks (CNN) and point cloud processing networks (such as PointNet++). Specifically, images collected by the camera, point cloud data generated by the lidar, and data from ultrasonic and infrared sensors are all input into the deep learning model for processing to extract feature information that is critical to the execution of subsequent tasks.
[0090] Data input and application of deep learning models:
[0091] In general, data from different sensors will first be preprocessed and standardized to ensure data consistency and quality. Then, these data will be passed as input to the deep learning and feature extraction module. In an embodiment of the present invention, the processing process includes convolution operations on image data and spatial feature learning of point cloud data.
[0092] Image data processing:
[0093] In this embodiment, the image data is processed by a convolutional neural network (CNN). CNN can automatically extract low-level features (such as edges and corners) and high-level features (such as the shape and texture of objects) from images. After the image data passes through the convolution layer, pooling layer, and fully connected layer of CNN, a high-level representation of the environment is finally obtained.
[0094] Specifically, deep convolutional neural networks such as ResNet-101 are used to extract image features from camera data. In this implementation, the input image data (in is the height of the image, is the width of the image, is the number of channels) will undergo a convolution operation to extract discriminative features. The output features of the image can be expressed as:
[0095] ;
[0096] in, is the extracted image features, is the dimension of the feature, representing the high-level semantic information in the image.
[0097] Processing of LiDAR data:
[0098] For point cloud data from LiDAR, this embodiment uses a point cloud processing network such as PointNet++ to extract features. LiDAR generates point cloud data by scanning the surrounding environment. (in is the number of points in the point cloud, each point is represented by three coordinates (x, y, z)) will pass through the PointNet++ network, which can effectively capture the geometric features in the point cloud, especially in irregular and sparse point cloud data, and can accurately extract important information.
[0099] Data processing of ultrasonic and infrared sensors:
[0100] For ultrasonic and infrared sensor data, although these sensors provide low spatial resolution, they are essential for detecting close-range obstacles. Generally, the sensor data is first denoised and filtered, and then features are extracted through rule-based algorithms. For ultrasonic sensors, the feature extraction process usually determines whether there is an obstacle based on the sensor's distance measurement and the obstacle's reflection intensity.
[0101] Training and optimization of deep learning models:
[0102] In some embodiments, the deep learning and feature extraction module is not only used for feature extraction, but also includes model training and optimization. By training the convolutional neural network and point cloud processing network with a large amount of environmental data, the deep learning model can gradually improve its ability to perceive the environment.
[0103] Formula 4: Loss function optimization
[0104] ;
[0105] in, is the loss function, are model parameters, is the input data (such as image data or point cloud data), is the true label, is the prediction result of the model, is the number of samples. By continuously optimizing the loss function, the trained model can better adapt to different environments and task requirements.
[0106] Feature preparation before data fusion:
[0107] After being processed by the deep learning and feature extraction modules, the data from each sensor has been converted into features with highly abstract information. , LiDAR features The features of the sensors and other sensors are uniformly formatted and ready for use by the multimodal data fusion module.
[0108] In some embodiments, data fusion adopts a weighted average or attention-based strategy, which enables the robot to automatically adjust the weight of each sensor data as needed in different tasks. For example, in an environment with strong light interference, the weight of image data may be lower, while LiDAR data provides more reliable spatial information.
[0109] In this way, the deep learning and feature extraction module not only realizes the feature extraction of different sensor data, but also provides important support for data fusion through intelligent algorithms.
[0110] In this embodiment, the deep learning and feature extraction module successfully converts the data of different sensors into high-level features that can be understood by the machine through advanced deep learning methods such as convolutional neural networks (CNN) and PointNet++. These features provide accurate information support for subsequent multimodal data fusion and robot control decision-making. Through the training and optimization of the deep learning model, the module can make effective judgments in different environments, improving the robot's perception and decision-making capabilities in complex environments.
[0111] Multimodal data fusion module, which is used to intelligently fuse feature data from different sensors to generate more accurate and comprehensive environmental perception information;
[0112] In this embodiment, the multimodal data fusion module fuses data from different sensors through algorithms such as weighted averaging and self-attention mechanisms (such as Transformer) to achieve intelligent environmental perception. Specifically, the module not only directly adds the features extracted by each sensor, but also automatically determines the importance of each sensor in a specific environment through intelligent weight allocation and fusion strategies, and dynamically adjusts according to task requirements.
[0113] Data input and reception of preprocessing results:
[0114] In this embodiment, the image features and LiDAR point cloud features It is obtained after deep learning and feature extraction module processing. The features of each sensor data are extracted in a separate network model and standardized so that the sensor data can be compatible and effectively fused.
[0115] Fusion method: self-attention mechanism and weighted fusion:
[0116] Output and subsequent processing of fusion features:
[0117] The fused features The fused features are fed into the subsequent control and decision modules to provide the robot with comprehensive environmental perception. Specifically, these fused features provide accurate information support for the subsequent robot behavior decision, path planning and obstacle avoidance modules. The high accuracy and robustness of the fused features ensure that the robot can perform tasks in complex and dynamic environments.
[0118] Optimization and dynamic adjustment of data fusion:
[0119] In some complex environments, the robot needs to adjust the weights of each sensor in real time. For this situation, the present invention adopts a dynamic adjustment mechanism based on feedback. This mechanism dynamically adjusts the sensor weights according to the real-time feedback of the robot during the task execution. For example, in a low-light environment, the weight of the image sensor will be reduced, while the weight of the infrared sensor will be increased accordingly to ensure that the detection of obstacles will not be affected.
[0120] In this way, the robot can flexibly adjust the weight of the sensors according to different environmental conditions, thereby improving the accuracy of environmental perception. This adjustment mechanism is not limited to the process of sensor data fusion, but also applies to the dynamic optimization of control decisions, further enhancing the robot's adaptive ability.
[0121] In this embodiment, the multimodal data fusion module intelligently fuses feature data from different sensors through the self-attention mechanism and weighted average method. This process enables the robot to fully and accurately perceive the surrounding environment and provide reliable information support for subsequent behavioral decisions and task execution. The application of the self-attention mechanism enables the system to dynamically adjust the weight of the data according to environmental conditions and task requirements, thereby ensuring the effectiveness and accuracy of data fusion under different circumstances. Through this module, the robot demonstrates strong perception and decision-making capabilities in complex and dynamic environments, providing technical support for the realization of efficient programming education robots.
[0122] The robot control and behavior decision module uses a deep reinforcement learning algorithm to make decisions based on the integrated environmental perception information and control the robot to perform tasks;
[0123] In this embodiment, the robot control and behavior decision module combines a variety of technical means, especially the deep reinforcement learning (DRL) algorithm, to optimize the robot's behavior decision. Deep reinforcement learning allows the robot to adjust its behavior strategy based on real-time feedback in a complex environment, thereby achieving autonomous learning and task optimization. Specifically, the robot uses the deep reinforcement learning model to select the action that best suits the current situation based on the environmental perception information obtained through the multimodal data fusion module.
[0124] Input data and module connection:
[0125] In general, the input of the robot control and behavior decision module is the fused feature data from the previous module. This data has been processed by the deep learning model, and after multimodal data fusion, it can fully express the overall picture of the environment. The data input includes image features , LiDAR point cloud features , and other related features of sensors. These fused environmental perception information not only includes the spatial layout around the robot, but also involves dynamic information such as the type, position, and speed of obstacles.
[0126] Once these feature data are processed by the aforementioned multimodal data fusion module, the robot control and behavior decision module will make behavioral decisions based on these data using reinforcement learning strategies to control the robot to perform corresponding tasks, such as obstacle avoidance and path planning.
[0127] Applications of Reinforcement Learning Algorithms:
[0128] In this embodiment, the robot control and behavior decision module uses a deep reinforcement learning (DRL) algorithm to optimize the robot's behavior strategy through the interaction between the agent and the environment. Specifically, when performing a task, the robot will select an action (such as moving forward, turning, stopping, etc.) based on the current state (for example, the current environment and position of the robot). The behavior selection is completed through Q learning or proximal policy optimization (PPO) algorithm.
[0129] Specifically, the deep Q learning (DQN) algorithm, as a common method of reinforcement learning, can and actions taken To estimate its future return value, and continuously adjust its behavior strategy based on this return value. The Q value update formula is as follows:
[0130] ;
[0131] in: :This function represents the time The robot is in state When you select Action The Q-value of the action is the expected future reward of the action in this state.
[0132] :Robot at all times The state it is in. This is usually a description of the environment, which may include the robot's position, speed, direction, and input from other sensors (such as image features, LiDAR point cloud features, etc.).
[0133] : At time The action selected by the robot. It can be a behavior performed by the robot in the environment, such as moving forward, turning, stopping, etc.
[0134] : Learning rate, usually in The range indicates the weight of the newly acquired experience in updating the Q value. The higher the learning rate, the greater the impact of new information on the update of the Q value.
[0135] :The robot is performing actions The reward obtained from the environment after the task is completed (i.e., immediate reward). This reward can be positive (such as successfully completing a task) or negative (such as encountering an obstacle).
[0136] : Discount factor, usually in In the range, it indicates the degree of discount on future rewards. A larger value means the robot places more emphasis on future rewards, while a smaller value means a greater preference for short-term rewards.
[0137] :The robot performs an action After that, the next state.
[0138] :This part indicates that in the next state All possible actions The action with the largest Q value in the game. It can also be understood as the expected reward of the robot's optimal behavior in the next state.
[0139] In addition, the Proximal Policy Optimization (PPO) algorithm uses the policy gradient method in the optimization process to directly optimize the probability distribution of the robot's action selection. Specifically, PPO measures the relative value of an action by calculating the advantage function, and then adjusts the strategy so that the robot can continuously improve the effectiveness of its behavioral decision-making when performing tasks.
[0140] Feedback mechanism for behavioral decision-making:
[0141] Under this feedback mechanism, algorithms such as Q learning and PPO enable robots to make increasingly accurate decisions in various tasks through repeated learning and updating, thereby improving the efficiency and accuracy of robots in performing tasks.
[0142] Path planning and obstacle avoidance decision:
[0143] Based on the environmental perception information provided by the aforementioned multimodal data fusion module, the robot control and behavior decision module is not only responsible for selecting behavioral actions, but also involves complex path planning and obstacle avoidance decisions. Path planning is the basis for the robot to complete the task, which determines the optimal path for the robot to reach the target from the starting point.
[0144] In this embodiment, the robot control and behavior decision module performs path planning using the A* algorithm or the D*Lite algorithm. The A* algorithm is based on a heuristic search strategy and calculates the cost of each candidate path to find the shortest path from the starting point to the target. The total cost function of the A* algorithm is It is expressed by the following formula:
[0145] ;
[0146] in: Is a node the total cost of is the cost from the starting point to the current node; is the estimated cost from the current node to the target node.
[0147] Behavioral decision-making and actual execution:
[0148] Finally, the robot control and behavior decision module issues specific control commands through the actuators based on the control instructions and path planning results generated by the deep reinforcement learning algorithm to guide the robot's behavior. For example, if it decides to move forward, the control system will send instructions to the drive motor; if it decides to turn, the control system will adjust the direction of the robot chassis.
[0149] During the execution process, the control module continuously obtains new information from sensor feedback and adjusts the robot's behavior in real time according to changes in the environment. Through this real-time feedback mechanism, the robot can efficiently perform tasks in complex and dynamic environments.
[0150] In this embodiment, the robot control and behavior decision module can make accurate decisions based on the environmental perception information provided by the aforementioned multimodal data fusion module through deep reinforcement learning and path planning technology. Through reinforcement learning algorithms such as deep Q learning or PPO, the robot continuously optimizes the decision-making strategy to improve the efficiency and accuracy of task execution. At the same time, combined with the A* algorithm or D*Lite algorithm for path planning, it is ensured that the robot can avoid obstacles in real time and select the optimal path when performing tasks. This module is the core component of the present invention and provides technical support for the robot to efficiently perform tasks in complex environments.
[0151] The path planning and obstacle avoidance module uses a path planning algorithm to generate the optimal path based on environmental perception information and target tasks, and avoids obstacles in real time during the path execution process;
[0152] In this embodiment, the path planning and obstacle avoidance module not only relies on traditional path planning algorithms, such as the A* algorithm or the D*Lite algorithm, but also combines intelligent decision-making from deep learning and reinforcement learning, and can adjust the path planning according to real-time feedback to ensure the safety of the robot and efficient execution of the task.
[0153] Environmental perception and reception of input data:
[0154] Generally, the input of the path planning and obstacle avoidance module comes from the aforementioned multimodal data fusion module. The fusion features provided by this module It includes environmental perception information from different sensors (such as cameras, LiDAR, ultrasonic sensors, etc.). This information describes the robot's current position, target position, and surrounding obstacles and dynamic environment. After processing, the environmental information provided by the sensor forms a unified spatial description, providing an accurate data basis for path planning and obstacle avoidance decisions.
[0155] Specifically, the fusion features It contains information such as the spatial distribution of obstacles, the type and shape of obstacles, and the target location. These data will be passed as input to the path planning and obstacle avoidance module to guide subsequent path planning and real-time obstacle avoidance decisions.
[0156] Application of path planning algorithm:
[0157] The core task of path planning is to generate the optimal path from the starting point to the target point. In this embodiment, the path planning adopts the A* algorithm or the D*Lite algorithm, wherein the A* algorithm is a classic heuristic search algorithm that finds the shortest path from the starting point to the target by calculating the total cost of each node. The D*Lite algorithm, as a dynamic path planning algorithm, is suitable for situations where there are real-time changes in the environment (such as obstacle movement, path blockage, etc.). Unlike the A* algorithm, the D*Lite algorithm can update the path in real time during the path execution process to ensure that the robot can respond quickly when the environment changes. By maintaining the cost of a set of nodes and adjusting the path based on real-time feedback, the algorithm can dynamically adjust the path planning during execution to avoid new obstacles.
[0158] Real-time obstacle avoidance decision-making:
[0159] On the basis of path planning, real-time obstacle avoidance is an indispensable part of robot task execution. Generally, when executing path planning, the robot may encounter sudden obstacles or environmental changes, causing the original planned path to no longer apply. Therefore, the path planning and obstacle avoidance module is not only responsible for generating the initial path, but also needs to make dynamic adjustments based on real-time sensor data and feedback mechanisms.
[0160] In this embodiment, the obstacle avoidance decision is combined with a deep reinforcement learning (DRL) algorithm. When the robot performs a task, real-time sensor data (such as the appearance of obstacles, distance changes, etc.) will be used as input to drive the reinforcement learning model to adjust the robot's behavior. For example, the robot may encounter a sudden obstacle during the execution of path planning. At this time, the control system will select the most appropriate action (such as turning, deceleration, backing up, etc.) based on the current state feedback (such as the distance and size of the obstacle) and update the path in real time.
[0161] Q-learning or Proximal Policy Optimization (PPO) algorithms in reinforcement learning can enable the robot to adapt to different obstacle avoidance scenarios through repeated trials and adjustments during the training process. Specifically, the robot and actions taken To estimate its future returns and continuously optimize decision-making strategies based on the returns.
[0162] Through reinforcement learning, the robot can optimize its obstacle avoidance strategy according to real-time changes in the environment to maximize the success rate of task completion.
[0163] Dynamic path adjustment and adaptive optimization:
[0164] In this embodiment, the path planning and obstacle avoidance module performs path planning by integrating the A* algorithm or the D*Lite algorithm, and combines the deep reinforcement learning algorithm to make real-time obstacle avoidance decisions. This module can generate efficient path planning based on the environmental perception information provided by the multimodal data fusion module, and adjust the path in real time to avoid obstacles during the task execution. Through the feedback mechanism of reinforcement learning, the robot can adaptively adjust the strategy in a constantly changing environment to ensure the smooth completion of the task. The implementation of this module provides powerful navigation and task execution capabilities for programming education robots.
[0165] Adaptive optimization module, which continuously adjusts decision-making strategies and control algorithms based on feedback from the robot's task execution to improve the robot's performance in complex environments;
[0166] In this embodiment, the adaptive optimization module is optimized through the following two main mechanisms: online learning and transfer learning.
[0167] Online learning enables the robot to continuously accumulate experience and adjust its behavior strategy according to real-time environmental changes while performing tasks. Specifically, after each task execution step, the robot will update its decision model based on the completion of the task (such as successful obstacle avoidance, reaching the target, etc.). Whenever the robot successfully completes a task, it will receive a positive reward; if the task fails or encounters an obstacle, it will receive a negative reward. This reward feedback will affect the adjustment strategy of the adaptive optimization module, thereby optimizing the robot's behavioral decision.
[0168] Transfer learning enables robots to quickly apply what they have learned from one task to other similar tasks.
[0169] Optimization objectives and adjustment mechanisms:
[0170] In some embodiments, the optimization objectives of the adaptive optimization module mainly include:
[0171] Task success rate: Ensure that the robot can complete the specified task efficiently and avoid unnecessary failures or erroneous behaviors.
[0172] Execution efficiency: Improve the speed and accuracy of robot task execution and reduce the time to complete tasks.
[0173] Path planning and obstacle avoidance efficiency: Optimize the speed of path planning in complex environments and reduce repeated adjustments and misoperations that may occur during obstacle avoidance.
[0174] Specifically, the optimization goal is usually defined by a reward function, which combines the feedback information during the robot's task execution. In general, the design of the reward function follows the following principles:
[0175] Give positive rewards for successful task execution;
[0176] Give negative rewards for failure or wrong behavior;
[0177] For behaviors such as path planning and obstacle avoidance, rewards are given for successfully executing the shortest path and minimizing the risk of collision.
[0178] Through this reward function, the robot can continuously adjust its behavior strategy when performing tasks to maximize the success rate and efficiency of the tasks.
[0179] Adjust strategies and optimize parameters:
[0180] In some embodiments, a deep Q-learning (DQN) or proximal policy optimization (PPO) algorithm is applied to the adaptive optimization module. Specifically, the Q-learning algorithm guides the robot's behavioral decisions by estimating the Q value of each action. The robot's goal is to maximize long-term rewards, so it adjusts its Q value based on the feedback of each action to learn the optimal action strategy.
[0181] In addition, the PPO algorithm directly optimizes the policy function of the robot's behavior decision-making through the policy gradient method. The PPO algorithm can maintain the stability of policy updates during the optimization process and avoid instability caused by excessive policy changes.
[0182] Real-time feedback for adaptive optimization:
[0183] In this embodiment, the adaptive optimization module adjusts the optimization strategy through real-time feedback. When the robot is performing a task, it dynamically adjusts the strategy based on the environmental information fed back from the sensor and the execution effect of the current behavior. Every time the robot performs an action, the feedback information (such as whether the obstacle is successfully avoided, whether the target is reached, etc.) will be input into the optimization algorithm as a reward signal to update the robot's behavior strategy.
[0184] In this embodiment, the adaptive optimization module continuously optimizes the robot's decision-making strategy during task execution through reinforcement learning algorithms such as deep Q learning (DQN) or proximal policy optimization (PPO). The module can adjust control parameters and optimize models based on real-time feedback from multiple aspects such as task success rate, execution efficiency, and obstacle avoidance. Through the combination of online learning and transfer learning, the robot can quickly adapt to different tasks and improve the efficiency and intelligence level of its overall task execution. The adaptive optimization module enables the robot to continuously improve in complex and dynamic environments, ultimately achieving more accurate and efficient task execution.
[0185] Programming education platform, which provides students with a programming interface, allowing them to control robots to perform tasks through graphical programming or text programming, and provides real-time feedback and scoring based on the results of task execution;
[0186] In this embodiment, the programming education platform is closely connected with the previous modules (such as sensor data processing, deep learning and feature extraction, path planning and obstacle avoidance, etc.). The robot performs tasks through the programming interface provided by the platform, and gives real-time feedback to students. At the same time, the robot's execution process is used as a visual feedback for programming learning. The platform is not only a control interface, but also a tool for students to learn and understand programming control logic, robot behavior and task planning.
[0187] Platform architecture and module functions:
[0188] The graphical programming module provides students with a graphical programming interface, using block programming (such as Blockly or Scratch), so that students can drag and drop different functional blocks to splice program logic and complete task settings. In this way, students do not need to understand complex programming syntax, but only drag and drop and connect modules to program, which lowers the programming threshold and is especially suitable for beginners and primary school students.
[0189] Task and feedback module:
[0190] In this embodiment, the task and feedback module is one of the core parts of the programming education platform, which is mainly responsible for the design of tasks, feedback on the task execution process, and the evaluation of the final results. Students create tasks through this module and assign tasks to robots for execution. Task types can be simple motion control or complex tasks such as path planning, obstacle avoidance, or target grasping. After the task is set, students upload the task script through the platform and let the robot execute it.
[0191] Specifically, after the student submits the program and starts the robot to perform the task, the task and feedback module will generate real-time feedback based on the robot's behavior. The feedback includes whether the robot successfully completes the task, the time of task execution, the optimality of path planning, the effectiveness of obstacle avoidance strategy, etc. The feedback results will be displayed in real time on the student's interface to help students understand whether their programming logic is correct and whether adjustments are needed.
[0192] The feedback of task execution can be presented in visual and textual ways. For example, the platform can display the robot's real-time video stream or graphically present the robot's path execution process and the completion status of the task. In addition, the platform can also display key parameters of the robot's execution, such as the number of collisions, path distance and other information.
[0193] Mission Scoring System:
[0194] The task scoring system is another important module of the platform, which is mainly used to evaluate the quality and efficiency of students' task completion. Based on the program submitted by the student, the platform will score the task execution according to a set of predefined standard scoring rules. The scoring criteria can be set according to the type and goal of the task, for example:
[0195] Task completion: whether the robot successfully executes all task steps; execution efficiency: the shortest time and path required to complete the task; control accuracy: whether the robot moves accurately according to the predetermined path and whether there is any deviation; error rate: the number of errors made by the robot during the task.
[0196] In some embodiments, the scoring system can also adjust the scoring rules according to different learning progress and task difficulty. Through the task scoring system, students can not only get feedback on the completion of the task, but also understand in which areas they need to improve according to the scoring results.
[0197] The feedback results of the task scoring system are displayed in intuitive scores or star ratings, and detailed execution analysis and improvement suggestions are provided to help students improve their programming skills and task understanding.
[0198] Real-time robot feedback interface:
[0199] In some embodiments, the programming education platform also includes a real-time robot feedback interface, which transmits the robot's execution status and environmental perception information in real time through communication with the robot control system. This allows students to intuitively understand the robot's performance in the actual execution process through real-time data and image feedback while writing programs.
[0200] During the execution process, students can also manually adjust the robot's behavior and change the execution strategy through the control buttons on the interface, further deepening their understanding of programming logic and robot control.
[0201] Collaboration between the platform and the aforementioned modules:
[0202] In this embodiment, the programming education platform provides a complete programming environment for students by closely cooperating with the robot control and behavior decision module, the path planning and obstacle avoidance module, etc. When students program through the platform, the robot control module will receive the program written by the students and pass the task instructions to the subsequent behavior decision module and path planning module to ensure that the robot performs the task according to the instructions set by the students.
[0203] In this embodiment, the programming education platform provides an intuitive and interactive programming learning environment by integrating graphical programming and text programming modules, combined with task execution feedback and scoring systems. Through real-time feedback and task scoring systems, the platform not only helps students understand the principles of robot control, but also promotes the improvement of students' programming skills. At the same time, the seamless connection between the platform and the aforementioned modules (such as robot control and behavior decision-making, path planning and obstacle avoidance modules) enables students to better understand robot behavior in actual programming tasks and quickly master complex programming and control skills.
[0204] A control method for a programming educational robot based on artificial intelligence, comprising:
[0205] S1. Sensor data collection:
[0206] First, the robot collects environmental data through multiple sensors (such as cameras, LiDAR, ultrasonic sensors, etc.), including RGB images, depth images, 3D point cloud data, obstacle distances, etc. These data provide necessary input for subsequent environmental perception and decision-making.
[0207] S2. Sensor data preprocessing and feature extraction:
[0208] By preprocessing the raw sensor data (such as denoising, standardization, etc.) and using deep learning models (such as convolutional neural networks CNN, PointNet++, etc.), useful features are extracted. Image data extracts features such as shape and texture, LiDAR data extracts the location and spatial structure of obstacles, and ultrasonic and infrared sensor data extracts distance information.
[0209] S3. Multimodal data fusion:
[0210] Through the self-attention mechanism or weighted average algorithm, the features from different sensors are weighted and fused. In this way, the robot can comprehensively utilize the advantages of different sensors to obtain more accurate and comprehensive environmental perception information.
[0211] S4. Deep reinforcement learning behavior decision-making:
[0212] Based on deep reinforcement learning algorithms (such as deep Q learning or proximal policy optimization), the robot selects appropriate behaviors (such as moving forward, turning, avoiding obstacles, etc.) according to the integrated environmental perception information. Through the reward and punishment mechanism, the robot continuously optimizes its behavior strategy during the execution of the task.
[0213] S5. Path planning and obstacle avoidance:
[0214] Through the A* algorithm or D*Lite algorithm for path planning, the robot generates the optimal path based on the current task and environmental information. At the same time, the robot will make obstacle avoidance decisions based on real-time sensor data to ensure that it avoids collisions and chooses the safest path during execution.
[0215] S6. Task execution and feedback evaluation:
[0216] The robot performs corresponding actions according to the task requirements and evaluates the task completion status through feedback during the execution process (such as whether it successfully avoids obstacles, whether it completes the target task, etc.). During the execution process, the robot may adjust the path or action according to environmental changes.
[0217] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A programming educational robot control system based on artificial intelligence, characterized in that: include: The hardware module includes a plurality of sensor components for collecting environmental data, wherein the sensors include: Camera: used to collect RGB images and depth images of the environment; LiDAR: used to generate 3D point cloud data of the environment; Ultrasonic sensor: used for close-range obstacle detection; Infrared sensor: used for obstacle detection in low-light or dark environments; The sensor data processing module is used to pre-process the environmental data collected by the sensor, remove noise and perform feature extraction to facilitate subsequent deep learning processing; Deep learning and feature extraction module, which is used to further process the data features from different sensors using deep neural networks to extract important environmental perception features; Multimodal data fusion module, which is used to intelligently fuse feature data from different sensors to generate more accurate and comprehensive environmental perception information; The robot control and behavior decision module uses a deep reinforcement learning algorithm to make decisions based on the integrated environmental perception information and control the robot to perform tasks; The path planning and obstacle avoidance module uses a path planning algorithm to generate the optimal path based on environmental perception information and target tasks, and avoids obstacles in real time during the path execution process; Adaptive optimization module, which continuously adjusts decision-making strategies and control algorithms based on feedback from the robot's task execution to improve the robot's performance in complex environments; Programming education platform, which provides students with a programming interface, allowing them to control robots to perform tasks through graphical programming or text programming, and provides real-time feedback and scoring based on the results of task execution; The deep learning and feature extraction module includes: Extracting environmental features from the image data using a convolutional neural network, wherein the convolutional neural network is trained to extract deep features related to obstacles, targets, and terrain; Use PointNet++ point cloud processing network to extract features from LiDAR point cloud data to obtain more accurate spatial perception information; The multimodal data fusion method is as follows: in, is the feature from the i-th sensor, d is the dimension of the extracted feature, α i is the weight corresponding to sensor i, indicating the weight coefficient of the sensor, weight α i Through dynamic calculation of the self-attention mechanism, the contribution of each sensor in multimodal fusion is adjusted. N is the total number of sensors. It is the fused feature vector, which combines the features of all sensors.
2. The artificial intelligence-based programming education robot control system according to claim 1 is characterized in that: The multimodal data fusion module adopts a self-attention mechanism or a Transformer network to perform weighted fusion on features from different sensors, combining time series information and spatial information to improve the perception accuracy and robustness of the system.
3. The artificial intelligence-based programming education robot control system according to claim 1 is characterized in that: The robot control and behavior decision-making module uses a deep reinforcement learning algorithm to select and execute corresponding control actions based on the integrated environmental perception information and task requirements, and continuously optimizes the decision-making strategy through a reward feedback mechanism to ensure that the robot can complete the task efficiently and intelligently.
4. The artificial intelligence-based programming education robot control system according to claim 3 is characterized in that: The deep reinforcement learning algorithm adopts deep Q learning or proximal strategy optimization algorithm, and rewards and punishes according to the actual effect of the robot in performing tasks, thereby optimizing the robot's behavioral decision-making ability.
5. The artificial intelligence-based programming education robot control system according to claim 1 is characterized in that: The path planning and obstacle avoidance module adopts the A* algorithm or the D*Lite algorithm, and dynamically adjusts the path planning in combination with real-time environmental information, so that the robot can efficiently avoid obstacles and complete the task objectives when performing tasks.
6. A control method for a programming education robot based on artificial intelligence, according to the control system for a programming education robot based on artificial intelligence according to any one of claims 1 to 5, characterized in that: include: S1. Collect environmental data through sensors, including RGB images, depth images, LiDAR point cloud data, ultrasonic sensor data, and infrared sensor data; S2. Preprocessing the environmental data and performing feature extraction on the preprocessed data, wherein the feature extraction includes extracting features from the image data using a convolutional neural network and extracting features from the LiDAR data using a point cloud processing network; S3, using multimodal data fusion technology to intelligently fuse feature information from different sensors to generate fused environmental perception information; S4. Based on the environmental perception information, use a deep reinforcement learning algorithm to make robot behavior decisions, select appropriate actions and control the robot to execute; S5. Based on the current state and target position of the robot, the path planning algorithm is used to generate the optimal path, and the obstacle avoidance algorithm is used to ensure the safety of the robot during the path execution; S6. Based on the feedback from task execution, use adaptive optimization algorithms to adjust decision strategies and control models to improve robot execution efficiency and task completion rate.
7. The control method of a programming educational robot based on artificial intelligence according to claim 6, characterized in that: The deep reinforcement learning algorithm is optimized based on the value functions of states and actions, evaluates the pros and cons of different actions through Q-value functions or policy functions, and updates policies based on real-time feedback.
8. The control method of a programming educational robot based on artificial intelligence according to claim 6, characterized in that: The path planning algorithm includes an A* algorithm or a D*Lite algorithm, which generates the shortest path by evaluating the cost from the current position to the target position and adjusts the path in real time to avoid obstacles.
9. The control method of a programming educational robot based on artificial intelligence according to claim 6, characterized in that: The adaptive optimization module adjusts the robot control strategy, reward function and action strategy based on the task execution results through online learning and transfer learning technology to adapt to new tasks and environments.
Citation Information
Patent Citations
Multi-mode navigation system of education robot in complex environment
CN118010009A