Robot visual guidance method and system based on artificial intelligence
By generating contextual feature vectors and using a visual dynamic field network to predict the optimal walking action, the problem of quadrupedal bionic robots relying on 3D reconstruction in complex terrains was solved. Stable walking was achieved on terrains such as home floors, grass, and gravel, improving the system's autonomy and adaptability.
Patent Information
- Application Number
- CN202511349793.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-22
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-09-22
AI Technical Summary
Existing vision-guided quadruped bionic robots rely on precise 3D geometric reconstruction in complex or unstructured terrains, which limits their robustness and autonomous walking ability, especially in home environments, grass, gravel, and other terrains where stable walking is difficult.
An AI-based robot vision guidance method is adopted. By acquiring the current image and motion state, a context feature vector is generated. The visual dynamic field generation network is used to predict the optimal walking action, bypassing the reconstruction of a precise 3D map, and realizing forward planning and closed-loop control.
It significantly expands the applicability of quadrupedal bionic robots, enabling them to handle complex terrain, improve the robustness and adaptability of the system, and realize the ability to adjust walking strategies according to ground characteristics, adapting to environments such as home floors, grass, and gravel.
Smart Images

Figure CN120839804A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of interdisciplinary technology of artificial intelligence and robotics, specifically to a robot vision guidance method and system based on artificial intelligence. Background Technology
[0002] With the rapid development of robotics technology, visual guidance technology has become crucial for enabling robots to handle complex and unstructured tasks. By equipping robots with visual sensors (such as cameras), they can perceive their working environment and autonomously adjust their movements. This allows them to replace traditional methods that rely on pre-programmed paths or remote control in complex environments such as home environments, field inspections, and disaster search and rescue, greatly enhancing the system's autonomy and adaptability.
[0003] However, existing robot vision-guided technologies still face numerous challenges when traversing complex, unstructured terrain. A mainstream approach relies on accurate 3D modeling and reconstruction of the working environment, acquiring complete 3D geometric information of the terrain through vision or depth sensors and generating path planning based on this. This method is highly dependent on the structure and material properties of the environment. When dealing with highly unstructured terrain such as furniture, gravel, grass, or shallow water, which lacks stable geometric features or is dynamically variable, establishing an accurate 3D model becomes extremely difficult, and the robustness and accuracy of the reconstruction algorithm are hard to guarantee. This severely limits the application capabilities of quadrupedal bionic robots in these real-world field scenarios.
[0004] Furthermore, existing visual guidance systems often lack a deep understanding of the task context, typically simplifying robot walking movements into purely geometric motion while ignoring the vastly different physical interactions between different gaits and terrains. For example, walking on gravel may cause gravel to shift or splash, while on grass, insufficient friction may lead to slipping. Traditional methods cannot model and predict these complex dynamic visual changes closely related to gait and terrain, making them ill-suited for adaptive tasks requiring real-time adjustments to walking strategies based on ground characteristics.
[0005] Furthermore, from the perspective of control strategies, many traditional methods place excessive demands on the accuracy of the system model or rely on open-loop execution plans. Such systems are highly sensitive to camera calibration errors, robot kinematic errors, and minor environmental disturbances (such as ground subsidence or foot slippage). If the initial model is flawed or interference occurs during execution, task failure is highly likely. Even when some systems introduce closed-loop feedback, their control methods are often passive responses based on current errors, lacking the ability to proactively plan for the consequences of actions. For example, they cannot predict beforehand that an action might cause severe shaking of the robot, thus affecting the system's robustness and adaptability in dynamic and uncertain environments. Summary of the Invention
[0006] The technical problem this invention aims to solve is that existing visual guidance technologies for quadrupedal bionic robots, especially in complex or unstructured terrains, typically rely on accurate 3D geometric reconstruction of the environment (such as constructing point cloud maps). When faced with environments like home floors, grass, gravel, or water surfaces—terrains with weak textures, dynamics, or deformability—accurate and robust 3D reconstruction becomes extremely difficult, or even impossible, thus limiting the robot's autonomous walking and adaptability in these challenging environments.
[0007] In view of this, the present invention aims to provide a robot vision guidance method and system based on artificial intelligence to solve the above-mentioned technical problems.
[0008] The first aspect of this invention provides a robot vision guidance method based on artificial intelligence, the method comprising:
[0009] Acquire the current image of the terrain in front of the quadrupedal bionic robot and the preset mission target image to guide the quadrupedal bionic robot to walk safely;
[0010] The current motion state of the quadrupedal bionic robot and contextual information including current gait information and terrain information in front are obtained.
[0011] The context information is encoded to generate a context feature vector;
[0012] Within the local motion space of the quadrupedal bionic robot, multiple candidate walking actions are sampled and generated based on the current motion state of the quadrupedal bionic robot.
[0013] For each candidate walking action, the current image of the terrain in front of the quadrupedal bionic robot, the current motion state of the quadrupedal bionic robot, the candidate walking action, and the context feature vector are input into a preset visual dynamic field generation network to generate a predicted visual dynamic field.
[0014] Based on the current image of the terrain in front of the quadrupedal bionic robot and the preset mission target image for guiding the quadrupedal bionic robot to walk safely, a desired visual dynamic field is determined.
[0015] An optimal walking action is determined from the plurality of candidate walking actions by comparing each of the predicted visual dynamic fields with the expected visual dynamic field.
[0016] Control the quadrupedal bionic robot to perform the optimal walking action.
[0017] In one specific embodiment, the step of encoding the context information specifically includes:
[0018] The current gait information and the terrain information ahead are respectively embedded into features to obtain a gait feature vector and a terrain feature vector;
[0019] The gait feature vector and the terrain feature vector are fused to generate the context feature vector.
[0020] Preferably, the visual dynamic field generation network includes at least one feature modulation layer. In the step of inputting the current image of the terrain in front of the quadrupedal bionic robot, the current motion state of the quadrupedal bionic robot, the candidate walking action, and the context feature vector into a preset visual dynamic field generation network, the context feature vector is used to generate modulation parameters. The modulation parameters perform affine transformation on the intermediate feature maps in the visual dynamic field generation network to achieve dynamic modulation of the predicted visual dynamic field.
[0021] Specifically, for the first in the network Intermediate feature map of the layer First, it consists of a subnetwork Based on context feature vectors Generate modulation parameters, i.e., scaling factors. and bias factor : ;
[0022] Subsequently, the intermediate feature map Perform an affine transformation to obtain the modulated feature map. : ;
[0023] in, This represents element-wise multiplication. Subnetwork Trainable parameters.
[0024] In one specific embodiment, the step of determining a desired visual dynamic field is as follows:
[0025] An optical flow estimation algorithm is used to calculate the pixel displacement between the current image of the terrain in front of the quadrupedal bionic robot and the preset target image for guiding the quadrupedal bionic robot to walk safely, thereby generating the desired visual dynamic field.
[0026] Preferably, the step of comparing each predicted visual dynamic field with the desired visual dynamic field specifically includes:
[0027] Calculate the difference measure between each predicted visual dynamic field and the expected visual dynamic field within a preset path key region;
[0028] And select the candidate walking action with the smallest difference metric as the optimal walking action.
[0029] In one specific embodiment, the visual dynamic field generation network is trained through self-supervised learning, and the training process includes:
[0030] Collect training data that includes terrain images before the action, contextual information, the walking action, and terrain images after the action;
[0031] Based on the terrain image before the action and the terrain image after the action, a ground truth visual dynamic field is generated, and the visual dynamic field generation network is trained based on the difference between the ground truth visual dynamic field and the visual dynamic field predicted by the visual dynamic field generation network.
[0032] Preferably, the steps for generating a truth visual dynamic field are as follows:
[0033] An optical flow estimation algorithm is used to calculate the actual pixel displacement between the image before the action and the image after the action, thereby generating the ground truth visual dynamic field.
[0034] In one specific embodiment, both the predicted visual dynamic field and the desired visual dynamic field are two-dimensional vector fields, and their size corresponds to the size of the current image of the terrain in front of the quadrupedal bionic robot. Each vector in the two-dimensional vector field represents the predicted or desired visual motion trend of the corresponding terrain pixel.
[0035] Preferably, the current gait information is a gait identifier in a preset gait library, and the terrain information ahead is a terrain category or attribute obtained through image segmentation or classification.
[0036] A second aspect of the present invention provides an artificial intelligence-based robot vision guidance system, the system comprising:
[0037] An information acquisition module is used to acquire the current image of the terrain in front of the quadrupedal bionic robot, the preset task target image for guiding the quadrupedal bionic robot to walk safely, the current motion state of the quadrupedal bionic robot, and contextual information including current gait information and terrain information in front.
[0038] A context encoding module is used to encode the context information to generate a context feature vector;
[0039] An action generation module is used to sample and generate multiple candidate walking actions based on the current motion state of the quadrupedal bionic robot within the local motion action space of the quadrupedal bionic robot.
[0040] A visual consequence prediction module is configured with a preset visual dynamic field generation network, which is used to generate a predicted visual dynamic field for each of the plurality of candidate walking actions, based on the current motion state of the quadrupedal bionic robot, the candidate walking action, and the context feature vector.
[0041] An action decision module is used to determine a desired visual dynamic field based on the current image of the terrain in front of the quadrupedal bionic robot and the preset task target image for guiding the quadrupedal bionic robot to walk safely, and to determine an optimal walking action from the plurality of candidate walking actions by comparing each predicted visual dynamic field with the desired visual dynamic field.
[0042] A robot control module is used to control the quadrupedal bionic robot to perform the optimal walking action.
[0043] This invention provides a robot vision guidance method and system based on artificial intelligence. It has the following beneficial effects:
[0044] 1. This invention generates control commands by directly predicting the dynamic consequences of a robot's walking movements in visual space and comparing them with the desired visual target, bypassing the intermediate step of accurately reconstructing a 3D map of the terrain. Since both the predicted and desired visual dynamic fields are defined and calculated in a 2D image space, this method does not rely on a precise 3D model, thus significantly expanding the applicability of quadrupedal bionic robots. This enables them to handle guidance tasks in complex terrains that are difficult to accurately model in 3D, such as home environments, grass, gravel, or shallow water.
[0045] 2. This invention introduces a contextual information encoding and modulation mechanism, enabling the system to understand the current walking scenario. By encoding the current gait information and the terrain information ahead into a contextual feature vector, and using this vector to dynamically modulate the intermediate feature map of the visual dynamic field generation network, the prediction result not only contains kinematic information but also reflects the physical effects of specific gait interactions with specific terrain. This allows the method to handle more complex adaptive walking tasks, such as adjusting the visual expectation of the foot placement based on whether the ground is hard or soft, thereby enabling applications such as walking smoothly across home environments, grass, gravel, or wading through water.
[0046] 3. This invention employs a closed-loop control strategy based on forward planning, enhancing the system's robustness and adaptability. In each control cycle, the system generates multiple candidate walking actions through sampling and predicts the visual consequences of each action, then selects the optimal walking action for execution. This iterative process of continuously predicting and making decisions based on the current image of the terrain in front of the quadrupedal bionic robot enables the system to compensate in real time for execution deviations caused by body sway, camera calibration errors, or minor environmental changes, without relying on an absolutely accurate world model or open-loop execution plan. Attached Figure Description
[0047] Figure 1 This is a functional block diagram of a robot vision guidance system according to an embodiment of the present invention;
[0048] Figure 2 This is a flowchart of a robot vision guidance method according to an embodiment of the present invention.
[0049] Among them, 10 is the information acquisition module; 20 is the context encoding module; 30 is the action generation module; 40 is the visual consequence prediction module; 50 is the action decision module; and 60 is the robot control module. Detailed Implementation
[0050] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0051] See attached document Figure 1 , Figure 1 This is a functional block diagram of a robot vision guidance system according to an embodiment of the present invention. The present invention provides an artificial intelligence-based robot vision guidance system, which may include: an information acquisition module 10, a context encoding module 20, an action generation module 30, a visual consequence prediction module 40, an action decision module 50, and a robot control module 60.
[0052] In one specific embodiment, the information acquisition module 10 is used to collect basic data required for system operation from the quadrupedal bionic robot's walking path. Specifically, the information acquisition module 10 acquires the current image of the terrain in front of the quadrupedal bionic robot at the current moment through a configured image sensor (e.g., an RGB-D camera), and reads a preset task target image guiding the quadrupedal bionic robot to walk safely from the storage unit. Simultaneously, the information acquisition module 10 acquires the robot's current motion state from the robot controller or state estimator, the current state including the robot's pose information and its joint states. The information acquisition module 10 also acquires contextual information containing current gait information and terrain information in front of the robot through sensors or preset information.
[0053] The context encoding module 20, whose input is connected to the output of the information acquisition module 10, is used to receive context information and process it into a fixed-dimensional context feature vector.
[0054] The motion generation module 30 has an input terminal for receiving the current state of the robot and for sampling and generating multiple candidate walking actions based on the current state of the robot within the local motion space of the quadrupedal bionic robot.
[0055] The visual consequence prediction module 40 is configured with a preset visual dynamic field generation network. The inputs of the visual consequence prediction module 40 are respectively used to receive the current image of the terrain in front of the quadrupedal bionic robot and the current motion state of the quadrupedal bionic robot from the information acquisition module 10, the context feature vector from the context encoding module 20, and multiple candidate walking actions from the action generation module 30. For each candidate walking action, the visual consequence prediction module 40 outputs a corresponding predicted visual dynamic field.
[0056] The action decision module 50 receives inputs from the information acquisition module 10, including a current image of the terrain in front of the quadrupedal bionic robot and a preset task target image for guiding the quadrupedal bionic robot to walk safely, as well as multiple predicted visual dynamic fields from the visual consequence prediction module 40. The action decision module 50 first determines a desired visual dynamic field based on the current image of the terrain in front of the quadrupedal bionic robot and the preset task target image for guiding the quadrupedal bionic robot to walk safely. Then, by comparing each predicted visual dynamic field with the desired visual dynamic field, it determines the optimal walking action from multiple candidate walking actions.
[0057] The robot control module 60, whose input end is connected to the output end of the motion decision module 50, is used to receive the optimal walking action and convert the optimal walking action into motion commands that can be executed by the robot's underlying controller, so as to control the quadrupedal bionic robot to perform the optimal walking action.
[0058] In a specific implementation scenario, the system of the present invention can be deployed on a hardware platform. The hardware platform includes, but is not limited to:
[0059] A quadrupedal bionic robot equipped with an inertial measurement unit (IMU) and joint torque sensors;
[0060] An RGB-D camera fixed to the robot's head or body is used to capture images of the terrain ahead;
[0061] And a computing server (which may be onboard to the robot or a remote server) equipped with a high-performance graphics processing unit (GPU) for performing the intensive computing tasks in the method of the present invention.
[0062] The system and method of this invention can run in a software environment. The software environment includes: an operating system, such as Linux; a robot operating system, such as the Robot Operating System (ROS), for handling robot communication and control; a deep learning framework, such as PyTorch or TensorFlow, for building, training, and deploying visual dynamic field generation networks; and a computer vision library, such as OpenCV, for performing operations such as image processing and optical flow calculation.
[0063] See attached document Figure 2 , Figure 2 This is a flowchart of a robot vision guidance method according to an embodiment of the present invention. In a complete task execution cycle, the method first initializes by acquiring the current image of the terrain in front of the quadrupedal bionic robot and the current motion state of the quadrupedal bionic robot. Subsequently, the system enters a closed-loop control loop. In each control loop, the system first calculates the desired visual dynamic field at the current moment based on the latest current image of the terrain in front of the quadrupedal bionic robot and a fixed task target image for guiding the quadrupedal bionic robot to walk safely. The system acquires the current motion state and walking context information of the quadrupedal bionic robot in parallel and encodes the context information. Based on the current motion state of the quadrupedal bionic robot, the system samples and generates a set of candidate walking actions covering the local motion space. Next, the system uses a visual dynamic field generation network to predict a visual consequence for each candidate walking action, i.e., the predicted visual dynamic field. By quantifying and comparing all predicted visual dynamic fields with the desired visual dynamic field, the system decides on the optimal walking action. Finally, the robot executes this optimal walking action, causing changes in the current image of the scene in front and the terrain in front of the quadrupedal bionic robot, and the system then enters the next control loop. This process iterates until the difference between the current image of the terrain in front of the quadrupedal bionic robot and the preset target image for guiding the quadrupedal bionic robot to walk safely meets the preset termination condition.
[0064] In a specific embodiment, the task initialization and target definition steps of the method are as follows.
[0065] The task target image is a pre-defined digital image with the same resolution and color space as the current image of the terrain in front of the quadrupedal bionic robot. The task target image remains unchanged throughout the task execution. It is a digital representation of the visual state expected of the robot when it reaches the target area or achieves the desired path. In different applications, the task target image can be a photograph taken by a camera representing a safe and traversable area, or a rendering generated based on navigation path points. This task target image is loaded into the system's storage unit before the method begins execution.
[0066] At the beginning of each control loop, the system needs to determine a desired visual dynamic field. The desired visual dynamic field is a two-dimensional vector field with dimensions identical to the current image of the terrain in front of the quadrupedal bionic robot and the preset target image for guiding the robot's safe walking. The desired visual dynamic field is calculated using an optical flow estimation algorithm, whose inputs are the current image of the terrain in front of the quadrupedal bionic robot and the preset target image for guiding the robot's safe walking.
[0067] Specifically, the optical flow estimation algorithm calculates a displacement vector for each pixel in the image. For example, for the current image of the terrain in front of the quadrupedal bionic robot, the coordinates are... For each pixel, the algorithm calculates the corresponding pixel coordinates within the preset target image for guiding the quadrupedal bionic robot to walk safely. Therefore, a displacement vector is obtained. The displacement vectors of all pixels together constitute the desired visual dynamic field. This vector field provides an instantaneous, pixel-level, quantitative motion target for subsequent action decision-making steps, where each vector represents the desired visual motion direction and speed of the corresponding terrain pixel in order to achieve the target state.
[0068] In each control loop, the system performs state perception and context encoding in parallel. The current motion state of the quadrupedal bionic robot specifically refers to its spatial pose in a predefined world coordinate system. This pose information is calculated using the forward kinematics state estimation model within the robot controller, based on readings from the joint encoders and IMU, and expressed as a coordinate system containing position coordinates. and pose description (e.g., quaternions) The numerical vector form of is provided to the system.
[0069] The context information consists of two parts: current gait information and terrain information ahead. The current gait information is a discrete identifier corresponding to a specific gait in a pre-defined gait library. When the robot switches gaits, the system is configured to use the corresponding gait identifier. For example, when the robot adopts a trot, the current gait information is set to a specific integer ID representing that gait.
[0070] Terrain information refers to the category or attribute of the terrain that the robot is currently treading or is about to tread. This information is obtained by analyzing the current image of the terrain in front of the quadrupedal bionic robot provided by the information acquisition module 10. In one specific embodiment, a pre-trained image segmentation network (e.g., Mask R-CNN) or image classification network (e.g., ResNet) processes the current image of the terrain in front of the quadrupedal bionic robot, identifies the main terrain located on the robot's path, and outputs its category label, such as "grass" or "gravel road".
[0071] After acquiring current gait information and terrain information ahead, the context encoding module 20 performs an encoding process to generate a context feature vector. This process includes two stages: feature embedding and feature fusion.
[0072] During the feature embedding stage, the integer ID of the current gait information is input into a gait embedding layer, which is a trainable lookup table used to map discrete gait IDs to a low-dimensional, dense gait feature vector. Similarly, the category labels of the foreground terrain information (also represented as integer IDs) are input into a terrain embedding layer, which maps them to a terrain feature vector. .
[0073] During the feature fusion stage, gait feature vectors and terrain feature vector Concatenating the features in a dimensional plane creates a combined feature vector. This combined feature vector is then fed into a multilayer perceptron (MLP) network. The MLP network contains at least one hidden layer and a non-linear activation function (e.g., ReLU). Through forward propagation, it ultimately outputs a single contextual feature vector that incorporates gait and terrain information. This context feature vector This will be used as a conditional input for subsequent visual consequence prediction.
[0074] See attached document Figure 2After completing state perception and context encoding, the system performs forward-looking action planning and decision-making. This step begins with the action generation module 30 sampling and generating multiple candidate walking actions within the local motion space of the quadrupedal bionic robot, based on the robot's current state. The local motion space defines a set of minute displacements of the robot's main body from its current pose. Each candidate walking action is represented as a six-dimensional vector. The first three terms represent the translation increment in the world coordinate system, and the last three terms represent the attitude rotation increment. In a specific embodiment, a set containing uniform random sampling is generated within a preset translation and rotation range. The set of distinct candidate walking actions is represented as .
[0075] Subsequently, the visual consequence prediction module 40 receives a set of candidate walking actions. For each candidate walking action in the set... (in From 1 to (index), the module will display the current image of the terrain in front of the quadrupedal bionic robot. Current motion state of the quadrupedal bionic robot Candidate walking movements and context feature vectors The inputs are fed into a pre-defined visual dynamic field generation network. The network performs a forward propagation calculation for each input combination and outputs a corresponding predicted visual dynamic field. This process is executed in batch mode, generating data at once. The candidate walking actions correspond to A predicted visual dynamic field .
[0076] Finally, the action decision module 50 executes the process of determining the optimal walking action. This process involves quantifying and comparing each predicted visual dynamic field. The desired visual dynamic field determined in the preceding steps The difference is used to achieve this. To focus the comparison on the image region directly related to the path, the system first defines a preset path key region. This region can be a binary mask. , where a pixel with a value of 1 corresponds to the projection position of the ground area that the robot is about to step on in front of it in the current image of the terrain in front of the quadrupedal bionic robot.
[0077] The difference metric is calculated within the critical region of the path. For the first... The difference measure of each candidate walking action. By calculating the predicted visual dynamic field and the expected visual dynamic field The L1 distance between them is obtained, and the specific calculation formula is as follows: ;
[0078] in, Represents pixel coordinates It is the value of the mask at the corresponding pixel; Represents the L1 norm; For the summation operator, let represent the summation of all pixel coordinates within the image. Perform a traversal and summation; Represents a normalization factor; Let be a two-dimensional vector, representing the vector formed by the first... The predicted visual dynamic field generated by each candidate walking action in pixels Displacement vector at; Let be a two-dimensional vector representing the desired visual dynamic field. The formula calculates the average L1 norm of the difference between the corresponding vectors of the two vector fields within the critical region of the path.
[0079] After calculating all Difference measurement of candidate walking actions Then, the action decision module 50 selects the candidate walking action with the smallest difference metric as the optimal walking action. Specifically, the index of the optimal walking action is first determined. :
[0080] ;
[0081] Subsequently, based on this index Determine the optimal walking motion:
[0082] ;
[0083] This optimal walking action The robot walking action is determined to be the one in the current control loop that is most likely to enable the visual transformation of the current terrain image into the preset target image for guiding the quadrupedal bionic robot to walk safely.
[0084] See attached document Figure 2 The optimal walking action is determined in the action decision module 50. The robot control module 60 then receives this optimal walking motion and converts it into motion commands executable by the robot's underlying controller. Optimal walking motion This is an incremental pose relative to the robot's current pose. The robot control module 60 first superimposes the incremental pose with the robot's current absolute pose to calculate a new target absolute pose. Then, the module uses the robot's inverse kinematics model to convert the new target absolute pose into desired angles or torques for each of the robot's leg joints. These desired values are then sent to the robot's servo controllers, driving the robot to perform minute movements. Depending on the robot's specific control architecture, this may involve position control, velocity control, or a more advanced model predictive control (MPC) mode to ensure precise motion and compliant interaction with the ground.
[0085] After the robot completes the optimal walking motion, the physical state of its body and the terrain in front of it will change. The information acquisition module 10 then re-acquires the current image of the terrain in front of the quadrupedal bionic robot. At the same time, the robot's current state, i.e., the real-time pose of the main body, is also updated through the robot controller. This latest perception data is used to start the next control loop.
[0086] The method of this invention operates in a closed-loop iterative manner. The system continuously senses the current state, predicts the consequences of actions, decides the optimal walking action and executes it, then senses again until a preset task termination condition is met. The task termination condition is typically based on the visual difference between the current image of the terrain in front of the quadrupedal bionic robot and a preset task target image guiding the quadrupedal bionic robot to walk safely. For example, the minimum difference metric calculated within the critical region of the path. (Right now If the value remains below a preset threshold (e.g., average pixel displacement error less than 0.05 pixels) and the amount of motion of the robot body in multiple consecutive control cycles is less than a preset small motion threshold (e.g., translation increment less than 0.1 mm, rotation increment less than 0.1 degrees), the system determines that it has reached the target position or completed path tracking and terminates the control cycle.
[0087] In one specific embodiment, the implementation of the context encoding module 20 includes two main stages: a feature embedding stage and a feature fusion stage.
[0088] During the feature embedding stage, the module encodes discrete current gait information and forward terrain information through two independent embedding layers. The current gait information is represented as an integer ID, which is input into a pre-defined gait embedding layer. The gait embedding layer is essentially a trainable lookup table that maps each unique gait ID to a fixed-dimensional floating-point gait feature vector. For example, if there are 10 different gaits, the embedding layer can be configured to map each gait ID to a 32-dimensional vector. Similarly, the foreground terrain information, represented as another integer ID (representing the terrain category), is input to a pre-defined terrain embedding layer. This layer is also a trainable lookup table that maps each unique terrain category ID to a fixed-dimensional floating-point terrain feature vector. For example, if 15 different terrains are identified, the embedding layer can map each terrain ID to a 64-dimensional vector. These two embedding layers are learned and optimized end-to-end throughout the system training process.
[0089] During the feature fusion stage, the context encoding module 20 receives gait feature vectors generated from the gait embedding layer. and terrain feature vectors generated from the terrain embedding layer First, these two feature vectors are concatenated along their dimensions to form a combined feature vector. For example, if the gait feature vector... It is a 32-dimensional terrain feature vector If the dimensions are 64, the concatenated combined feature vector will be 96 dimensions. This combined feature vector is then input into a Multilayer Perceptron (MLP) network. An MLP network consists of multiple fully connected layers; for example, it may contain two or three hidden layers, each followed by a batch normalization layer and a non-linear activation function (e.g., ReLU, RectifiedLinearUnit) to enhance the network's expressive power and training stability. The first hidden layer maps the 96-dimensional input to 128 dimensions, and the second hidden layer maps the 128-dimensional input to 64 dimensions. The output layer of the MLP network is a fully connected layer that maps the output of the previous hidden layer to the final, fixed-dimensional context feature vector. For example, contextual feature vectors It can be 32-dimensional or 64-dimensional. This context feature vector. It densely encodes comprehensive information on gait and terrain involved in the current walk, which is used to adjust the behavior of the visual dynamic field generation network in the subsequent visual consequence prediction module 40.
[0090] In one specific embodiment, the visual consequence prediction module 40 includes a preset visual dynamic field generation network. The core function of the network is to receive the current image of the terrain in front of the quadrupedal bionic robot, as well as additional conditional information representing the robot's state, candidate walking actions, and walking context, and output a two-dimensional predicted visual dynamic field with the same resolution as the current image of the terrain in front of the quadrupedal bionic robot.
[0091] The visual dynamic field generation network adopts an encoder-decoder structure, similar to the U-Net architecture, and incorporates a conditional input mechanism.
[0092] Encoder section:
[0093] The encoder consists of a series of convolutional layers, batch normalization layers, activation functions (such as ReLU), and downsampling layers (such as max pooling layers).
[0094] Its main input is the current image of the terrain in front of the quadrupedal bionic robot provided by the information acquisition module 10. The current image of the terrain in front of the quadrupedal bionic robot is subjected to feature extraction and spatial dimension reduction through multiple convolutional blocks.
[0095] At different levels of the encoder, the robot's current state (6-dimensional pose vector), candidate walking actions (6-dimensional increment vector), and contextual feature vector (e.g., 32-dimensional) are processed and incorporated into the image features. Specifically, to more precisely integrate conditional information into the image features, this invention employs an adaptive modulation mechanism. These three low-dimensional vectors represent the current motion state of the quadrupedal bionic robot. Candidate walking movements and context feature vectors First, the conditions are concatenated along the dimensional lines to form a comprehensive conditional vector. Then, this comprehensive conditional vector is input into a separate conditional embedding network (usually a small multilayer perceptron network).
[0096] This conditional embedding network generates dynamic modulation parameters for the feature maps of different layers in the encoder based on the comprehensive conditional vector. More specifically, for the first layer of the encoder... Intermediate feature map of the layer A subnetwork (The dedicated multilayer perceptron associated with this layer) is based on the context feature vector In addition to other conditional information, modulation parameters are generated, namely, scaling factors for the feature map of this layer. and bias factor :
[0097] ;
[0098] here, Subnetwork Trainable parameters. and The dimension will be designed to match the feature map. The number of channels matches.
[0099] Subsequently, the intermediate feature map A affine transformation is performed channel by channel to obtain the modulated feature map. :
[0100] ;
[0101] in, This represents element-wise multiplication. The modulated feature map... This will serve as the input to the next convolutional layer. Through this mechanism, the network can dynamically adjust its extraction and processing of visual features based on the robot's real-time state, the action to be performed, and the walking context, thereby achieving conditional feature learning at different levels of abstraction.
[0102] The decoder also consists of a series of convolutional layers, batch normalization layers, and activation functions, and includes upsampling layers (such as transposed convolutions or bilinear interpolation followed by convolutions).
[0103] The decoder receives the compressed feature representation of the encoder's last layer output and upsamples it step by step to restore the spatial resolution of the image.
[0104] After each upsampling step, the decoder concatenates the upsampled feature map with the corresponding image feature map from the encoder, passed through skip connections. These skip connections help preserve the detailed information of the image captured by the encoder (such as terrain edges and textures), fusing it with the high-level features reconstructed by the decoder, avoiding the loss of fine-grained features during downsampling, thereby improving the accuracy of the predicted dynamic field.
[0105] Output layer:
[0106] The final layer of the decoder is a convolutional layer with two output channels. These two channels represent each pixel in the horizontal direction. ) and vertical direction ( The predicted displacement components on ).
[0107] This output is the predicted visual dynamic field, which is exactly the same as the current image resolution of the terrain in front of the input quadrupedal bionic robot. Since the shift can be any real number, the output layer typically does not use an activation function, or uses linear activation.
[0108] The visual dynamics field generation network is trained end-to-end using a large amount of training data. This training data includes various robot actions, corresponding initial images, robot states, contextual information, and the actual pixel displacements observed after each action (i.e., the true visual dynamics field). The network's optimization objective is to minimize the difference between the predicted and true visual dynamics fields (e.g., using L1 or L2 loss), enabling it to accurately predict the visual consequences of any candidate walking action under given conditions. In practical training, commonly used loss functions include L1 or L2 loss, typically focusing on pixels within critical regions of the path.
[0109] In one specific embodiment, the visual dynamic field generation network of the present invention is trained using a self-supervised learning paradigm. Its core lies in automatically generating labeled data for model training through the actual interaction between the robot and its environment, thereby avoiding the time-consuming and expensive manual annotation process. The collection of training data is fundamental to this process.
[0110] See attached document Figure 2 The training data collection process mainly involves the robot autonomously exploring or executing a series of predefined basic actions, and recording key information at each step. The specific process is as follows:
[0111] Robot behavior generation:
[0112] The robot is programmed to perform a series of exploratory walking movements. These movements can be small incremental displacements randomly sampled from the robot's motion space, or they can be pre-programmed, taught trajectories covering a wide range of motion patterns. For example, the robot can perform small translations (along the X, Y, and Z axes) and rotations (around the X, Y, and Z axes) to explore various visual changes in its movement across different terrains. In some embodiments, these exploratory movements can be remotely taught by a human operator or guided by a low-level controller with a given coarse objective.
[0113] Diversity is key to data collection, ensuring that the robot can tread on different types of terrain, use different gaits, and perform actions in different initial postures, so that the collected data can represent the various real-world walking scenarios that the robot may encounter.
[0114] Data point records:
[0115] Before and after each action, the system records a complete data sample. Specifically, for each action performed by the robot, the system records the following before the action is executed:
[0116] Current image of the terrain in front of the quadrupedal bionic robot RGB image acquired by information acquisition module 10.
[0117] Current motion state of the quadrupedal bionic robot The precise pose (position and orientation) of the robot's main body in the world coordinate system.
[0118] Context information : The gait ID currently in use and the terrain category ID obtained through image analysis or preset.
[0119] Execute action The robot's actual performance relative to The incremental pose. This action is the candidate walking action provided as input to the visual dynamic field generation network during training.
[0120] After the action is performed, the system will capture images for the next moment. .
[0121] Self-supervised label generation (real visual dynamic field):
[0122] This is the core of self-supervised training. For each data point, instead of manually labeling the predicted target, we automatically generate the true (or "ground truth") visual dynamic field using an optical flow estimation algorithm. .
[0123] Specifically, a high-precision optical flow estimation algorithm (e.g., deep learning-based RAFT (RecurrentAll-PairsFieldTransforms) or LiteFlowNet, or the traditional Farneback algorithm) is applied to two consecutive frames of images. and The algorithm will calculate the image. Each pixel in the array is used to match... The two-dimensional displacement vector that occurs at the corresponding point in the vector.
[0124] The pixel-level displacement vector field output by the optical flow algorithm is considered as the result of the action. Later by arrive The real visual dynamic field between .this These will be used as training labels to supervise the predictions of the visual dynamic field generation network. .
[0125] Using the above method, each training sample includes: ( These data are stored as a large dataset for offline training or online fine-tuning of the visual dynamic field generation network. The goal of training is to enable the network to learn the mapping between the current image of the terrain in front of the quadrupedal bionic robot, the robot's state, candidate walking actions, and contextual information, and the real visual dynamic field.
[0126] In one specific embodiment, the visual dynamic field generation network (i.e., the core component of the visual consequence prediction module 40) used in this invention is trained in a self-supervised manner. This training paradigm utilizes actual interaction data between the robot and its environment to automatically generate the "truth labels" required for training, thereby greatly reducing the reliance on manual annotation and enabling the model to learn from large-scale real-world experience.
[0127] The truth generation process is the core of self-supervised training. For each training sample obtained through the aforementioned data acquisition steps ( The system uses an optical flow estimation algorithm to calculate the corresponding true visual dynamic field. .
[0128] The specific steps are as follows:
[0129] Input image pair: Images of the robot before it performs the action. and the image after the action is performed It is provided as input to the optical flow estimation algorithm.
[0130] Optical flow calculation: Image analysis using optical flow algorithms and The pixel movement between pixels affects the image. Each pixel in Calculate a two-dimensional vector This indicates that the pixel is from Move the position in the middle to The displacement required at the corresponding position.
[0131] Generate realistic visual dynamic fields All these pixel-level displacement vectors together constitute the real visual dynamic field. This venue is a... Two-dimensional vector fields with the same resolution, where Represents pixels The displacement vector actually observed.
[0132] To improve the accuracy and robustness of the ground truth, the system can employ advanced deep learning optical flow models, such as RAFT (RecurrentAll-PairsFieldTransforms) or LiteFlowNet. These models can provide more accurate optical flow estimates under conditions of large displacement and complex textures.
[0133] In calculation At this time, masks can be further applied. Only optical flow within critical regions of the path is calculated to ensure that the training process is more focused on task-related visual changes.
[0134] This was automatically generated using an optical flow algorithm. This refers to the "target output" or "truth label" used during network training.
[0135] Network optimization:
[0136] In obtaining a large amount of ( After training with data consisting of , the visual dynamic field generation network is optimized by minimizing the difference between its predicted values and the true values.
[0137] Network Input: During the training phase, for each training sample, the visual dynamic field generator network receives the following input:
[0138] Current image of the terrain in front of the quadrupedal bionic robot ;
[0139] Current motion state of the quadrupedal bionic robot (Encoded as a vector);
[0140] Actual actions performed (Encoded as a vector);
[0141] Context feature vector (Based on context encoding module 20) (Generation). These inputs are integrated into the network's encoder-decoder structure to generate the predicted visual dynamic field.
[0142] Network output: The network output is the predicted visual dynamic field. Its structure and resolution are consistent with the real visual dynamic field. Totally consistent.
[0143] Loss Function: To quantify the difference between the prediction and the true value, the system employs a loss function. Commonly used loss functions are L1 loss (Mean Absolute Error, MAE) or L2 loss (Mean Squared Error, MSE), especially at the pixel level.
[0144] ;
[0145] Alternatively, use L2 loss:
[0146] ;
[0147] in, Represents pixel coordinates It is the value of the mask at the corresponding pixel, ensuring that the loss calculation is focused on the relevant area; This represents the total loss value for the current training batch or a single training sample. Represents the pixel coordinates in the image; Represents the coordinates of all pixels within the image. A summation operator for iterative summation; Represents a normalization factor; This indicates that the network generated by the visual dynamic field is at the current time step. Predicted, in pixels Two-dimensional displacement vector at the location; This represents the pixel-level optical flow estimation obtained from actual image changes using an optical flow estimation algorithm. The true (ground truth) two-dimensional displacement vector at that location; Represents the L1 norm; This represents the square of the L2 norm, also known as the square of the Euclidean distance. L1 loss is generally more robust during training because it penalizes outliers (such as optical flow estimation errors) less.
[0148] Optimizer: The loss function's gradient is computed using the backpropagation algorithm, and then an optimizer (e.g., Adam, SGDwithMomentum) is used to update all trainable parameters in the network (including convolutional kernels, batch normalization parameters, MLP weights, and embedding layer weights in the context encoding module). The optimizer's goal is to iteratively adjust the network parameters to minimize the total loss. .
[0149] Training Objective: Through the self-supervised training process described above, the Visual Dynamics Field Generation Network learns to accurately predict the displacement of each pixel in an image, given the current visual environment, the robot's own state, the upcoming action, and the walking context. This enables the network to learn efficiently on large-scale datasets without costly manual annotation, thus providing reliable visual consequence prediction capabilities for subsequent forward-looking action planning and decision-making.
[0150] In one specific embodiment, the method and system provided by the present invention are applied to tasks involving quadrupedal bionic robots traversing complex terrain. For example, in field search and rescue or inspection, the robot needs to cross a gravel area from flat ground to reach the opposite grassland. This embodiment will be illustrated using this example.
[0151] Application scenario: Quadrupedal bionic robot traversing areas with loose rocks
[0152] Mission objective: To stably and safely traverse the rocky terrain ahead, avoiding slips or instability caused by uneven ground or slippery materials, and ultimately reaching the target area smoothly. Due to the highly unstructured and dynamic uncertainties of rocky terrain, traditional methods based on fixed gait sequences or pre-built geometric path planning are difficult to effectively adapt to its complex characteristics.
[0153] System Configuration:
[0154] Robot: A quadrupedal bionic robot equipped with a high-performance onboard computing unit (such as the NVIDIA Jetson series).
[0155] Camera: A high-resolution RGB-D camera, fixed to the robot's head, provides current images and depth information of the terrain in front of the quadrupedal bionic robot from a forward-looking perspective.
[0156] Workbench: A mixed terrain environment consisting of flat ground, gravel, and grass.
[0157] Specific implementation steps:
[0158] Initialization and task definition:
[0159] Before the mission began, the robot was positioned on flat ground, facing the gravel area to be traversed.
[0160] Expected visual dynamic field Generation: In this embodiment, the task objective is defined by a target image, which shows the visual scene that the robot should observe at the target point (e.g., the grass opposite) after successfully traversing a gravel area. Based on the current image of the terrain in front of the quadrupedal bionic robot and this target image, the system calculates the desired visual dynamic field using an optical flow estimation algorithm or a pre-trained target dynamic field generation model. .this It describes the ideal motion pattern (i.e., visual optical flow) that pixels in the current field of view should undergo in order to move smoothly and linearly from the current position to the target position.
[0161] Path critical region mask Definition: For terrain traversal tasks, the path-critical region is defined as the projection of the ground area directly in front of the robot onto the current image of the terrain in front of the quadrupedal bionic robot. This mask can be dynamically generated by analyzing depth information or using a pre-trained segmentation network, ensuring that the decision-making process focuses on the terrain regions that directly affect the robot's stability.
[0162] Information Acquisition and Context Encoding:
[0163] Information Acquisition Module 10: Real-time acquisition of RGB images of the terrain ahead. And obtain the current precise pose of the robot body. .
[0164] Context encoding module 20:
[0165] Current gait information: Based on the current task stage, the robot's gait is set to "slow and stable gait", and its ID is encoded as a gait feature vector. .
[0166] Forward terrain information: Analyze the RGB image of the foreground terrain using an image segmentation network. The main terrain within the critical area of the path was identified as "gravel," and its ID was encoded into a terrain feature vector. .
[0167] By concatenating and using a multilayer perceptron (MLP), these two feature vectors are fused to generate the final context feature vector. .this It includes background information on "walking on gravel with a slow, steady gait".
[0168] Forward-looking action planning and decision-making:
[0169] Motion generation module 30: Based on the robot's current pose Generate a series of candidate walking actions These candidate walking actions are typically small six-dimensional pose increments that encompass the conditioning movements required to maintain stability while traversing complex terrain, such as:
[0170] Forward translation at different speeds and in different directions (exploring forward strategies).
[0171] Vertical translation of the body's center of gravity (lowering the center of gravity to increase stability).
[0172] Lateral translation of the body (fine-tuning the landing point).
[0173] Minor adjustments to body posture (pitch, tilt, and yaw to adapt to ground undulations). For example, 20 candidate walking movements can be sampled, including 10 forward movements at different speeds, 5 movements involving lowering or raising the center of gravity, and 5 minor adjustments to body posture.
[0174] Visual Consequence Prediction Module 40: For each candidate walking action Compare it with the RGB image of the terrain in front. Current motion state of the quadrupedal bionic robot and context feature vectors These are input together into a pre-trained visual dynamic field generation network. The network is for each... Output a corresponding predicted visual dynamic field For example, if A relatively fast movement on a gravel road is what the network (which has mastered the unstable nature of gravel through self-supervised learning) might predict as a sudden, irregular pixel shift in the field of vision, representing visual instability caused by slipping or swaying.
[0175] Action Decision Module 50:
[0176] For each candidate walking action Calculate its difference measure . By comparing the predicted visual dynamic fields With the pre-set desired visual dynamic field In the critical area of the path The average L1 distance within the range is obtained.
[0177] ;
[0178] Choose to The smallest candidate walking action is selected as the optimal walking action. .this Within the current control cycle, this refers to the minute movements that best enable the robot to produce a visual effect that closely approximates the ideal, smooth forward movement. For example, the system might discard a fast movement that predicts violent shaking and instead choose a movement that lowers the center of gravity and moves at a slower pace, because its predicted visual consequences best match the expectation of smooth forward movement.
[0179] Action execution and iteration:
[0180] Robot control module 60: Receives the optimal walking motion This is converted into torque or angle commands for each joint that can be executed by the robot's whole-body controller, and then drives the robot to perform that tiny movement.
[0181] Iteration: After the robot completes its action, the information acquisition module 10 immediately acquires new terrain images ahead and updates the robot's state. The system then enters the next control loop, repeating the above steps until the task termination condition is met.
[0182] Mission termination conditions:
[0183] Visual convergence: The optimal difference measure over multiple consecutive control cycles. (i.e., the smallest) If the value remains below a preset minimum threshold (e.g., 0.01 pixels), it indicates that the visual difference between the current image of the terrain in front of the quadrupedal bionic robot and the target image is negligible.
[0184] Motion convergence: Simultaneously, the robot body executes [actions] over multiple consecutive cycles. If the norm (e.g., L2 norm) remains below a small threshold (e.g., 0.01 mm and 0.1 degrees), it indicates that the robot has essentially stopped moving.
[0185] Through the above embodiments, the method of the present invention enables a quadrupedal bionic robot to adaptively and stably traverse complex, unstructured terrain without relying on high-precision 3D environment maps and complex dynamic models, solely through visual feedback and forward-looking predictive learning of action consequences. The self-supervised training paradigm ensures that the model can continuously learn and improve from actual walking experience, thereby greatly enhancing the system's adaptability and robustness.
Claims
1. A robot vision guidance method based on artificial intelligence, characterized in that, Includes the following steps: Acquire the current image of the terrain in front of the quadrupedal bionic robot and the preset mission target image to guide the quadrupedal bionic robot to walk safely; The current motion state of the quadrupedal bionic robot and contextual information including current gait information and terrain information in front are obtained. The context information is encoded to generate a context feature vector; Within the local motion space of the quadrupedal bionic robot, multiple candidate walking actions are sampled and generated based on the current motion state of the quadrupedal bionic robot. For each candidate walking action, the current image of the terrain in front of the quadrupedal bionic robot, the current motion state of the quadrupedal bionic robot, the candidate walking action, and the context feature vector are input into a preset visual dynamic field generation network to generate a predicted visual dynamic field. Based on the current image of the terrain in front of the quadrupedal bionic robot and the preset mission target image for guiding the quadrupedal bionic robot to walk safely, a desired visual dynamic field is determined. An optimal walking action is determined from the plurality of candidate walking actions by comparing each of the predicted visual dynamic fields with the expected visual dynamic field. Control the quadrupedal bionic robot to perform the optimal walking action.
2. The robot vision guidance method based on artificial intelligence according to claim 1, characterized in that, The step of encoding the context information specifically includes: The current gait information and the terrain information ahead are respectively embedded into features to obtain a gait feature vector and a terrain feature vector; The gait feature vector and the terrain feature vector are fused to generate the context feature vector.
3. The robot vision guidance method based on artificial intelligence according to claim 1, characterized in that, The visual dynamic field generation network contains at least one feature modulation layer. In the step of inputting the current image of the terrain in front of the quadrupedal bionic robot, the current motion state of the quadrupedal bionic robot, the candidate walking action, and the context feature vector into a preset visual dynamic field generation network, the context feature vector is used to generate modulation parameters. The modulation parameters perform affine transformation on the intermediate feature map in the visual dynamic field generation network to dynamically modulate the predicted visual dynamic field.
4. The robot vision guidance method based on artificial intelligence according to claim 1, characterized in that, The specific steps for determining a desired visual dynamic field are as follows: An optical flow estimation algorithm is used to calculate the pixel displacement between the current image of the terrain in front of the quadrupedal bionic robot and the preset target image for guiding the quadrupedal bionic robot to walk safely, thereby generating the desired visual dynamic field.
5. The robot vision guidance method based on artificial intelligence according to claim 1, characterized in that, The step of comparing each of the predicted visual dynamic fields with the expected visual dynamic field specifically includes: Calculate the difference measure between each predicted visual dynamic field and the expected visual dynamic field within a preset path key region; And select the candidate walking action with the smallest difference metric as the optimal walking action.
6. The robot vision guidance method based on artificial intelligence according to claim 1, characterized in that, The visual dynamic field generation network is trained using a self-supervised learning method, and the training process includes: Collect training data that includes terrain images before the action, contextual information, the walking action, and terrain images after the action; Based on the terrain image before the action and the terrain image after the action, a ground truth visual dynamic field is generated, and the visual dynamic field generation network is trained based on the difference between the ground truth visual dynamic field and the visual dynamic field predicted by the visual dynamic field generation network.
7. The robot vision guidance method based on artificial intelligence according to claim 6, characterized in that, The specific steps for generating a truth-based visual dynamic field are as follows: An optical flow estimation algorithm is used to calculate the actual pixel displacement between the terrain image before the action and the terrain image after the action, and to generate the ground truth visual dynamic field.
8. The robot vision guidance method based on artificial intelligence according to claim 1, characterized in that, Both the predicted visual dynamic field and the expected visual dynamic field are two-dimensional vector fields, and their size corresponds to the size of the current image of the terrain in front of the quadrupedal bionic robot. Each vector in the two-dimensional vector field represents the predicted or expected visual motion trend of the corresponding terrain pixel.
9. The robot vision guidance method based on artificial intelligence according to claim 1, characterized in that, The current gait information is a gait identifier in a preset gait library, and the terrain information ahead is a terrain category or attribute obtained through image segmentation or classification.
10. A robot vision guidance system based on artificial intelligence, characterized in that, include: An information acquisition module is used to acquire the current image of the terrain in front of the quadrupedal bionic robot, the preset task target image for guiding the quadrupedal bionic robot to walk safely, the current motion state of the quadrupedal bionic robot, and contextual information including current gait information and terrain information in front. A context encoding module is used to encode the context information to generate a context feature vector; An action generation module is used to sample and generate multiple candidate walking actions based on the current motion state of the quadrupedal bionic robot within the local motion action space of the quadrupedal bionic robot. A visual consequence prediction module is configured with a preset visual dynamic field generation network, which is used to generate a predicted visual dynamic field for each of the plurality of candidate walking actions, based on the current image of the terrain in front of the quadrupedal bionic robot, the current motion state of the quadrupedal bionic robot, the candidate walking action, and the context feature vector. An action decision module is used to determine a desired visual dynamic field based on the current image of the terrain in front of the quadrupedal bionic robot and the preset task target image for guiding the quadrupedal bionic robot to walk safely, and to determine an optimal walking action from the plurality of candidate walking actions by comparing each predicted visual dynamic field with the desired visual dynamic field. A robot control module is used to control the quadrupedal bionic robot to perform the optimal walking action.
Citation Information
Patent Citations
Method and device for automatic wiring based on visual guidance
CN115629066A
Cotton field robot fusion navigation path generation method
CN120333428A
Discrepancy detection apparatus and methods for machine learning
US20150148953A1
Vision guided robot path programming
US20190047145A1
Obstacle recognition method for autonomous robots
US20200225673A1