Robot obstacle avoidance and navigation method based on multi-modal fusion and visual language model
Through multimodal fusion and visual language model, robots can achieve efficient and accurate obstacle avoidance and navigation in complex environments, solving the shortcomings of perception and decision-making in traditional methods and improving operational efficiency and safety.
Patent Information
- Application Number
- CN202510780915.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-12
AI Technical Summary
The obstacle avoidance and navigation methods of traditional mobile robots in complex dynamic environments have problems such as incomplete perception, inaccurate decision-making, and insufficient computing resources, resulting in low operating efficiency and poor security in complex environments.
Multimodal fusion and visual language model are adopted, through the fusion of camera images, lidar point clouds and IMU data, combined with multi-layer Transformer modules and visual language models, semantic maps and action strategies are generated to achieve real-time environmental perception and decision-making.
It improves the perception accuracy and decision-making efficiency of robots in complex environments, can deal with dynamic obstacles in real time, ensure safe operation, and support efficient task execution in multiple scenarios.
Smart Images

Figure CN120293156A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of mobile robots, and particularly relates to a robot obstacle avoidance and navigation method based on multi-modal fusion and vision language model. Background Art
[0002] With the rapid development of artificial intelligence and robot technology, mobile robots, with their high degree of automation and flexibility, are being applied more and more widely and deeply in many fields such as industry, service, and medical care. In industrial production, mobile robots can efficiently complete tasks such as material handling and assembly, significantly improving production efficiency and reducing labor costs; in the service field, they can provide services such as guiding tours and cleaning, bringing convenience to people's lives; in medical scenarios, mobile robots can assist medical staff in tasks such as drug delivery and patient transfer, reducing the workload of medical staff.
[0003] However, the traditional methods relied on by current mobile robots for obstacle avoidance and navigation have many limitations, seriously restricting their application effectiveness in complex dynamic environments, which are specifically manifested as follows: Traditional mobile robots usually use a single sensor to achieve environmental perception, such as lidar or camera commonly. Lidar obtains the distance information of surrounding objects by emitting laser beams and measuring the time of reflected light. Although it can accurately measure the distance, it lacks the ability to perceive semantic information such as the color and texture of objects; the camera is good at capturing visual information of images, but in complex lighting conditions such as low light and strong light interference, its imaging quality will be severely affected, resulting in deviations in environmental perception. In a complex dynamic environment, there are often objects of various shapes, materials, and motion states. A single sensor is difficult to comprehensively and accurately perceive this information, thus unable to provide a reliable environmental cognition basis for the robot and increasing the risk of the robot colliding with obstacles.
[0004] Traditional obstacle avoidance and navigation algorithms mainly make decisions based on simple rules or pre-set map information, lacking in-depth understanding of semantic information. When faced with natural language instructions, traditional algorithms are difficult to effectively associate the semantic content in the instructions with actual environmental information and cannot accurately understand the true intention of the instructions. For example, when an operator gives a natural language instruction such as "Go to the meeting room and avoid the moving cleaning cart", traditional algorithms may not be able to accurately identify the location of the "meeting room" and the motion state of the "cleaning cart", resulting in a slow and error-prone decision-making process and unable to quickly and accurately plan a suitable path.
[0005] Traditional methods often require a lot of computing resources when dealing with environmental perception and decision-making tasks. Due to the limited computing power of the robot itself, many complex computing tasks have to rely on cloud servers to complete. However, it takes a certain amount of time to transmit data between the robot and the cloud, which will cause high latency in the system. In a dynamic environment, environmental information changes rapidly. This high latency makes it impossible for the robot to respond to environmental changes in a timely manner, making it difficult to meet the needs of real-time obstacle avoidance and navigation, which seriously affects the robot's operating efficiency and safety.
[0006] In recent years, multimodal visual language models (such as Qwen) have demonstrated strong capabilities in image understanding, semantic reasoning, and cross-modal interaction. Multimodal visual language models can simultaneously process information from multiple modalities, such as images and texts. By learning from a large amount of data, they have acquired rich semantic knowledge and reasoning capabilities. They can not only accurately understand information such as objects and scenes in images, but also make reasonable inferences and decisions based on text instructions.
[0007] However, how to organically combine these advanced model capabilities with multimodal sensor fusion technology and successfully apply them to obstacle avoidance and navigation tasks of mobile robots is still a key issue that needs to be solved in the current field of robotics. By fusing multimodal visual language models with multimodal sensors, it is expected to give full play to the advantages of both, effectively overcome the limitations of traditional methods, and provide more efficient and accurate solutions for obstacle avoidance and navigation of mobile robots in complex dynamic environments. Summary of the invention
[0008] In view of the deficiencies in environmental perception, decision making and real-time control in the prior art, the purpose of the present invention is to provide a robot obstacle avoidance and navigation method based on multimodal fusion and visual language model, which combines the multimodal visual language model with the multimodal sensor fusion technology to enhance the perception, decision-making and real-time response capabilities of mobile robots in complex dynamic environments.
[0009] The present invention achieves the above-mentioned purpose through the following technical solutions: A robot obstacle avoidance and navigation method based on multimodal fusion and visual language model includes the following steps: Collect different modal data in real time through multiple sensors, and perform time synchronization and normalization processing; Extract features of different modal data and fuse them through an intermediate layer, where the intermediate layer is a graph neural network that fuses image features, point cloud features, and IMU features to generate a unified perception output; Input the fused multi-modal data into a vision-language model, and generate a semantic map of the environment using the semantic segmentation and object detection results of the model; among them, use a vision encoder to encode the visual input to generate a visual feature vector, use a language encoder to encode the language input to generate a language feature vector, and implement the interaction between the visual feature and the language feature through a cross-attention mechanism to generate cross-modal features f cross , which is expressed as the following formula:
[0010] Among them, α i,j represents the attention weight, which is used to measure the f lang relationship between the i -th element in the language feature vector f vis,j and the j -th element in the visual feature vector f vis,j is the j -th element in the visual feature vector, W Q 、W K 、 W V is a learnable weight matrix, d is a scaling factor, T represents the transpose of the matrix (Transpose); Use a multi-layer Transformer module to process the cross-modal features to generate visual analysis results; Combine the natural language instructions and the visual analysis results to generate an action strategy; Convert the generated action strategy into a control signal executable by the robot to achieve closed-loop control.
[0011] According to a robot obstacle avoidance and navigation method based on multi-modal fusion and vision-language model provided by the present invention, when performing time synchronization processing, for the image acquisition device, lidar, and IMU device, use a PPS pulse signal or a shared clock source for coarse time alignment; For the modality data that is not strictly synchronized, use an interpolation or resampling method based on timestamps to align the discrete time points through a linear interpolation formula, which is expressed as:
[0012] Among them, x(t) is the interpolation data at time t , x k andx k+1 for adjacent timestamps t k and t k+1 the raw data at For the mode with non - linear time drift, the Kalman filter is used to estimate the time offset Δt, and the timestamp is dynamically adjusted through the state update equation, which is expressed as:
[0013] where K k is the Kalman gain t ref is the reference time t imu,k is the IMU timestamp
[0014] According to a robot obstacle avoidance and navigation method based on multi - modal fusion and visual language model provided by the present invention, for image features, a pre - trained CNN model is used to extract features from the images captured by the image acquisition device to obtain high - level semantic feature vectors; Global average pooling is performed on the output of the intermediate layer or fully - connected layer of the CNN model to compress the three - dimensional feature map C×H×W into a one - dimensional feature vector C×1, and its calculation formula is:
[0015] where F c,i,j is the activation value of the c - th channel of the feature map at the position (i,j) and f img is the compressed image feature vector
[0016] According to a robot obstacle avoidance and navigation method based on multi - modal fusion and visual language model provided by the present invention, for point cloud features, the PointNet or PointNet++ network is used to extract geometric features from the point cloud data collected by the lidar, specifically including: The original point cloud data is represented as an N×3 matrix, where N is the number of points in the point cloud, and each row contains three - dimensional coordinates (x, y, z); Feature embedding is performed on each point through a multi - layer perceptron MLP to generate high - dimensional feature vectors, and an alignment transformation of the feature space is performed through the T - Net network. The calculation formula of the transformation matrix T is:
[0017] where P is the original point cloud matrix P' is the transformed point cloud matrix, MLPθ a multi-layer perceptron with parameter θ; Use max pooling to aggregate the features of all points to generate a global point cloud feature vector f lidar , expressed as:
[0018] where, f n is the feature vector of the nth point, R D is the feature dimension.
[0019] According to a robot obstacle avoidance and navigation method based on multi-modal fusion and vision language model provided by the present invention, for IMU features, for the accelerometer data a( t ) = a x ( t ), a y ( t ), a z ( t )] T and gyroscope data ω ( t ) = ω x ( t ), ω y ( t ), ω z ( t )] T , perform the following processing: Use a complementary filter or a Kalman filter to denoise the raw data and estimate the attitude angles, including the pitch angle θ , roll angle ϕ , yaw angle ψ , and the update formula of its complementary filter is:
[0020] where, θ est is the estimated value, representing the estimate of a certain parameter or state at the current moment, α is the learning rate or weighting factor, controlling the balance between new information and the old estimate, ω y ( t ) is the angular velocity, Δ t is the sampling time interval, a x ( t) is the acceleration, ||a( t )|| is the magnitude of the acceleration; Perform time integration on the gyroscope data to generate the rotation quaternion q(t) = [q0(t), q1(t), q2(t), q3(t)] T , and its quaternion differential equation is:
[0021] where ⊗ is the quaternion multiplication operator.
[0022] Combine the attitude angle, rotation quaternion, and the magnitude of the acceleration ||a( t )|| into the IMU feature f imu : .
[0024] According to a robot obstacle avoidance and navigation method based on multimodal fusion and visual language model provided by the present invention, when using a graph neural network to fuse the extracted image features, point cloud features, and IMU features, it includes: Define the graph structure, regard the features of each modality as a node in the graph, where the image feature corresponds to node A, the point cloud feature corresponds to node B, the IMU feature corresponds to node C, the edges between the nodes represent the correlation between modalities, and the edge weights are determined by the correlation between modalities; Initialize the node features, the initial features of each node are generated by the feature extractor of the corresponding modality. Among them, the features of node A are extracted by the CNN model, the features of node B are extracted by the PointNet or PointNet++ network, and the features of node C are generated by the complementary filter or Kalman filter; Update the node features through the message passing mechanism, expressed as the following formula:
[0025] where is the feature of node l in the i th layer, is the set of neighbor nodes of node i , W and are learnable weight matrices, σ is the activation function; After multiple rounds of message passing, use the global pooling operation to aggregate the features of all nodes into a global feature vector, and this global feature vector is the output of multimodal fusion.
[0026] According to a robot obstacle avoidance and navigation method based on multimodal fusion and visual language model provided by the present invention, when the image featuresf img , point cloud features f lidar and IMU features f imu After fusion, generate fused multi-modal features f fused : Take the fused multi-modal features f fused as visual input, and at the same time take the user's natural language instructions U as language input to construct the input pair of the vision-language model ([[]] f fused , U ); Use a vision encoder to encode the visual input f fused to generate a visual feature vector f vis , expressed by the following formula: f vis = Vis enc ( f fused ; θ vis ) where θ vis are the vision encoder parameters; Adopt an object detection algorithm to perform object bounding box localization and category classification on the object to generate object detection results; combine the visual feature vector f vis and the object detection results to generate a semantic map of the environment through spatial relationship modeling.
[0027] According to a robot obstacle avoidance and navigation method based on multi-modal fusion and vision-language model provided by the present invention, use a language encoder to encode the language input U to generate a language feature vector f lang , expressed by the following formula: f lang = Lang enc ( U ; θ lang ) where θ lang are the language encoder parameters.
[0028] According to a robot obstacle avoidance and navigation method based on multi-modal fusion and vision-language model provided by the present invention, use a multi-layer Transformer module for cross-modal features fcross Perform progressive enhancement and output a semantic map M sem and the target position T The visual analysis results of are expressed by the following formula: f enhanced = Transformer N dec ( f cross ; θ cross ) M sem = Decoder map ( f enhanced ; θ map ) T = Detector( f enhanced ; θ det ) A = Policy( f enhanced ; θ policy ) Among them, f enhanced is the enhanced cross-modal feature vector, Transformer N dec is a module composed of N-layer Transformer decoders, Decoder map is the semantic map decoder, Detector is the target detector, Policy is the policy generation module, θ cross and θ map and θ det and θ policy are the parameters of the cross-modal Transformer decoder, semantic map decoder, target detector, and policy generation module respectively; Based on the semantic map M sem, perform semantic annotation on different regions in the environment; According to the target position T , mark specific targets in the semantic map; Combine the action policy A and update the semantic map in real time to reflect the dynamic changes in the environment.
[0029] A robot obstacle avoidance and navigation method based on multi-modal fusion and vision language model provided by the present invention, the action strategy A includes: In the navigation task, infer the target position according to the semantic map and generate a sequence of path points: Based on cross-modal features f cross , use the A* algorithm or Dijkstra algorithm to search for the optimal path in the semantic map M sem, and generate a sequence of path points M. The generated sequence of path points M is expressed as {( x 1, y 1),( x 2, y 2),…,( xn , yn )}, where ([[]] xi , yi ) is the coordinate of the i th path point; The obstacle avoidance rule formed by using the dynamic window method to plan the path in the obstacle avoidance task: When a dynamic obstacle is detected, adjust the path through the local path replanning algorithm to generate a new sequence of path points M'.
[0030] It can be seen that the present invention proposes a robot obstacle avoidance and navigation method based on multi-modal fusion and vision language model. Compared with the traditional robot obstacle avoidance and navigation technology, it has significant and multi-faceted beneficial effects, providing strong support for the efficient operation of robots in various complex scenarios. The specific manifestations are as follows: The present invention integrates camera images, lidar point clouds, and IMU data, giving full play to the advantages of each sensor: Camera images provide rich visual information, enabling the robot to identify features such as the color, shape, and texture of objects; Lidar point clouds accurately measure the distance and position of objects, constructing a detailed three-dimensional environment model; IMU data reflects the motion state of the robot in real time, including information such as acceleration and angular velocity. Through the fusion processing of multi-modal data, the robot can perceive the environment from multiple dimensions, construct a more comprehensive and accurate environmental perception, and effectively cope with various challenges in complex dynamic environments. For example, in a home environment, the robot can accurately identify different objects such as furniture, appliances, and pets, as well as their positions and motion states, providing a solid data foundation for obstacle avoidance and navigation.
[0031] The method of the present invention can real-time sense the movement trajectory and speed of dynamic obstacles. The camera image can capture the appearance changes of the obstacles, the lidar point cloud can accurately measure their position changes, and the IMU data can assist in judging the movement trend of the obstacles. By comprehensively analyzing these multi-modal data, the robot can predict the action path of dynamic obstacles in advance and timely adjust its own movement strategy, effectively avoiding collisions and ensuring operation safety. For example, in a logistics warehouse, when a staff member or a forklift suddenly appears on the robot's traveling path, the robot can quickly sense and make an avoidance action.
[0032] The present invention combines natural language instructions and visual analysis results to generate an action strategy. The visual language model deeply analyzes the camera image, identifies visual information such as objects, scenes, signs, etc. in the environment, and at the same time accurately analyzes the intentions, goals and constraints in the natural language instructions. By fusing natural language and visual information, the robot can better adapt to different scenarios. Based on the semantic map, the visual analysis results can real-time update the environmental information, and the natural language instructions can adjust the task goals according to the actual situation. For example, in a shopping mall environment, the original task of the robot is to go to a certain store, but on the way, it encounters the situation that the store location has changed or new obstacles have been added. By discovering the changes through visual analysis and combining with the possible natural language correction instructions of the user, the robot can quickly adjust the action strategy and re-plan the path to ensure the smooth completion of the task.
[0033] In the obstacle avoidance task, the present invention uses the dynamic window method to plan the path and form obstacle avoidance rules. The dynamic window method searches for feasible speed combinations in the robot's speed space according to the robot's kinematic constraints and the current environmental information, and generates a smooth and safe obstacle avoidance path. Combining the environmental perception data provided by the multi-modal sensor fusion, the robot can real-time sense the positions, speeds and movement trends of the surrounding obstacles, dynamically adjust the window size and search range, and quickly plan the optimal obstacle avoidance path.
[0034] In the navigation task, the present invention infers the target position according to the semantic map and generates a sequence of path points. Through the understanding of visual information by the visual language model, the robot can accurately identify the landmark objects and areas in the semantic map, and thus accurately infer the target position. Based on the target position, the robot can generate a reasonable sequence of path points to guide itself to reach the destination smoothly. For example, in a hospital environment, the robot knows the locations of different departments according to the semantic map, and combined with the natural language instructions of the patient, it can plan the optimal path from the current position to the target department, improving the accuracy and efficiency of navigation.
[0035] The method of the present invention has strong versatility and adaptability and can be widely applied to various scenarios such as home service, logistics transportation, and industrial production. In the home service scenario, the robot can complete tasks such as cleaning, delivering items, and accompanying, bringing convenience to family life; in the logistics transportation scenario, the robot can efficiently carry goods in complex environments such as warehouses and workshops, improving the efficiency and automation level of logistics transportation; in the industrial production scenario, the robot can assist workers in completing tasks such as material handling and assembly, reducing the labor intensity of workers. By applying in different scenarios, the method of the present invention can give full play to the role of the robot and create greater value for enterprises and society.
[0036] In summary, a robot obstacle avoidance and navigation method based on multi-modal fusion and visual language model of the present invention significantly improves the perception, decision-making, and execution capabilities of the robot in complex environments through multi-dimensional perception fusion, natural language interaction support, and efficient real-time control, provides strong support for the wide application and intelligent development of the robot, and has broad market prospects and important social significance.
[0037] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Description of the Drawings
[0038] Figure 1 is a flowchart of an embodiment of a robot obstacle avoidance and navigation method based on multi-modal fusion and visual language model of the present invention.
[0039] Figure 2 is a flowchart of time synchronization processing in an embodiment of a robot obstacle avoidance and navigation method based on multi-modal fusion and visual language model of the present invention.
[0040] Figure 3 is a flowchart of extracting image features in an embodiment of a robot obstacle avoidance and navigation method based on multi-modal fusion and visual language model of the present invention.
[0041] Figure 4 is a flowchart of extracting point cloud features in an embodiment of a robot obstacle avoidance and navigation method based on multi-modal fusion and visual language model of the present invention.
[0042] Figure 5 is a flowchart of extracting IMU features in an embodiment of a robot obstacle avoidance and navigation method based on multi-modal fusion and visual language model of the present invention.
[0043] Figure 6 is a flowchart of outputting the global feature vector of multi-modal fusion in an embodiment of a robot obstacle avoidance and navigation method based on multi-modal fusion and visual language model of the present invention.
[0044] Figure 7 This is a flowchart for generating a semantic map in an embodiment of a robot obstacle avoidance and navigation method based on multimodal fusion and vision language model of the present invention.
[0045] Figure 8 This is a flowchart for generating visual analysis results in an embodiment of a robot obstacle avoidance and navigation method based on multimodal fusion and vision language model of the present invention.
[0046] Figure 9 This is a system schematic diagram of an embodiment of a robot obstacle avoidance and navigation method based on multimodal fusion and vision language model of the present invention. Detailed implementation manners
[0047] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without making creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.
[0048] The mention of "embodiment" in this document means that the specific features, structures or characteristics described in connection with the embodiment can be included in at least one embodiment of the present application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0049] See Figure 1 , this embodiment provides a robot obstacle avoidance and navigation method based on multimodal fusion and vision language model, and the method includes the following steps: Step S1, collecting different modal data in real time through multiple sensors, and performing time synchronization processing and normalization processing; Step S2, extracting the features of different modal data and fusing them through an intermediate layer, where the intermediate layer is a graph neural network for fusing image features, point cloud features and IMU features to generate a unified perception output; Step S3: Input the fused multi-modal data into the vision-language model, and use the semantic segmentation and object detection results of the model to identify objects (such as tables and chairs) in the image and their spatial relationships, generate a semantic map of the environment, and mark the passable areas and obstacles. Among them, use a vision encoder to encode the visual input to generate a visual feature vector, use a language encoder to encode the language input to generate a language feature vector, implement the interaction between the visual feature and the language feature through a cross-attention mechanism to generate cross-modal features, and use a multi-layer Transformer module to process the cross-modal features to generate visual analysis results; Step S4: Combine the natural language instructions and the visual analysis results to generate an action strategy, including using the dynamic window method to plan a path in the obstacle avoidance task to form an obstacle avoidance rule, and inferring the target position based on the semantic map and generating a sequence of path points in the navigation task. For example, the user issues an instruction through voice or text: "Take me to the living room", and the vision-language model combines the semantic map to infer the position of the living room, generates a sequence of path points, and uses the dynamic window method (DWA) to optimize the local obstacle avoidance performance.
[0050] Step S5: Convert the generated action strategy into a control signal executable by the robot, and drive the robot to move along the planned path to achieve closed-loop control. Among them, continuously collect new sensor data and update the path planning in real time.
[0051] In the above step S1, use an image acquisition device (such as an RGB camera) to capture the environmental image with a resolution of 1920×1080; use a lidar to collect depth data with a range of 0~10 meters; use an IMU to collect acceleration and angular velocity data; among them, collect environmental data in real time through multiple sensors such as cameras, lidars, and IMUs (inertial measurement units), and the environmental data includes visual data (images), depth data (lidar point clouds), and motion state data (IMU acceleration and angular velocity).
[0052] In the above step S1, perform time synchronization and format conversion on the collected multi-modal data to ensure the consistency of different sensor data in time and space, which is specifically achieved through timestamp alignment and normalization processing.
[0053] Among them, as Figure 2 shown, when performing time synchronization processing, for the image acquisition device, lidar, and IMU device, use the PPS pulse signal or a shared clock source for coarse time alignment; For the modal data that is not strictly synchronized, use the interpolation or resampling method based on timestamps to align the discrete time points through a linear interpolation formula, expressed as:
[0054] Among them,x(t) is the interpolated data at time t , x k and x k+1 are adjacent timestamps t k and t k+1 are the original data at For modalities with non-linear time drift (such as IMU), the Kalman filter is used to estimate the time offset Δt, and the timestamp is dynamically adjusted through the state update equation, expressed as:
[0055] where K k is the Kalman gain t ref is the reference time t imu,k is the IMU timestamp
[0056] In the above step S2, as Figure 3 shown, for image features, a pre-trained CNN model (such as ResNet or EfficientNet) is used to extract features from the images captured by the image acquisition device to obtain high-level semantic feature vectors Global average pooling is performed on the output of the intermediate layer or fully connected layer of the CNN model to compress the three-dimensional feature map C×H×W into a one-dimensional feature vector C×1, and its calculation formula is:
[0057] where F c,i,j is the activation value of the c-th channel of the feature map at position (i,j) f img is the compressed image feature vector
[0058] In the above step S2, as Figure 4 shown, for point cloud features, the PointNet or PointNet++ network is used to extract geometric features from the point cloud data collected by the lidar, specifically including: The original point cloud data is represented as an N×3 matrix, where N is the number of points in the point cloud, and each row contains three-dimensional coordinates (x, y, z); Feature embedding is performed on each point through a multi-layer perceptron MLP to generate high-dimensional feature vectors, and the feature space is aligned and transformed through the T-Net network. The calculation formula for the transformation matrix T is:
[0059] Among them, P is the original point cloud matrix, P' is the transformed point cloud matrix, and MLP θ is a multi-layer perceptron with parameters θ; Use max pooling to aggregate the features of all points to generate a global point cloud feature vector f lidar , expressed as:
[0060] Among them, f n is the feature vector of the nth point, R D is the feature dimension.
[0061] In the above step S2, as Figure 5 shown, for the IMU feature, for the accelerometer data a( t ) = a x ( t ), a y ( t ), a z ( t )] T and the gyroscope data ω ( t ) = ω x ( t ), ω y ( t ), ω z ( t )] T , perform the following processing: Use a complementary filter or a Kalman filter to denoise the original data and estimate the attitude angles, including the pitch angle θ , roll angle ϕ , yaw angle ψ , and the update formula of its complementary filter is:
[0062] Among them, θ est is the estimated value, representing the estimate of a certain parameter or state at the current moment, α is the learning rate or weighting factor, controlling the balance between new information and the old estimate, ω y ( t ) is the angular velocity, Δt is the sampling time interval, a x ( t ) is the acceleration, and ||a( t )|| is the acceleration magnitude; Perform time integration on the gyroscope data to generate the rotation quaternion q(t) = [q0(t), q1(t), q2(t), q3(t)] T , and its quaternion differential equation is:
[0063] where ⊗ is the quaternion multiplication operator.
[0064] Combine the attitude angle, rotation quaternion, and acceleration magnitude ||a( t )|| into the IMU feature f imu :
[0065] Through the above steps, high-level semantic features of camera images, geometric features of lidar point clouds, and motion state features of IMUs are extracted respectively, providing high-quality and structured feature representations for subsequent multi-modal feature fusion.
[0066] In the above step S2, as Figure 6 shown, when using a graph neural network to fuse the extracted image features, point cloud features, and IMU features, it includes: Define the graph structure, regarding the features of each modality as a node in the graph, where the image features correspond to node A, the point cloud features correspond to node B, and the IMU features correspond to node C. The edges between the nodes represent the correlations between modalities, and the edge weights are determined by the correlations between modalities; Initialize the node features. The initial features of each node are generated by the feature extractor of the corresponding modality. Among them, the features of node A are extracted by the CNN model, the features of node B are extracted by the PointNet or PointNet++ network, and the features of node C are generated by a complementary filter or a Kalman filter; Update the node features through the message passing mechanism, expressed as the following formula:
[0067] where is the feature of node l in the i th layer, is the set of neighbor nodes of node i , W and are learnable weight matrices, σis an activation function; After multiple rounds of message passing, a global pooling operation is used to aggregate the features of all nodes into a global feature vector, which is the output of multimodal fusion.
[0068] In the above step S3, as Figure 7 shown, after fusing the image features f img , point cloud features f lidar and IMU features f imu a fused multimodal feature f fused is generated: Taking the fused multimodal feature f fused as the visual input, and at the same time taking the user's natural language instruction U as the language input, an input pair for building a vision-language model (such as Qwen) is constructed( f fused , U ); Using a vision encoder (such as ViT or CLIP) to encode the visual input f fused to generate a visual feature vector f vis , which is expressed by the following formula: f vis = Vis enc ( f fused ; θ vis ) where θ vis are the parameters of the vision encoder; Adopting an object detection algorithm (such as YOLOv8 or Faster R-CNN) to perform object bounding box localization and class classification on the object to generate object detection results; combining the visual feature vector f vis and the object detection results, a semantic map of the environment is generated through spatial relationship modeling.
[0069] As Figure 8 shown, using a language encoder to encode the language input U to generate a language feature vector f lang , which is expressed by the following formula: f lang = Lang enc ( U ; θlang ) Among them, θ lang are the parameters of the language encoder; Convert the user's natural language instruction U (such as "Take me to the kitchen") into a sequence of tokens U = u 1, u 2,…, uL , and generate the language feature vector f lang .
[0070] Implement the interaction between visual features f vis and language features f lang through the cross-attention mechanism to generate cross-modal features f cross , expressed as the following formula:
[0071] Among them, α i,j represents the attention weight, which is used to measure the degree of association between the f lang -th element in the language feature vector i and the f vis,j -th element in the visual feature vector j , f vis,j is The j -th element in the visual feature vector, W Q 、W K 、 W V is a learnable weight matrix, d is a scaling factor, T represents the transpose of the matrix (Transpose), that is, ( W Q f lang ) T represents the transpose of the matrix W Q f lang .
[0072] The above formula measures the correlation between language features and visual features through dot product operations. The softmax function converts these correlations into a probability distribution, i.e., attention weights, which can accurately capture the complex interaction relationships between the visual and language modalities. The generated attention weights are used to perform weighted summation on the visual features, thereby integrating language information into the visual features to generate cross-modal features. These cross-modal features can effectively integrate the information of both modalities, enabling the generated cross-modal features to contain both visual and language semantic information. And by adjusting W Q 、W K 、W V these matrices, the model can learn the optimal mapping relationship between different modal features, thereby improving the quality of cross-modal features. The basic idea of the cross-attention mechanism can be easily extended to handle more modal cases. For example, if it is necessary to add the audio modality, only the corresponding query, key, and value matrices need to be introduced, and the attention weights are calculated and feature fusion is performed in a similar manner, making the model highly scalable.
[0073] Using multiple layers of Transformer modules to gradually enhance the cross-modal features f cross and output the semantic map M sem and the visual analysis results of the target position T , which are expressed by the following formula: f enhanced =Transformer N dec ( f cross ; θ cross ) M sem =Decoder map ( f enhanced ; θ map ) T =Detector( f enhanced ; θ det ) A=Policy( f enhanced ; θ policy ) where f enhancedis the enhanced cross-modal feature vector, Transformer N dec denotes a module composed of N-layer Transformer decoders, Decoder map is the semantic map decoder, Detector is the target detector, Policy is the policy generation module, θ cross 、 θ map 、 θ det 、 θ policy are the parameters of the cross-modal Transformer decoder, semantic map decoder, target detector, and policy generation module respectively; Based on the semantic map M sem, semantic annotation is performed on different regions in the environment (such as "road", "obstacle", "building", etc.); according to the target position T , specific targets (such as "pedestrian in front", "vehicle on the left", etc.) are marked in the semantic map; combined with the action policy A, the semantic map is updated in real time to reflect the dynamic changes in the environment.
[0074] In this embodiment, the action policy A includes: Sequence of path points: Based on the cross-modal feature f cross , the A* algorithm or Dijkstra algorithm is used to search for the optimal path in the semantic map M sem, and a sequence of path points M is generated. The generated sequence of path points M is represented as {( x 1, y 1),( x 2, y 2),…,( xn , yn )}, where ([[]] xi , yi ) is the coordinate of the i th path point; Obstacle avoidance rule: When a dynamic obstacle is detected, the path is adjusted through a local path replanning algorithm to generate a new sequence of path points M'. For example, when an obstacle is detected in front, the obstacle avoidance rule can include: decelerating or stopping to avoid collision; adjusting the robot's posture to bypass the obstacle; according to the distribution of the obstacles, selecting a suitable obstacle avoidance path.
[0075] Through the above steps, the effective fusion of image, point cloud, and IMU features is achieved, and the semantic segmentation and target detection capabilities of the vision-language model are used to generate a high-precision semantic map, target position, and action policy, significantly improving the semantic understanding and task planning capabilities in complex environments.
[0076] In this embodiment, it further includes: using a lightweight model to accelerate the vision-language model and deploying the optimized vision-language model on edge devices. For example, using TensorRT to accelerate the Qwen model to reduce the inference latency, and deploying the optimized model on NVIDIA Jetson devices to achieve low-latency real-time control.
[0077] In practical applications, a robot obstacle avoidance and navigation method proposed in this embodiment, which is based on multi-modal fusion (fusing camera images, lidar point clouds, and IMU data) and a vision-language model (such as Qwen), realizes the deep combination of visual features and language features through the vision-language model, completes the semantic understanding of the environment and task planning, and adopts a series of scientific and effective steps during the inference process.
[0078] The vision-language model uses a vision encoder to perform semantic segmentation on images, and can accurately label the object categories in the images, such as shelves, storage locations, conveyor lines, etc. At the same time, it uses object detection algorithms (such as YOLO or Faster R-CNN) to accurately identify the positions and sizes of specific targets. In a complex environment, this ability enables the robot to clearly identify various objects and avoid misjudging the environment. For example, in a logistics warehouse, the robot can accurately identify the specific positions and layouts of the shelves, as well as the status of the goods on the conveyor line, providing detailed environmental information for subsequent task execution.
[0079] Combining the results of semantic segmentation and object detection, the vision-language model can generate a semantic map of the environment, which details information such as the workshop layout and obstacle distribution. This semantic map not only contains the geometric information of the environment but also assigns specific semantic meanings to each area and object. For example, in a factory workshop, the semantic map can clearly show the layout information such as the shelves on the left, the conveyor line on the right, and the buffer racks and conveyor lines in the center of the workshop, enabling the robot to understand the environmental structure from a macroscopic level and providing strong support for task planning and path navigation.
[0080] The language encoder parses the user instructions and can quickly extract key information to clarify the target location and specific requirements of the task. For example, for the instruction "Go to conveyor line No. 1 to load 1 bin", the language encoder can accurately parse the target location as "conveyor line No. 1", thus ensuring that the robot can accurately understand the user's intention and avoid task failure caused by incorrect instruction understanding.
[0081] The visual language model's ability to parse language instructions enables it to flexibly adapt to various forms of natural language instructions. Users can issue tasks in concise and clear language without following specific formats or rules. For example, users can say "deliver the goods to the warehouse door" or "go to the office to get the documents", and the robot can accurately understand and perform the corresponding tasks, greatly improving the user experience and reducing the difficulty of operation.
[0082] Using the cross-attention mechanism, the visual language model can accurately match language instructions with semantic maps. For example, based on the information in the semantic map that the "conveyor line" is located in the center of the workshop, combined with the language instruction "go to conveyor line 1", the model can generate a path plan starting from the current position, bypassing obstacles, and reaching the conveyor line interface along the channel, so that the robot can combine language instructions with the actual environment to generate a more reasonable and efficient path.
[0083] While the robot is performing a task, the environment may change, such as new obstacles appearing or the target location moving. The cross-modal reasoning capability of the visual language model enables the robot to perceive environmental changes in real time and dynamically adjust the path based on the new information. For example, when the robot encounters a temporarily placed obstacle on the way to the target location, the model can quickly replan the path, bypass the obstacle and continue to move forward, ensuring the successful completion of the task.
[0084] In the decision-making stage, the visual language model can output specific action strategies, including waypoint sequences and obstacle avoidance rules. For example, the waypoint sequence [(x1, y1), (x2, y2), (x3, y3)] provides a clear route for the robot to follow, and the obstacle avoidance rule "adjust the local path when a dynamic obstacle is detected" ensures that the robot can respond in time to avoid collision when encountering an obstacle. These specific action strategies provide detailed guidance for the robot's operation, enabling the robot to perform tasks according to the predetermined plan.
[0085] The decision-making process fully considers various factors in the actual scene, such as obstacle distribution, target location, etc. Through reasonable path planning and obstacle avoidance rules, the robot can complete tasks safely and efficiently in complex environments. For example, in the above example of "avoiding obstacles in front and going to the door", the generated path point sequence [(0, 0), (1, 1), (3, 2), (5, 0)] can guide the robot to bypass the box and reach the door smoothly, improving the success rate of task execution.
[0086] Using lightweight models and hardware acceleration techniques can effectively reduce the inference latency of vision-language models. Under the premise of ensuring model accuracy, lightweight models reduce the number of model parameters and computational complexity, improving the inference speed. Hardware acceleration techniques then use specialized hardware devices (such as GPUs, TPUs, etc.) to accelerate the model, further shortening the inference time. This enables the robot to respond to environmental changes in real-time and adjust its action strategy promptly.
[0087] In a complex and dynamic environment, the operation of a robot requires a high degree of stability and real-time performance. Through real-time optimization processing, the vision-language model can ensure that the robot can still make quick and accurate decisions when facing a rapidly changing environment. For example, in a logistics warehouse, the movement of goods and personnel is relatively frequent, and the robot needs to perceive these changes in real-time and adjust its actions. Real-time optimization processing enables the robot to maintain a stable operating state in this dynamic environment, improving the reliability and stability of the system.
[0088] In summary, a multi-modal fusion robot obstacle avoidance and navigation method based on a vision-language model in this embodiment significantly improves the robot's obstacle avoidance and navigation capabilities in complex environments through precise semantic understanding, intelligent language instruction parsing, cross-modal reasoning fusion, scientific decision-making generation, and real-time optimization processing, providing strong support for the wide application and intelligent development of robots, and having broad market prospects and important practical value.
[0089] As Figure 9 shown, this embodiment also provides a robot obstacle avoidance and navigation system based on multi-modal fusion and a vision-language model, including the following modules: Data acquisition module: Used to obtain camera images, lidar point clouds, and IMU data in real-time.
[0090] Preprocessing module: Performs time synchronization and format conversion on multi-modal data.
[0091] Multi-modal fusion module: Extracts the features of different modal data and fuses them through an intermediate layer.
[0092] Vision-language model module: Responsible for image analysis and semantic reasoning.
[0093] Decision generation module: Generates an action strategy in combination with language instructions.
[0094] Real-time optimization module: Uses lightweight models and hardware acceleration techniques to improve real-time performance.
[0095] Action execution module: Converts the strategy into a control signal and drives the robot.
[0096] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0097] The above embodiments are only the preferred embodiments of the present invention, and the scope of protection of the present invention cannot be limited thereby. Any non-substantial changes and substitutions made by those skilled in the art on the basis of the present invention belong to the scope of protection required by the present invention.
Claims
1. A robot obstacle avoidance and navigation method based on multimodal fusion and vision language model, characterized in that, It includes the following steps: Collect different modality data in real time through multiple sensors, and perform time synchronization processing and normalization processing; Extract the features of different modality data and fuse them through an intermediate layer, where the intermediate layer is a graph neural network for fusing image features, point cloud features, and IMU features to generate a unified perception output; Input the fused multi-modal data into a vision-language model, and use the semantic segmentation and object detection results of the model to generate a semantic map of the environment; among them, use a vision encoder to encode the visual input to generate a visual feature vector, use a language encoder to encode the language input to generate a language feature vector, and implement the interaction between the visual feature and the language feature through a cross-attention mechanism to generate cross-modal features f cross , which is expressed by the following formula: Among them, α i,j represents the attention weight, which is used to measure the f lang degree of association between the i th element in the language feature vector f vis,j and the j th element in the visual feature vector, f vis,j is the j th element in the visual feature vector, W Q 、W K 、W V is the learnable weight matrix, d is the scaling factor, T represents the transpose of the matrix; Use a multi-layer Transformer module to process cross-modal features to generate visual analysis results; Generate an action strategy by combining natural language instructions and visual analysis results; Convert the generated action strategy into a control signal executable by the robot to achieve closed-loop control.
2. The method according to claim 1, wherein: When performing time synchronization processing, for the image acquisition device, lidar, and IMU device, use the PPS pulse signal or a shared clock source for coarse time alignment; For modality data that is not strictly synchronized, use an interpolation or resampling method based on timestamps to align discrete time points through a linear interpolation formula, expressed as: wherein, x(t) is the interpolated data at t time, x k and x k+1 are the adjacent timestamps t k and t k+1 are the original data at; For modalities with non-linear time drift, use a Kalman filter to estimate the time offset Δt, and dynamically adjust the timestamp through the state update equation, expressed as: Among them, K k is the Kalman gain, t ref is the reference time, t imu,k is the IMU timestamp.
3. The method according to claim 1, wherein: For image features, use a pre-trained CNN model to extract features from the images captured by the image acquisition device to obtain high-level semantic feature vectors; Perform global average pooling on the output of the intermediate layer or fully connected layer of the CNN model to compress the three-dimensional feature map C×H×W into a one-dimensional feature vector C×1, and its calculation formula is: in, F c,i,j The cth channel of the feature map is at position (i,j) The activation value at f img is the compressed image feature vector.
4. The method according to claim 1, wherein: For point cloud features, use the PointNet or PointNet++ network to extract geometric features from the point cloud data collected by the lidar, specifically including: Represent the original point cloud data as an N×3 matrix, where N is the number of points in the point cloud, and each row contains three-dimensional coordinates (x, y, z); Perform feature embedding on each point through a multi-layer perceptron MLP to generate high-dimensional feature vectors, and perform an alignment transformation on the feature space through the T-Net network, and the calculation formula of its transformation matrix T is: Among them, P is the original point cloud matrix, P' is the transformed point cloud matrix, and MLP θ is a multi-layer perceptron with parameters θ; Aggregate the features of all points using max pooling to generate a global point cloud feature vector f lidar , expressed as: Among them, f n is the feature vector of the nth point, R D is the feature dimension.
5. The method according to claim 1, wherein: For the IMU features, for the accelerometer data a( t ) = a x ( t ), a y ( t ), a z ( t )] T and the gyroscope data ω ( t ) = ω x ( t ), ω y ( t ), ω z ( t )] T , the following processing is performed: Denoise the original data and estimate the attitude angles, including the pitch angle, using a complementary filter or a Kalman filter θ , roll angle ϕ , and yaw angle ψ . The update formula for its complementary filter is as follows: wherein, θ est is an estimated value, representing the estimation of a certain parameter or state at the current moment, α is the learning rate or weighting factor, controlling the balance between new information and the old estimate, ω y ( t ) is the angular velocity, Δ t is the sampling time interval, a x ( t ) is the acceleration, ||a( t )|| is the acceleration magnitude; Integrate the gyroscope data over time to generate the rotation quaternion q(t) = [q0(t), q1(t), q2(t), q3(t)] T , and its quaternion differential equation is: wherein, ⊗ is the quaternion multiplication operator; Combine the attitude angle, rotation quaternion, and acceleration magnitude ||a( t )|| into an IMU feature f imu .
6. The method according to claim 1, characterized in that When using a graph neural network to fuse the extracted image features, point cloud features, and IMU features, it includes: Define a graph structure, regard the features of each modality as a node in the graph, where the image feature corresponds to node A, the point cloud feature corresponds to node B, and the IMU feature corresponds to node C, and the edges between the nodes represent the correlation between modalities, and the edge weights are determined by the correlation between modalities; Initialize the node features, and the initial features of each node are generated by the corresponding modality feature extractor. Among them, the features of node A are extracted by the CNN model, the features of node B are extracted by the PointNet or PointNet++ network, and the features of node C are generated by a complementary filter or a Kalman filter; Update the node features through a message passing mechanism, expressed as the following formula: Among them, is the l feature of the node in the i layer, is the set of neighbor nodes of the node i , W and are learnable weight matrices, σ is an activation function; After multiple rounds of message passing, a global pooling operation is used to aggregate the features of all nodes into a global feature vector, which is the output of multimodal fusion.
7. The method according to any one of claims 1 to 6, characterized in that: When fusing image features f img , point cloud features f lidar and IMU features f imu to generate fused multi-modal features f fused : The fused multi-modal features f fused are used as visual inputs, while the user's natural language instructions U are used as language inputs to construct the input pairs of the vision-language model( f fused , U ); Encode the visual input using a visual encoder f fused to generate a visual feature vector f vis , expressed by the following formula: f vis =Vis enc ( f fused ; θ vis ) Among them, θ vis are visual encoder parameters; Use a target detection algorithm to locate the target bounding box and classify the target category, generating target detection results; combine the visual feature vector f vis and the target detection results to generate a semantic map of the environment through spatial relationship modeling.
8. The method according to claim 7, characterized in that: Encode the language input U using a language encoder to generate a language feature vector f lang , which is expressed by the following formula: f lang =Lang enc ( U ; θ lang ) Among them, θ lang are the parameters of the language encoder.
9. The method according to claim 8, characterized in that: Using a multi-layer Transformer module to gradually enhance cross-modal features f cross and output a semantic map M sem and the visual analysis results of the target position T are expressed by the following formula: f enhanced =Transformer N dec ( f cross ; θ cross ) M sem =Decoder map ( f enhanced ; θ map ) T = Detector( f enhanced ; θ det ) A = Policy( f enhanced ; θ policy ) Among them, f enhanced is the enhanced cross-modal feature vector, Transformer N dec is a module representing a Transformer decoder composed of N layers, Decoder map is the semantic map decoder, Detector is the target detector, Policy is the policy generation module, θ cross 、 θ map 、 θ det 、 θ policy are the parameters of the cross-modal Transformer decoder, semantic map decoder, target detector, and policy generation module, respectively; Based on the semantic map M sem, perform semantic annotation on different regions in the environment; according to the target location T , mark specific targets in the semantic map; combine with action strategy A to update the semantic map in real time to reflect the dynamic changes in the environment.
10. The method according to claim 9, characterized in that: The action strategy A includes: In the navigation task, infer the target location based on the semantic map and generate a sequence of path points: based on cross-modal features f cross , use the A* algorithm or Dijkstra algorithm to search for the optimal path in the semantic map M sem and generate a sequence of path points M. The generated sequence of path points M is represented as {( x 1, y 1),( x 2, y 2),…,( xn , yn )}, where ([[]] xi , yi ) is the coordinate of the i th path point; An obstacle avoidance rule formed by using the dynamic window method to plan a path in an obstacle avoidance task: when a dynamic obstacle is detected, the path is adjusted through a local path replanning algorithm to generate a new path point sequence M'.
Citation Information
Patent Citations
Natural language navigation method based on radar and vision multi-mode fusion
CN113156419A
Self-adaptive multi-mode sensor fusion method and system for robot navigation and obstacle avoidance
CN118443000A
Multi-robot collaborative navigation method and system based on visual language large model
CN119756375A
Cited By
Robot control method and device based on physical constraint embedding, equipment and medium
CN120862691A
Mechanical arm control method and device based on large model, equipment and storage medium
CN120901973A
Method and device for controlling robot arm based on large model, equipment and storage medium
CN120901973B
Robot navigation method and device, electronic equipment and storage medium
CN120991880A
Power robot control method and system based on visual voice action model
CN121043156A