Visual guidance robot active path detection method

Through the visually guided robot active path detection method, multimodal sensors and deep learning models are used to build cost maps and adjust paths, solving the problem that traditional methods cannot cope in complex environments in real time, and achieving efficient and autonomous navigation of robots in complex terrain environments.

CN120088648APending Publication Date: 2025-06-03WUXI INSTITUTE OF TECHNOLOGY

Patent Information

Application Number
CN202510175100.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

Traditional robot path planning methods are not capable of expressing themselves in complex and changing environments, especially in dynamic environments, where paths cannot be adjusted in real time to cope with environmental changes.

Method used

The active path detection method of vision-guided robots is adopted to obtain visual, lidar and depth data in real time through multimodal sensors, and preprocess and target detection using convolutional neural networks and deep learning models, build cost maps and adjust real-time travel paths.

Benefits of technology

It realizes effective obstacle avoidance of static and dynamic obstacles by robots in complex terrain environments, can dynamically respond to ground changes and environmental factors, and improves the robot's adaptability and autonomy in shallow wading environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088648A_ABST
    Figure CN120088648A_ABST
Patent Text Reader

Abstract

The invention discloses a vision-guided robot active path detection method, and relates to the technical field of robots, and the method comprises the steps: obtaining visual modal data, laser radar modal data and depth modal data of a detection area of a robot in real time through a multi-modal sensor; preprocessing the visual modal data, the laser radar modal data and the depth modal data by using a convolutional neural network; a plurality of terrain type regions are separated from the preprocessed visual modal data, the preprocessed laser radar modal data and the preprocessed depth modal data by applying a color space conversion strategy; performing target detection on the plurality of terrain type regions through a deep learning model, and extracting candidate regions of terrain obstacles; pixel segmentation is carried out on the candidate region of the terrain obstacle, and the position and the region boundary of the terrain obstacle are determined; and adjusting the real-time advancing path of the robot according to the constructed cost map.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of robotics technology. Specifically, it relates to a method for actively detecting the path of a robot guided by vision. Background Art

[0002] With the continuous development of technology, the application of robotics technology has been continuously expanding in many fields. Especially in terms of automation, intelligence, and autonomous navigation, robots are gradually being applied to more complex environments, such as disaster rescue, exploration, military reconnaissance, etc. Especially in complex terrain environments, the challenges faced by robots are more diverse and severe. A complex terrain environment generally refers to a terrain with rough ground, irregular obstacles, variable weather and geographical conditions, and possible dynamic change factors.

[0003] Traditional robot path planning methods often rely on static environment maps or relatively simplified sensor inputs. Although they can solve some simple problems, when faced with complex and changing environments, these methods are inadequate.

[0004] In the prior art, the Chinese invention patent with the publication number CN116339342A discloses an autonomous path planning and obstacle avoidance system and method. This patent proposes a path planning method based on a pre-constructed map. Although it can effectively plan paths, it has poor adaptability in dynamic environments and cannot adjust paths in real time to cope with environmental changes. The Chinese invention patent with the publication number CN118565498A discloses a path planning method and device for a mobile robot. This patent performs well in dealing with static obstacles, but has insufficient response ability to sudden obstacles in dynamic environments.

[0005] In view of the above problems, no effective solution has been proposed yet. Summary of the Invention

[0006] In view of this, the present invention provides a method for actively detecting the path of a robot guided by vision to solve the above technical problems.

[0007] To achieve the above object, a method for actively detecting the path of a robot guided by vision provided by the present invention includes the following steps: S101: Real-time obtain visual modality data, lidar modality data, and depth modality data of the detection area of the robot through a multi-modal sensor; S102: Use a convolutional neural network to preprocess the visual modality data, the lidar modality data, and the depth modality data; wherein, the preprocessing includes image denoising processing and edge detection processing; S103: Apply a color space conversion strategy to separate multiple terrain type regions from the preprocessed visual modality data, preprocessed lidar modality data, and preprocessed depth modality data; S104: Perform object detection on the multiple terrain type regions through a deep learning model to extract candidate regions of terrain obstacles; wherein, the terrain obstacles include static obstacles and dynamic obstacles; S105: Perform pixel segmentation on the candidate regions of the terrain obstacles to determine the positions and regional boundaries of the terrain obstacles; optimize the regional boundaries of the terrain obstacles through morphological processing; wherein, the morphological processing includes dilation processing and erosion processing; S106: Construct a cost map based on the regional boundaries of the terrain obstacles; and adjust the real-time travel path of the robot according to the constructed cost map.

[0008] As a further improvement method of the present invention: Further, in the cost map, the cost weight of the regional boundary of the terrain obstacle is higher than the weight threshold; the adjusting the real-time travel path of the robot according to the constructed cost map includes: Generate a preliminary path according to the constructed cost map and the motion ability constraints of the robot; Use a particle swarm optimization algorithm to smooth the preliminary path to obtain a reference path; Adjust the real-time travel path of the robot according to the reference path.

[0009] Further, the static obstacles include one or more of the following obstacles: pits, protrusions, cracks, gullies, steep slopes, cliffs, bogs, and holes; The terrain type regions include one or more of the following regions: water surface region, grassland region, sandy region, muddy region, snowy region, ice surface region, gravel region, forest region, vegetation region, tidal flat region, rocky region, marsh region, and farmland region.

[0010] Further, the detection area includes a shoal environment area; the multiple terrain type regions include a water surface region, a sandy region, a muddy region, a gravel region, a vegetation region, a tidal flat region, and a rocky region; the preprocessed visual modality data includes first visual modality data and first visual modality mask data; The applying a color space conversion strategy to separate multiple terrain type regions from the preprocessed visual modality data, preprocessed lidar modality data, and preprocessed depth modality data includes: Apply the described color space conversion strategy to process the first visual modality data and the first visual modality mask data, obtaining the second visual modality data and the second visual modality mask data; Input the second visual modality data and the second visual modality mask data into the DNSNet model to obtain a reference visual image with shadows removed; wherein, the DNSNet model includes a shadow segmentation module, a dual-channel Transformer module, a dual-channel attention block, a local window attention mechanism, and a noise suppression attention aggregation module; Optimize the reference visual image with shadows removed according to the preprocessed lidar modality data and the preprocessed depth modality data; Separate the water area, sand area, mud area, gravel area, vegetation area, tidal flat area, and rock area from the optimized reference visual image.

[0011] Furthermore, the step of inputting the second visual modality data and the second visual modality mask data into the DNSNet model to obtain a reference visual image with shadows removed includes: Input the second reference visual modality data and the second visual modality mask data into the shadow segmentation module to obtain sub-modal data and shadow mask data ; wherein, the shadow segmentation module extracts underlying features through linear projection; sequentially obtains embedded features through depthwise separable convolution and residual convolution; performs downsampling through multiple first convolution operations, and then performs upsampling through multiple first transposed convolution operations; finally obtains sub-modal data and shadow mask data ; Input the sub-modal data and the shadow mask data into the dual-channel Transformer module to obtain an intermediate feature map; wherein, the dual-channel Transformer module adopts an encoder-decoder architecture, and the first encoding layer corresponds to the first decoding layer; the dual-channel Transformer module embeds the local window attention mechanism; Input the intermediate feature map into the dual-channel attention block, and obtain an enhanced feature map through tensor addition; wherein, the dual-channel attention block processes the input intermediate feature map through two first convolution layers and a ReLU activation function; Input the enhanced feature map into the noise suppression attention aggregation module to obtain the denoised feature map; wherein, the noise suppression attention aggregation module generates an encoded feature representation through a second encoder, preserves key features while gradually reducing the spatial size and number of channels of the image; performs reconstruction through a second decoder to restore the encoded feature map to an image of the same size as the enhanced feature map; and performs noise reduction by combining the denoising autoencoder mechanism. Combine the sub-modal data output by the shadow segmentation module and the shadow mask data 、the intermediate feature map output by the dual-channel Transformer module, the enhanced feature map output by the dual-channel attention block, and the denoised feature map output by the noise suppression attention aggregation module to determine the reference visual image after removing the shadow.

[0012] Further, the detection area includes a shoal environment area; the terrain obstacles include pits, protrusions, and dynamic obstacles in multiple terrain type areas; wherein, the protrusions in the multiple terrain type areas include reefs, shipwrecks, fishing reefs, and sand dunes; the dynamic obstacles in the multiple terrain type areas include fishing fences, boats, floating objects, and animals. Construct the cost map based on the regional boundaries of the terrain obstacles; generate a preliminary path according to the constructed cost map and the motion ability constraints of the robot, including: According to the motion ability constraints of the robot, assign respective cost weights to each terrain type area. Increase the cost weights of the regional boundaries of the pits, protrusions, and dynamic obstacles in the multiple terrain type areas to construct the cost map. Generate the preliminary path according to the constructed cost map and the motion ability constraints of the robot by using path planning based on triangulation; wherein, the path planning algorithm based on triangulation is the TA (Triangle A) algorithm based on triangulation or the greedy algorithm based on triangulation.

[0013] Further, the path planning algorithm based on triangulation is the TA algorithm based on triangulation; the step of generating the preliminary path according to the constructed cost map and the motion ability constraints of the robot by using path planning based on triangulation includes: Determine the starting position and the ending position, and add the starting position to the open list. Select the node with the minimum value of the cost function from the open list as the current node for expansion; wherein, the value of the cost function includes the actual cost value from the starting position to the current node and the estimated cost value from the current node to the ending position. During the node expansion process, in combination with the motion ability constraints of the robot, nodes in adjacent areas that meet the motion ability of the robot are used as expandable nodes; For each expandable node, calculate the value of its cost function and update its parent node information; among them, if the expandable node is already in the open list, compare the new value of the cost function with the original value of the cost function; if the new value of the cost function is smaller, update the information of the expandable node; if the expandable node is not in the open list, add it to the open list; Repeat the node expansion process until the end position is found or the open list is empty; If the end position is found, based on the recorded parent node information, backtrack to generate a preliminary path from the start position to the end position.

[0014] Furthermore, based on the interaction interface between the robot and the user, the working state, real-time travel path, and task status information of the robot are displayed in real time; wherein, the user remotely controls the robot by sending instructions to the robot.

[0015] Furthermore, after adjusting the real-time travel path of the robot based on the cost map constructed, the method further includes: Convert the adjusted real-time travel path into real-time motion instructions for the robot; wherein, the real-time motion instructions are used to adjust the speed, steering angle, and acceleration of the robot's travel; Control the travel process of the robot according to the real-time motion instructions.

[0016] Furthermore, the multi-modal sensor includes a camera, a depth camera, and a lidar; through the multi-modal sensor, visual modal data, lidar modal data, and depth modal data are obtained in real time, including: Use the camera to obtain the visual modal data in real time; Use the depth camera to obtain the depth modal data in real time; Use the lidar to obtain the lidar modal data in real time.

[0017] Based on the embodiments provided in the present application, an active path detection method for a robot based on visual guidance is implemented. Through visual guidance analysis, the robot can effectively cope with the changes in the shallow water wading environment, improving its autonomy and efficiency. In other words, in a complex terrain environment, the robot not only needs to avoid static obstacles, but also needs to be able to dynamically respond to ground changes, sudden obstacles, and unpredictable environmental factors. Especially in a shallow water wading environment such as a shoal, the active path detection method based on visual guidance provided in the present application can actively explore the shallow water wading environment and adjust the path in real time. Without a pre-constructed map, it can make autonomous decisions based on visual information, avoid deep water areas, and adjust the traveling path, thus greatly enhancing the adaptability of the robot in the shallow water wading environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 FIG. is a flowchart of an alternative active path detection method for a robot based on visual guidance according to an embodiment of the present application; Figure 2 FIG. is a flowchart of another alternative active path detection method for a robot based on visual guidance according to an embodiment of the present application; The realization, functional features, and advantages of the objectives of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0020] Optionally, as Figure 1 shown, the present application provides an active path detection method for a robot based on visual guidance, including: S101, obtaining visual modality data, lidar modality data, and depth modality data of the detection area of the robot in real time through a multi-modal sensor; Among them, visual modality data mainly obtains image data through cameras, including RGB images, etc. For example, using a camera carried by a drone to take ground images, from which visual features such as terrain texture and color can be extracted for identifying different landform types, such as mountains, plains, rivers, forests, etc. In addition, object clues in the images, such as buildings and roads, can also be used to assist in perceiving the terrain environment. LiDAR modality data is obtained by LiDAR, which can emit lasers and receive reflected signals to obtain three-dimensional point cloud data of the terrain. These data can accurately reflect the geometric shape of the terrain, including information such as height, slope, and unevenness, which is crucial for constructing a detailed terrain model. Depth modality data is obtained by a depth camera, and each pixel contains depth information, that is, the distance of the point from the camera. The depth image can intuitively represent the distance relationship of the scene and help construct the three-dimensional structure of the scene.

[0021] S102, Use a convolutional neural network to preprocess the visual modality data, LiDAR modality data, and depth modality data; among them, the preprocessing includes image denoising processing and edge detection processing; The preprocessing can enhance the recognizability of terrain features.

[0022] S103, Apply a color space conversion strategy to separate multiple terrain type regions from the preprocessed visual modality data, preprocessed LiDAR modality data, and preprocessed depth modality data; Among them, the color space conversion strategy can include the strategy of converting from RGB to HSV or Lab color space. Thus, the shoal area can be recognized more clearly. Use image enhancement techniques, such as histogram equalization, to optimize the image to ensure that the water surface and ground features are clearer.

[0023] S104, Perform object detection on the multiple terrain type regions through a deep learning model to extract candidate regions of terrain obstacles; among them, terrain obstacles include static obstacles and dynamic obstacles; Among them, the deep learning model can be Faster R-CNN; S105, Perform pixel segmentation on the candidate regions of terrain obstacles to determine the location and regional boundaries of terrain obstacles; through morphological processing, optimize the regional boundaries of terrain obstacles; among them, morphological processing includes dilation processing and erosion processing; Among them, U-Net or Mask R-CNN can be used to perform pixel segmentation on the candidate regions of terrain obstacles; S106, Based on the regional boundaries of terrain obstacles, construct a cost map; according to the constructed cost map, adjust the real-time travel path of the robot.

[0024] Based on the embodiments provided in this application, a method for active path detection of a robot based on visual guidance is implemented. Through visual guidance analysis, the robot can effectively respond to changes in the shallow water wading environment, improving its autonomy and efficiency. In other words, in a complex terrain environment, the robot not only needs to avoid static obstacles but also be able to dynamically respond to ground changes, sudden obstacles, and unpredictable environmental factors. Especially in a shallow water wading environment such as a shoal, the active path detection method based on visual guidance provided in this application can actively explore the shallow water wading environment and adjust the path in real time. Without a pre-constructed map, it can make autonomous decisions based on visual information, avoid deep water areas, and adjust the travel path, thus greatly enhancing the adaptability of the robot in the shallow water wading environment.

[0025] Furthermore, in the cost map, the cost weight of the regional boundary of the terrain obstacle is higher than the weight threshold; according to the constructed cost map, the real-time travel path of the robot is adjusted, including: Generate a preliminary path according to the constructed cost map and the motion ability constraints of the robot; Among them, the motion ability constraints of the robot include the maximum climbing angle and the ability to cross water surfaces; Use the particle swarm optimization algorithm to smooth the preliminary path to obtain a reference path, so as to reduce the complexity and redundancy of the path and ensure that the robot can travel efficiently and stably; Adjust the real-time travel path of the robot according to the reference path.

[0026] Furthermore, the static obstacles include one or more of the following obstacles: pits, protrusions, cracks, gullies, steep slopes, cliffs, mud bogs, and holes; the protrusions include but are not limited to sand dunes, stone piles, tree roots, etc.

[0027] The terrain type areas include one or more of the following areas: water surface area, grassland area, sandy area, muddy area, snowy area, ice surface area, gravel area, forest area, vegetation area, tidal flat area, rocky area, marsh area, and farmland area.

[0028] Furthermore, the detection area includes a shallow water environment area; the multiple terrain type areas include a water surface area, a sandy area, a muddy area, a gravel area, a vegetation area, a tidal flat area, and a rocky area; the preprocessed visual modality data includes first visual modality data and first visual modality mask data; Optionally, as Figure 2 shown, applying a color space conversion strategy, multiple terrain type areas are separated from the preprocessed visual modality data, the preprocessed lidar modality data, and the preprocessed depth modality data, including: S201. Apply a color space conversion strategy to process the first visual modality data and the first visual modality mask data to obtain the second visual modality data and the second visual modality mask data; Among them, applying a color space conversion strategy to process data can specifically include: Step 1. Prepare data: an RGB image, that is, the first visual modality data; and its corresponding binary mask, that is, the first visual modality mask data. The RGB image contains color information, and the binary mask is used to specify the regions of interest in the image.

[0029] Step 2. Read the image and the mask: Use an image processing tool or programming library to read the RGB image and the binary mask, ensuring that both have the same size and resolution.

[0030] Step 3. Color space conversion: Convert the RGB image from the RGB color space to the HSV or Lab color space. The HSV color space consists of hue, saturation, and value, while the Lab color space consists of lightness, red-green axis, and yellow-blue axis. The purpose of the conversion is to more easily process color information in the new color space.

[0031] Step 4. Apply the mask: Use the binary mask to select the regions of interest in the HSV or Lab image. This step ensures that only the regions specified by the mask are processed, ignoring other regions.

[0032] Step 5. Process the regions of interest: In the new color space, adjust the color attributes as needed. For example, the hue can be adjusted to change the color, the saturation can be adjusted to enhance or weaken the vividness of the color, and the value can be adjusted to change the brightness of the color. Among them, adjust the hue: If it is necessary to convert the color of a specific region from one hue to another. For example, convert all red regions to blue regions. In the HSV color space, the hue range of red is usually 0 - 10 degrees, and the hue range of blue is 110 - 130 degrees. You can achieve this conversion by finding the pixels in the red region and changing their hue values to the hue values of blue. Adjust the saturation and value: Similarly, the saturation and value can be adjusted to enhance or weaken the vividness and brightness of the color. For example, increasing the saturation can make the color more vivid, and increasing the value can make the color brighter.

[0033] Step 6. Save and use the processed image.

[0034] S202. Input the second visual modality data and the second visual modality mask data into the DNSNet model to obtain a reference visual image after removing shadows; among them, the DNSNet model includes a shadow segmentation module, a dual-channel Transformer module, a dual-channel attention block, a local window attention mechanism, and a noise suppression attention aggregation module; S203. Optimize the reference visual image after shadow removal according to the preprocessed lidar modality data and the preprocessed depth modality data; Among them, the data provided by the lidar and the depth sensor can be used to improve the quality of the visual image that has already had shadows removed. The lidar modality data provides the three-dimensional structure information of the environment, and the depth modality data provides the depth information of each pixel point. The combination of these two can more accurately understand the geometric structure of the scene. Through this information, the visual image can be further optimized, such as improving the details of the image, enhancing the contrast, correcting the color, etc., making the image more realistic and useful. Specifically, it can include the following steps: Data fusion: Integrate the lidar data and the depth data with the visual image data. This can be achieved in various ways, such as feature-level fusion, decision-level fusion, etc. For example, a multi-modal VoxelNet like MVX-Net can be used, which enhances the lidar point cloud with semantic image features and achieves the fusion of image and lidar features in early learning to realize accurate 3D object detection. Shadow removal: Use deep learning methods to remove the shadows in the visual image again. For example, a Transformer model like SpA-Former can be used, which can adaptively process the shadows projected on different semantic regions and has strong generalization ability. Image optimization: Optimize the image after shadow removal using the fused data. This may include using specific algorithms to enhance specific aspects of the image, such as guiding the object query to regress to a position closer to the real 3D bounding box through the triangular geometric loss while geometrically retaining the properties of the 2D bounding box.

[0035] S204. Separate the water area, sand area, mud area, gravel area, vegetation area, tidal flat area, and rock area from the optimized reference visual image.

[0036] Further, input the second visual modality data and the second visual modality mask data into the DNSNet model to obtain the reference visual image after shadow removal, including: Input the second reference visual modality data and the second visual modality mask data into the shadow segmentation module to obtain the sub-modal data and the shadow mask data ; Among them, the shadow segmentation module extracts the underlying features through linear projection; ; is the underlying feature; represents the real number field, indicating that each element in the feature map is a real number; is the number of channels; is the height; is the width; Obtain the embedded features through depthwise separable convolution and residual convolution in sequence; Perform downsampling through multiple first convolution operations, and then perform upsampling through multiple first transposed convolution operations; wherein, the first convolution operation can be a 4×4 convolution operation; the first transposed convolution operation can be a 2×2 transposed convolution operation; Finally, obtain the sub-modal data and the shadow mask data through a single second convolution operation; the second convolution operation can be a 3×3 convolution operation; Input the sub-modal data and the shadow mask data into a dual-channel Transformer module to obtain an intermediate feature map; wherein, the dual-channel Transformer module adopts an encoder-decoder architecture, and the first encoding layer corresponds to the first decoding layer; the dual-channel Transformer module embeds a local window attention mechanism; Input the intermediate feature map into a dual-channel attention block, and obtain an enhanced feature map through tensor addition; wherein, the dual-channel attention block processes the input intermediate feature map through two first convolutional layers and a ReLU activation function; For example, the first convolutional layer can be a 3×3 convolutional layer; is the feature map after two 3×3 convolutions and ReLU activation; is the intermediate feature map; wherein, a channel attention mechanism is applied to highlight the importance of key channels; is the feature map after the channel attention mechanism; is the global average pooling operation, which is used to reduce the spatial dimension of the feature map; Apply a spatial attention mechanism to focus on the shadow-affected area and extract finer details; is the feature map after the spatial attention mechanism; Obtain an enhanced feature map through tensor addition .

[0037] The enhanced feature map is input into the noise suppression attention aggregation module to obtain the denoised feature map. Among them, the noise suppression attention aggregation module generates an encoded feature representation through a second encoder, while gradually reducing the spatial size and number of channels of the image and retaining key features; performs reconstruction through a second decoder to restore the encoded feature map to an image of the same size as the enhanced feature map; and combines the denoising autoencoder mechanism to perform noise reduction. Combined with the sub-modal data output by the shadow segmentation module and the shadow mask data 、the intermediate feature map output by the dual-channel Transformer module, the enhanced feature map output by the dual-channel attention block, and the denoised feature map output by the noise suppression attention aggregation module to determine the reference visual image after removing the shadow.

[0038] Furthermore, the detection area includes the shoal environment area; the terrain obstacles include pits, protrusions, and dynamic obstacles in multiple terrain type areas. Among them, the protrusions in multiple terrain type areas include reefs, shipwrecks, fishing reefs, and sand dunes; the dynamic obstacles in multiple terrain type areas include fishing fences, boats, floating objects, and animals. Among them, boats include fishing boats, leisure boats, commercial ships, and military ships; floating objects include branches, plastic garbage, oil drums, and other sundries; animals include fish, birds, mammals, and amphibians. The dynamic obstacles can also include but are not limited to natural factors such as water currents, winds, swells, and undercurrents; human activities such as swimmers, rowers, fishermen, shore pedestrians, and other water sports participants; dynamic obstacle groups such as schools of fish, flocks of ducks, crowds, and fleets of ships.

[0039] Based on the regional boundaries of the terrain obstacles, a cost map is constructed; according to the constructed cost map and the motion ability constraints of the robot, a preliminary path is generated, including: According to the motion ability constraints of the robot, respective cost weights are assigned to each terrain type area. The cost weights of the regional boundaries of the pits, the cost weights of the regional boundaries of the protrusions, and the cost weights of the regional boundaries of the dynamic obstacles in multiple terrain type areas are increased to construct a cost map. Specifically, in combination with the motion ability constraints of the robot, different cost weights are assigned to different terrain types. For example, for muddy areas and steep slopes that are difficult for the robot to pass through, higher costs are assigned, while for flat and easily passable areas, lower costs are assigned. The cost map constructed in this way can comprehensively reflect the complexity of the environment and the motion constraints of the robot.

[0040] According to the constructed cost map and the motion ability constraints of the robot, a preliminary path is generated using triangulation-based path planning.

[0041] In the embodiments of the present application, the path planning algorithm based on triangulation may include, but is not limited to, the TA algorithm based on triangulation and the greedy algorithm based on triangulation.

[0042] Further, taking the path planning algorithm based on triangulation as the TA algorithm based on triangulation as an example, according to the constructed cost map and the motion ability constraints of the robot, a preliminary path is generated by using the path planning based on triangulation, including: Determine the starting position and the ending position, and add the starting position to the open list; Select the node with the smallest value of the cost function from the open list as the current node for expansion; wherein, the value of the cost function includes the actual cost value from the starting position to the current node and the estimated cost value from the current node to the ending position; Among them, the actual cost value includes factors such as path length, soil water content, and height change of the path. The weights of these factors can be adjusted according to actual needs and the importance of the cost function to balance the relationship between different constraints. The estimated cost value also considers factors such as path length, soil water content, and height change. By expanding the node search area, the algorithm can select paths with relatively redundant lengths to achieve relative optimality of other constraints. During the node expansion process, in combination with the motion ability constraints of the robot, the nodes in the adjacent areas that meet the motion ability of the robot are used as expandable nodes to avoid generating paths that the robot cannot actually pass; For each expandable node, calculate the value of its cost function and update its parent node information; Specifically, the actual cost value from the starting point to the current node can be expressed as: Among them, is the path length; is the soil water content; is the path height change; 、 、 are the weight coefficients; The estimated cost value from the current node to the ending position can be expressed as: Among them, is the estimated path length; is the estimated soil water content; is the estimated height change; Among them, if the expandable node is already in the open list, compare the value of the new cost function with the value of the original cost function; if the value of the new cost function is smaller, update the information of the expandable node; if the expandable node is not in the open list, add it to the open list; Repeat the node expansion process until the end position is found or the open list is empty; If the end position is found, generate a preliminary path from the start position to the end position by backtracking according to the recorded parent node information.

[0043] If the open list is empty but the end position is not found, it means that a feasible path cannot be found under the current cost map and motion ability constraints, and the environmental information needs to be re-evaluated or the motion ability constraints of the robot need to be adjusted.

[0044] Optionally, the path can be further adjusted according to the motion characteristics of the robot. For example, adjust the curvature of the path to ensure that it is within the maximum turning radius of the robot; optimize the height change of the path to make it meet the climbing ability of the robot, etc. The path can be continuously adjusted and optimized iteratively until a certain optimization goal is met or a predetermined number of iterations is reached. During the optimization process, it is still necessary to ensure that the path conforms to the constraints of the cost map and the motion ability constraints of the robot.

[0045] Furthermore, based on the interaction interface between the robot and the user, the working state, real-time travel path, and task status information of the robot are displayed in real time; among them, the user sends instructions to the robot to remotely control the robot.

[0046] Furthermore, after adjusting the real-time travel path of the robot according to the constructed cost map, the method further includes: Convert the adjusted real-time travel path into the real-time motion instructions of the robot; among them, the real-time motion instructions are used to adjust the speed, steering angle, and acceleration of the robot's travel; Control the travel process of the robot according to the real-time motion instructions.

[0047] Furthermore, the multi-modal sensor includes a camera, a depth camera, and a lidar; through the multi-modal sensor, visual modal data, lidar modal data, and depth modal data are obtained in real time, including: Use the camera to obtain visual modal data in real time; Use the depth camera to obtain depth modal data in real time; Use the lidar to obtain lidar modal data in real time.

[0048] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.

Claims

1. A vision-guided robot active path detection method, characterized in that: include: S101: Obtain visual modal data, laser radar modal data, and depth modal data of the detection area of ​​the robot in real time through a multimodal sensor; S102: using a convolutional neural network, preprocessing the visual modality data, the lidar modality data, and the depth modality data; wherein the preprocessing includes image denoising processing and edge detection processing; S103: applying a color space conversion strategy to separate a plurality of terrain type regions from the preprocessed visual modality data, the preprocessed lidar modality data, and the preprocessed depth modality data; S104: performing target detection on the multiple terrain type areas through a deep learning model to extract candidate areas of terrain obstacles; wherein the terrain obstacles include static obstacles and dynamic obstacles; S105: performing pixel segmentation on the candidate area of ​​the terrain obstacle to determine the position and area boundary of the terrain obstacle; optimizing the area boundary of the terrain obstacle through morphological processing; wherein the morphological processing includes dilation processing and erosion processing; S106: constructing a cost map based on the regional boundaries of the terrain obstacles; and adjusting the real-time travel path of the robot according to the constructed cost map.

2. The vision-guided robot active path detection method according to claim 1, characterized in that: In the cost map, the cost weight of the regional boundary of the terrain obstacle is higher than the weight threshold; and adjusting the real-time travel path of the robot according to the constructed cost map includes: Generate a preliminary path based on the constructed cost map and the motion capability constraints of the robot; Using a particle swarm optimization algorithm to smooth the preliminary path to obtain a reference path; The real-time travel path of the robot is adjusted according to the reference path.

3. The vision-guided robot active path detection method according to claim 1, characterized in that: The static obstacles include one or more of the following obstacles: pits, protrusions, cracks, gullies, steep slopes, cliffs, swamps and potholes; The terrain type area includes one or more of the following areas: water surface area, grassland area, sand area, mud area, snow area, ice area, gravel area, woodland area, vegetation area, tidal flat area, rocky area, swamp area and farmland area.

4. The vision-guided robot active path detection method according to claim 2, characterized in that: The detection area includes a shoal environment area; multiple terrain type areas include a water surface area, a sand area, a mud area, a gravel area, a vegetation area, a tidal flat area and a rock area; The preprocessed visual modality data includes first visual modality data and first visual modality mask data; The color space conversion strategy is applied to separate multiple terrain type areas from the preprocessed visual modality data, the preprocessed lidar modality data, and the preprocessed depth modality data, including: Applying the color space conversion strategy to process the first visual modality data and the first visual modality mask data to obtain second visual modality data and second visual modality mask data; Inputting the second visual modality data and the second visual modality mask data into a DNSNet model to obtain a reference visual image after removing the shadow; wherein the DNSNet model includes a shadow segmentation module, a dual-channel Transformer module, a dual-channel attention block, a local window attention mechanism, and a noise suppression attention aggregation module; Optimizing the reference visual image after removing shadows according to the preprocessed LiDAR modal data and the preprocessed depth modal data; The water surface area, sand area, mud area, gravel area, vegetation area, tidal flat area and rock area are separated from the optimized reference visual image.

5. The vision-guided robot active path detection method according to claim 4, characterized in that: The step of inputting the second visual modality data and the second visual modality mask data into the DNSNet model to obtain a reference visual image after removing the shadow comprises: The second reference visual modality data With the second visual modality mask data Input to the shadow segmentation module to obtain sub-modal data and shadow mask data ; The shadow segmentation module extracts the underlying features through linear projection; obtains embedded features through depthwise separable convolution and residual convolution in turn; performs downsampling through multiple first convolution operations, and then performs upsampling through multiple first transposed convolution operations; finally obtains submodal data through a second convolution operation and shadow mask data ; The submodal data and the shadow mask data Input the dual-channel Transformer module to obtain an intermediate feature map; wherein the dual-channel Transformer module adopts an encoder-decoder architecture, and the first encoding layer corresponds to the first decoding layer; the dual-channel Transformer module is embedded with the local window attention mechanism; Input the intermediate feature map into the dual-channel attention block, and obtain the enhanced feature map by tensor addition; wherein the dual-channel attention block processes the input intermediate feature map through two first convolutional layers and a ReLU activation function; The enhanced feature map is input into the noise suppression attention aggregation module to obtain a denoised feature map; wherein the noise suppression attention aggregation module generates an encoded feature representation through a second encoder, gradually reducing the spatial size and the number of channels of the image while retaining key features; reconstructs through a second decoder to restore the encoded feature map to an image of the same size as the enhanced feature map; and performs noise reduction in combination with a denoising autoencoder mechanism; Combined with the sub-modal data output by the shadow segmentation module and shadow mask data , the intermediate feature map output by the dual-channel Transformer module, the enhanced feature map output by the dual-channel attention block, and the denoised feature map output by the noise suppression attention aggregation module to determine a reference visual image after removing shadows.

6. The vision-guided robot active path detection method according to claim 4, characterized in that: The detection area includes a shoal environment area; the terrain obstacles include pits, protrusions and dynamic obstacles in multiple terrain type areas; wherein the protrusions in multiple terrain type areas include reefs, shipwrecks, fishing reefs and sand dunes; the dynamic obstacles in multiple terrain type areas include fishing traps, ships, floating objects and animals; Based on the regional boundary of the terrain obstacle, the cost map is constructed; and according to the constructed cost map and the motion capability constraint of the robot, a preliminary path is generated, including: According to the motion capability constraints of the robot, each terrain type area is assigned a respective cost weight; Adding cost weights of area boundaries of pits, cost weights of area boundaries of convexities, and cost weights of area boundaries of dynamic obstacles in multiple terrain type areas to construct the cost map; According to the constructed cost map and the motion capability constraints of the robot, the preliminary path is generated using triangulation-based path planning; wherein the triangulation-based path planning algorithm is a triangulation-based TA algorithm or a triangulation-based greedy algorithm.

7. The vision-guided robot active path detection method according to claim 6, characterized in that: The triangulation-based path planning algorithm is the triangulation-based TA algorithm; The method of generating the preliminary path by using triangulation-based path planning based on the constructed cost map and the robot's motion capability constraints includes: Determine a starting position and an end position, and add the starting position to an open list; Selecting a node with the smallest cost function value from the open list as the current node for expansion; wherein the value of the cost function includes an actual cost value from the starting position to the current node and an estimated cost value from the current node to the end position; In the process of node expansion, in combination with the motion capability constraint of the robot, nodes in adjacent areas that meet the motion capability of the robot are used as expandable nodes; For each scalable node, calculate the value of its cost function and update its parent node information; if the scalable node is already in the open list, compare the new cost function value with the original cost function value; if the new cost function value is smaller, update the information of the scalable node; if the scalable node is not in the open list, add it to the open list; Repeat the node expansion process until the end position is found or the open list is empty; If the end point is found, a preliminary path from the start point to the end point is generated backtrackingly according to the recorded parent node information.

8. The vision-guided robot active path detection method according to claim 1, characterized in that: Based on the interactive interface between the robot and the user, the working status, real-time travel path and task status information of the robot are displayed in real time; wherein, the user remotely controls the robot by sending instructions to the robot.

9. The vision-guided robot active path detection method according to claim 1, characterized in that: After adjusting the real-time travel path of the robot according to the constructed cost map, the method further includes: Converting the adjusted real-time travel path into real-time motion instructions of the robot; wherein the real-time motion instructions are used to adjust the speed, steering angle and acceleration of the robot; The moving process of the robot is controlled according to the real-time motion instruction.

10. The vision-guided robot active path detection method according to claim 1, characterized in that: The multimodal sensor includes a camera, a depth camera and a laser radar; The real-time acquisition of visual modality data, lidar modality data, and depth modality data by multimodal sensors includes: Acquiring the visual modality data in real time using the camera; Acquiring the depth modality data in real time using the depth camera; The laser radar modal data is acquired in real time using the laser radar.

Citation Information

Patent Citations

  • Autonomous path planning and obstacle avoidance system and method

    CN116339342A

  • Path planning method and device for mobile robot

    CN118565498A

Cited By

  • Unmanned aerial vehicle inspection path autonomous planning method based on machine vision

    CN120800389A

  • Mobile robot control device giving consideration to field navigation patrol and transportation load

    CN120821235A

  • A mobile robot control device that takes into account field navigation patrol and transportation of a load

    CN120821235B

  • AGV environment sensing method and system based on laser beams

    CN120927009A

  • Intelligent logistics warehouse guide line visual detection method based on deep learning

    CN120953620A