A robot three-dimensional environment perception method and device based on deep visual learning

Through deep visual learning methods, synchronized visual data streams are generated and multi-scale feature fusion is performed, which solves the problem of insufficient deep semantics and spatial structure understanding of three-dimensional environmental perception in existing technologies, and achieves high-precision and robust environmental perception suitable for autonomous navigation and object recognition.

CN120564156BActive Publication Date: 2025-09-26DONGGUAN XINBAIREN ROBOT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511057338.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-09-26
Estimated Expiration
2045-07-30

AI Technical Summary

Technical Problem

Existing technologies lack the ability to comprehensively understand the deep semantics and spatial structure of the environment in complex three-dimensional environment perception, resulting in insufficient accuracy and robustness of perception results in dynamic environments or occluded scenes.

Method used

Through a method based on deep visual learning, visual data is collected to generate a synchronized visual data stream, spatial structural features and preliminary semantic segmentation maps are extracted, geometric semantic feature maps are generated, and multi-scale feature tensors are optimized through loss functions. A three-dimensional geometric model is constructed, and fusion optimization is performed with the geometric semantic feature maps. Finally, structured perception results are output through a verification framework.

Benefits of technology

It achieves high-precision perception of complex three-dimensional environments, improves perception accuracy and robustness, adapts to dynamic environments, and improves the accuracy and practicality of autonomous navigation and object recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120564156B_ABST
    Figure CN120564156B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of target detection technology, specifically a method and device for robot three-dimensional environment perception based on deep visual learning, the method comprising: collecting visual data based on a perception system carried by the robot and synchronously processing it, extracting spatial structural features and a preliminary semantic segmentation map of the synchronized visual data stream, combining the spatial structural features and the preliminary semantic segmentation map to generate a geometric semantic feature map; extracting local, regional, and global features of the geometric semantic feature map and optimizing them, performing three-dimensional modeling processing based on a multi-scale feature tensor to generate a three-dimensional geometric model; fusing and optimizing the structural information of the three-dimensional geometric model with the geometric semantic feature map, and performing verification processing based on a verification framework to obtain three-dimensional environment perception data. This application solves the defects of the existing technology in complex three-dimensional environment perception, such as insufficient accuracy and limited deep semantic understanding capabilities, by fusing deep visual learning with visual data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of target detection technology, and in particular to a robot three-dimensional environment perception method and device based on deep vision learning. Background Art

[0002] Robotic environmental perception is a core technology in the field of intelligent robotics, with applications in multiple scenarios, including autonomous navigation, object recognition, and scene understanding. Existing technologies typically use a variety of sensors, such as lidar, cameras, and ultrasonic sensors, to collect environmental data and process it using traditional computer vision techniques or simple machine learning algorithms. These methods have achieved some success in two-dimensional environmental perception. For example, lidar-based simultaneous localization and mapping (SLAM) technology can generate environmental maps from point cloud data and, combined with basic image processing algorithms, implement simple obstacle detection and path planning.

[0003] However, existing technologies have significant limitations when dealing with complex three-dimensional environments, namely a lack of comprehensive understanding of the environment's deep semantics and spatial structure. This shortcoming results in insufficient accuracy and robustness in complex scenarios (such as dynamic environments or those with occlusion), limiting the performance of robots in practical applications.

[0004] Existing research in 3D environmental perception focuses primarily on processing or simply fusing single-modal data. For example, RGB-D data processing techniques based on depth cameras can extract depth information from the environment, but they cannot effectively integrate the spatiotemporal correlations between multi-source data. This makes it difficult to address the real-time perception needs of dynamic environments and capture the deep semantics of the environment. Consequently, existing technologies struggle to achieve high-precision environmental modeling and scene analysis in practical applications. Summary of the Invention

[0005] The main purpose of the present invention is to provide a robot three-dimensional environment perception method and device based on deep visual learning, aiming to overcome the limitations of the existing technology in the perception of complex three-dimensional environments, especially the technical problem of insufficient understanding of the deep semantics and spatial structure of the environment.

[0006] In order to solve the above-mentioned problems, the present invention proposes a robot three-dimensional environment perception method based on deep vision learning, which includes:

[0007] Collecting visual data based on a perception system carried by the robot, and synchronously processing the visual data to obtain a synchronized visual data stream;

[0008] Extracting spatial structural features and a preliminary semantic segmentation map of the synchronized visual data stream, and combining the spatial structural features and the preliminary semantic segmentation map to generate a geometric semantic feature map;

[0009] Extracting local, regional, and global features of the geometric semantic feature map, and optimizing them based on a loss function to generate a multi-scale feature tensor;

[0010] Performing three-dimensional modeling processing according to the multi-scale feature tensor to generate a three-dimensional geometric model;

[0011] fusing and optimizing the structural information of the three-dimensional geometric model with the geometric semantic feature map to output a structured perception result;

[0012] The structured perception results are verified based on the verification framework to obtain three-dimensional environmental perception data.

[0013] Furthermore, the step of collecting visual data based on the perception system carried by the robot and synchronously processing the visual data to obtain a synchronized visual data stream includes:

[0014] Performing data separation processing on the visual data collected by the robot perception system to obtain multiple separated data;

[0015] Mapping the separated data to a three-dimensional spatial reference system based on the coordinate system transformation matrix, and performing time alignment on the separated data to obtain visual data;

[0016] Acquiring current posture information of the robot, and optimizing the visual data according to the current posture information to obtain optimized visual data;

[0017] The data transmission order in the optimized visual data is adjusted based on a preset priority to generate a synchronous visual data stream including three-dimensional point cloud, image sequence, motion trajectory and distance information.

[0018] Furthermore, the step of extracting the spatial structural features and the preliminary semantic segmentation map of the synchronized visual data stream, and combining the spatial structural features and the preliminary semantic segmentation map to generate a geometric semantic feature map includes:

[0019] Performing dual-path decomposition processing on the synchronized visual data stream to obtain geometric branch data and semantic branch data;

[0020] Identifying high-density areas in the geometric branch data, merging similar points in the high-density areas into geometric clusters, and fitting the geometric clusters based on surface fitting technology to obtain spatial structural features;

[0021] Segmenting the semantic branch data based on a pre-trained deep convolutional neural network to obtain a preliminary semantic segmentation map;

[0022] Performing heterogeneous graph fusion processing on the spatial structure features and the preliminary semantic segmentation map to obtain an initial geometric semantic feature map;

[0023] Obtaining projection errors between a three-dimensional point cloud and an image sequence in a synchronized visual data stream, and constructing a bias correction matrix based on the projection errors;

[0024] The initial geometric semantic feature map is subjected to deviation correction processing according to the deviation correction matrix to obtain a geometric semantic feature map.

[0025] Furthermore, the step of extracting local, regional and global features of the geometric semantic feature map and optimizing based on a loss function to generate a multi-scale feature tensor includes:

[0026] Spatially partitioning the geometric semantic feature map to generate an initial feature set containing local, regional and global information;

[0027] Initializing a deep neural network, performing feature extraction processing on the initial feature set according to the deep neural network, and generating a hierarchical feature tensor;

[0028] Constructing a loss function consisting of object classification loss, boundary detection loss, and spatial positioning loss, and optimizing the hierarchical feature tensor according to the loss function to obtain an optimized feature tensor;

[0029] The local, regional and global scale features in the optimized feature tensor are subjected to feature fusion processing to generate a multi-scale feature tensor.

[0030] Furthermore, the step of performing three-dimensional modeling processing according to the multi-scale feature tensor to generate a three-dimensional geometric model includes:

[0031] hierarchically decoding the local, regional, and global features in the multi-scale feature tensor, extracting spatial information corresponding to the local features, regional features, and global features, respectively, to generate a spatial representation;

[0032] Performing registration processing on the point cloud data according to the spatial representation to obtain a preliminary three-dimensional point cloud model;

[0033] Acquire a neighborhood relationship between points in the preliminary three-dimensional point cloud model, construct a topological map of the point cloud, and construct a topological point cloud model based on the topological map;

[0034] Performing boundary prediction processing on the topological point cloud model based on a boundary prediction algorithm to obtain a boundary enhanced point cloud model;

[0035] The boundary-enhanced point cloud model is segmented, high-semantic regions are modeled preferentially, and global modeling is completed through incremental expansion to generate a three-dimensional geometric model.

[0036] Furthermore, the step of fusing and optimizing the structural information of the three-dimensional geometric model with the geometric semantic feature map and outputting a structured perception result includes:

[0037] Extracting point cloud features of the three-dimensional geometric model, and spatially registering the point cloud features with pixel-level semantic labels of the semantic feature map to obtain an aligned feature set;

[0038] Analyzing and processing the point cloud geometric information and semantic labels contained in the alignment feature set based on a graph convolutional network to construct a spatial and semantic association graph between objects;

[0039] Acquire multi-view data of the robot perception system, project node features of the space and semantic association graph into the two-dimensional feature space of each view, and perform weighted fusion on the multi-view features to generate a fused feature tensor;

[0040] Identifying an occluded area in the fused feature tensor and performing depth completion processing on the occluded area to obtain a completed feature field, wherein the occluded area includes an area where point clouds are missing or semantically discontinuous;

[0041] The point cloud in the completed feature field is segmented into multiple object instances through a clustering algorithm, and the local features of each object instance are extracted based on a 3D convolutional network.

[0042] Global feature aggregation is performed on the multiple local features, and the output includes the three-dimensional coordinates, direction vector and point cloud representation of the object to form a structured perception result.

[0043] Furthermore, the step of verifying the structured perception results based on the verification framework to obtain three-dimensional environment perception data includes:

[0044] Mapping multiple data in the synchronized visual data stream to a unified coordinate system, and calculating the data deviation of the visual data stream in the unified coordinate system to generate a verification matrix;

[0045] Performing confidence modeling processing on the structured perception result according to the verification matrix to obtain a confidence distribution map;

[0046] Identifying areas in the confidence distribution map that are below a preset threshold, grouping low-confidence points into continuous deviation areas through cluster analysis, and generating a deviation area annotation set;

[0047] Performing local optimization processing on the structured perception result according to the deviation area annotation set to obtain an optimized perception result;

[0048] The optimized perception results are subjected to structured coding processing to obtain three-dimensional environment perception data.

[0049] The present invention also proposes a robot three-dimensional environment perception device based on deep visual learning, comprising:

[0050] An acquisition module is used to collect visual data based on the perception system carried by the robot, and synchronously process the visual data to obtain a synchronized visual data stream;

[0051] an extraction module, configured to extract spatial structural features and a preliminary semantic segmentation map of the synchronized visual data stream, and combine the spatial structural features and the preliminary semantic segmentation map to generate a geometric semantic feature map;

[0052] An optimization module, configured to extract local, regional, and global features of the geometric semantic feature map and perform optimization based on a loss function to generate a multi-scale feature tensor;

[0053] A modeling module, configured to perform three-dimensional modeling processing based on the multi-scale feature tensor to generate a three-dimensional geometric model;

[0054] A fusion module, configured to fuse and optimize the structural information of the three-dimensional geometric model with the geometric semantic feature map, and output a structured perception result;

[0055] The verification module is used to verify the structured perception results based on the verification framework to obtain three-dimensional environment perception data.

[0056] The present invention further provides a computer device, comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned method when executing the computer program.

[0057] The present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the above method when executed by a processor.

[0058] Compared with the prior art, this application has the following beneficial effects:

[0059] This application proposes a method and device for robotic three-dimensional environment perception based on deep visual learning. By collecting and synchronously processing visual data to generate a synchronized visual data stream, this method effectively integrates data from multiple sensors (such as lidar and cameras), improving the robustness and real-time performance of data processing. By extracting spatial structural features and preliminary semantic segmentation maps from the synchronized visual data streams and combining them to generate a geometric semantic feature map, this method achieves a comprehensive understanding of the deep semantics and spatial structure of the environment. Compared to existing solutions that rely solely on RGB-D data processing or traditional SLAM (Simultaneous Local Area Mapping) technology, this method can more accurately capture the geometric and semantic information of complex three-dimensional environments, improving perception accuracy. Furthermore, by extracting local, regional, and global features from the geometric semantic feature map and generating a multi-scale feature tensor based on loss function optimization, the method enhances the comprehensiveness and multi-scale adaptability of feature representation. A three-dimensional geometric model is constructed and optimized through fusion with the geometric semantic feature map to output structured perception results. Verification of the structured perception results through a verification framework ensures the reliability and consistency of the perception data, making this method more accurate and practical in application scenarios such as autonomous navigation, object recognition, and scene understanding.

[0060] In summary, the method of this application solves the shortcomings of existing technologies in complex three-dimensional environment perception, such as insufficient accuracy, poor robustness, and limited deep semantic understanding capabilities, through the innovative design of deep visual learning and data fusion. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0062] The structures, proportions, sizes, etc. depicted in the drawings of this specification are only used to match the contents disclosed in the specification so as to facilitate understanding and reading by persons familiar with this technology. They are not intended to limit the conditions under which this application can be implemented, and therefore have no substantive technical significance. Any structural modifications, changes in proportional relationships, or adjustments in size, without affecting the efficacy and objectives that can be achieved by this application, should still fall within the scope of the technical contents disclosed in this application.

[0063] Figure 1 This is a schematic diagram of the steps of a robot three-dimensional environment perception method based on deep vision learning in one embodiment of the present invention;

[0064] Figure 2This is a schematic block diagram of the structure of a robot three-dimensional environment perception device based on deep vision learning according to an embodiment of the present invention;

[0065] Figure 3 is a schematic block diagram of the structure of a computer device according to an embodiment of the present invention;

[0066] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION

[0067] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0068] Those skilled in the art will appreciate that, unless expressly stated otherwise, the singular forms "a", "an", "above", and "the" used herein may also include plural forms. It should be further understood that the term "comprising" used in the specification of the present invention refers to the presence of features, integers, steps, operations, elements, modules, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, modules, components, and / or groups thereof. It should be understood that when an element is said to be "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be intermediate elements. In addition, "connected" or "coupled" as used herein may include wireless connections or wireless couplings. The term "and / or" used herein includes all or any module and all combinations of one or more associated listed items.

[0069] Those skilled in the art will understand that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by those skilled in the art in the art to which this invention belongs. It should also be understood that terms such as those defined in common dictionaries should be understood to have meanings consistent with their meanings in the context of the prior art and, unless specifically defined as such, will not be interpreted in an idealized or overly formal sense.

[0070] Reference Figure 1 , an embodiment of the present invention provides a robot three-dimensional environment perception method based on deep visual learning, comprising the following steps:

[0071] S1: collecting visual data based on the perception system carried by the robot, and synchronously processing the visual data to obtain a synchronized visual data stream;

[0072] In step S1, the robot's onboard perception system integrates multiple sensors to collect visual data. These sensors may include high-resolution RGB-D cameras, CMOS cameras, or CCD cameras. For example, an RGB-D camera captures color images and depth information, generating high-resolution RGB-D visual data that captures the texture of the environment and object outlines. During data collection, a multi-threaded parallel acquisition mechanism optimizes data throughput, ensuring efficient and real-time data transmission from each sensor. For example, the RGB-D camera captures images at 30 frames per second, and the LiDAR generates point clouds at 10 Hz. To avoid data accumulation and latency, multi-threading technology is used to parallelize the data streams from different sensors. Each sensor data frame is assigned a timestamp, and Network Time Protocol (NTP) or GPS time synchronization modules are used to maintain timestamp accuracy within milliseconds. A spatial calibration algorithm is also used to align data from different sensors, for example, by projecting the RGB-D camera's point cloud data onto the image plane to achieve spatial consistency. In another embodiment, visual data collection incorporates a wide-angle and zoom lens switching mechanism to dynamically adjust the viewing angle coverage. For example, in open areas, the camera switches to a wide-angle lens to cover a wider range of the environment. When detailed observation of an object is required (such as identifying a small object on a table), the camera switches to a zoom lens to obtain higher-resolution details. This increases the flexibility of data acquisition and adapts to the needs of different environments. The output synchronized visual data stream includes a 3D point cloud, RGB-D image, and motion trajectory, all organized with a unified timestamp and spatial coordinate system.

[0073] S2: extracting spatial structural features and a preliminary semantic segmentation map of the synchronized visual data stream, and combining the spatial structural features and the preliminary semantic segmentation map to generate a geometric semantic feature map;

[0074] In step S2, a dual-path preprocessing framework is employed to extract spatial structural features and a preliminary semantic segmentation map. The geometric path processes 3D point cloud and motion data to capture the spatial geometry of the environment. Specifically, the geometric path groups the point cloud data into voxel grids through voxel clustering, reducing computational complexity while preserving spatial structural information. Surface fitting algorithms (such as RANSAC or least squares) are used to extract smooth geometric features, such as planes, edges, or curved surfaces, from the voxels. These features describe the basic geometry of objects in the environment, such as the edges of a wall, tabletop, or chair. Meanwhile, the semantic path processes RGB-D image data to generate a preliminary semantic segmentation map. RGB-D images contain both color (RGB) and depth (D) information, providing rich visual cues for semantic segmentation. The semantic path uses a pretrained convolutional neural network (CNN), such as ResNet or DeepLabv3+, to perform pixel-level classification on the RGB-D image, generating a preliminary semantic segmentation map. The segmentation map assigns an object class label, such as table, chair, or person, to each pixel in the image. To improve segmentation accuracy, depth information can be used to optimize the RGB image segmentation results, for example, by filtering out background areas using the depth map. After extracting spatial structural features and a preliminary semantic segmentation map, these two pieces of information are integrated into a geometric semantic feature map using a heterogeneous graph fusion algorithm. Geometric and semantic features are represented as a graph structure: nodes represent objects or regions in the environment (e.g., table, chair), and edges represent the spatial relationships between these objects (e.g., table is next to chair). Specifically, features such as planes and edges extracted from the geometric path are associated with the object category labels of the semantic path. For example, a table plane feature might be associated with the table category label, and a chair edge feature with the chair category label. Subsequently, a heterogeneous graph is constructed, in which each node contains geometric information (e.g., plane normal vector) and semantic information (e.g., object category), while edges are defined by spatial distance or relative positional relationships (e.g., proximity or inclusion). In another embodiment, to enhance adaptability to dynamic environments, the fusion process incorporates a spatiotemporal attention mechanism. This mechanism dynamically adjusts the weights of nodes and edges by analyzing temporal and spatial information in a multi-frame data stream, thereby enhancing the representation of dynamic objects. Furthermore, the projective transformation between the point cloud and the RGB-D image can correct for spatial deviations between the two modal data. For example, due to sensor noise, the point cloud data is not fully aligned with the depth information of the RGB-D image. The cross-modal calibration module uses the camera's intrinsic and extrinsic parameters to perform a projective transformation to ensure that the planar features in the point cloud are spatially aligned with the table labels in the image. The final output is a geometric semantic feature map that integrates the spatial structure information (such as planes and edges) and semantic information (such as object categories) of the environment, and represents objects and their spatial relationships through a heterogeneous graph structure.

[0075] S3: extracting local, regional, and global features of the geometric semantic feature map, and optimizing based on a loss function to generate a multi-scale feature tensor;

[0076] In step S3, the geometric semantic feature map is a composite data structure that combines spatial structural features and preliminary semantic segmentation information. To extract multi-scale features from this feature map, namely local features (such as object edge details), regional features (such as surface texture or shape), and global features (such as the spatial layout of different objects in the environment), a specialized neural network model is designed. A large NND (Neural Network for Depth) model is initialized. This model uses a hybrid architecture that combines a 3D convolutional network, a Transformer, and a graph neural network (GNN). In its implementation, the NND model uses a dynamic branch selection mechanism to adaptively allocate computational resources based on the semantic density of the input feature map. For example, in an indoor environment, the feature map contains dense semantic information (such as areas with densely packed objects like furniture and walls) and sparse semantic information (such as open areas on the floor). Based on the distribution of semantic density, the dynamic branch selection mechanism prioritizes more computational resources for dense regions, ensuring feature extraction accuracy in complex areas while reducing computational overhead in sparse regions, thereby optimizing overall efficiency. To implement this mechanism, the model first preprocesses the geometric semantic feature map, calculating the semantic density of each region (for example, by using the distribution density of object categories in the semantic segmentation map). It then dynamically adjusts the depth and width of feature extraction branches based on a density threshold. This adaptive mechanism not only improves computational efficiency but also ensures comprehensive and accurate feature extraction. During the feature extraction process, separate processing modules are designed for local, regional, and global features. The local feature extraction module is primarily based on a 3D convolutional network, using small convolution kernels (e.g., 3×3×3) to capture subtle geometric details such as object edges and corners. The regional feature extraction module uses larger convolution kernels or pooling operations to focus on the overall shape or material characteristics of an object's surface, such as identifying the flat surface of a bookshelf or the texture of wood. The global feature extraction module uses the Transformer's global attention mechanism to analyze the spatial relationships between different objects in the environment. To enhance the contextual relevance of deep features, the model introduces residual connections and cross-layer skip structures. These structures allow shallow features (such as edge details) to be directly transferred to the deep network for fusion with global features, thereby avoiding information loss and improving feature representation. During the optimization phase, the NND model uses a multi-task loss function to simultaneously optimize object classification, boundary detection, and spatial localization. By weightedly combining the losses of multiple subtasks, the model ensures a balance between the different tasks. By jointly optimizing these tasks, the model generates a multi-scale feature tensor that contains rich 3D environmental information. The feature tensor includes local details (such as geometric information about object edges), regional features (such as semantic information about object surfaces), and the global layout (such as the spatial relationships between objects in the environment).

[0077] S4: performing three-dimensional modeling processing according to the multi-scale feature tensor to generate a three-dimensional geometric model;

[0078] In step S4, information is extracted from the multi-scale feature tensor and processed for 3D modeling. Geometric properties in space are defined using neural networks (such as multi-layer perceptrons (MLPs), offering greater flexibility and accuracy. Specifically, Neural Radiance Fields (NeRFs) are combined with point cloud registration. NeRF models a continuous geometric representation of the environment by learning the radiance and volumetric density of spatial points. Point cloud registration aligns discrete point cloud data collected by the perception system to enhance the model's topological consistency. During the modeling process, the multi-scale feature tensor is input into a multi-layer perceptron (MLP) network, which is tasked with predicting spatial occupancy probabilities and object boundaries in the environment. Occupancy probabilities reflect whether a point in space is occupied by an object, while object boundaries accurately demarcate the boundaries between different objects. To ensure the accuracy of the model output, geometric consistency constraints are introduced in the optimization process. By comparing the geometric representation generated by the model with the topological structure of the point cloud data, the model output is forced to conform to the geometric characteristics of the real environment. Furthermore, the model is optimized through a game between a generator and a discriminator. The generator generates a 3D geometric model based on the feature tensor, while the discriminator evaluates the degree of match between the generated model and the real environment. Through this adversarial training, the model can better reconstruct the geometric shape of occluded objects, and the final output 3D geometric model contains the precise position, orientation and shape information of objects in the environment.

[0079] S5: fusing and optimizing the structural information of the 3D geometric model with the geometric semantic feature map, and outputting a structured perception result;

[0080] In step S5, the goal of fusion optimization is to integrate these two pieces of information and generate more consistent perception results through a semantically driven reasoning process. Specifically, a graph convolutional network (GCN) is used to model the spatial and semantic relationships between objects. A GCN constructs a graph structure, representing objects or regions in the environment as nodes, with edges between nodes representing spatial adjacency or semantic correlation. For example, in an indoor scene, a table and a chair might be modeled as nodes with strong semantic associations because they often appear in the same functional area (such as a restaurant). A GCN propagates and aggregates node features through multiple layers of convolution operations, ensuring that each object's features rely not only on its own geometric and semantic information but also incorporate contextual information from surrounding objects, thereby generating globally consistent perception results. This global consistency is particularly important in complex scenes, as the recognition and localization of a single object may be affected by local occlusion or noise. Global reasoning can correct these errors by leveraging contextual relationships. Furthermore, a multi-view fusion strategy is employed to correct positioning errors in the 3D geometric model by integrating visual data (such as RGB-D images) collected from different viewpoints by the robot perception system. For example, imagine a robot perceiving a table in an office setting. From one viewpoint, it may only capture a portion of the table's point cloud data due to occlusion, resulting in an incomplete 3D geometric model. Multi-view fusion can supplement the missing geometric information with RGB-D images from other viewpoints, while also incorporating semantic information (such as the texture or color of the table's surface) to further confirm the object's category and boundaries. The fusion process typically involves coordinate transformation and data alignment, ensuring that data from different viewpoints are integrated in a unified coordinate system, thereby generating a more accurate 3D geometric model and semantic annotations. In another embodiment, depth completion technology is used to predict the position and shape of occluded objects. For example, in a warehouse filled with clutter, a robot may only see a portion of a box's surface, while the rest of the box is obscured by other objects. By analyzing the geometric features and semantic information of the visible portion and combining it with a depth completion algorithm, the complete shape and position of the box can be inferred. Depth completion technology is based on deep learning models, such as generative adversarial networks (GANs) or variational autoencoders (VAEs), which are trained to learn prior knowledge of object shape to generate reasonable completion results. The introduction of the occlusion perception module significantly improves the robustness of the system, enabling it to generate reliable perception results even in complex and crowded scenes. The output structured perception results include the object category, three-dimensional coordinates (x, y, z), direction (expressed in Euler angles), and point cloud representation.

[0081] S6: Verify the structured perception results based on the verification framework to obtain three-dimensional environmental perception data.

[0082] In step S6, a cross-validation mechanism is designed to enhance the robustness and accuracy of perception results. Specifically, the validation framework utilizes multiple data types collected by the robot's perception system, including visual data (such as RGB images or depth images), point cloud data (from LiDAR or depth cameras), and motion data (from the robot's own IMU or encoders). Through collaborative analysis of these data, the reliability of perception results across different dimensions is ensured. During implementation, the structured perception results from step S5 are matched and compared with visual data. For example, point cloud data can provide high-precision 3D spatial position information, useful for verifying the accuracy of object positions in 3D geometric models. Visual data, on the other hand, excels at identifying object categories and can be used to calibrate the category information in semantic segmentation results. In one exemplary scenario, assume that the robot perceives a table in an indoor environment. Step S5 generates a 3D geometric model of the table and labels it as a table. In step S6, the validation framework extracts the point cloud data and verifies whether the 3D coordinates of the table surface are consistent with the geometric model. Furthermore, the visual data analyzes the texture and shape characteristics of the table to confirm the correct category labeling. If the point cloud data indicates positional drift on the table surface, or if the visual data suggests the object is more likely a chair, the verification framework flags this result as low confidence, triggering further processing. Furthermore, a deviation detection mechanism is designed, specifically using Bayesian filtering to analyze the confidence of the perception results. Bayesian filtering is a probabilistic inference method that integrates visual observations and dynamically updates the confidence of perception results. For example, in the aforementioned table scene, Bayesian filtering calculates the joint confidence of the table's position and category based on the observation probabilities of the point cloud and visual data. If the confidence falls below a certain threshold (for example, due to unreliable visual data due to insufficient lighting), a potential anomaly is identified, such as positional drift or misclassification. At this point, the verification framework triggers a local data recollection mechanism, returning to step S1 to recollect visual data to correct the error. This feedback mechanism ensures that the system can adaptively optimize perception results in complex and dynamic environments. Verified structured perception results are output as 3D visualizations, containing rich structured data such as the object's category, position, orientation, and morphology. This data not only provides accurate environmental information for robot navigation, but also supports more advanced tasks such as object grasping or path planning. This real-time performance is particularly important for real-time robot navigation and obstacle avoidance.

[0083] In one embodiment, the steps of collecting visual data based on the perception system carried by the robot and synchronously processing the visual data to obtain a synchronized visual data stream include:

[0084] Performing data separation processing on the visual data collected by the robot perception system to obtain multiple separated data;

[0085] Mapping the separated data to a three-dimensional spatial reference system based on the coordinate system transformation matrix, and performing time alignment on the separated data to obtain visual data;

[0086] Acquiring current posture information of the robot, and optimizing the visual data according to the current posture information to obtain optimized visual data;

[0087] The data transmission order in the optimized visual data is adjusted based on a preset priority to generate a synchronous visual data stream including three-dimensional point cloud, image sequence, motion trajectory and distance information.

[0088] In the above-described embodiment, during the visual data separation and processing phase, signal demultiplexing and channel isolation algorithms are employed to effectively process heterogeneous data, decomposing the raw signals into independent data streams. Through an embedded signal preprocessing module, these data are formatted into standardized data streams, retaining the timestamp and sensor identifier for each modality. The separated visual data undergoes temporal-spatial synchronization and calibration to map them to a unified 3D spatial reference system and achieve time alignment. For temporal synchronization, a high-precision timestamp alignment algorithm is employed to eliminate temporal deviations caused by sampling frequency or clock drift between sensors through timestamp interpolation and clock offset correction. For example, the sampling time of an RGB-D camera may have microsecond offsets. An interpolation algorithm is used to align the timestamps of the point cloud data to the time axis of the image frame to ensure temporal consistency. For spatial synchronization, a coordinate transformation matrix is ​​constructed based on sensor extrinsic calibration to map the data from each modality into a unified 3D spatial reference system. For example, in a warehouse environment, shelf images perceived by a robot (from an RGB-D camera) are mapped to a robot-centered 3D reference system, where each point on the shelf has a consistent coordinate representation in space. In addition, the system adopts a multi-threaded parallel computing framework to accelerate the data alignment process and avoids data loss through a dynamic buffer management mechanism, thereby generating time-space aligned visual data.

[0089] In one embodiment, the step of extracting the spatial structural features and the preliminary semantic segmentation map of the synchronized visual data stream, and combining the spatial structural features and the preliminary semantic segmentation map to generate a geometric semantic feature map includes:

[0090] Performing dual-path decomposition processing on the synchronized visual data stream to obtain geometric branch data and semantic branch data;

[0091] Identifying high-density areas in the geometric branch data, merging similar points in the high-density areas into geometric clusters, and fitting the geometric clusters based on surface fitting technology to obtain spatial structural features;

[0092] Segmenting the semantic branch data based on a pre-trained deep convolutional neural network to obtain a preliminary semantic segmentation map;

[0093] Performing heterogeneous graph fusion processing on the spatial structure features and the preliminary semantic segmentation map to obtain an initial geometric semantic feature map;

[0094] Obtaining projection errors between a three-dimensional point cloud and an image sequence in a synchronized visual data stream, and constructing a bias correction matrix based on the projection errors;

[0095] The initial geometric semantic feature map is subjected to deviation correction processing according to the deviation correction matrix to obtain a geometric semantic feature map.

[0096] In the above-described embodiment, the input synchronous visual data stream is processed using a dual-path decomposition framework, decomposing the data into geometric and semantic branch data. The decomposition process utilizes a modality separation algorithm to assign point cloud data to the geometric branch. Voxel gridding is then used to organize the data into a 3D voxel representation, preserving spatial coordinates and density information. Simultaneously, the RGB image and depth image are assigned to the semantic branch and normalized and channel-aligned to ensure compatibility for subsequent feature extraction. A timestamp alignment mechanism ensures temporal consistency across modal data, and a modality weight assigner dynamically adjusts processing priorities based on data characteristics. Next, a density-based spatial clustering algorithm is used to identify high-density regions within the voxel grid, and a region growing algorithm is used to merge adjacent voxels into geometric clusters. Each geometric cluster is then fitted with a local quadratic surface to generate a smooth surface representation, preserving boundary and curvature information. A distance-weighted strategy adjusts fitting weights based on voxel density to enhance the representation of complex geometric structures. Simultaneously, the semantic branch data is segmented using a pre-trained deep convolutional neural network to generate a preliminary semantic segmentation map. RGB and depth images are input into a convolutional network with an encoder-decoder architecture. The encoder extracts deep features, while the decoder restores spatial resolution. The edge enhancement module weights boundary region predictions using a gradient operator to improve segmentation accuracy. Subsequently, the spatial structural features and the preliminary semantic segmentation map are integrated using a heterogeneous graph fusion algorithm to generate an initial geometric semantic feature map. A heterogeneous graph is constructed, where nodes represent geometric clusters and semantic regions, and edges reflect spatial adjacency and semantic associations. The graph convolutional network updates node features through multi-layer propagation, combining surface parameters and class labels to generate a fused feature vector. A multi-hop neighborhood aggregation mechanism captures complex relationships, and an edge weight matrix adjusts the fusion ratio of geometry and semantics. The initial geometric semantic feature map is optimized through cross-modal bias correction. Based on the projection error between the point cloud and the image, the algorithm constructs a bias correction matrix. It then adjusts spatial position and label deviations using geometric constraints (such as normal vector consistency) and semantic constraints (such as class probability distribution). An iterative optimization framework alternately updates the correction matrix and feature map until convergence, generating the final geometric semantic feature map after correction.

[0097] In one embodiment, the step of extracting local, regional, and global features of the geometric semantic feature map and optimizing based on a loss function to generate a multi-scale feature tensor includes:

[0098] Spatially partitioning the geometric semantic feature map to generate an initial feature set containing local, regional and global information;

[0099] Initializing a deep neural network, performing feature extraction processing on the initial feature set according to the deep neural network, and generating a hierarchical feature tensor;

[0100] Constructing a loss function consisting of object classification loss, boundary detection loss, and spatial positioning loss, and optimizing the hierarchical feature tensor according to the loss function to obtain an optimized feature tensor;

[0101] The local, regional and global scale features in the optimized feature tensor are subjected to feature fusion processing to generate a multi-scale feature tensor.

[0102] In the above embodiment, the geometric semantic feature map is spatially partitioned to generate an initial feature set containing local, regional and global information. This process is achieved by designing a hybrid decomposition network. The network combines three-dimensional convolution operations and adaptive pooling mechanisms to divide the feature map into feature sub-maps of different scales. High-resolution convolution kernels are used to capture edge and texture details for local features. Regional features use medium-scale convolution kernels to extract object surface and shape information. Global features use global average pooling to compress spatial dimensions to extract environmental layout information. During the decomposition process, the semantically guided attention mechanism dynamically adjusts the weights of features at each scale according to the semantic segmentation distribution of the feature map, ensuring that areas with high semantic density obtain more computing resources, thereby generating a hierarchical initial multi-scale feature set. Next, a deep neural network is initialized based on the initial multi-scale feature set to perform feature extraction processing and generate a hierarchical feature tensor. The network adopts a hybrid architecture combining a 3D convolutional network, a Transformer, and a graph neural network. It first performs a preliminary encoding of the feature set using a 3D convolutional module to generate a high-dimensional feature representation. The Transformer module then captures long-range dependencies between features to enhance contextual relevance, while the graph neural network further enriches the feature representation by modeling the topological relationships of the feature subgraph through node aggregation and edge updates. A dynamic branch selector design enables the network to adaptively allocate computational resources based on the scale distribution of the feature set, prioritizing regions of high semantic complexity. Residual connections and a cross-layer skip structure ensure the fluidity and expressiveness of deep features, resulting in a hierarchical feature tensor containing multi-level and multi-dimensional features. During the optimization phase, a multi-task loss function consisting of an object classification loss, a boundary detection loss, and a spatial localization loss is constructed to optimize the hierarchical feature tensor to produce an optimized feature tensor. Specifically, the hierarchical feature tensor is fed into a multi-task prediction head, which generates classification, boundary, and localization predictions, respectively. The loss for each task is calculated by comparing the predictions with the ground-truth labels. This head assesses the degree of feature sharing between tasks, thereby optimizing the feature representation across tasks. An adaptive weight regulator dynamically balances the importance of each loss term. During the optimization process, gradient normalization and momentum decay strategies are used to ensure training stability. The resulting optimized feature tensor has higher feature discrimination and task adaptability. Finally, the local, regional, and global scale features in the optimized feature tensor are implemented through a cross-scale feature fusion network. This network designs a multi-branch fusion framework, aligns the scales of the feature tensor through upsampling and downsampling operations, and then uses feature splicing and channel attention mechanisms to achieve interactive fusion of cross-scale features. The adaptive fusion controller dynamically adjusts the fusion weights according to the semantic density and geometric complexity of the feature tensor, prioritizing the enhancement of feature regions with high information content. The multi-layer perceptron further compresses the feature dimensions to reduce redundant information. The final output multi-scale feature tensor integrates complete information at local, regional, and global scales, and has a compact and efficient three-dimensional environment feature representation.

[0103] In one embodiment, the step of performing three-dimensional modeling processing based on the multi-scale feature tensor to generate a three-dimensional geometric model includes:

[0104] hierarchically decoding the local, regional, and global features in the multi-scale feature tensor, extracting spatial information corresponding to the local features, regional features, and global features, respectively, to generate a spatial representation;

[0105] Performing registration processing on the point cloud data according to the spatial representation to obtain a preliminary three-dimensional point cloud model;

[0106] Acquire a neighborhood relationship between points in the preliminary three-dimensional point cloud model, construct a topological map of the point cloud, and construct a topological point cloud model based on the topological map;

[0107] Performing boundary prediction processing on the topological point cloud model based on a boundary prediction algorithm to obtain a boundary enhanced point cloud model;

[0108] The boundary-enhanced point cloud model is segmented, high-semantic regions are modeled preferentially, and global modeling is completed through incremental expansion to generate a three-dimensional geometric model.

[0109] In the above-described embodiment, a multi-layer perceptron network performs feature mapping on information at different scales within the feature tensor, hierarchically decoding local features (such as detailed textures of point clouds), regional features (such as local shapes of objects), and global features (such as overall scene layout) to generate a continuous spatial representation. This spatial representation incorporates geometric information and incorporates semantic information. In its specific implementation, a volume disparity-based sampling strategy prioritizes dense sampling of highly semantically significant regions, and gradient descent is used to optimize the weights of the multi-layer perceptron, ensuring that the spatial representation accurately captures the geometric structure and semantic distribution of the environment. Based on the generated spatial representation, point cloud data is registered to construct a preliminary 3D point cloud model. This process utilizes a point cloud registration algorithm to achieve geometric alignment. Regions with high spatial occupancy probability are extracted from the spatial representation to generate a virtual point cloud. Subsequently, an iterative closest point algorithm is used to perform point-by-point matching between the actual point cloud data and the virtual point cloud, calculating the spatial transformation matrix for registration. To improve the robustness of the registration, a local feature matching mechanism is introduced to optimize matching accuracy by comparing the normal and curvature features of the point clouds. A random sampling consensus algorithm is also used to eliminate outliers. This process preprocesses the actual point cloud to remove noise and samples high-probability points from the spatial representation to ensure the accuracy of the registration results. Subsequently, a topological map of the point cloud is constructed by analyzing the neighborhood relationships of each point in the point cloud. A graph cut algorithm is used to identify and repair topologically broken areas, ensuring surface continuity and the absence of self-intersections. Laplace smoothing techniques are combined to geometrically constrain the point cloud surface to maintain curvature continuity. A semantic-based weighting factor is introduced to prioritize the topological structure of highly semantically significant regions. This process iteratively optimizes the position and connectivity of points to generate a topologically consistent point cloud model. Furthermore, geometric and semantic features related to object boundaries are extracted from the topologically consistent point cloud model. A multi-layer perceptron is used to regress and predict the boundary probability of each point, marking high-probability boundary points and enhancing their feature representation. An attention-based feature aggregation method is employed to focus on local features in boundary regions. A cross-scale feature fusion strategy combines global information from multi-scale feature tensors to improve the accuracy of boundary predictions, resulting in a boundary-enhanced point cloud model. High-semantic regions are prioritized for modeling, and a volumetric geometry reconstruction algorithm is used to construct the geometric structure of key objects. This is then gradually expanded to low-semantic regions to generate a complete geometric model. Voxel-based meshing technology is used to convert the point cloud into a voxel representation, and stereo projection is used to optimize the connectivity between voxels. An incremental update mechanism is also introduced to dynamically adjust the model's topology to accommodate new data. This process, through region segmentation, prioritized modeling, and incremental expansion, ensures that the resulting 3D geometric model accurately captures the details of high-semantic regions while fully covering the entire scene.

[0110] In one embodiment, the step of fusing and optimizing the structural information of the three-dimensional geometric model with the geometric semantic feature map and outputting a structured perception result includes:

[0111] Extracting point cloud features of the three-dimensional geometric model, and spatially registering the point cloud features with pixel-level semantic labels of the semantic feature map to obtain an aligned feature set;

[0112] Analyzing and processing the point cloud geometric information and semantic labels contained in the alignment feature set based on a graph convolutional network to construct a spatial and semantic association graph between objects;

[0113] Acquire multi-view data of the robot perception system, project node features of the space and semantic association graph into the two-dimensional feature space of each view, and perform weighted fusion on the multi-view features to generate a fused feature tensor;

[0114] Identifying an occluded area in the fused feature tensor and performing depth completion processing on the occluded area to obtain a completed feature field, wherein the occluded area includes an area where point clouds are missing or semantically discontinuous;

[0115] The point cloud in the completed feature field is segmented into multiple object instances through a clustering algorithm, and the local features of each object instance are extracted based on a 3D convolutional network.

[0116] Global feature aggregation is performed on the multiple local features, and the output includes the three-dimensional coordinates, direction vector and point cloud representation of the object to form a structured perception result.

[0117] In the above-described embodiment, a multimodal feature alignment algorithm is used to achieve unified cross-modal representation based on the coordinates and surface normals of point cloud data, as well as the spatial distribution and category labels in the semantic feature map. Specifically, the 3D geometric model can be voxelized to segment the point cloud data into a regular voxel grid to reduce computational complexity. A multi-layer perceptron (MLP) is then used to map the local geometric features of the point cloud to the same feature space as the semantic feature map. The correspondence between the point cloud coordinates and the pixels of the semantic feature map is calculated using volumetric geometry constraints. An attention mechanism is then introduced to align the feature scales of the different modalities, ultimately generating an aligned feature set that describes both geometric and semantic consistency. This process continuously adjusts the alignment parameters through a cyclic optimization process to ensure the accuracy of the feature set. Using the point cloud geometric information and semantic labels in the aligned feature set, a graph structure is constructed, in which nodes represent point cloud clusters or objects in the scene, and edges represent spatial adjacency or semantic correlation between objects. The Euclidean distances and direction vectors between objects are calculated based on the point cloud coordinates to generate an initial spatial adjacency matrix. Simultaneously, a category correlation matrix is ​​generated based on the semantic labels. Node features are then aggregated and updated through a multi-layer graph convolutional network. Each convolution layer fuses geometric properties (such as position and orientation) and semantic properties (such as category) of nodes, and dynamically adjusts edge weights to control the strength of neighborhood information transmission. After multiple rounds of iterative convolution, the graph structure effectively captures the spatial and semantic relationships between objects, ultimately generating a semantic-geometric relationship graph through global pooling. Leveraging multi-view data provided by the robot's perception system (such as depth maps or RGB images), the node features of the relationship graph are projected into a two-dimensional feature space at different viewpoints through a projection transformation, generating view-dependent local feature descriptors. Fusion of these multi-view features utilizes a transform encoder, using a self-attention mechanism to capture inter-view feature correlations and correcting projection errors through cross-view geometric constraints. After multiple rounds of attention computation, the fused features undergo dimensionality reduction to generate a high-dimensional fused feature tensor containing globally consistent geometric and semantic information. The point cloud distribution and semantic labels in the fused feature tensor are analyzed, and occluded areas are identified by detecting areas of sudden changes in point cloud density or semantic labels. A generative adversarial network (GAN) is used to predict the underlying geometric structure of the occluded area. The generator is responsible for generating the completed point cloud features, while the discriminator optimizes the consistency of the completed features with the true point cloud distribution, while also combining semantic labels to constrain the category consistency of the completed area. In terms of execution logic, the boundaries of the occluded area are first detected, and then the point cloud features are completed layer by layer through a multi-scale generative network. Finally, a continuous completed feature field is generated through semantic-guided optimization. This feature field can fully describe the geometric and semantic distribution of the scene.Finally, a density-based clustering algorithm (such as DBSCAN) is used to segment the point cloud in the completed feature field into multiple object instances and assign a semantic label to each instance. A three-dimensional convolutional network is used to extract the local features of each object instance, capturing its geometric and semantic characteristics. Multiple local features are fused through a global feature aggregation module, and a multi-task learning framework is used to simultaneously optimize the object's category prediction, position regression, and morphological reconstruction. The final output is a structured perception result containing the object's three-dimensional coordinates, direction vector, and point cloud representation.

[0118] In one embodiment, the step of verifying the structured perception results based on the verification framework to obtain three-dimensional environment perception data includes:

[0119] Mapping multiple data in the synchronized visual data stream to a unified coordinate system, and calculating the data deviation of the visual data stream in the unified coordinate system to generate a verification matrix;

[0120] Performing confidence modeling processing on the structured perception result according to the verification matrix to obtain a confidence distribution map;

[0121] Identifying areas in the confidence distribution map that are below a preset threshold, grouping low-confidence points into continuous deviation areas through cluster analysis, and generating a deviation area annotation set;

[0122] Performing local optimization processing on the structured perception result according to the deviation area annotation set to obtain an optimized perception result;

[0123] The optimized perception results are subjected to structured coding processing to obtain three-dimensional environment perception data.

[0124] In the above-described embodiments, data streams from different modalities, such as visual data, point cloud data, and motion data, are unified into a common three-dimensional coordinate system through projective mapping. For example, two-dimensional visual data is projected into three-dimensional space using camera intrinsics and extrinsic parameters, or point cloud data collected by different sensors is aligned using a point cloud registration algorithm. During this process, spatial consistency errors between the data modalities in the unified coordinate system are calculated, such as the projection deviation between the visual data and the point cloud data, or the degree of temporal mismatch between the motion data and the three-dimensional geometric model. These deviations are quantified using a consistency analysis algorithm into a verification matrix that records the distribution of spatial and temporal deviations. A Bayesian filtering algorithm is used to iteratively update the observation probability of the visual data and the prior model. For example, the semantic information provided by the visual data, the geometric information provided by the point cloud data, and the temporal information provided by the motion data are input into the Bayesian model to generate confidence scores for each 3D point and its semantic label. These scores reflect the credibility of the perception results in different regions, ultimately forming an intuitive confidence distribution map, where high-confidence areas indicate reliable perception results, while low-confidence areas may contain errors or missing data. The confidence distribution map is analyzed to identify regions below a preset threshold, and cluster analysis is used to generate a set of deviation region annotations. This process uses a threshold segmentation algorithm to extract from the distribution map a set of points with confidence below a certain threshold, such as regions with confidence below 0.7. These low-confidence points are then grouped into contiguous deviation regions using a clustering algorithm (such as DBSCAN or K-means). Combined with the geometric semantic feature map from the structured perception results, each deviation region is annotated with its spatial location, semantic category, and possible error sources (such as sensor noise or occlusion). Guided by the deviation region annotation set, the structured perception results are locally optimized through local data re-collection and feature fusion. For example, the sensor parameters or acquisition path of the robot perception system can be adjusted based on the spatial location and semantic information of the deviation region, and visual data covering the deviation region can be re-collected. This newly collected data is timestamped and aligned with the original data stream to form a local enhanced data stream. Feature extraction is then performed on the local enhanced data stream, such as extracting spatial structural features from the point cloud or a semantic segmentation map from the visual data, and then integrated with the corresponding region in the original perception results using a weighted fusion algorithm. During the fusion process, a multi-scale feature tensor optimization mechanism can be employed to ensure a smooth transition between new and old data, thereby updating the 3D geometric model and geometric semantic feature map. The updated 3D geometric model and geometric semantic feature map are then encoded using volumetric geometry analysis and a graph neural network, integrating point cloud data, semantic labels, and motion trajectory information into a unified structured data format. This data structure not only contains information such as object category, spatial position, orientation, and morphology, but also preserves the correlation between visual data.

[0125] Reference Figure 2 , a robot three-dimensional environment perception device based on deep visual learning, comprising:

[0126] The acquisition module 100 is used to collect visual data based on the perception system carried by the robot, and synchronously process the visual data to obtain a synchronized visual data stream;

[0127] An extraction module 200 is configured to extract spatial structural features and a preliminary semantic segmentation map of the synchronized visual data stream, and combine the spatial structural features and the preliminary semantic segmentation map to generate a geometric semantic feature map;

[0128] An optimization module 300 is used to extract local, regional and global features of the geometric semantic feature map and optimize it based on a loss function to generate a multi-scale feature tensor;

[0129] A modeling module 400 is configured to perform three-dimensional modeling processing based on the multi-scale feature tensor to generate a three-dimensional geometric model;

[0130] A fusion module 500 is configured to fuse and optimize the structural information of the 3D geometric model with the geometric semantic feature map, and output a structured perception result;

[0131] The verification module 600 is used to verify the structured perception results based on the verification framework to obtain three-dimensional environment perception data.

[0132] Reference Figure 3 In the embodiment of the present application, a computer device is also provided. The computer device may be a server, and its internal structure may be as follows: Figure 3 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer design is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data such as a robot three-dimensional environment database. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a robot three-dimensional environment perception method based on deep visual learning is implemented.

[0133] An embodiment of the present application also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, a robot three-dimensional environment perception method based on deep visual learning is implemented, comprising the following steps: collecting visual data based on a perception system carried by the robot, synchronously processing the visual data to obtain a synchronized visual data stream; extracting spatial structural features and a preliminary semantic segmentation map of the synchronized visual data stream, and combining the spatial structural features and the preliminary semantic segmentation map to generate a geometric semantic feature map; extracting local, regional and global features of the geometric semantic feature map, and optimizing it based on a loss function to generate a multi-scale feature tensor; performing three-dimensional modeling processing based on the multi-scale feature tensor to generate a three-dimensional geometric model; fusing and optimizing the structural information of the three-dimensional geometric model with the geometric semantic feature map to output a structured perception result; and verifying the structured perception result based on a verification framework to obtain three-dimensional environment perception data.

[0134] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When executed, the computer program can include the processes of the above-described method embodiments. Any reference to memory, storage, database, or other media provided herein and used in the embodiments may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double-speed SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct RAMbus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM).

[0135] The above description is only a preferred embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the contents of the present invention description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. A robot three-dimensional environment perception method based on deep visual learning, characterized in that: include: Collecting multiple visual data based on a perception system carried by the robot, and synchronously processing the visual data to obtain a synchronized visual data stream; Extracting spatial structural features and a preliminary semantic segmentation map of the synchronized visual data stream, and combining the spatial structural features and the preliminary semantic segmentation map to generate a geometric semantic feature map; Extracting local, regional, and global features of the geometric semantic feature map, and optimizing them based on a loss function to generate a multi-scale feature tensor; Performing three-dimensional modeling processing according to the multi-scale feature tensor to generate a three-dimensional geometric model; fusing and optimizing the structural information of the three-dimensional geometric model with the geometric semantic feature map to output a structured perception result; Verify the structured perception results based on the verification framework to obtain three-dimensional environmental perception data; The steps of collecting visual data based on the perception system carried by the robot and synchronously processing the visual data to obtain a synchronized visual data stream include: Performing data separation processing on the visual data collected by the robot perception system to obtain multiple separated data; Mapping the separated data to a three-dimensional spatial reference system based on the coordinate system transformation matrix, and performing time alignment on the separated data to obtain visual data; Acquiring current posture information of the robot, and optimizing the visual data according to the current posture information to obtain optimized visual data; Adjusting the data transmission order in the optimized visual data based on a preset priority to generate a synchronized visual data stream including a three-dimensional point cloud, an image sequence, a motion trajectory, and distance information; The step of extracting the spatial structural features and the preliminary semantic segmentation map of the synchronized visual data stream, and combining the spatial structural features and the preliminary semantic segmentation map to generate a geometric semantic feature map comprises: Performing dual-path decomposition processing on the synchronized visual data stream to obtain geometric branch data and semantic branch data; Identifying high-density areas in the geometric branch data, merging similar points in the high-density areas into geometric clusters, and fitting the geometric clusters based on surface fitting technology to obtain spatial structural features; Segmenting the semantic branch data based on a pre-trained deep convolutional neural network to obtain a preliminary semantic segmentation map; Performing heterogeneous graph fusion processing on the spatial structure features and the preliminary semantic segmentation map to obtain an initial geometric semantic feature map; Obtaining projection errors between a three-dimensional point cloud and an image sequence in a synchronized visual data stream, and constructing a bias correction matrix based on the projection errors; The initial geometric semantic feature map is subjected to deviation correction processing according to the deviation correction matrix to obtain a geometric semantic feature map.

2. A robot three-dimensional environment perception method based on deep visual learning according to claim 1, characterized in that: The step of extracting local, regional, and global features of the geometric semantic feature map and optimizing based on a loss function to generate a multi-scale feature tensor includes: Spatially partitioning the geometric semantic feature map to generate an initial feature set containing local, regional and global information; Initializing a deep neural network, performing feature extraction processing on the initial feature set according to the deep neural network, and generating a hierarchical feature tensor; Constructing a loss function consisting of object classification loss, boundary detection loss, and spatial positioning loss, and optimizing the hierarchical feature tensor according to the loss function to obtain an optimized feature tensor; The local, regional and global scale features in the optimized feature tensor are subjected to feature fusion processing to generate a multi-scale feature tensor.

3. The robot three-dimensional environment perception method based on deep visual learning according to claim 1 is characterized in that: The step of performing three-dimensional modeling processing according to the multi-scale feature tensor to generate a three-dimensional geometric model includes: hierarchically decoding the local, regional, and global features in the multi-scale feature tensor, extracting spatial information corresponding to the local features, regional features, and global features, respectively, to generate a spatial representation; Performing registration processing on the point cloud data according to the spatial representation to obtain a preliminary three-dimensional point cloud model; Acquire a neighborhood relationship between points in the preliminary three-dimensional point cloud model, construct a topological map of the point cloud, and construct a topological point cloud model based on the topological map; Performing boundary prediction processing on the topological point cloud model based on a boundary prediction algorithm to obtain a boundary enhanced point cloud model; The boundary-enhanced point cloud model is segmented, high-semantic regions are modeled preferentially, and global modeling is completed through incremental expansion to generate a three-dimensional geometric model.

4. The robot three-dimensional environment perception method based on deep visual learning according to claim 1 is characterized in that: The step of fusing and optimizing the structural information of the three-dimensional geometric model with the geometric semantic feature map and outputting a structured perception result includes: Extracting point cloud features of the three-dimensional geometric model, and spatially registering the point cloud features with pixel-level semantic labels of the geometric semantic feature map to obtain an aligned feature set; Analyzing and processing the point cloud geometric information and semantic labels contained in the alignment feature set based on a graph convolutional network to construct a spatial and semantic association graph between objects; Acquire multi-view data of the robot perception system, project node features of the space and semantic association graph into the two-dimensional feature space of each view, and perform weighted fusion on the multi-view features to generate a fused feature tensor; Identifying an occluded area in the fused feature tensor and performing depth completion processing on the occluded area to obtain a completed feature field, wherein the occluded area includes an area where point clouds are missing or semantically discontinuous; The point cloud in the completed feature field is segmented into multiple object instances through a clustering algorithm, and the local features of each object instance are extracted based on a 3D convolutional network. Global feature aggregation is performed on the multiple local features, and the output includes the three-dimensional coordinates, direction vector and point cloud representation of the object to form a structured perception result.

5. The robot three-dimensional environment perception method based on deep visual learning according to claim 1 is characterized in that: The step of verifying the structured perception results based on the verification framework to obtain three-dimensional environment perception data includes: Mapping multiple data in the synchronized visual data stream to a unified coordinate system, and calculating the data deviation of the visual data stream in the unified coordinate system to generate a verification matrix; Performing confidence modeling processing on the structured perception result according to the verification matrix to obtain a confidence distribution map; Identifying areas in the confidence distribution map that are below a preset threshold, grouping low-confidence points into continuous deviation areas through cluster analysis, and generating a deviation area annotation set; Performing local optimization processing on the structured perception result according to the deviation area annotation set to obtain an optimized perception result; The optimized perception results are subjected to structured coding processing to obtain three-dimensional environment perception data.

6. A robot three-dimensional environment perception device based on deep visual learning, applied to the robot three-dimensional environment perception method according to any one of claims 1 to 5, characterized in that: include: An acquisition module is used to collect visual data based on the perception system carried by the robot, and synchronously process the visual data to obtain a synchronized visual data stream; an extraction module, configured to extract spatial structural features and a preliminary semantic segmentation map of the synchronized visual data stream, and combine the spatial structural features and the preliminary semantic segmentation map to generate geometric semantic features; An optimization module, configured to extract local, regional, and global features of the geometric semantic feature map and perform optimization based on a loss function to generate a multi-scale feature tensor; A modeling module, configured to perform three-dimensional modeling processing based on the multi-scale feature tensor to generate a three-dimensional geometric model; A fusion module, configured to fuse and optimize the structural information of the three-dimensional geometric model with the geometric semantic feature map, and output a structured perception result; The verification module is used to verify the structured perception results based on the verification framework to obtain three-dimensional environment perception data.

7. A computer device, characterized in that: The method comprises a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 5 when executing the computer program.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Semantic mapping method based on visual SLAM and two-dimensional semantic segmentation

    CN111462135A

  • Indoor dynamic SLAM method under multi-source semantic perception

    CN116007607A