Robot integrating multi-modal scene recognition and automatic surveying and mapping functions and method thereof

Through intelligent building surveying and mapping robots integrating multimodal sensors and advanced algorithms, the problems of low identification accuracy and insufficient surveying and mapping efficiency in the existing technology are solved, and high-precision three-dimensional map construction and intelligent decision-making are realized, which improves the efficiency and quality of building construction.

CN120063305APending Publication Date: 2025-05-30ZHENGZHOU UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510142687.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing intelligent surveying and mapping robots have low recognition accuracy in complex environments, their perception systems are susceptible to light, occlusion and noise, and lack multimodal data fusion and intelligent decision-making capabilities, resulting in insufficient surveying and mapping efficiency and accuracy.

Method used

Design a robot that integrates multimodal scene recognition and automatic mapping functions, adopts multi-sensor fusion technology, including lidar, depth camera and voice module, and combines composite exploration algorithms and SLAM algorithms to realize high-precision three-dimensional map construction and semantic map analysis.

Benefits of technology

It improves the accuracy and efficiency of construction, reduces costs and safety risks, promotes the digital and intelligent development of the construction industry, and realizes a more efficient, accurate and intelligent building surveying and mapping verification system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120063305A_ABST
    Figure CN120063305A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of house construction, in particular to a robot integrating multi-mode scene recognition and automatic surveying and mapping functions and a method thereof. Comprising an execution module, a decision module, a sensing module, a data processing module and a voice module. The execution module is used for driving the robot to move; the decision module is used for controlling work of the execution module; the sensing module is used for detecting the environment on the outer side; the data processing module is used for processing the data detected by the sensing module; and the voice module is used for receiving external voice and playing a processing result through the voice. According to the robot integrating the multi-modal scene recognition and automatic surveying and mapping functions and the method thereof, the defects in the aspects of verification speed, precision, automation degree and the like in the prior art can be overcome, and through the combination of the robot, three-dimensional reconstruction and a CAD model comparison technology, the accuracy of verification is improved. A more efficient, accurate and intelligent building surveying and mapping verification system is constructed, so that the surveying and mapping efficiency and quality of building engineering are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of building construction, and particularly to a robot with multi-modal scene recognition and automatic mapping functions and a method thereof. Background Art

[0002] Intelligent building mapping robots are currently mainly used in the processes such as in the initial stage of house and building construction, in the roughcast, decoration testing and inspection.

[0003] In the roughcast stage, intelligent mapping robots can conduct detailed mapping and recording of the building structure. They can capture every detail, including the thickness of walls, the positions of doors and windows, the layout of pipes and wires, etc. These data are crucial for subsequent decoration design and construction, because they can ensure that the decoration plan fits perfectly with the building structure and avoid unnecessary troubles and losses during the decoration process.

[0004] In the decoration testing stage, intelligent mapping robots can conduct real-time monitoring and evaluation of the decoration effect. They can capture every detail change during the decoration process, including color changes, material textures, furniture placement, etc. By analyzing these data, the robot can timely discover problems existing in the decoration process and propose corresponding solutions, thus ensuring that the decoration quality reaches the expected standard.

[0005] In the later inspection and maintenance, intelligent mapping robots also play an important role. They can conduct regular inspections and maintenance of the building, and timely discover potential safety hazards, such as wall cracks, pipe leaks, etc. Through timely repair and maintenance, the service life of the building can be extended and people's life and property safety can be guaranteed.

[0006] In short, intelligent building mapping robots play an important role throughout the life cycle of houses and buildings. They not only improve the efficiency and accuracy of mapping and design, but also provide strong support for subsequent decoration, inspection and maintenance. With the continuous development of technology, the application of intelligent mapping robots in the construction field will be more and more extensive.

[0007] Compared with traditional manual mapping, intelligent building mapping robots have unique advantages, such as high measurement accuracy, time saving, low cost, high efficiency, etc., and have been widely used in the field of building mapping. Currently, intelligent building mapping robots mainly rely on single radar scanning, and generate other maps such as point cloud maps through radar scanning of the surrounding environment. Thus, an internal contour map of a building is generated.

[0008] Intelligent surveying and mapping robots adopt advanced sensor technologies and automated measurement processes, and can complete large-scale surveying and mapping tasks in a short time. Compared with traditional surveying and mapping that requires manual operation of instruments or relies on equipment such as total stations for high-precision measurement, the surveying and mapping efficiency of intelligent surveying and mapping robots has undoubtedly been greatly improved.

[0009] In terms of control, existing intelligent surveying and mapping robots mainly rely on single program instructions. Although they can achieve some simple direction control and form a modeling map through autonomous cruising, it is only a single-modal human-machine interaction and lacks true intelligence.

[0010] Current intelligent surveying and mapping robots need to have the ability to perceive and adapt to complex environments. However, in actual applications, the perception systems of robots are often interfered by environmental factors such as light, occlusion, and noise, resulting in a decrease in perception accuracy and stability.

[0011] The deficiencies of existing intelligent surveying and mapping robots in the prior art are as follows:

[0012] 1. Many existing construction robots only have a single function, such as only conducting surveying and mapping or only conducting scene recognition, lacking the ability of multi-modal fusion. This requires users to cooperate with multiple devices, increasing the operation complexity and cost.

[0013] 2. Traditional scene recognition technologies have low recognition accuracy in complex environments and are easily affected by factors such as light and perspective, resulting in inaccurate generated data and affecting subsequent construction decisions.

[0014] 3. Existing technologies often have bottlenecks in processing multi-source data and cannot effectively fuse data collected by different sensors, resulting in information loss or redundancy and affecting data analysis and decision support.

[0015] 4. The system usually only relies on a single input method (such as voice or image) and lacks diverse interaction options. This limits the user experience and reduces the flexibility and accuracy of human-machine interaction.

[0016] 5. Traditional scene recognition technologies have low recognition accuracy in complex environments and are easily affected by factors such as light and perspective, resulting in inaccurate generated data and affecting subsequent construction decisions.

[0017] 6. Current surveying and mapping methods mostly rely on manual operation or step-by-step processing, which is time-consuming and difficult to meet the needs of rapid building design and construction, reducing the overall project efficiency.

[0018] 7. Many existing construction robots lack intelligent decision-making capabilities and rely on preset programs for operation, unable to flexibly respond to on-site changes and complex situations.

[0019] 8. Currently, the designs of many construction robots do not take into account the complexity of various construction environments and lack adaptability in different industries or scenarios, resulting in poor performance in actual applications.

[0020] Therefore, a robot and its method integrating multi-modal scene recognition and automatic mapping functions are designed to provide another technical solution to the above technical problems. Summary of the Invention

[0021] Based on this, in order to solve the above technical problems, it is necessary to provide a robot and its method integrating multi-modal scene recognition and automatic mapping functions to solve the technical problems raised in the above background technology. To solve the above technical problems, the present invention adopts the following technical solutions:

[0022] A robot integrating multi-modal scene recognition and automatic mapping functions includes an execution module, a decision-making module, a perception module, a data processing module, and a voice module;

[0023] The execution module is used to drive the robot to move;

[0024] The decision-making module is used to control the operation of the execution module;

[0025] The perception module is used to detect the external environment;

[0026] The data processing module is used to process the data detected by the perception module;

[0027] The voice module is used to receive external voices and play the processing results through voice.

[0028] As a preferred embodiment of the robot integrating multi-modal scene recognition and automatic mapping functions provided by the present invention, the execution module is a four-wheel drive sports chassis, and the four-wheel drive sports chassis is composed of a chassis, a power supply, a motor driver, a motor, and wheels. Wheels are rotatably connected to the four end corners of the chassis. Motors are fixed at positions corresponding to the wheels inside the chassis. The output ends of the motors are connected to the wheels. A power supply is fixed inside the chassis. A voltage conversion module is fixed at the bottom of the battery inside the chassis. A motor driver is fixed at one end of the power supply on the top of the chassis.

[0029] As a preferred embodiment of the robot integrating multi-modal scene recognition and automatic mapping functions provided by the present invention, the decision-making module is a main control of the Robot Operating System (ROS), and the ROS main control is installed inside the chassis and at one end of the motor driver.

[0030] As a preferred embodiment of the robot with multi-modal scene recognition and automatic mapping functions provided by the present invention, the sensing module is a sensor, the sensor is lower than the top of the chassis, and the sensor is at least composed of a lidar sensor and a depth camera.

[0031] A detection method for a robot with multi-modal scene recognition and automatic mapping functions. Through the detection of the environment by the sensing module, the decision-making module controls the execution module to drive the robot to move;

[0032] When moving, the robot is guided to move towards the target point and explore the unknown area through a composite exploration algorithm;

[0033] Feature points are extracted from the indoor weak texture environment through the Simultaneous Localization and Mapping (SLAM) algorithm.

[0034] As a preferred embodiment of the detection method for the robot with multi-modal scene recognition and automatic mapping functions provided by the present invention, the composite exploration algorithm construction method is as follows:

[0035] Obtain omnidirectional environmental information through sensors;

[0036] Perform image recognition and semantic segmentation through deep learning technology to extract key information in the environment;

[0037] Fuse lidar data and image data to construct a high-precision three-dimensional map and semantic map;

[0038] According to the position of the target point, divide the map into half-planes pointing to the target point to limit the exploration area;

[0039] Within the divided area, find all boundary points through search algorithms such as Depth-First Search (DFS);

[0040] Construct boundary lines to form the boundary contour of the exploration area;

[0041] Score the boundary points through an evaluation function, and select the boundary point with the highest score as the best boundary point;

[0042] Perform vector synthesis on the best boundary point and the target point to generate a composite exploration point;

[0043] Through global path planning algorithms such as D*, according to the composite exploration point and the current map information, plan the optimal path from the current position of the robot to the composite exploration point;

[0044] During the movement of the robot, algorithms such as the Dynamic Window Approach (DWA) are used to adjust the robot's motion trajectory in real time;

[0045] Through positioning algorithms such as Cartographer, the position information of the robot is updated in real time, and the map is updated in real time according to the updated position of the robot.

[0046] As a preferred implementation manner of the detection method of the robot with multi-modal scene recognition and automatic mapping functions provided by the present invention, the SLAM algorithm is constructed as follows:

[0047] Through improved algorithms such as the Line Segment Detector (LSD), line features are extracted from the image;

[0048] Through algorithms such as Oriented FAST and Rotated BRIEF (ORB), point features are extracted from the image;

[0049] Through point feature matching, corresponding points in two frames of images are found, and preliminary pose estimation is performed through algorithms such as the Perspective-n-Point (PnP) algorithm;

[0050] Through line feature matching, and using the matched line features as additional constraints, the pose is iteratively optimized;

[0051] Loop closure detection is performed by minimizing the reprojection error to determine whether the robot has returned to a previously visited location;

[0052] According to the pose estimation results and sensor data, a three-dimensional map or a two-dimensional grid map of the environment is constructed;

[0053] Through known CAD building feature data and laser ranging data, the constructed three-dimensional map or two-dimensional grid map is cross-checked and corrected to obtain corrected mapping data and map.

[0054] It can be undoubtedly seen that through the above technical solutions of the present application, the technical problems to be solved by the present application can surely be solved.

[0055] At the same time, through the above technical solutions, the present invention has at least the following beneficial effects:

[0056] 1. The robot and its method with multimodal scene recognition and automatic surveying and mapping functions provided by the present invention can solve the deficiencies of the prior art in aspects such as verification speed, accuracy, and automation level. By combining the robot, 3D reconstruction, and CAD model comparison technologies, a more efficient, accurate, and intelligent building surveying and verification system is constructed to improve the surveying and mapping efficiency and quality of construction projects.

[0057] 2. The intelligent building robot of the present invention, through its multimodal scene recognition and automatic surveying and mapping functions, not only improves the accuracy and efficiency of building construction, reduces costs and safety risks, but also promotes the digital and intelligent development of the construction industry, and has important practical application value and broad market prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0059] Figure 1 is the framework diagram of the present invention;

[0060] Figure 2 is the algorithm flowchart of the present invention;

[0061] Figure 3 is the schematic diagram of the main architecture of the robot of the present invention;

[0062] Figure 4 is the electrical connection schematic diagram of the signal processing end of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0063] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the following will further elaborate on the present invention in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0064] In order to enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in combination with the drawings.

[0065] It should be noted that, without conflict, the embodiments in the present invention and the features and technical solutions in the embodiments can be combined with each other.

[0066] It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0067] Example 1:

[0068] A robot integrating multi-modal scene recognition and automatic mapping functions, including an execution module, a decision-making module, a perception module, a data processing module, and a voice module;

[0069] The execution module is used to drive the robot to move;

[0070] The execution module is a four-wheel drive sports chassis, which is composed of a chassis, a power supply, a motor driver, a motor, and wheels. Wheels are rotatably connected to the four end corners of the chassis. Motors are fixed at positions corresponding to the wheels inside the chassis. The output ends of the motors are connected to the wheels. A power supply is fixed inside the chassis. A voltage conversion module is fixed at the bottom of the battery inside the chassis, so as to be able to supply power to the motor through the power supply and adjust the output voltage of the battery through the voltage conversion module. A motor driver is fixed at one end of the power supply on the top of the chassis, and then the motor driver is used to control the operation of the motor.

[0071] Thus, when in use, a T-head wire is led out from the power supply to connect the voltage conversion module and the power switch first, and then multiple T-head wires with the same supply voltage as the demand module are respectively led out from the voltage conversion module to connect to the motor, the tram driver, etc. for power supply. And a dedicated wire is led out from the ROS main control to connect to the motor driver through a dedicated wire, and the motor driver is connected to the motor.

[0072] The decision-making module is used to control the operation of the execution module;

[0073] The decision-making module is a ROS main control, which is installed inside the chassis and at one end of the motor driver, so as to be able to control the ROS main control through the ROS main control. Then, a T-head wire is led out from the voltage conversion module to connect for power supply, a dedicated wire is led out from the ROS main control to connect to the motor driver, and then the robot movement is controlled through the motor driver.

[0074] The perception module is used for detecting the external environment;

[0075] The perception module is a sensor, which is lower than the top of the chassis and above the front-wheel motor. The sensor is at least composed of a lidar sensor and an OAK-D-Pro depth camera. Thus, the surrounding environment can be continuously scanned outside through the lidar sensor to generate a reflection map and depth, and the depth camera can generate a grayscale map and obtain more details by using structured light. During power supply, a T-head wire is led out from the voltage conversion module to connect to the lidar and is connected to the depth camera through a Y-type adapter for power supply. And in terms of signals, a data transmission line is led out from the lidar through an Ethernet adapter, and the depth camera leads out a data transmission line through a Y-type adapter.

[0076] The data processing module is used to process the data detected by the sensing module.

[0077] The voice module is used to receive external voices and play the processing results through voice.

[0078] It may also include Jeston AGX Origin 64GB AI and iFlytek AIUI R818 microphone array. During power supply, a T-head wire is led out from the voltage transformation module for connection and power supply. In the signal, a dedicated wire is used to connect the microphone array development kit to the microphone board through a 40-pin cable. At the UAC interface, a dedicated wire is used to connect to the Jeston AGX Origin USB interface for audio output. The external power amplifier is connected to its reference signal interface. Finally, a dedicated wire is used to connect the Jeston AGX Origin to the serial port interface of the microphone array development kit, and the data transmission lines led out by the previous lidar and sensing camera are connected to the Jeston AGX Origin 64GB computing module to complete the signal connection. And the relevant connections of the previous microphone array development kit.

[0079] Furthermore, by installing a microphone array development kit in the middle of the robot, the spatial filtering characteristics of the microphone array can be utilized to form a directional pickup beam through the angle positioning of waking people, and suppress the noise outside the beam, improving the far-field pickup quality.

[0080] Embodiment 2

[0081] Reference Figures 1-4 , on the basis of the above Embodiment 1, a method is disclosed.

[0082] By processing the data collected by multiple sensors and the pose estimation results, the present invention can quickly obtain the 3D map and 2D grid map of the construction site, reduce the manual operation time, and greatly improve the speed of surveying and verification.

[0083] By combining the 3D reconstruction and the comparison technology of the preset CAD model, the present invention can accurately identify the subtle differences in the building structure and significantly improve the verification accuracy.

[0084] By means of the autonomous mobile robot, the present invention realizes the automatic surveying and comparison at the construction site, reduces the human intervention, improves the automation degree of the system, and ensures the reliability and consistency of the verification results.

[0085] By introducing the multi-modal sensor fusion technology, the present invention enables the robot to simultaneously utilize various sensing means such as vision, lidar, and inertial sensors to ensure its real-time and accurate navigation in complex environments.

[0086] By introducing a self-learning navigation algorithm, the present invention enables the robot to autonomously adjust the path planning and navigation strategy, enhancing its autonomous decision-making ability in complex construction site scenarios.

[0087] A lidar and a depth camera are fixed in front of the robot. The lidar can generate a reflection map and a depth map by continuously scanning the surrounding environment, while the depth camera can generate a grayscale map and obtain more details using structured light. Subsequently, the edges detected from these images are calibrated through an algorithm until satisfactory parameters are obtained. Then, SLAM is used for three-dimensional reconstruction. If there is still a large error between the result of the SLAM three-dimensional reconstruction and the actual result, the PNP algorithm needs to be introduced to optimize and reduce the error. For the simple case with a large number of points, a linear line transformation is performed, and for the case with fewer points, the Efficient Perspective-n-Point (EPnP) algorithm is used. Based on the initial value estimation, a more accurate result is obtained through iterative optimization. After PnP optimization, in order to minimize the reprojection error, Bundle Adjustment (BA) optimization is performed and corrected by loop detection.

[0088] Through the detection of the environment by the perception module, the decision-making module controls the execution module to drive the robot to move.

[0089] The robot is guided to move towards the target point and explore unknown areas through a composite exploration algorithm. The construction method is as follows:

[0090] The principle of composite exploration involves region segmentation, boundary point search and evaluation, composite exploration point design, path planning and obstacle avoidance, as well as mapping and updating. First, the exploration area is restricted by dividing the map into half-planes pointing to the target point. Then, all boundary points are searched within the divided area and a boundary line is constructed. Next, the best boundary point is determined through an evaluation function that comprehensively considers the boundary length and the distances to the robot and the target point. After that, the best boundary point and the target point are vectorially combined to generate a composite exploration point, guiding the robot to move towards the target point and explore unknown areas. During this process, the D* algorithm is used for global path planning, the DWA algorithm is used for local obstacle avoidance, and the Cartographer algorithm is used for real-time positioning and map updating.

[0091] First, integrate sensors such as lidar, cameras, ultrasonic sensors, etc. to obtain rich environmental information. Then, through the fusion of multi-modal data, such as using deep learning techniques for image recognition and semantic segmentation, and combining lidar data for map construction, high-precision mapping and semantic map construction are achieved. On this basis, aiming at the characteristics of the building environment, improve the design of composite exploration points and introduce a target traction mechanism to improve the exploration efficiency. At the same time, considering multi-floor environments and dynamic obstacles, optimize the path planning and obstacle avoidance algorithms.

[0092] In addition, the project also focuses on human-computer interaction. Through the interface, it is convenient to set target points and planning tasks, such as inspection, cleaning, etc. The application scenarios are extended to indoor navigation, delivery, and surveying and inspection in outdoor construction sites, industrial parks, etc., giving full play to the advantages of the composite exploration principle in actual projects.

[0093] 1. Environmental perception and data acquisition

[0094] Configure lidar, cameras, ultrasonic sensors, etc. to ensure that environmental information can be obtained omnidirectionally.

[0095] Calibrate each sensor to ensure the consistency and accuracy of the data.

[0096] Use deep learning techniques for image recognition and semantic segmentation to extract key information in the environment.

[0097] Fuse lidar data with image data to construct high-precision 3D maps and semantic maps.

[0098] 2. Region segmentation and boundary establishment

[0099] Region segmentation: According to the position of the target point, divide the map into half-planes pointing to the target point to limit the exploration area.

[0100] Boundary point search: In the segmented area, use appropriate search algorithms (such as DFS, BFS) to find all boundary points.

[0101] Construct boundary lines to form the boundary contour of the exploration area.

[0102] 3. Boundary point evaluation and selection

[0103] Evaluation function design: Design an evaluation function, comprehensively considering factors such as boundary length, the distance from the robot to the boundary point, and the distance from the boundary point to the target point.

[0104] Let B be the set of boundary points, r be the current position of the robot, g be the position of the target point, b i be the boundary point i, L be the boundary length, d(r, b i ) be the distance from the robot to the boundary point i, d(b i, g) is the distance from the boundary point i to the target point, and the evaluation function E(b i ) is defined accordingly:

[0105]

[0106] where w 1 , w 2 and w 3 are weight coefficients used to balance the influence of different factors on the evaluation result, and their values need to be adjusted according to the actual application scenario. represents the reciprocal of the distance from the robot to the boundary point. The closer the distance, the higher the score. represents the reciprocal of the distance from the boundary point to the target point. The closer the distance, the higher the score. represents the reciprocal of the boundary length. The shorter the boundary length, the higher the score.

[0107] Determination of the best boundary point: Use the evaluation function to score the boundary points, and select the boundary point with the highest score as the best boundary point.

[0108] 4. Design of Composite Exploration Points and Path Planning

[0109] Generation of composite exploration points: Vectorially combine the best boundary point and the target point to generate composite exploration points.

[0110] The composite exploration points should be located on the path between the current position of the robot and the target point and be biased towards the unknown area to guide the robot to explore.

[0111] Global path planning: Use global path planning algorithms such as the D* algorithm to plan the optimal path from the current position of the robot to the composite exploration point according to the composite exploration points and the current map information.

[0112] Local obstacle avoidance: During the movement of the robot, use local obstacle avoidance algorithms such as the DWA algorithm to adjust the movement trajectory of the robot in real time to avoid collisions with obstacles.

[0113] 5. Real-time Positioning and Map Update

[0114] Real-time positioning: Use positioning algorithms such as Cartographer to update the position information of the robot in real time.

[0115] Map update: According to the new information explored by the robot, update the map in real time, including information such as newly added obstacles and channels.

[0116] Map update: According to the new information explored by the robot, update the map in real time, including information such as newly added obstacles and channels, that is, the continuous change of the error between the laser points of the lidar and the points on the positioning map. The objective function of Scan Matching (SM) is used to update the position information:

[0117]

[0118] Among them, E(x) is the error function, representing the total error in the current pose. x is the current pose of the robot, usually including position and orientation. z t,i is the measurement value of the i-th laser point at the t-th moment. m(x, z t,i ) is the point in the map corresponding to the laser point z t,i at the pose x. d 2 is the square of the distance, used to calculate the error between the laser point and the map point. w i is the weight, used to adjust the contribution of different laser points to the total error.

[0119] 6. Multi-modal scene recognition

[0120] Multi-modal data fusion:

[0121] Combine multi-modal data (such as images, lidar data, etc.) to identify and analyze the characteristics of the building environment. According to the environmental characteristics, improve the design of composite exploration points, introduce a target traction mechanism, and improve the exploration efficiency.

[0122] Path planning and obstacle avoidance optimization:

[0123] Consider the influence of multi-floor environments and dynamic obstacles, and optimize the path planning and obstacle avoidance algorithms. For example, in a multi-floor environment, a cross-floor path planning algorithm needs to be designed; in a dynamic obstacle environment, the response speed and accuracy of the obstacle avoidance algorithm need to be improved.

[0124] Extract feature points in the indoor weak texture environment through the SLAM algorithm. The construction method of the SLAM algorithm is as follows:

[0125] 1. Image preprocessing and feature extraction

[0126] 1.1 Improve the LSD algorithm to extract line features

[0127] Histogram equalization of the input image: Perform histogram equalization on the input binocular image to enhance the contrast of the image, especially the contrast in the weak texture area.

[0128] Length suppression strategy: Introduce a length threshold in the LSD algorithm to filter out too short or too long line segments, and retain line segments of moderate length as line features.

[0129] Extract line features: Apply the improved LSD algorithm to extract line features from the image.

[0130] The improved LSD algorithm has the following steps:

[0131] First, a length threshold is introduced to filter out too short or too long line segments and retain the line segments related to the building structure. For the extracted line segment L i , its length is l i , and the width of the image is W. By setting l min and l max for screening, the following effect is achieved: l min ≤ l i ≤ l max

[0132] Then, the Density-Based Spatial Clustering of Applications with Noise (DBSCAN) method with noise is used to group line segments with similar directions and adjacent in space. For the previously identified line segment L i , its neighborhood N(L i ) is defined as:

[0133] N(L i ) = {L j ∣ |θ i - θ j | < θ th and min(d end , d start ) < d th}

[0134] Where θ i and θ j are the directions of the line segments, θ th is the angle threshold, d start and d end are the starting and ending data in the adjacent line segments in space, min(d end , d start ) is the minimum value in the adjacent line segments in space, and d th is the spatial distance threshold.

[0135] Based on the definition, eligible line segments can be screened according to |N(L i )| ≥ minPts. Among them, the meaning of |N(L i )| is the number of line segments contained in the neighborhood of the line segment L i , and minPts is a variable that can be adjusted independently, indicating the minimum number of points that need to be contained in the neighborhood of a point to mark this point as a core point.

[0136] From this, the core line segments can be obtained, and then gradually extended to the line segments in their neighborhoods to form a cluster:

[0137] C = {L j ∣ L j∈ N(L i ) or L j ∈ N(L k ) and L k ∈ C}

[0138] Among them, the line segment L j is all within the neighborhood of the line segment L i or L k . The line segments not included in the cluster are marked as noise.

[0139] Due to the limitation of the angle threshold, there will still be a problem of insufficient accuracy in the initially obtained line segment clusters. Therefore, it is necessary to further perform Hough transform fitting on the line segments in each clustering cluster to generate a more accurate straight line. In the image space, it is changed to the polar coordinate system to establish the parameter space H(ρ, θ). For all the line segment endpoints on the cluster C, calculate their straight lines in the parameter space:

[0140]

[0141] Among them, δ(x) is the Dirac function, ρ represents the length, θ represents the angle, and x i , y i represent the coordinates of the line segment endpoints in the cluster C. Then, find the peak value in the parameter space H(ρ, θ):

[0142]

[0143] Among them, ρ * , θ * are the length and angle of the peak point, and argmax is to find the maximum value of the independent variable. Use the peak value to generate the fitting curve:

[0144] x cosθ * + y sinθ * = ρ *

[0145] Take the fitting curve as the line feature for subsequent processing. Among them, x and y are the variables of the fitting curve.

[0146] 1.2 Point feature extraction

[0147] Select a feature point extraction algorithm: Extract point features from the image, such as ORB, etc.

[0148] Descriptor calculation: Calculate descriptors for each feature point for subsequent feature matching.

[0149] In buildings, walls and the edges of the ground are common. The boundaries formed by the edges will intersect with each other, which can also be called corner points. The ORB algorithm is used for its recognition project. It combines FAST feature detection and BRIEF descriptor. The specific steps are as follows:

[0150] FAST detects corner points by comparing the brightness differences between a pixel and its neighboring pixels. For a pixel p in an image, if it satisfies the condition:

[0151] I(p) > I(q i ) + t or I(p) < I(q i ) - t,

[0152] then p is considered a corner point. Here, I(p) is the grayscale value of point p, and t is the threshold. To further screen corner points, the Harris corner response function is used:

[0153] R = det(M) - k·(trace(M)) 2

[0154] where det(M) = λ 1 λ 2 and trace(M) = λ 1 + λ 2 , M is the second-order matrix of the image, k is an empirical constant that needs to be adjusted according to the effect. λ 1 , λ 2 are the eigenvalues of M, and the pixel points on the image can be classified into straight lines, planes, and corner points. When both λ 1 and λ 2 are relatively large and approximately equal, it can be considered a corner point.

[0155] Then, the BRIEF descriptor is calculated for each feature point, and a binary descriptor is generated by comparing pixel pairs within the neighborhood of the feature point:

[0156]

[0157] where x j and y j are any pair of pixel points, I(x j ) and I(y j ) are the pixel sums within a certain window. If I(x j ) < I(y j ), the corresponding position of the descriptor is set to 1, otherwise it is set to 0. Finally, the binary vector descriptor is generated: d = [d 1 , d 2 , …, d n for subsequent feature matching.

[0158] 2. Feature Matching and Initial Pose Estimation

[0159] 2.1 Point Feature Matching

[0160] Feature matching: Use the descriptor to match between feature points and find the corresponding points in two frames of images.

[0161] Initial pose estimation: Based on the matched feature points, perform initial pose estimation using epipolar geometry or the PnP algorithm.

[0162] Find the pair of matched feature points by calculating the similarity between the feature descriptors of the current frame and the historical frames. Here, the cosine similarity is adopted:

[0163]

[0164] where d i · d' j represents the inner product between two pixel points, and ||d i ||||d' j || represents the 2-norm of the two pixel matrices, i.e., the modulus length. The cosine distance of the image is calculated according to the angle between two pixel points. Filter out the pair of matched feature points N match :

[0165] N match ={(d i , d′ j ) | s(d i , d′ j ) < τ}

[0166] where the threshold of τ needs to be adjusted multiple times according to the actual scenario, and the number of N match needs to be greater than the preset threshold N threshold of the number of pairs of matched points. In addition, the pairs of feature points also need to meet certain spatial distribution conditions. By calculating its fundamental matrix F, judge whether F satisfies x' j Fx i = 0 for verification, and finally obtain the valid pairs of matched points that meet the requirements.

[0167] 2.2 Line feature matching and constraint

[0168] Line feature matching: Transfer and use the principle of point feature matching, and perform line feature matching according to information such as the endpoints, lengths, and directions of line segments.

[0169] Introduce line feature constraints: Take the matched line features as additional constraints and add them to the optimization problem of pose estimation.

[0170] 3. Pose optimization by combining IMU data

[0171] IMU data preprocessing: Perform preprocessing such as denoising and calibration on the data of the Inertial Measurement Unit (IMU).

[0172] ​Tight coupling fusion: Tightly couple and fuse IMU data with visual data, and use factor graph optimization method for joint pose estimation to improve the accuracy of pose estimation.

[0173] Pose optimization: According to the fused data, iteratively optimize the pose to obtain a more accurate pose estimation.

[0174] The acceleration a obtained by the IMU k and the angular acceleration ω k as well as the 2D feature points z extracted by point-line detection k are input as part of the factor graph. Additionally, there is pose data obtained from multiple sensors:

[0175]

[0176] where p k is the position, v k is the velocity, and q k is the attitude. Integrating the IMU data gives the factor error:

[0177]

[0178] where, x k , x k+1 are two adjacent points in the 3D space point set, t is the time, g is the gravitational acceleration, and R is the rotation matrix. The visual data uses z k to describe the projection relationship between the 3D point and the 2D feature point, and find a rotation matrix R and a translation of Using the camera reprojection function π(·), the specific model is:

[0179]

[0180] The resulting visual factor error:

[0181]

[0182] where, X i is the actually observed image coordinate. In addition to the input generated by obtaining data in the forward direction, there is also the loop factor error generated by loop detection:

[0183]

[0184] Finally, minimize the above three parts of error factors, that is:

[0185]

[0186] Among them, ∥·∥ represents the Euclidean norm of a vector. Through the nonlinear least squares method or other optimization algorithms, the pose of the camera and the positions of the 3D points can be iteratively updated to minimize this error function. Specifically, the optimization algorithm calculates the gradients of the error function with respect to the pose and the positions of the 3D points, and updates the pose and the positions of the 3D points along the negative direction of the gradients until the error function converges to a local minimum. In the above framework, as long as the results can be derived into a general residual factor, various other sensors can be easily added to it.

[0187] 4. Further optimization by fusing lidar data

[0188] 4.1 Lidar data preprocessing

[0189] Point cloud data denoising: Denoise the point cloud data obtained by the lidar to remove invalid points or noise points.

[0190] Point cloud registration: Register the point cloud data of the current frame with the point cloud data of the previous frame to find the relative transformation between them.

[0191] 4.2 Fusion of lidar and visual data

[0192] External calibration: Use an automatic calibration method based on the scan matching objective function to externally calibrate the lidar and the depth camera to obtain their relative poses.

[0193] Data fusion: Fusion the pose estimation results of the lidar with the pose estimation results obtained by visual-IMU fusion. Recursive Bayesian filters or other fusion algorithms can be used to fuse these data.

[0194] Final pose estimation: Obtain the final pose estimation result based on the fused data.

[0195] 5. Loop detection and global optimization

[0196] Loop detection: Use minimizing the reprojection error for loop detection to determine whether the robot has returned to a previously visited location.

[0197] According to the matching method in the feature matching step, among the point pairs and line pairs obtained after successful matching, there may still be point lines that are incorrectly matched successfully near the threshold. The pose can be adjusted by comparison to improve the accuracy, and the minimum value of the reprojection error can be calculated:

[0198]

[0199] Among them, T t is the pose of the robot at time t, T h is the pose of the historical frame, x i ,xj are two points in three-dimensional space, X i and X j are the actually observed image coordinates.

[0200] Global optimization: If a loop closure is detected, use the global optimization algorithm to globally optimize the entire map and trajectory to eliminate the cumulative error.

[0201] 6. Map construction and update

[0202] Map construction: According to the pose estimation results and sensor data, process through the Cartographer positioning algorithm to construct a three-dimensional map or a two-dimensional grid map of the environment.

[0203] Map update: As the robot moves and new sensor data is acquired, continuously update the map information, and use the scan matching objective function to compare the laser points with the map points to achieve the update of the position information.

[0204] 7. Experimental verification and evaluation

[0205] Dataset testing: Test the performance of the algorithm on publicly available datasets or datasets collected by oneself, and evaluate the positioning accuracy and robustness of the algorithm.

[0206] Practical application verification: Apply the algorithm to actual scenarios (such as indoor navigation, etc.) to verify the actual effect of the algorithm.

[0207] Achieved effects:

[0208] 1. Improved construction accuracy and efficiency:

[0209] In terms of construction accuracy, intelligent building robots can achieve multi-modal perception and recognition of the construction site by integrating advanced sensors and algorithms, including the fusion processing of various information sources such as vision and laser. This multi-modal perception ability enables the robot to accurately identify the positions, shapes, and relationships between building elements in a complex and changing construction environment, thus avoiding the errors caused by human factors in traditional surveying methods. At the same time, the built-in high-precision surveying system of the robot can collect three-dimensional data of the construction site in real time and generate accurate point cloud models, providing reliable data support for construction. These high-precision surveying data not only help improve the quality of construction but also provide a scientific basis for subsequent construction adjustment and optimization.

[0210] In terms of construction efficiency, intelligent building robots demonstrate powerful automated operation capabilities. The robots can autonomously complete multiple construction tasks, such as wall spraying, floor grinding, and steel bar binding. This not only reduces the labor intensity of manual work but also increases the construction speed. Through multi-robot linkage and intelligent scheduling systems, the robots can efficiently cooperate on the construction site, further enhancing the construction efficiency. In addition, the robots can be flexibly adjusted according to construction requirements, such as adjusting the spraying speed and grinding intensity, to adapt to different construction scenarios. This flexibility enables the robots to maintain an efficient and stable operation state when facing complex and changing construction tasks.

[0211] 2. Reduced labor costs and labor intensity:

[0212] After integrating multi-modal scene recognition and automatic surveying and mapping functions, intelligent building robots have significantly reduced labor costs and labor intensity, which is mainly reflected in their highly automated and intelligent operation mode and efficient and accurate surveying and mapping capabilities.

[0213] By integrating multi-modal sensors and advanced algorithms, intelligent building robots can real-time sense and analyze the construction environment, accurately identify the positions, shapes, and attributes of building elements, and thus autonomously complete multiple construction tasks including wall spraying, floor grinding, and steel bar binding. This highly automated operation mode greatly reduces manual intervention and dependence on skilled workers, thus effectively reducing labor costs. At the same time, intelligent building robots have high-precision surveying and mapping capabilities, can real-time collect three-dimensional data of the construction site, and generate accurate point cloud models, providing reliable data support for construction. This high-precision surveying and mapping not only improves the construction quality but also makes the construction process more standardized and reduces rework and material waste caused by construction errors, further reducing costs.

[0214] 3. Enhanced environmental adaptability and flexibility:

[0215] After integrating multi-modal scene recognition and automatic surveying and mapping functions, the environmental adaptability and flexibility of intelligent building robots have been significantly improved. These robots can use advanced sensors and algorithms to real-time sense and analyze complex environmental information on the construction site, including various factors such as light, temperature, humidity, terrain, and landform. Based on this information, the robots can automatically adjust the operation strategy, optimize the surveying and mapping accuracy and operation efficiency, to adapt to different construction conditions and requirements.

[0216] In addition, the multi-modal scene recognition technology also enables the robot to quickly respond when facing emergencies or environmental changes, adjust the operation path or strategy, and ensure the continuity and stability of the construction process. This high level of environmental adaptability and flexibility not only improves the construction efficiency and quality, but also enables the robot to maintain an efficient and stable operation state in various complex scenarios, providing strong technical support for the intelligent transformation of the construction industry and diverse construction needs.

[0217] 4. Improved the real-time and accuracy of data:

[0218] After integrating multi-modal scene recognition and automatic mapping functions, intelligent construction robots have significantly improved the real-time and accuracy of data. By using multi-modal sensors to capture complex information on the construction site in real time and combining advanced algorithms for rapid processing and analysis, the robot can instantly generate high-precision and highly reliable data. This real-time and accurate data not only provides a scientific basis for construction decisions, but also ensures the precise control of the construction process, thus effectively improving the overall quality and efficiency of construction projects. In addition, the real-time update and accuracy of the data also promote immediate adjustment and optimization during the construction process, further enhancing the flexibility and response speed of the project.

[0219] 5. Promoted the digital transformation of the construction industry:

[0220] After integrating multi-modal scene recognition and automatic mapping functions, intelligent construction robots have significantly promoted the digital transformation of the construction industry. They can capture and process complex data on the construction site in real time, realizing digital and intelligent control of the construction process. This technological innovation not only improves the construction efficiency and precision, but also promotes the innovation of building design and construction methods, accelerating the transformation pace of the construction industry towards digitalization and intelligence, and injecting new impetus into the sustainable development of the construction industry.

[0221] 6. Achieved environment-friendly design:

[0222] First of all, through high-precision mapping and multi-modal sensing technologies, these robots can accurately capture and analyze various information of the building environment, including structure, materials, lighting, temperature, etc. This ability enables designers to have a more comprehensive understanding of the building environment, thus formulating more reasonable, energy-saving and environmentally friendly design schemes.

[0223] Secondly, intelligent building robots also demonstrate environmental friendliness during the construction process. They can precisely control the construction process and material usage, reducing waste and pollution. Through automatic surveying and positioning technologies, robots can accurately perform tasks such as material cutting, assembly, and installation, avoiding material waste and errors caused by improper manual operations in traditional construction methods. In addition, robot construction can also reduce the emissions of pollutants such as noise and dust, causing less interference and damage to the environment surrounding the construction site.

[0224] More importantly, the use of intelligent building robots also promotes the development of the construction industry towards a more environmentally friendly and sustainable direction. By improving construction efficiency and precision, robots can shorten the construction period, reduce energy consumption and carbon emissions. At the same time, robot construction can also reduce the dependence on human resources, thus alleviating the pressure on the environment and society exerted by the construction industry.

[0225] The preferred embodiments of the present invention disclosed above are only used to help illustrate the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the present invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the present invention, so that those skilled in the art can well understand and utilize the present invention. The present invention is only limited by the claims and their full scope and equivalents.

Claims

1. A robot integrating multimodal scene recognition and automatic mapping functions, characterized in that: It includes execution module, decision-making module, perception module, data processing module and speech module; The execution module is used to drive the robot to move; The decision module is used to control the work of the execution module; The sensing module is used to detect the external environment; The data processing module is used to process the data detected by the perception module; The voice module is used to receive external voice and play the processing result through voice.

2. A robot integrating multimodal scene recognition and automatic mapping functions according to claim 1, characterized in that: The execution module is a four-wheel drive motion chassis, which consists of a chassis, a power supply, a motor driver, a motor and wheels. The four end corners of the chassis are rotatably connected to wheels, a motor is fixed inside the chassis at a position corresponding to the wheels, the output end of the motor is connected to the wheels, a power supply is fixed inside the chassis, a transformer module is fixed inside the chassis and at the bottom of the battery, and a motor driver is fixed on the top of the chassis and at one end of the power supply.

3. The robot integrating multimodal scene recognition and automatic mapping functions according to claim 1, characterized in that: The decision module is a robot operating system main control, and the robot operating system main control is installed inside the chassis and located at one end of the motor driver.

4. The robot integrating multimodal scene recognition and automatic mapping functions according to claim 1, characterized in that: The perception module is a sensor, which is lower than the top of the chassis and consists of at least a lidar sensor and a depth camera.

5. A detection method for a robot integrating multimodal scene recognition and automatic mapping functions, used for a robot integrating multimodal scene recognition and automatic mapping functions as described in claims 1 to 4, characterized in that: Through the detection of the environment by the perception module, the decision module controls the execution module to drive the robot to move; When moving, the robot is guided to move to the target point and explore unknown areas through a composite exploration algorithm; Feature points are extracted in indoor weak-texture environments through simultaneous positioning and mapping algorithms.

6. The method for detecting a robot integrating multimodal scene recognition and automatic mapping functions according to claim 5, characterized in that: The composite exploration algorithm is constructed as follows: Obtain all-round environmental information through sensors; Use deep learning technology to perform image recognition and semantic segmentation to extract key information in the environment; Fuse lidar data with image data to build high-precision 3D maps and semantic maps; According to the location of the target point, the map is divided into half planes pointing to the target point to limit the exploration area; In the segmented area, all boundary points are found through the depth-first search algorithm; Construct boundary lines to form the boundary outline of the exploration area; The boundary points are scored by the evaluation function, and the boundary point with the highest score is selected as the best boundary point; Perform vector synthesis of the best boundary point and the target point to generate a composite exploration point; Through the D* global path planning algorithm, the optimal path from the robot's current position to the composite exploration point is planned according to the composite exploration point and the current map information; During the robot's movement, the dynamic window algorithm is used to adjust the robot's motion trajectory in real time; Through the Cartographer positioning algorithm, the robot's location information is updated in real time, and the map is updated in real time based on the robot's location.

7. The method for detecting a robot integrating multimodal scene recognition and automatic mapping functions according to claim 5, characterized in that: The SLAM algorithm is constructed as follows: Through the improved line segment detector algorithm, line features are extracted from the image; Extract point features from images using directional fast and rotational binary robust independent basic feature algorithms; Through point feature matching, the corresponding points in the two frames are found, and the preliminary pose estimation is performed through the perspective n-point method algorithm; Through line feature matching, the matched line features are used as additional constraints to iteratively optimize the pose; Loop detection is performed by minimizing the reprojection error to determine whether the robot has returned to the place it has passed before; Build a 3D map or 2D grid map of the environment based on the pose estimation results and sensor data; Through the known CAD building feature data and laser ranging data, the constructed three-dimensional map or two-dimensional grid map is cross-checked and corrected to obtain the corrected surveying and mapping data and map.

Citation Information

Cited By

  • Alternating current adapter intelligent ordered control algorithm and device with real-time state sensing function

    CN120628217A

  • Construction process inspection track generation method and device based on SLAM (Simultaneous Localization and Mapping)

    CN120970630A