Method, apparatus, device and medium for constructing three-dimensional semantic map in indoor scenario
By integrating semantic labels from Deeplabv3+ into SLAM systems, the method creates robust pixel-level 3D semantic maps for indoor environments, enabling improved navigation and interaction capabilities for robots.
Patent Information
- Application Number
- CN202210316142.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-28
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-03-28
AI Technical Summary
Existing indoor mobile service robots lack object semantic information in an unstructured environment and cannot meet the needs of semantic navigation, interaction and crawling. Most of the existing maps are pure geometric structure information.
The three-dimensional semantic map construction algorithm based on visual SLAM and deep learning semantic segmentation is adopted. Through pixel coordinate consistency data association, the semantic labels output by the Deeplabv3+ semantic segmentation algorithm are fused into the three-dimensional map of the sparse direct method visual odometer (DSO), and the pixel-level three-dimensional semantic map construction is realized.
The construction of pixel-level three-dimensional semantic maps has been realized, which improves the intelligence level of the robot in the indoor environment, and provides the help of semantic-based navigation, interaction and grabbing functions.
Smart Images

Figure CN114782530B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of visual navigation and image processing, and in particular to a method, device, equipment and medium for constructing a three-dimensional semantic map in an indoor scene. Background Art
[0002] With the continuous development of robot technology, the demand for indoor mobile service robots shows an upward trend. However, the prerequisite for indoor mobile service robots to enter home applications on a large scale is to have intelligent environmental perception and understanding capabilities. One of the key technologies is that the robot can have the ability to build a semantic map. Currently, the maps relied on by robots for positioning and navigation in unstructured environments are mainly grid maps, topological maps, etc., mostly pure geometric structure information, lacking object semantic information in the environment, and unable to meet the future scenario requirements of indoor mobile service robots. Therefore, the semantic map, as the core technology of intelligent three-dimensional perception, has recently received extensive attention. Summary of the Invention
[0003] In order to solve the above technical problems, the present invention proposes a method, device, equipment and medium for constructing a three-dimensional semantic map in an indoor scene. Based on a three-dimensional semantic map construction algorithm of visual SLAM and deep learning semantic segmentation, the semantic labels output by the Deeplabv3+ semantic segmentation algorithm are fused into the three-dimensional map constructed by a visual SLAM system based on a sparse direct method visual odometer (DSO) through a data association method of pixel coordinate consistency, realizing the construction of a pixel-level three-dimensional semantic map.
[0004] In order to achieve the above object, the technical solution of the present invention is as follows:
[0005] A method for constructing a three-dimensional semantic map in an indoor scene, comprising the following steps:
[0006] Obtain an indoor scene map;
[0007] Input the indoor scene map into a visual SLAM system to sense the indoor environment and extract a three-dimensional point cloud map; at the same time, input the indoor scene map into a preset semantic segmentation model to predict the semantic label of each pixel point and obtain a semantic segmentation label map;
[0008] Based on the corresponding relationship of each pixel point in the point cloud map and the semantic segmentation label map, extract the semantic information of the pixel points from the semantic segmentation label map and synchronously map it to the three-dimensional point cloud map to obtain a pixel-level three-dimensional semantic map.
[0009] Preferably, inputting the indoor scene map into a visual SLAM system to sense the indoor environment and extract a three-dimensional point cloud map specifically includes the following steps:
[0010] Run the DSO algorithm to obtain the camera pose and the depth value of the pixel points;
[0011] Based on the obtained pixel depth values and the camera internal parameters, obtain the position of the pixel in the camera coordinate system with the camera as the reference origin;
[0012] According to the camera pose, calculate the position of the pixel in the standard coordinate system;
[0013] Calculate the position of each pixel point in the standard coordinate system, and establish a three-dimensional point cloud map of the indoor scene.
[0014] Preferably, the construction process of the preset semantic segmentation model:
[0015] Select common objects in the indoor scene from the public dataset, extract them to form a new dataset, and preprocess the dataset to divide the data into a training sample set and a test sample set;
[0016] Input the training sample set into the DeepLabv3+ network model for model training to obtain a preliminary model;
[0017] Input the test sample set into the preliminary model for testing, adjust the original hyperparameters according to the test results until the error of the prediction result of the preliminary model meets the preset threshold, and output the current model as the semantic segmentation model.
[0018] Preferably, use mIoU as the evaluation index to evaluate the performance of the prediction result.
[0019] Preferably, the public dataset includes ADE20K, COCO, and Pascal.
[0020] Preferably, the common objects in the indoor scene include desks, doors, people, vases, bookcases, floors, monitors, armchairs, boxes, walls, table lamps, chairs, whiteboards, curtains, glass, hanging paintings, clocks, tables, sofas, and plants.
[0021] Preferably, it further includes the following steps:
[0022] Locate the boundary of the object through the contour detection method, and learn to predict the distance and direction from the boundary to the inside of the object;
[0023] Replace the semantic label of the pixel point at the object boundary with the semantic label of the pixel point inside the object.
[0024] A three-dimensional semantic map construction device for an indoor scene, including: an acquisition module, a first extraction module, a second extraction module, and a composition module, where,
[0025] The acquisition module is used to acquire an indoor scene map;
[0026] The first extraction module is configured to receive an indoor scene map and extract a three-dimensional point cloud map based on a visual SLAM system using sparse direct method visual odometry;
[0027] The second extraction module is configured to receive an indoor scene map and predict the semantic label of each pixel point based on a preset semantic segmentation model to obtain a semantic segmentation label map;
[0028] The mapping module is configured to extract the semantic information of the pixel points from the semantic segmentation label map based on the corresponding relationship between the pixel points in the point cloud map and the semantic segmentation label map, and synchronously map it to the three-dimensional point cloud map to obtain a pixel-level three-dimensional semantic map.
[0029] A computer device includes: a memory for storing a computer program; a processor for implementing the three-dimensional semantic map construction method in an indoor scene as described in any one of the above when executing the computer program.
[0030] A readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the three-dimensional semantic map construction method in an indoor scene as described in any one of the above.
[0031] Based on the above technical solutions, the beneficial effects of the present invention are: The present invention studies a three-dimensional semantic map construction algorithm based on visual SLAM and deep learning semantic segmentation for an actual indoor environment. By using a data association method based on pixel coordinate consistency, the semantic labels output by the Deeplabv3+ semantic segmentation algorithm are fused into the three-dimensional map constructed by a visual SLAM system based on sparse direct method visual odometry (DSO), realizing the construction of a pixel-level three-dimensional semantic map. This algorithm has good robustness and will provide assistance for mobile robots to realize functions such as semantic-based navigation, interaction, and grasping, effectively improving their intelligent level. Description of the Drawings
[0032] Figure 1 is a schematic flowchart of a three-dimensional semantic map construction method in an indoor scene in an embodiment;
[0033] Figure 2 is a schematic diagram of the principle of a three-dimensional semantic map construction method in an indoor scene in an embodiment;
[0034] Figure 3 is a schematic diagram of the principle of a semantic segmentation boundary optimization method in an embodiment;
[0035] Figure 4 is a comparison diagram of semantic segmentation effects in an embodiment, where a is an indoor scene map; b is the segmentation effect diagram of the Deeplabv3+ algorithm; c is the segmentation effect diagram of the Deeplabv3+ optimized algorithm;
[0036] Figure 5 It is a comparison diagram of depth information, point cloud information, and semantic segmentation information after processing the same frame of image in an embodiment;
[0037] Figure 6 It is a comparison diagram of the robot's running trajectory before and after optimization in an embodiment;
[0038] Figure 7 It is a schematic structural diagram of a three-dimensional semantic map construction device in an indoor scene in an embodiment;
[0039] Figure 8 It is a structural block diagram of a computer device in an embodiment. Specific embodiments
[0040] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention.
[0041] As Figure 1 shown, Figure 1 It is a schematic flow chart of a three-dimensional semantic map construction method in an indoor scene. An embodiment of the present application provides a three-dimensional semantic map construction method in an indoor scene, which is applied to a mobile robot and specifically includes the following steps:
[0042] Step S1, obtain an indoor scene map;
[0043] Step S2, input the indoor scene map into a visual SLAM system to sense the indoor environment and extract a three-dimensional point cloud map; at the same time, input the indoor scene map into a preset semantic segmentation model to predict the semantic label of each pixel point and obtain a semantic segmentation label map;
[0044] Step S3, based on the corresponding relationship of each pixel point in the point cloud map and the semantic segmentation label map, extract the semantic information of the pixel points from the semantic segmentation label map and synchronously map it to the three-dimensional point cloud map to obtain a pixel-level three-dimensional semantic map.
[0045] For robots to achieve positioning and navigation in unstructured environments, the maps they rely on are mainly grid maps, topological maps, etc., mostly pure geometric structure information, lacking object semantic information in the environment, and unable to meet the requirements of future scenarios such as semantic navigation, interaction, and grasping for indoor mobile service robots. For the actual indoor environment, a three-dimensional semantic map construction algorithm based on visual SLAM and deep learning semantic segmentation is studied. By using a data association method based on pixel coordinate consistency, the semantic labels output by the Deeplabv3+ semantic segmentation algorithm are fused into the three-dimensional map constructed by a visual SLAM system based on a sparse direct method visual odometry (DSO), realizing the construction of a pixel-level three-dimensional semantic map. This algorithm has good robustness and will provide assistance for mobile service robots to achieve functions such as semantic-based navigation, interaction, and grasping, effectively improving their intelligence level.
[0046] As Figure 2 shown, the specific principle of establishing a three-dimensional semantic map in this embodiment: An indoor scene image (RGB format) is simultaneously input into a visual SLAM system and a preset Deeplabv3+ semantic segmentation model. The semantic segmentation model will predict the semantic labels of each pixel point in the image. The semantic labels are also a 2D image in form, namely the semantic segmentation label map, which has the same resolution as the input indoor scene image, and there is a one-to-one correspondence between each pixel point. Therefore, when a pixel point p(u, v) is converted into a three-dimensional point P c through the depth value d and the camera internal parameter K, and transformed into a three-dimensional map point P w through the camera pose transformation T, according to the coordinate consistency principle, the semantic information of this pixel point is extracted from the corresponding semantic segmentation label map, and then used as the semantic attribute value of this three-dimensional point to be synchronously mapped into the three-dimensional point cloud map to construct a pixel-level three-dimensional semantic map.
[0047] In the method for constructing a three-dimensional semantic map in the indoor scene of an embodiment, a process of inputting the indoor scene image into the visual SLAM system to perceive the indoor environment and extract the three-dimensional point cloud map is also provided, which specifically includes the following steps:
[0048] Run the DSO algorithm to obtain the camera pose T and the pixel point depth value d;
[0049] According to the obtained pixel point depth value d and the camera internal parameter K, obtain the position P c of a pixel point p(u, v) in the camera coordinate system with the camera as the reference origin;
[0050] According to the camera pose T, calculate the position P w (X W , Y W , ZW ) The formula is as follows:
[0051]
[0052] In the formula, the internal parameter R is the rotation matrix and t is the translation vector;
[0053] Calculate the position of each pixel point in the standard coordinate system and establish a three-dimensional point cloud map of the indoor scene.
[0054] In the method for constructing a three-dimensional semantic map in the indoor scene of an embodiment, a construction process of a semantic segmentation model is further provided, which specifically includes the following steps:
[0055] Select common objects in the indoor scene from the public dataset, extract them to form a new dataset, and preprocess the dataset to divide the data into a training sample set and a test sample set;
[0056] Input the training sample set into the DeepLabv3+ network model for model training to obtain a preliminary model;
[0057] Input the test sample set into the preliminary model for testing, adjust the original hyperparameters according to the test results until the error of the prediction result of the preliminary model meets the preset threshold, and output the current model as the semantic segmentation model.
[0058] In this embodiment, in order to improve the semantic mapping accuracy and quality in the indoor environment, the Deeplabv3+ network model is optimized. Specifically, 20 types of common objects in the indoor scene are selected from three public datasets of ADE20K, COCO, and Pascal VOC. The 20 types of objects include desks, doors, people, vases, bookcases, floors, monitors, armchairs, boxes, walls, table lamps, chairs, whiteboards, curtains, glass, hanging paintings, clocks, tables, sofas, and plants. Extract the 20 types of objects to form a new dataset, with a total of 18,000 pictures, of which 15,000 are used for training and 3,000 are used for testing.
[0059] When training the DeepLabv3+ network model, it is necessary to set the hyperparameters for model training. Considering the characteristics of the Deeplabv3+ algorithm and the characteristics of the new dataset, adjustments are made based on the original hyperparameters given in the Deeplabv3+ algorithm paper. The final obtained hyperparameters are shown in Table 1:
[0060] Table 1 Hyperparameter Configuration
[0061]
[0062] The training and testing of the DeepLabv3+ network model were carried out on the Ubuntu18.04 system with an Intel E5-2678 processor, and a total of 160,000 iterations were trained. After training, it was tested on the test set, and the mIoU was used as the evaluation index to evaluate the performance of the prediction results.
[0063] In the method for constructing a three-dimensional semantic map in the indoor scene of an embodiment, a semantic segmentation boundary optimization process is also provided, which specifically includes the following steps:
[0064] Locate the boundary of the object through the contour detection method, and learn to predict the distance and direction from the boundary to the inside of the object;
[0065] Replace the semantic label of the pixel at the object boundary with the semantic label of the pixel inside the object.
[0066] In this embodiment, aiming at the problem that the DSO direct method visual SLAM is more sensitive to the pixels at the object boundary position and the current semantic segmentation algorithms, including Deeplabv3+, usually have insufficiently fine semantic segmentation at the object boundary. Through theoretical analysis and combined with actual observations, it is found that in semantic segmentation, for an object, the semantic segmentation result of its internal image is usually accurate, but the closer to the boundary, the less accurate. To solve this problem, a model-independent semantic segmentation boundary optimization method (Boundary Refinement) is proposed on the basis of Deeplabv3+, and its principle is as Figure 3 shown. First, locate the boundary of the object through the contour detection method, and learn to predict the distance and direction from the boundary to the inside of the object, and then replace the semantic label at the boundary with the semantic label of the internal pixel, so as to reduce the segmentation error at the boundary and improve the segmentation quality.
[0067] To verify the effect of the semantic segmentation boundary optimization algorithm, actual indoor office scene pictures were also selected for comparative testing at the same time, and the effects are compared as Figure 4 shown. The segmentation object is a stool as an example. It can be found in the comparison diagram that the optimized algorithm has a more refined boundary. The benchmark version of the algorithm and the version proposed in this paper using the semantic segmentation boundary optimization algorithm were respectively tested and visually compared and analyzed. The results are shown in Table 2: After using semantic boundary optimization (Boundary Refinement), the accuracy of the algorithm is improved by 1.2 percentage points, and there is basically no impact on the model parameters and running time.
[0068] Table 2 Semantic segmentation test results
[0069]
[0070] To verify the effectiveness of the semantic mapping algorithm in the indoor scenario proposed in this paper, the robot was moved around the meeting room for one week, and pictures of the entire experimental environment were collected in real time, with approximately 2,000 pictures collected in total. The collected indoor scene pictures were imported into the algorithm. While obtaining the semantic information of the pictures, the point cloud position information of the pictures could also be obtained. The above information was finally output as a semi-dense three-dimensional semantic map with semantic information on the three-dimensional map through the semantic data association algorithm. Two evaluation metrics are proposed in this paper. On the one hand, the number of recognized categories in the semantic map reflects the richness of the reference object information and can also indirectly evaluate the effectiveness of the algorithm. As shown in Table 3, the recognition rate of the categories of the algorithm in this paper can reach 100%, and the recognition rate of the instances can reach nearly 75%, and rich instance information in the environment can be extracted.
[0071] Table 3 Recognition Effect of Semantic Map
[0072]
[0073] On the other hand, whether the semantic map can form a closed-loop route consistent with the actual walking trajectory reflects whether the algorithm can perform point cloud matching and correction. As Figure 5 shows the depth information, point cloud information, and semantic segmentation information during the processing of each frame of the picture. The point cloud map and the semantic segmentation map can correspond regularly to the objects; Figure 6 shows the trajectory before closed-loop optimization and the running trajectory of the robot after closed-loop optimization. This experiment shows that the semantic map construction framework and optimization method proposed in this paper can have good point cloud segmentation and semantic recognition in the indoor environment, and can reconstruct the three-dimensional semantic map of the indoor environment and automatically generate the running path of the robot.
[0074] The embodiment of the present application also provides a three-dimensional semantic map construction device in the indoor scenario. The specific implementation manner is the same as the implementation manner and the achieved technical effects recorded in the embodiment of the three-dimensional semantic map construction method in the indoor scenario above, and some contents will not be elaborated.
[0075] As Figure 7 shown, a three-dimensional semantic map construction device 100 in the indoor scenario is provided. The device includes: an acquisition module 110, a first extraction module 120, a second extraction module 130, and a composition module 140, where,
[0076] The acquisition module 110 is configured to acquire an indoor scene map;
[0077] The first extraction module 120 is configured to receive the indoor scene map and extract a three-dimensional point cloud map based on a visual SLAM system of a sparse direct method visual odometer;
[0078] The second extraction module 130 is configured to receive an indoor scene map, predict the semantic label of each pixel point based on a preset semantic segmentation model, and obtain a semantic segmentation label map.
[0079] The composition module 140 is configured to extract the semantic information of the pixel points from the semantic segmentation label map based on the corresponding relationship between the pixel points in the point cloud map and the semantic segmentation label map, and synchronously map it to the three-dimensional point cloud map to obtain a pixel-level three-dimensional semantic map.
[0080] The devices and modules illustrated in the above embodiments may be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.
[0081] As Figure 8 shown, an embodiment of the present application further provides a computer device 200, which includes at least one memory 210, at least one processor 220, and a bus 230 connecting different platform systems. Among them,
[0082] The memory 210 may include a readable medium in the form of volatile memory, such as a random access memory (RAM) 211 and / or a cache memory 212, and may further include a read-only memory (ROM) 213.
[0083] Among them, the memory 210 further stores a computer program, and the computer program can be executed by the processor 220, so that the processor 220 executes the steps of the three-dimensional semantic map construction method in the indoor scene of the embodiment of the present application. The specific implementation manner is consistent with the implementation manner and the achieved technical effects described in the embodiments of the three-dimensional semantic map construction method in the indoor scene above, and some contents will not be repeated.
[0084] The memory 210 may further include a utility 214 having at least one program module 215. Such program modules 215 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each of these examples or some combination thereof may include the implementation of a network environment.
[0085] Correspondingly, the processor 220 can execute the above computer program and can also execute the utility 214.
[0086] The bus 230 can represent one or more of several types of bus structures, including a memory bus or a memory controller, a peripheral bus, an accelerated graphics port, a processor, or a local bus using any of the various bus structures.
[0087] The computer device 200 can also communicate with one or more external devices 240 such as a keyboard, a pointing device, a Bluetooth device, etc., and can also communicate with one or more devices capable of interacting with the computer device 200, and / or communicate with any device (such as a router, a modem, etc.) that enables the computer device 200 to communicate with one or more other computing devices. Such communication can be carried out through the input / output interface 250. Moreover, the computer device 200 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 260. The network adapter 260 can communicate with other modules of the computer device 200 through the bus 230. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in conjunction with the computer device 200, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage platforms, etc.
[0088] The embodiment of the present application also provides a readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0089] Obtain an indoor scene map;
[0090] Input the indoor scene map into a visual SLAM system to sense the indoor environment and extract a three-dimensional point cloud map; at the same time, input the indoor scene map into a preset semantic segmentation model to predict the semantic label of each pixel point and obtain a semantic segmentation label map;
[0091] Based on the correspondence relationship of each pixel point between the point cloud map and the semantic segmentation label map, extract the semantic information of the pixel points from the semantic segmentation label map and synchronously map it to the three-dimensional point cloud map to obtain a pixel-level three-dimensional semantic map.
[0092] A computer-readable medium includes permanent and non-permanent, removable and non-removable media that can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media such as modulated data signals and carrier waves.
[0093] The above are only the preferred embodiments of the embodiments of the present application and are not used to limit the embodiments of the present application. For those skilled in the art, various changes and modifications can be made to the embodiments of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the embodiments of the present application shall be included in the protection scope of the embodiments of the present application.
Claims
1. A method for constructing a three-dimensional semantic map in an indoor scene, characterized in that It includes the following steps: Obtain an indoor scene map; Input the indoor scene map into a visual SLAM system to perceive the indoor environment and extract a three-dimensional point cloud map; at the same time, input the indoor scene map into a preset semantic segmentation model to predict the semantic label of each pixel point and obtain a semantic segmentation label map. It also includes the following steps: Locate the boundary of the object through a contour detection method, and learn to predict the distance and direction from the boundary to the inside of the object; Replace the semantic label of the pixel point at the object boundary with the semantic label of the pixel point inside the object. Based on the correspondence between each pixel point in the point cloud map and the semantic segmentation label map, extract the semantic information of the pixel points from the semantic segmentation label map and synchronously map it to the three-dimensional point cloud map to obtain a pixel-level three-dimensional semantic map. Among them, the construction process of the preset semantic segmentation model: Select common objects in the indoor scene from the public dataset, extract them to form a new dataset, and preprocess the dataset to divide the data into a training sample set and a test sample set. Input the training sample set into the DeepLabv3+ network model for model training to obtain a preliminary model; Input the test sample set into the preliminary model for testing, and adjust the original hyperparameters according to the test results until the error of the prediction result of the preliminary model meets the preset threshold, and output the current model as the semantic segmentation model.
2. The method for constructing a three-dimensional semantic map in an indoor scene according to claim 1, wherein Input the indoor scene map into a visual SLAM system to perceive the indoor environment and extract a three-dimensional point cloud map, specifically including the following steps: Run the DSO algorithm to obtain the camera pose and pixel point depth value; Based on the obtained pixel point depth value and camera internal parameters, obtain the position of the pixel in the camera coordinate system with the camera as the reference origin; According to the camera pose, calculate the position of the pixel in the standard coordinate system; Calculate the position of each pixel point in the standard coordinate system and establish a three-dimensional point cloud map of the indoor scene.
3. The method for constructing a three-dimensional semantic map in an indoor scene according to claim 1, characterized in that, Use mIoU as an evaluation index to evaluate the performance of the prediction result.
4. The method for constructing a three-dimensional semantic map in an indoor scene according to claim 1, characterized in that The public dataset includes ADE20K, COCO, and Pascal.
5. The method for constructing a three-dimensional semantic map in an indoor scene according to claim 1, wherein The common objects in the indoor scene include desks, doors, people, vases, bookcases, floors, monitors, boxes, walls, table lamps, chairs, whiteboards, curtains, glass, hanging paintings, clocks, tables, sofas, and plants.
6. A three-dimensional semantic map construction device for indoor scenarios, characterized in that It includes: An acquisition module, a first extraction module, a second extraction module, and a composition module, where The acquisition module is used to obtain an indoor scene map; The first extraction module is used to receive the indoor scene map and extract a three-dimensional point cloud map based on the visual SLAM system of the sparse direct method visual odometer. The second extraction module is used to receive an indoor scene map, predict the semantic label of each pixel point based on a preset semantic segmentation model, and obtain a semantic segmentation label map. It further includes the following steps: locating the boundary of an object through a contour detection method, and learning to predict the distance and direction from the boundary to the interior of the object; replacing the semantic label of the pixel points at the object boundary with the semantic label of the pixel points inside the object. Among them, the construction process of the preset semantic segmentation model is as follows: selecting common objects in the indoor scene from a public dataset, extracting them to form a new dataset, and preprocessing the dataset to divide the data into a training sample set and a test sample set; inputting the training sample set into the DeepLabv3+ network model for model training to obtain a preliminary model; inputting the test sample set into the preliminary model for testing, adjusting the original hyperparameters according to the test results until the error of the prediction result of the preliminary model meets the preset threshold, and outputting the current model as the semantic segmentation model. The composition module is used to extract the semantic information of the pixel points from the semantic segmentation label map based on the corresponding relationship between the pixel points in the point cloud map and the semantic segmentation label map, and synchronously map it to the three-dimensional point cloud map to obtain a pixel-level three-dimensional semantic map.
7. A computer device, characterized in that, It includes: A memory for storing computer programs; A processor for implementing the three-dimensional semantic map construction method in the indoor scene according to any one of claims 1 to 5 when executing the computer program.
8. A readable storage medium, characterized in that, A computer program is stored on the readable storage medium, and when the computer program is executed by the processor, it implements the three-dimensional semantic map construction method in the indoor scene according to any one of claims 1 to 5.
Citation Information
Patent Citations
Indoor three-dimensional semantic map construction method
CN111340939A
Simultaneous localization and mapping system using illumination invariant image, and method for mapping pointcloud thereof
KR101988555B1