Multi-modal Wall Tile Recognition Method, System and Device
By combining three-dimensional point clouds and two-dimensional image data, using deformable convolution networks and improved anchor frame generation algorithms, the morphological similarity and environmental interference problems of wall tiles recognition in the prior art are solved, and high-precision and robust wall tiles recognition are achieved.
Patent Information
- Application Number
- CN202210944935.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-08
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-08-08
AI Technical Summary
The existing wall tiles identification methods are difficult to effectively distinguish objects with similar shapes, and environmental factors such as lighting are prone to misdetection and misdetection. The existing technology lacks semantic information and two-dimensional image data with two-dimensional image data with texture information are insufficient complementarity when combined.
The multimodal recognition method is adopted, combining three-dimensional point cloud data and two-dimensional image data, and RGB images are generated through semantic segmentation, rotation, rasterization, intensity map and depth map, and identification is used using a deformable convolutional network and an improved anchor box generation algorithm to adapt to wall brick targets of different shapes and scales.
It improves the accuracy and robustness of wall tiles identification, reduces noise interference, and can obtain high-quality identification results in complex environments, adapt to wall tiles of different shapes and scales, reducing calculation overhead.
Smart Images

Figure CN115457537B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of three-dimensional reconstruction and computer graphics, and particularly relates to a method and system for identifying wall tiles in an indoor scene. Background Art
[0002] How to perform three-dimensional description of the real world and realize the digitization of the scene is currently in increasing demand and application. For example: with the rapid development of sensor technology and the emergence of a large number of software processing tools, three-dimensional point cloud data has been popularized. Three-dimensional point cloud data realizes the digitization of the scene through a large number of three-dimensional point sets, and can densely and accurately represent the three-dimensional geometric shape of objects in the environment.
[0003] In the fields of building construction, interior decoration design, and three-dimensional scene reconstruction, the identification of wall bricks is a key step in improving the automation level of indoor modeling. Whoever completes this key link first will have a stronger competitive advantage. Existing brick identification methods are very difficult to effectively distinguish wall tiles, which are objects with very similar three-dimensional point cloud distribution patterns. At the same time, the method of only using two-dimensional image recognition is easily affected by environmental factors such as lighting, resulting in false detection and missed detection. This is mainly because although point cloud data contains the spatial geometric information of the target, it lacks semantic information; while image data, although lacking depth information, provides rich texture information. Therefore, point cloud data and image data have a certain information complementarity, and comprehensively using the two types of data helps to improve the performance of wall tile identification in complex scenes. Summary of the Invention
[0004] The present invention proposes a multi-modal method, system, and device for identifying wall tiles, which comprehensively utilizes the advantages of two-dimensional and three-dimensional data, can effectively overcome the noise on the wall surface, and obtain accurate identification results. At the same time, it tries to avoid the disadvantages of the two types of data as much as possible, reduce noise, and improve the robustness of the algorithm.
[0005] According to the first aspect of the embodiments of the present invention, a multi-modal method for identifying wall tiles is provided, including:
[0006] Obtain indoor three-dimensional point cloud data;
[0007] Perform semantic segmentation on the three-dimensional point cloud data to extract wall point clouds perpendicular to the horizontal plane;
[0008] Rotate the wall point clouds to the horizontal plane to obtain the wall point clouds under a top view, and record the rotation matrix;
[0009] Perform rasterization processing on the horizontal wall point clouds, calculate the intensity values in a Gaussian weighted manner, and generate an intensity map;
[0010] Normalize the height values of the horizontal wall point cloud, and use the normalized height values as the grayscale values of the depth map to generate a depth map;
[0011] Fuse the intensity map and the depth map into an RGB image according to the RGB channels;
[0012] Use a wall tile recognition model to recognize the RGB image. The feature extraction network of the wall tile recognition model includes a deformable convolutional network that enables the convolutional kernel to adapt to wall tile regions of different shapes. After the RGB image is input into the feature extraction network, a feature map is obtained. Use RPN to generate candidate boxes. The feature map and the candidate boxes are subjected to region of interest pooling and classification to obtain an image recognition result. When generating the candidate boxes, generate corresponding aspect ratios of anchor boxes according to the aspect ratio of the wall tile target so that anchor boxes of different sizes and scaling ratios can adapt to wall tile targets of different scales;
[0013] Generate a three-dimensional detection result of the wall tile region according to the recognition result of the wall tile recognition model and the rotation matrix.
[0014] According to the second aspect of the embodiments of the present invention, a multi-modal wall tile recognition system is provided, including:
[0015] A data acquisition module configured to acquire indoor three-dimensional point cloud data;
[0016] A data preprocessing module configured to perform semantic segmentation on the three-dimensional point cloud data, extract wall point clouds perpendicular to the horizontal plane, rotate the wall point clouds to the horizontal plane to obtain wall point clouds under a top view, record the rotation matrix, perform rasterization processing on the horizontal wall point clouds, calculate intensity values in a Gaussian weighted manner to generate an intensity map, normalize the height values of the horizontal wall point clouds, use the normalized height values as the grayscale values of the depth map to generate a depth map, and fuse the intensity map and the depth map into an RGB image according to the RGB channels;
[0017] A wall tile recognition model configured to recognize the RGB image. The feature extraction network of the wall tile recognition model includes a deformable convolutional network that enables the convolutional kernel to adapt to wall tile regions of different shapes. After the RGB image is input into the feature extraction network, a feature map is obtained. Use RPN to generate candidate boxes. The feature map and the candidate boxes are subjected to region of interest pooling and classification to obtain an image recognition result. When generating the candidate boxes, generate corresponding aspect ratios of anchor boxes according to the aspect ratio of the wall tile target so that anchor boxes of different sizes and scaling ratios can adapt to wall tile targets of different scales;
[0018] A three-dimensional detection result generation module for wall tile regions configured to generate a three-dimensional detection result of the wall tile region according to the recognition result of the wall tile recognition model and the rotation matrix.
[0019] According to a third aspect of an embodiment of the present invention, there is provided a multi-modal wall tile recognition device, including: a lidar for collecting indoor three-dimensional point cloud data; a processor; and a memory for storing instructions executable by the processor; wherein, when the processor is configured to execute the instructions in the memory, all or part of the steps of the method are implemented.
[0020] According to a fourth aspect of an embodiment of the present invention, there is provided a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, all or part of the steps of the method are implemented.
[0021] Existing recognition methods are very difficult to effectively distinguish some objects with very similar point cloud distribution patterns. And interference from environmental factors such as light is likely to cause the detection model to be difficult to accurately locate and identify the target. However, the present invention can effectively solve these problems. Even when processing point cloud data containing interference such as light and data loss, a very high recognition quality can still be obtained. The present invention is a multi-modal wall tile recognition method, which uses a deformable convolutional network to replace the traditional single regular convolution, so that the convolution kernel can adapt to different shapes and scales of wall tiles, and as much as possible extract features more relevant to the real target. At the same time, the RPN is optimized and the anchor box generation algorithm is improved, so that the generated anchor boxes are more in line with the scale of wall tiles, solving the problem that targets with various shapes such as long and narrow wall tiles and small wall tiles are likely to be missed, and improving the accuracy of the original anchor box positioning. The present invention has good robustness, small computational overhead and high recognition accuracy. This has very important significance for work such as building construction, interior decoration design and three-dimensional scene reconstruction. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings of the embodiments will be briefly introduced below.
[0023] Figure 1 It is a flowchart of a multi-modal wall tile recognition method provided by an embodiment of the present invention.
[0024] Figure 2 It is a schematic diagram of extracting wall point clouds provided by an embodiment of the present invention.
[0025] Figure 3 It is an intensity map, a depth map and an image after channel fusion provided by an embodiment of the present invention.
[0026] Figure 4 It is an improved feature extraction network diagram provided by an embodiment of the present invention.
[0027] Figure 5 It is a schematic diagram of the structure of a wall tile recognition model provided by an embodiment of the present invention.
[0028] Figure 6 The recognition result of the wall tile recognition model provided by an embodiment of the present invention.
[0029] Figure 7 Schematic diagram of the three-dimensional detection result provided by an embodiment of the present invention. Detailed implementation manners
[0030] Figure 1 A flowchart of a multi-modal wall tile recognition method is shown. The method shown below will be described in detail. Figure 1 The method shown will be described in detail.
[0031] Step 1: Obtain indoor three-dimensional point cloud data.
[0032] Step 2: Preprocess the three-dimensional point cloud data to obtain a preprocessing result that fuses the intensity map and the depth map, and record the rotation matrix from the three-dimensional point cloud to the two-dimensional image.
[0033] The detailed method for preprocessing the three-dimensional point cloud data is as follows:
[0034] Step 2.1, as Figure 2 shown, since the recognition object is a wall tile, semantic segmentation is performed on the three-dimensional point cloud of the roughcast house to extract the wall point cloud.
[0035] Step 2.2, the extracted wall point cloud is perpendicular to the horizontal plane. Rotate the wall point cloud to the xOy plane to be parallel to the horizontal plane to obtain the wall point cloud under top view, and record the rotation matrix M. In order to convert the image detection result in the subsequent steps back to the three-dimensional space, the rotation matrix is recorded.
[0036] Step 2.3, perform rasterization processing on the horizontal wall point cloud, calculate the intensity value in a Gaussian weighted manner, and generate an intensity map.
[0037] Step 2.4, normalize the Z value (height) of the horizontal wall point cloud, and use the normalized Z value as the gray value of the depth map to generate a depth map.
[0038] Step 2.5, as Figure 3 shown, fuse the intensity map and the depth map into an RGB image according to the RGB channels. Among them, the depth map is used as the R channel, the intensity map is used as the G channel, and the gray value of the B channel can be set to be the same as that of the R channel.
[0039] Step 3: Training of the wall tile recognition model.
[0040] Step 3.1, use the fused RGB image as a training sample, mark the wall tile area according to the format of the VOC2012 dataset to obtain the marked training sample.
[0041] Step 3.2: Use the training samples to train the constructed network model to obtain a wall tile recognition model. Due to the influence of doors and windows on the wall tile area, a deformable convolutional network is introduced to adapt to wall tile areas of different shapes. The improved feature extraction network is as shown in Figure 4 . Replace C3, C4, and C5 in the feature extraction network with deformable convolutional networks, so that the convolutional kernels can adapt to different morphologies and scales of wall tiles, and extract features that are more relevant to the real targets as much as possible.
[0042] Step 3.3: Input the image into the feature extraction network in Step 3.2 to obtain a feature map, and then use RPN to generate candidate boxes. In this step, optimize the way of generating anchor boxes according to the sizes of wall tile targets in the dataset. The optimization principle is to generate corresponding aspect ratios of anchor box aspect ratios according to the aspect ratios of the targets, and the scaling ratio is related to the size of the targets, so that anchor boxes of different sizes and scaling ratios can adapt to targets of different scales. In an example, generate 10 anchor boxes with five aspect ratios of [0.25, 0.5, 1.0, 2.0, 4.0] and two scaling ratios of [4, 8].
[0043] Step 3.4: Obtain the image detection results through ROI pooling and Classification based on the feature map and candidate boxes. The wall tile recognition network structure is as shown in Figure 5 .
[0044] Step 4: Input the preprocessed image into the trained wall tile recognition model for automatic recognition to obtain the recognition results. As shown in Figure 6 , the wall tile recognition model can accurately recognize the fused wall tile image. Since the pixel points of the images before and after fusion correspond to each other, the recognition results can be directly used.
[0045] Step 5: Generate the three-dimensional detection results of the wall tile area according to the recognition results and the corresponding transformation matrix.
[0046] Step 5.1: Remove redundancy from the results recognized by the model to obtain the position information on the image.
[0047] Step 5.2: According to the rotation matrix M recorded in Step 2.2, convert the position information on the image to obtain the spatial coordinate information of the three-dimensional point cloud.
[0048] Step 5.2: Since the wall tiles are on the wall and the thickness can be ignored, directly use the spatial coordinates of the wall tile area as the final three-dimensional detection results. Figure 7 For the final three-dimensional detection results.
[0049] In some examples, a multimodal wall tile recognition system is also provided, including a data acquisition module, a data preprocessing module, a wall tile recognition model, and a three-dimensional detection result generation module for the wall tile area. The data acquisition module is configured to acquire indoor three-dimensional point cloud data collected by a lidar. The data preprocessing module is configured to perform the above-mentioned step 2 to implement the preprocessing of the three-dimensional point cloud data. The preprocessed dataset is used to train the wall tile recognition model through the above-mentioned step 3. The preprocessed image is input into the trained wall tile recognition model for automatic recognition to obtain the recognition result. The three-dimensional detection result generation module for the wall tile area is configured to perform the above-mentioned step 5 to generate a three-dimensional detection result for the wall tile area according to the recognition result of the wall tile recognition model and the corresponding transformation matrix.
[0050] In some examples, a multimodal wall tile recognition device is also provided, including: a lidar, a memory, and a processor. The lidar collects indoor three-dimensional point cloud data. The memory stores instructions executable by the processor. The processor executes the instructions in the memory to implement all or part of the steps of the above method.
Claims
1. A multi-modal wall tile recognition method, characterized in that, including: Obtain indoor three-dimensional point cloud data; Perform semantic segmentation on the three-dimensional point cloud data to extract wall point clouds perpendicular to the horizontal plane; Rotate the wall point clouds to the horizontal plane to obtain wall point clouds under top view and record the rotation matrix; Perform rasterization processing on the horizontal wall point clouds, calculate intensity values in a Gaussian weighted manner, and generate an intensity map; Normalize the height values of the horizontal wall point clouds, use the normalized height values as the grayscale values of the depth map, and generate a depth map; Fuse the intensity map and the depth map into an RGB image according to the RGB channels; Use a wall tile recognition model to recognize the RGB image. The feature extraction network of the wall tile recognition model includes a deformable convolutional network that enables the convolutional kernel to adapt to wall tile regions of different shapes. After the RGB image is input into the feature extraction network, a feature map is obtained. The RPN is used to generate candidate boxes. The feature map and the candidate boxes are processed through region of interest pooling and classification to obtain an image recognition result. When generating the candidate boxes, corresponding aspect ratios of anchor boxes are generated according to the aspect ratio of the wall tile target, so that anchor boxes of different sizes and scaling ratios can adapt to wall tile targets of different scales; Generate a three-dimensional detection result of the wall tile region according to the recognition result of the wall tile recognition model and the rotation matrix.
2. The multimodal wall tile recognition method according to claim 1, wherein The specific generation method of the three-dimensional detection result is as follows: Remove redundancy from the recognition result of the wall tile recognition model to obtain position information on the image; Use the rotation matrix to convert the position information on the image to obtain the spatial coordinate information of the three-dimensional point cloud, and the spatial coordinates are the three-dimensional detection result.
3. A multimodal wall tile recognition system, characterized in that, including: A data acquisition module configured to obtain indoor three-dimensional point cloud data; A data preprocessing module configured to perform semantic segmentation on the three-dimensional point cloud data, extract wall point clouds perpendicular to the horizontal plane, rotate the wall point clouds to the horizontal plane to obtain wall point clouds under top view and record the rotation matrix, perform rasterization processing on the horizontal wall point clouds, calculate intensity values in a Gaussian weighted manner, generate an intensity map, normalize the height values of the horizontal wall point clouds, use the normalized height values as the grayscale values of the depth map, generate a depth map, and fuse the intensity map and the depth map into an RGB image according to the RGB channels; A wall tile recognition model configured to recognize the RGB image. The feature extraction network of the wall tile recognition model includes a deformable convolutional network that enables the convolutional kernel to adapt to wall tile regions of different shapes. After the RGB image is input into the feature extraction network, a feature map is obtained. The RPN is used to generate candidate boxes. The feature map and the candidate boxes are processed through region of interest pooling and classification to obtain an image recognition result. When generating the candidate boxes, corresponding aspect ratios of anchor boxes are generated according to the aspect ratio of the wall tile target, so that anchor boxes of different sizes and scaling ratios can adapt to wall tile targets of different scales; A three-dimensional detection result generation module for the wall tile region configured to generate a three-dimensional detection result of the wall tile region according to the recognition result of the wall tile recognition model and the rotation matrix.
4. The multimodal wall tile recognition system according to claim 3, characterized in that, The specific generation method of the three-dimensional detection result is as follows: the redundant information is removed from the recognition result of the wall tile recognition model to obtain the position information on the image; by using the rotation matrix, the position information on the image is converted to obtain the spatial coordinate information of the three-dimensional point cloud, and the spatial coordinate is the three-dimensional detection result.
5. A multi-modal wall tile recognition device, characterized in that, It includes: a lidar for collecting indoor three-dimensional point cloud data; a processor; a memory for storing instructions executable by the processor; wherein, when the processor is configured to execute the instructions in the memory, the method described in any one of claims 1-2 is implemented.
6. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the method described in any one of claims 1 to 2 is implemented.
Citation Information
Patent Citations
Three-dimensional target detection method and system based on image restoration
CN111079545A
Method of constructing indoor two-dimensional semantic map with wall corner as critical feature based on robot platform
US20220244740A1