Irregular room detection method, device, storage medium and program product
By combining the dual filtering and merging strategies of 3D space and 2D projection, the method of calculating the vertices of room layout is solved, and the error detection and multi-checking problems in the existing technology are solved, and accurate room detection for complex scenarios is achieved.
Patent Information
- Application Number
- CN202411977314.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-12-31
AI Technical Summary
Existing room detection methods have problems of mis-checking or multi-checking when dealing with complex shapes and irregular rooms, and lack the first perspective and real-time perception of dynamic environments.
An irregular room detection method is adopted to obtain plane detection data, combine the dual filtering and merging strategies of 3D space and 2D projection, and perform plane filtering and merging processing, calculate room layout vertices, and update the current room layout vertices.
It effectively removes noise and plane redundancy, improves the quality of processing results, can apply to complex scenarios such as irregular rooms, and achieves accurate room detection.
Smart Images

Figure CN119399190B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of room detection, and in particular, relates to a method, device, storage medium and program product for detecting irregular rooms. Background Art
[0002] Existing room detection methods include deep learning-based methods and SLAM-based methods. Deep learning-based methods use convolutional neural networks to extract features from images or point cloud data to identify planes and other structures. SLAM-based methods use sensors such as cameras or lidar to capture environmental information, build environmental maps and perform positioning.
[0003] The inventors found that the existing method had the following problems during the implementation of this embodiment:
[0004] 1. Although the deep learning method can provide relatively accurate detection results and the process is relatively clear, it requires a large amount of labeled data support when dealing with complex shapes, has poor adaptability to rooms with irregular shapes, and is prone to false detection or multiple detection.
[0005] 2. Traditional room layout detection methods are usually based on static and third-person perspective data, lack the first-person perspective and real-time perception of dynamic environments, and have certain limitations when dealing with multi-room layouts. Summary of the invention
[0006] In response to the problems existing in the prior art, the present invention provides an irregular room detection method, device, storage medium and program product, which at least partially solve the problem that the prior art cannot be applied to complex scenes such as irregular rooms due to limitations.
[0007] In a first aspect, an embodiment of the present disclosure provides a method for detecting an irregular room, comprising:
[0008] Obtain plane detection data of irregular rooms;
[0009] Performing plane filtering and merging processing on the detection data to obtain processed data, wherein the plane filtering and merging adopts double filtering and merging combining 3D space and 2D projection;
[0010] Filtering the processed data in 3D space to obtain filtered data;
[0011] Calculate room layout vertices based on the filtered data to obtain room layout vertices;
[0012] Update the current room layout vertices based on the obtained room layout vertices to complete the room detection.
[0013] Optionally, the acquiring of plane detection data of the irregular room includes:
[0014] Convert the acquired RGBD data into point cloud data, and detect the plane based on the point cloud data;
[0015] For each detected plane, calculate the maximum, minimum, mean, and standard deviation of its 3D coordinates;
[0016] Calculates the normal vector and center point of a plane based on the maximum, minimum, mean, and standard deviation of the 3D coordinates.
[0017] Optionally, the 2D projection filtering and merging in the double filtering and merging combining 3D space and 2D projection includes:
[0018] Project all wall plane point clouds to the 2D image coordinate system and filter out 2D outliers;
[0019] Calculate the 2D bounding box of each plane, and calculate the intersection and union ratio between each two planes based on the 2D bounding box;
[0020] Merge two planes whose intersection-over-union ratio is greater than the first set threshold.
[0021] Optionally, the 3D space filtering and merging in the dual filtering and merging combining 3D space and 2D projection includes:
[0022] By setting a second threshold, non-wall planes are filtered out by plane normal vectors based on prior knowledge;
[0023] The center point distance and the plane normal vector angle between the two wall planes are calculated, and if the center point distance and the normal vector angle meet a third set threshold, the two planes are merged.
[0024] Optionally, filtering the processed data in 3D space to obtain filtered data includes:
[0025] Use statistical methods to remove abnormal points, thereby filtering 3D outliers;
[0026] Removing planes whose areas are smaller than a fourth set threshold, thereby deleting invalid planes;
[0027] The spatial bounding box of each wall is calculated according to the wall plane normal vector to obtain a 3D bounding box, and the 3D bounding box is projected into the 2D image coordinate system.
[0028] Optionally, the calculating room layout vertices based on the filtered data to obtain the room layout vertices includes:
[0029] If there is only one wall plane, the vertex of the plane is returned as the room layout vertex;
[0030] If the number of planes is greater than 1, the relationship between the planes is determined based on their orientation, angle, and distance, and new room layout vertices are added based on the relationship between the planes.
[0031] Optionally, if the number of planes is greater than 1, the relationship between the planes is determined based on their orientation, angle, and distance, and room layout vertices are added based on the relationship between the planes, including:
[0032] If the two planes face the same direction and the included angle is less than the fifth set threshold, a distance judgment is performed; if the distance between the two planes is greater than the sixth set threshold, they are determined to be two non-identical and non-adjacent planes, and two new room layout vertices are added; if the distance between the two planes is not greater than the sixth set threshold, they are determined to be the same plane, the planes are merged, and one new room layout vertex is added; if the two planes only face the same direction and the included angle is not less than the fifth set threshold, they are determined to be non-identical and non-adjacent planes, and two new room layout vertices are added;
[0033] If the two planes are perpendicular, calculate the intersection point of the planes. If the intersection point is farther from both planes than the seventh set threshold, the two planes are determined to be perpendicular non-intersecting planes, and two new room layout vertices are added. Otherwise, they are perpendicular and intersecting planes, and the intersection point is used as the new room layout vertex.
[0034] Optionally, updating the current room layout vertices based on the obtained room layout vertices to complete the room detection includes:
[0035] Calculate the distance between each newly added room layout vertex and the historical room layout vertex;
[0036] If the distance is greater than the eighth set threshold, the newly added room layout vertex is confirmed to be a valid vertex, otherwise it is confirmed that the newly added room vertex coincides with the historical room vertex.
[0037] Optionally, after the step of updating the current room layout vertices based on the obtained room layout vertices, the step further includes:
[0038] Traverse the historical rooms. If the number of layout vertices of the historical room is less than 4, mark it as a suspected uncertain room.
[0039] If the number of layout vertices of the historical room is greater than or equal to 4, calculate whether the agent is in the room; if the agent is in the room, the room is the current room, otherwise, mark the room as a suspected uncertain room;
[0040] If the room type marked as a suspected uncertain room is a newly added type, a new room ID is assigned;
[0041] For rooms marked as suspected uncertain, if the room types are the same, find the nearest room and update it as the current room.
[0042] Optionally, the room type is identified using a deep convolutional neural network model.
[0043] In a second aspect, an embodiment of the present disclosure further provides an electronic device, the electronic device comprising:
[0044] at least one processor; and,
[0045] a memory communicatively connected to the at least one processor; wherein,
[0046] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the irregular room detection method described in any one of the first aspects.
[0047] In a third aspect, an embodiment of the present disclosure further provides a computer-readable storage medium, which stores computer instructions, and the computer instructions are used to enable a computer to execute any irregular room detection method described in the first aspect.
[0048] In a fourth aspect, an embodiment of the present disclosure further provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements any of the irregular room detection methods described in the first aspect.
[0049] The irregular room detection method, device, storage medium and program product provided by the present invention, wherein the irregular room detection method ensures that the generated room layout is coherent and reasonable by calculating the room layout vertices, thereby achieving the purpose of being applicable to complex scenes such as irregular rooms. The dual filtering and merging strategy of 3D space and 2D projection is combined to effectively remove noise and plane redundancy, and improve the quality of the processing results. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] The above and other objects, features and advantages of the present disclosure will become more apparent through a more detailed description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, wherein like reference numerals generally represent like components throughout the exemplary embodiments of the present disclosure.
[0051] Figure 1 A flow chart of an irregular room detection method provided by an embodiment of the present disclosure;
[0052] Figure 2 A principle block diagram of an irregular room detection device provided by an embodiment of the present disclosure;
[0053] Figure 3 A functional block diagram of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0054] The embodiments of the present disclosure are described in detail below with reference to the accompanying drawings.
[0055] It should be clear that the following embodiments of the present disclosure are described by specific specific examples, and those skilled in the art can easily understand other advantages and effects of the present disclosure from the contents disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all of the embodiments. The present disclosure can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present disclosure. It should be noted that the following embodiments and features in the embodiments can be combined with each other in the absence of conflict. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present disclosure.
[0056] It should be noted that various aspects of the embodiments within the scope of the appended claims are described below. It should be apparent that the aspects described herein may be embodied in a wide variety of forms, and any specific structure and / or function described herein is merely illustrative. Based on the present disclosure, it should be understood by those skilled in the art that an aspect described herein may be implemented independently of any other aspect, and two or more of these aspects may be combined in various ways. For example, any number of aspects described herein may be used to implement the device and / or practice the method. In addition, other structures and / or functionalities other than one or more of the aspects described herein may be used to implement this device and / or practice this method.
[0057] It should also be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present disclosure. The illustrations only show components related to the present disclosure rather than being drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component may be changed arbitrarily, and the component layout may also be more complicated.
[0058] Additionally, in the following description, specific details are provided to facilitate a thorough understanding of the examples. However, it will be understood by those skilled in the art that the aspects described may be practiced without these specific details.
[0059] By analyzing the existing technical solutions, we found that the difficulty of the existing technology lies in how to effectively handle plane detection, parameter update, filtering and merging in irregular rooms, as well as management between multiple rooms. Specifically, these difficulties include how to obtain effective and accurate wall planes in complex scenes based on the first-person perspective, how to accurately filter, merge and retain wall planes, how to calculate the current room layout vertices based on partial walls, how to manage multiple rooms, etc.
[0060] To address these issues, this embodiment designs a detection solution for the layout and type of irregular multi-rooms based on RGB data in the first perspective, and a detection method for the layout and type of irregular multi-rooms based on the first perspective and RGBD. The specific solution is as follows.
[0061] For ease of understanding, Figure 1 As shown, this embodiment discloses a non-regular room detection method, including
[0062] Step S101: Acquire plane detection data of an irregular room;
[0063] Plane detection: First, convert the RGBD data into point cloud data, and use the RANSAC algorithm to detect the plane.
[0064] The conversion of RGBD data into point cloud data is a process of converting data containing color information (RGB) and depth information (D) into a collection of points in three-dimensional space.
[0065] The steps to convert RGBD data into a point cloud include extracting point cloud coordinates from the depth image:
[0066] The pixel values in the depth image are used to calculate the coordinates of each pixel in three-dimensional space. This involves converting the depth value into coordinates (X, Y, Z) in the camera coordinate system.
[0067] Determine the camera's intrinsic and extrinsic parameters: The camera's intrinsic parameters include focal length, principal point, and distortion parameters, which are used to convert pixel coordinates into coordinates in the camera coordinate system.
[0068] The camera extrinsics include the position and orientation of the camera in the world coordinate system, which are used to convert the coordinates in the camera coordinate system into the coordinates in the world coordinate system.
[0069] Convert RGB image to point cloud color: For each point cloud coordinate, use the camera intrinsic and extrinsic parameters to correspond to the pixel coordinate on the RGB image.
[0070] Extract the color information of the corresponding pixel from the RGB image, such as RGB value or HSV value.
[0071] Correct and enhance colors: To improve color accuracy and robustness, you can correct and enhance colors, such as using a color correction matrix to correct color deviations, or using histogram equalization and contrast enhancement to enhance color contrast.
[0072] Combine point cloud coordinates and color information into a point cloud: Combine the coordinates and color information of each point into a point cloud, which can be achieved by storing the coordinates and color information in a structure or array.
[0073] The RANSAC (Random Sample Consensus) algorithm is an iterative method for estimating mathematical model parameters from a data set containing outliers and noise. In point cloud data processing, the RANSAC algorithm is widely used in plane detection, point cloud segmentation, registration, feature extraction, and shape fitting.
[0074] The core idea of the RANSAC algorithm is to estimate the model parameters by randomly selecting some samples in the data set and evaluating how well these parameters fit the remaining samples. The algorithm flow is as follows:
[0075] Random sampling: randomly select a certain number of sample points from the data set.
[0076] Model estimation: Use these sample points to estimate the parameters of the mathematical model. For example, in plane detection, the model parameters may include the plane’s normal vector and its offset from the origin.
[0077] Model evaluation: Use the estimated model parameters to evaluate the remaining data points, computing the reprojection error of each data point to the model.
[0078] Model optimization: Select the optimal model parameters based on the reprojection error.
[0079] Iteration: Repeat the above steps until a certain number of iterations is met or the number of inliers reaches a threshold.
[0080] Application of RANSAC algorithm in plane detection: In point cloud data, RANSAC algorithm can be used to detect and segment different geometric shapes, especially planes. The following are the specific applications of RANSAC algorithm in plane detection:
[0081] Plane segmentation: The RANSAC algorithm can be used to segment planes from point clouds. For example, the RansacPlane class in PCL (Point Cloud Library) can be used to segment planes from point clouds.
[0082] Iterative detection: By iteratively removing points on the detected plane and repeating the RANSAC algorithm, multiple planes in the point cloud can be detected.
[0083] Parameter calculation: For each detected plane, calculate the maximum, minimum, mean, and standard deviation of its 3D coordinates.
[0084] There are many ways to represent point cloud data in deep learning, each with its own specific advantages and applicable scenarios. The following are some common point cloud data representation methods:
[0085] Direct point cloud representation: directly taking point cloud data as input, each point contains coordinates (x, y, z) and possible other information (such as color, intensity, etc.). This method is simple and direct, allowing the network to learn complex relationships between points.
[0086] Voxelized representation: Convert point cloud data into a voxel grid, which is a regular 3D grid where each voxel represents a small cube in space. This method is easy to process using traditional convolutional neural networks because the voxelized data has a regular spatial structure.
[0087] Multi-view representation: 2D view: The point cloud is projected onto a 2D plane from multiple viewpoints (such as top view, front view, side view, etc.) to form a series of 2D images. These images can then be processed using a 2D CNN.
[0088] Graph-based representation: The point cloud is represented as a graph, where the points are nodes and the neighborhood relationships between points are edges. Graph convolutional networks can effectively process this structured data and capture the local connection patterns between points.
[0089] Hybrid representation: a combination of points and voxels: Some methods combine the advantages of point clouds and voxels, such as sampling points in a voxel grid and then using these points for deep learning processing, which can preserve precise location information while encoding rich scene context information.
[0090] Feature extraction representation: extract hand-crafted geometric features such as normals, curvatures, edge strengths, etc. from point clouds, and then use these features as input to the deep learning model.
[0091] Embedding representation: The point cloud is embedded into a continuous tensor field, which can be a multidimensional array where the position of each point corresponds to an element in the array. This approach allows the use of continuous mathematical operations and convolutions.
[0092] Transformation representation: Input and feature transformation: Transform the input point cloud (such as alignment, scaling, etc.) and transform the features to improve the performance and generalization ability of the model.
[0093] Compute the normal vector and center point of the plane.
[0094] Step S102: performing plane filtering and merging processing on the detection data to obtain processed data, wherein the plane filtering and merging adopts double filtering and merging combining 3D space and 2D projection;
[0095] 2D projection filtering and merging:
[0096] Project all wall plane point clouds to the 2D image coordinate system and filter out 2D outliers.
[0097] After projecting the point cloud to a 2D image, you can use outlier filtering methods to filter out 2D outliers. The following are some commonly used outlier filtering methods:
[0098] Statistical Outlier Removal: This method assumes that most points in the point cloud are inliers, while outliers are a minority. It calculates the distance between each point and the surrounding points, and then determines which points are outliers based on the standard deviation of these distances. Specific parameters include:
[0099] setMeanK: The number of domain points, usually set to 50.
[0100] setStddevMulThresh: The outlier threshold, usually set to 1.0, indicating that the distance is greater than 1 times the standard deviation.
[0101] Conditional Removal: This method removes all data points that meet a certain condition, such as based on the point’s intensity, color, or other attributes.
[0102] Calculate the 2D bounding box of each plane, and calculate the IoU (intersection over union) between the two planes. If the IoU is greater than the first set threshold, the two planes are considered to belong to the same wall and are merged.
[0103] Use a depth-first traversal algorithm to find all adjacent and similar wall planes, then merge these planes and update the merged plane parameters.
[0104] To find all adjacent and similar wall planes using a depth-first traversal algorithm, follow these steps:
[0105] Defining the similarity of wall planes: First, we need to define what are “similar” wall planes. In point cloud data processing, similarity may be based on several factors:
[0106] Similarity of normal vectors: Two planes are considered similar if the angle between their normal vectors is small.
[0107] Similarity by distance: Two planes can be considered similar if the distance between them is within a certain threshold.
[0108] Similarity of size: Two planes are considered similar if their areas or sizes are within a certain range.
[0109] Constructing a graph model: Each detected wall plane is considered as a node in a graph. If two planes satisfy the above similarity conditions, an edge is established between the corresponding two nodes. In this way, a graph is constructed in which nodes represent wall planes and edges represent similarities between planes.
[0110] Applying Depth-First Traversal Algorithm: Use the depth-first traversal algorithm to traverse the graph, starting from any node (wall plane), and visit all nodes connected by edges (similar wall planes). Applying the depth-first traversal algorithm can visit all connected nodes in the graph.
[0111] 3D spatial filtering and merging:
[0112] Filtering non-wall planes: By setting a second threshold, based on prior knowledge that walls are usually perpendicular to the ground, non-wall planes are filtered out by plane normal vectors.
[0113] By setting a second threshold, based on prior knowledge, non-wall planes are filtered out by plane normal vectors, including the following steps:
[0114] Calculate the normal vector of the plane: First, you need to estimate the normal vector of each plane from the point cloud data. This can be achieved by using the normal estimation method in the Point Cloud Library (PCL). For example, you can use the NormalEstimation class to calculate the normal of each point, and then average the point cloud of each plane to obtain the normal vector of the plane.
[0115] Define the normal vector characteristics of the wall plane: Based on prior knowledge, we know that the normal vector of a wall plane is usually perpendicular to the ground. Therefore, we can set a threshold, such as the angle with the vertical direction is less than 45 degrees, to determine whether a plane is a wall plane. This means that if the dot product of the normal vector of a plane and the vertical direction is greater than a certain value (for example, cos(45 degrees) = 0.707), the plane can be considered to be a wall plane.
[0116] Apply threshold to filter non-wall planes: Using the threshold defined above, non-wall planes can be filtered out. Specifically, for each plane, calculate the dot product of its normal vector and the vertical direction. If the dot product is less than the set threshold, the plane is considered not to be a wall plane and is removed from the wall plane set.
[0117] Merge wall planes: Calculate the center point distance and plane normal vector angle between two wall planes. If the distance and angle meet the third set threshold, the two planes are considered to belong to the same wall and merged.
[0118] Use a depth-first traversal algorithm to find all adjacent and similar wall planes, then merge these planes and update the merged plane parameters.
[0119] Step S103: filtering the processed data in 3D space to obtain filtered data;
[0120] Filter 3D outliers: Use statistical methods to remove abnormal points.
[0121] Using statistical methods to remove abnormal points or outliers in 3D point clouds, the following methods can be used:
[0122] Statistical outlier removal: It is a filtering method based on statistical principles. It analyzes the average distance from each point in the point cloud to the point in its neighborhood, assumes that these distances conform to the Gaussian distribution, and then determines which points are outliers based on the set standard deviation multiples. The specific steps are as follows:
[0123] Calculate the average distance of each point to the points in its neighborhood: First, for each point in the point cloud, calculate the average distance from it to its nearest neighbor.
[0124] Identify outliers: Then, based on the mean and standard deviation of the global distance, determine which points have an average distance outside the standard range. These points are considered outliers and are removed.
[0125] Radius outlier removal: removes points that have fewer neighbors within a sphere of a given radius. This method can be adjusted using two parameters: one is the minimum number of points that should be included in the sphere. The other parameter is used to calculate the radius of the sphere for neighboring points.
[0126] Delete invalid planes: Remove planes whose area is smaller than the fourth set threshold.
[0127] To remove planes whose area is smaller than the fourth set threshold and delete invalid planes, you can use the following steps:
[0128] Plane detection: First, detect each plane from the point cloud. This can be achieved by the RANSAC algorithm, which can fit a model from noisy data. In the PCL library, you can use SampleConsensusModelPlane to implement plane detection.
[0129] Calculate plane area: For each detected plane, calculate its area. This can be done by extracting the coordinates of the points within the plane and then using geometric methods to calculate the area of the plane.
[0130] Set area threshold: Set an area threshold to determine whether a plane is valid. If the area of a plane is smaller than this threshold, the plane is considered invalid.
[0131] Set an area threshold to determine whether a plane is valid. If the area of a plane is smaller than this threshold, the plane is considered invalid.
[0132] Calculate 3D bounding boxes: Calculate the spatial bounding box of each wall according to the wall plane normal vector to obtain the 3D bounding box, and project these 3D bounding boxes into the 2D image coordinate system.
[0133] Step S104: Calculate room layout vertices based on the filtered data to obtain room layout vertices;
[0134] Initialization: If there is only one wall plane, directly return the vertices of the plane as the room layout vertices.
[0135] Calculate wall plane relationships:
[0136] Sort the wall planes along the x-axis. Traverse each pair of adjacent wall planes and determine the relationship between the planes based on their orientation, angle, and distance. Add room layout vertices based on the relationship between the planes.
[0137] If the two planes face the same direction and the included angle is less than the fifth set threshold, the distance is determined: if the distance is greater than the sixth set threshold, they are two non-identical and non-adjacent planes, and two new room layout vertices are added. Otherwise, they are determined to be the same plane, the planes are merged, and one new room layout vertex is added; if the two planes only face the same direction and the included angle is not less than the fifth set threshold, they are non-identical and non-adjacent planes, and two new room layout vertices are added.
[0138] If the two planes are perpendicular, calculate the intersection point of the planes: if the intersection point is farther from both planes than the seventh set threshold, then they are perpendicular non-intersecting planes, and two new room layout vertices are added. Otherwise, they are perpendicular and intersecting planes, and the intersection point is used as the new room layout vertex.
[0139] Step S105: Update the current room layout vertices based on the obtained room layout vertices to complete the room detection.
[0140] For each newly added room layout vertex, the distance between it and the historical room layout vertex is calculated.
[0141] If the distance is greater than the eighth set threshold, the newly added room layout vertex is confirmed to be a valid vertex.
[0142] Otherwise, the newly added room vertex coincides with the historical room vertex, and the historical room layout vertex is updated.
[0143] Multiple rooms are managed as follows:
[0144] If the current number of rooms is 0, initialize the first room ID to 0.
[0145] Determine the position change of the intelligent agent:
[0146] If the agent's position has not changed, no processing is required.
[0147] If the agent's position changes, start processing suspected uncertain rooms.
[0148] Suspected uncertain room handling
[0149] Traverse the historical rooms: If the number of layout vertices of the historical room is less than 4, mark it as a suspected uncertain room. If the number of layout vertices of the historical room is greater than or equal to 4, calculate whether the agent is in the room. If the agent is in the room, the room is the current room. Otherwise, mark the room as a suspected uncertain room.
[0150] Dealing with suspected uncertain rooms:
[0151] If the current room type is a newly added type, a new room ID is assigned.
[0152] For suspected uncertain rooms of the same type, find the nearest room and update it as the current room.
[0153] Room type recognition: Through deep convolutional neural networks, data of different types of rooms are collected and labeled, and the collected and labeled data are used to complete deep convolutional neural network model training.
[0154] Based on RGB data, a deep convolutional neural network is used to collect and label different types of room data to complete model training. The model uses the ResNet series.
[0155] Based on RGB data, the deep convolutional neural network is used to collect and annotate different types of room data to complete model training. The process of using the ResNet series model can be divided into the following steps:
[0156] Data collection and annotation: First, we need to collect RGB image data. For the room type data, we can focus on the "scene.txt" file, which contains the scene type information, and the "annotation2Dfinal" and "annotation3Dfinal" files, which contain 2D segmentation and 3D bounding box annotation information, respectively, which is particularly important for room type recognition.
[0157] Model selection: Select a model from the ResNet series based on the task requirements. The ResNet series model is suitable for deep network training because its residual block design can effectively solve the gradient vanishing problem. You can choose ResNet18, ResNet34, ResNet50, ResNet101 or ResNet152, depending on the computing resources and task requirements.
[0158] Model training: Use deep convolutional neural networks to train the collected data. The core of the ResNet model is its residual block, which alleviates the gradient vanishing problem by skipping connections, allowing the network to learn features more deeply.
[0159] During the training process, you can use pre-trained weights for transfer learning, which can speed up the training and improve the generalization ability of the model. For example, you can fine-tune the ResNet model pre-trained on the ImageNet dataset.
[0160] Feature extraction and fusion: For RGB images, the ResNet model can be used to extract color and texture features. If depth information is also available, the improved two-stream convolutional recurrent neural network (Re-CRNN) can be used to fuse RGB features and depth features to improve the accuracy of object recognition. This fusion method can make full use of the potential feature information of RGB-D images and overcome the problem that existing literature only focuses on single-modal recognition results and ignores the complementary advantages of RGB images and depth images.
[0161] Model optimization and evaluation:
[0162] During the training process, different mechanisms can be designed to avoid overfitting, such as weight decay, batch normalization, etc. At the same time, the model performance can be optimized by adjusting hyperparameters such as learning rate and optimizer. Finally, the performance of the model is evaluated using a labeled test dataset to ensure that the model can effectively distinguish different types of rooms.
[0163] Through the above steps, the recognition model training of different types of rooms can be completed based on RGB data through a deep convolutional neural network. Among them, the ResNet series model is the preferred one due to its powerful feature extraction ability and effective gradient management.
[0164] When using the ResNet series model for training, the following strategies can be used to prevent overfitting: Data enhancement: Generate new training samples by transforming the original data, such as flipping, translating, and random cropping, and increase data diversity, thereby effectively alleviating the overfitting problem.
[0165] Dropout layer: Add a Dropout layer to the model to randomly "discard" some neurons, reduce the dependency between neurons, and enhance the generalization ability of the network. The Dropout rate is generally set to 0.5.
[0166] Regularization: Use L1 or L2 regularization to limit the size of model parameters and reduce the risk of overfitting by penalizing model weights.
[0167] Early Stopping: Monitor the performance of the validation set and stop training when the validation error of the model no longer improves in consecutive iterations to prevent overfitting.
[0168] Optimizer settings: Use a learning rate decay strategy or an adaptive learning rate optimizer to dynamically adjust the learning rate, and appropriately adjust the momentum and weight decay parameters to improve the stability and generalization ability of the model.
[0169] Simplify the model structure: reduce the number of layers and parameters of the model to make it simpler and reduce the risk of overfitting.
[0170] Use pre-trained models: Using models pre-trained on large datasets as a basis and then fine-tuning them on specific tasks can improve the generalization ability of the model.
[0171] Ensemble learning: Improve generalization ability by combining the prediction results of multiple models and reduce the risk of overfitting of a single model.
[0172] Dimensionality reduction: Use dimensionality reduction methods such as PCA and t-SNE to reduce the dimension of features and remove redundant features, thereby reducing the complexity of the model.
[0173] Noise processing: Remove noise data or outliers during data preprocessing to reduce the number of meaningless features learned by the model.
[0174] The above method can effectively prevent or alleviate the overfitting problem when training with the ResNet series model.
[0175] Data augmentation is an important technique used in deep learning to improve the generalization ability of models. The following are some common data augmentation operations:
[0176] Random flipping: including horizontal flipping and vertical flipping, which increases the diversity of data by flipping the image.
[0177] Random rotation: Rotate the image at random angles to simulate data from different perspectives.
[0178] Random translation: Perform random translation in a certain direction of the image to simulate data at different locations.
[0179] Random Scale: Randomly scale the image to simulate objects of different sizes.
[0180] Color transformation: including adjustment of brightness, contrast, saturation, and grayscale operations to simulate data under different lighting conditions.
[0181] Cut and Paste: Cut and paste images onto other images to simulate different compositions and backgrounds.
[0182] HSV Data Augmentation: Augment data by adjusting hue, saturation, and brightness in the HSV color space.
[0183] Affine transformation: includes operations such as rotation, translation, and scaling to simulate different geometric transformations.
[0184] Shear transformation: Shear the image to simulate different perspectives and poses.
[0185] Mosaic: stitches four images into one, so that the number and position of objects in each image vary.
[0186] CutMix: Split two images and combine them into one image.
[0187] RandomMix: Randomly select regions of two images to mix.
[0188] ResizeMix: Resize the images before mixing them.
[0189] ClassMix: Mix images of different categories together.
[0190] These data augmentation techniques can be used individually or in combination to generate more diverse training data and help the model generalize better to unknown data. Through these transformations, we can increase the size and diversity of the training dataset, allowing the model to generalize better to new, unseen data.
[0191] like Figure 2 As shown, this embodiment also discloses an irregular room detection device, including:
[0192] An acquisition module is used to acquire plane detection data of irregular rooms;
[0193] A filtering and merging module is used to perform plane filtering and merging processing on the detection data to obtain processed data, wherein the plane filtering and merging adopts double filtering and merging combining 3D space and 2D projection;
[0194] A 3D space filtering module is used to filter the processed data in the 3D space to obtain filtered data;
[0195] A vertex calculation module, used for calculating room layout vertices based on the filtered data to obtain room layout vertices;
[0196] The updating module is used to update the current room layout vertices based on the obtained room layout vertices to complete the room detection.
[0197] The electronic device disclosed in this embodiment includes a memory and a processor. The memory is used to store non-temporary computer-readable instructions. Specifically, the memory may include one or more computer program products, and the computer program product may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, a random access memory (RAM) and / or a cache memory (cache), etc. The non-volatile memory may include, for example, a read-only memory (ROM), a hard disk, a flash memory, etc.
[0198] The processor may be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device to perform desired functions. In one embodiment of the present disclosure, the processor is used to run the computer-readable instructions stored in the memory, so that the electronic device performs all or part of the steps of the irregular room detection method of each embodiment of the present disclosure.
[0199] Those skilled in the art should be able to understand that in order to solve the technical problem of how to obtain a good user experience, the present embodiment may also include well-known structures such as a communication bus and an interface, and these well-known structures should also be included in the protection scope of the present disclosure.
[0200] like Figure 3 The present invention provides a schematic diagram of the structure of an electronic device according to an embodiment of the present invention, which is suitable for implementing the electronic device in the embodiment of the present invention. Figure 3 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0201] like Figure 3 As shown, the electronic device may include a processing device (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) or a program loaded from a storage device into a random access memory (RAM). In the RAM, various programs and data required for the operation of the electronic device are also stored. The processing device, ROM, and RAM are connected to each other through a bus. The input / output (I / O) interface is also connected to the bus.
[0202] Typically, the following devices can be connected to the I / O interface: input devices such as sensors or visual information acquisition devices; output devices such as display screens; storage devices such as tapes, hard disks, etc.; and communication devices. The communication device allows the electronic device to communicate with other devices (such as edge computing devices) wirelessly or by wire to exchange data. Figure 3 An electronic device having various devices is shown, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.
[0203] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device, or installed from a ROM. When the computer program is executed by a processing device, all or part of the steps of the irregular room detection method of the embodiment of the present disclosure are executed.
[0204] For detailed description of this embodiment, reference may be made to the corresponding descriptions in the aforementioned embodiments, which will not be repeated here.
[0205] The computer-readable storage medium disclosed in this embodiment stores non-transitory computer-readable instructions. When the non-transitory computer-readable instructions are executed by a processor, all or part of the steps of the irregular room detection method of each embodiment of the present disclosure are executed.
[0206] The above-mentioned computer-readable storage media include, but are not limited to: optical storage media (e.g., CD-ROM and DVD), magneto-optical storage media (e.g., MO), magnetic storage media (e.g., magnetic tape or mobile hard disk), media with built-in rewritable non-volatile memory (e.g., memory card) and media with built-in ROM (e.g., ROM box).
[0207] For detailed description of this embodiment, reference may be made to the corresponding descriptions in the aforementioned embodiments, which will not be repeated here.
[0208] The basic principles of the present disclosure are described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, effects, etc. mentioned in the present disclosure are only examples and not limitations, and it cannot be considered that these advantages, strengths, effects, etc. are required by each embodiment of the present disclosure. In addition, the specific details disclosed above are only for the purpose of illustration and ease of understanding, and are not limitations. The above details do not limit the present disclosure to the necessity of adopting the above specific details to be implemented.
[0209] In the present disclosure, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between these entities or operations. The block diagrams of the devices, devices, equipment, and systems involved in the present disclosure are only illustrative examples and are not intended to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagram. As will be appreciated by those skilled in the art, these devices, devices, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including", "comprising", "having", etc. are open words, referring to "including but not limited to", and can be used interchangeably with them. The words "or" and "and" used here refer to the words "and / or" and can be used interchangeably with them, unless the context clearly indicates otherwise. The words "such as" used here refer to the phrase "such as but not limited to", and can be used interchangeably with them.
[0210] Additionally, as used herein, "or" used in a list of items beginning with "at least one" indicates a separate list, so that, for example, a list of "at least one of A, B, or C" means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). Furthermore, the word "exemplary" does not mean that the example described is preferred or better than other examples.
[0211] It should also be noted that in the system and method of the present disclosure, each component or each step can be decomposed and / or recombined. Such decomposition and / or recombination should be regarded as equivalent solutions of the present disclosure.
[0212] Various changes, substitutions, and modifications of the techniques described herein may be made without departing from the teachings defined by the appended claims. Furthermore, the scope of the claims of the present disclosure is not limited to the specific aspects of the processes, machines, manufactures, compositions of events, means, methods, and actions described above. Currently existing or later to be developed processes, machines, manufactures, compositions of events, means, methods, or actions that perform substantially the same functions or achieve substantially the same results as the corresponding aspects described herein may be utilized. Thus, the appended claims include such processes, machines, manufactures, compositions of events, means, methods, or actions within their scope.
[0213] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the aspects shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
[0214] The above description has been given for the purpose of illustration and description. In addition, this description is not intended to limit the embodiments of the present disclosure to the forms disclosed herein. Although multiple example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, changes, additions and sub-combinations thereof.
Claims
1. A method for detecting irregular rooms, characterized in that: include: Obtain plane detection data of irregular rooms; Performing plane filtering and merging processing on the detection data to obtain processed data, wherein the plane filtering and merging adopts double filtering and merging combining 3D space and 2D projection; Filtering the processed data in 3D space to obtain filtered data; Calculate room layout vertices based on the filtered data to obtain room layout vertices; Update the current room layout vertices based on the obtained room layout vertices to complete the room detection; The 2D projection filtering and merging in the dual filtering and merging combining 3D space and 2D projection includes: Project all wall plane point clouds to the 2D image coordinate system and filter out 2D outliers; Calculate the 2D bounding box of each plane, and calculate the intersection and union ratio between each two planes based on the 2D bounding box; Merge two planes whose intersection / union ratio is greater than a first set threshold; The updating of the current room layout vertices based on the obtained room layout vertices to complete the room detection includes: Calculate the distance between each newly added room layout vertex and the historical room layout vertex; If the distance is greater than the eighth set threshold, the newly added room layout vertex is confirmed to be a valid vertex, otherwise it is confirmed that the newly added room layout vertex coincides with the historical room vertex; After the step of updating the current room layout vertices based on the obtained room layout vertices, the method further includes: Traverse the historical rooms. If the number of layout vertices of the historical room is less than 4, mark it as a suspected uncertain room. If the number of layout vertices of the historical room is greater than or equal to 4, calculate whether the agent is in the room; if the agent is in the room, the room is the current room, otherwise, mark the room as a suspected uncertain room; If the room type marked as a suspected uncertain room is a newly added type, a new room ID is assigned; For rooms marked as suspected uncertain, if the room types are the same, find the nearest room and update it as the current room.
2. The irregular room detection method according to claim 1, characterized in that: The obtaining of plane detection data of the irregular room includes: Convert the acquired RGBD data into point cloud data, and detect the plane based on the point cloud data; For each detected plane, calculate the maximum, minimum, mean, and standard deviation of its 3D coordinates; Calculates the normal vector and center point of a plane based on the maximum, minimum, mean, and standard deviation of the 3D coordinates.
3. The irregular room detection method according to claim 2, characterized in that: The 3D space filtering and merging in the dual filtering and merging combining 3D space and 2D projection includes: By setting a second threshold, non-wall planes are filtered out by plane normal vectors based on prior knowledge; The center point distance and the plane normal vector angle between the two wall planes are calculated, and if the center point distance and the normal vector angle meet a third set threshold, the two planes are merged.
4. The irregular room detection method according to claim 2, characterized in that: The filtering of the processed data in the 3D space to obtain filtered data includes: Use statistical methods to remove abnormal points, thereby filtering 3D outliers; Removing planes whose areas are smaller than a fourth set threshold, thereby deleting invalid planes; The spatial bounding box of each wall is calculated according to the wall plane normal vector to obtain a 3D bounding box, and the 3D bounding box is projected into the 2D image coordinate system.
5. The irregular room detection method according to claim 1, characterized in that: The step of calculating the room layout vertices based on the filtered data to obtain the room layout vertices includes: If there is only one wall plane, the vertex of the plane is returned as the room layout vertex; If the number of planes is greater than 1, the relationship between the planes is determined based on their orientation, angle, and distance, and new room layout vertices are added based on the relationship between the planes.
6. The irregular room detection method according to claim 5, characterized in that: If the number of planes is greater than 1, the relationship between the planes is determined based on their orientation, angle, and distance, and new room layout vertices are added based on the relationship between the planes, including: If the two planes face the same direction and the included angle is less than the fifth set threshold, a distance judgment is performed; if the distance between the two planes is greater than the sixth set threshold, they are determined to be two non-identical and non-adjacent planes, and two new room layout vertices are added; if the distance between the two planes is not greater than the sixth set threshold, they are determined to be the same plane, the planes are merged, and one new room layout vertex is added; if the two planes only face the same direction and the included angle is not less than the fifth set threshold, they are determined to be non-identical and non-adjacent planes, and two new room layout vertices are added; If the two planes are perpendicular, calculate the intersection point of the planes. If the intersection point is farther from both planes than the seventh set threshold, the two planes are determined to be perpendicular non-intersecting planes, and two new room layout vertices are added. Otherwise, they are perpendicular and intersecting planes, and the intersection point is used as the new room layout vertex.
7. The irregular room detection method according to claim 1, characterized in that: The room type is identified using a deep convolutional neural network model.
8. An electronic device, characterized in that: The electronic device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the irregular room detection method described in any one of claims 1-7.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the irregular room detection method described in any one of claims 1-7.
10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the irregular room detection method described in any one of claims 1-7 is implemented.
Citation Information
Patent Citations
Recognizing method of wall plane profile
CN102750553A
Method for acquiring room layout plan and electronic equipment
CN113269877A
Algorithm for automatically identifying room internal partition data according to building three-dimensional model
CN118070396A