Indoor building outline recognition method based on image super-resolution reconstruction
By introducing an hourglass network of channel attention module and multi-scale module, combined with Real-ESRGAN network for image super-resolution reconstruction, the problem of object scale differences and insufficient image clarity in indoor building house contour recognition is solved, and higher recognition accuracy and accuracy are achieved.
Patent Information
- Application Number
- CN202510886381.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-06-30
AI Technical Summary
The prior art is difficult to accurately identify objects of different scales in the identification of indoor building house profiles, and the image clarity is insufficient, resulting in poor recognition accuracy.
The image super-resolution reconstruction processing method is adopted, and the channel attention module and multi-scale module are introduced through the improvement of the hourglass network, and the image super-resolution reconstruction is combined with the Real-ESRGAN network to improve the image clarity, and the model is optimized using the multi-task loss function.
It improves the accuracy and accuracy of indoor building contour recognition, especially in low-resolution image processing, which significantly improves the accuracy and recognition effect of object positioning.
Smart Images

Figure CN120411748B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image recognition, and in particular relates to an indoor building outline recognition method based on image super-resolution reconstruction processing. Background Art
[0002] When recognizing the outlines of indoor buildings, achieving accurate floor plan recognition requires solving two difficult problems. First, floor plan recognition requires precise identification of room types, walls, doors, and windows, involving objects of different scales. These objects vary greatly in size and features, posing a significant challenge to the recognition task. Second, floor plan images may contain noise, and the low resolution results in blurred edges and outlines of elements such as walls, doors, and windows in the floor plan, which affects the accuracy of object positioning during recognition.
[0003] Early traditional methods for processing floor plans relied on rules and low-level image processing. These methods extracted the floor plan structure by considering the spatial relationships between elements. While easy to implement and requiring minimal hardware, they also suffered from low accuracy and poor generalization. With the development of deep learning, research has focused on models based on convolutional neural networks and graph neural networks. These models offer advantages over traditional methods in terms of recognition accuracy and generalization, but suffer from limited image clarity.
[0004] Therefore, based on the above problems, when performing indoor building outline recognition, how to accurately identify the floor plan and improve image recognition accuracy is an urgent problem that needs to be solved for those skilled in the art. Summary of the Invention
[0005] In order to solve the problems existing in the background technology, the present invention adopts the following technical solutions:
[0006] The method for recognizing indoor building outlines based on image super-resolution reconstruction processing comprises the following steps:
[0007] Obtain an original building floor plan dataset, perform data enhancement processing on the original building floor plan data, and obtain an expanded building floor plan dataset;
[0008] The hourglass network is improved. The hourglass network has an encoder-decoder structure. A channel attention module is introduced into the residual module of the hourglass network. After the decoder upsampling process, a multi-scale module is introduced. The improved hourglass network is trained using an expanded building floor plan dataset and optimized using the Adam optimizer. The training objective is to minimize the multi-task loss function. Finally, a multi-task floor plan recognition model is obtained.
[0009] Build and train an image super-resolution model based on the Real-ESRGAN network. The trained image super-resolution model can perform super-resolution reconstruction on planar images.
[0010] When recognizing building floor plans, the resolution of the building floor plan to be recognized is first detected. If the resolution is higher than 600×600, super-resolution reconstruction is not performed. If it is lower than 600×600, super-resolution reconstruction is performed. The processed building floor plan is then recognized using the floor plan recognition model to obtain the recognition result.
[0011] Furthermore, the method for training the improved hourglass network using the expanded building floor plan dataset is as follows:
[0012] First, the convolutional neural network (CNN) is used to extract features from the expanded building plan.
[0013] The extracted features enter the encoder for processing. The encoder includes a four-layer network. Each layer of the network includes a maximum pooling layer and multiple residual modules. The encoding process of each layer of the network can be expressed by the following formula:
[0014] ;
[0015] in, Indicates the The encoding features of the layer network, It is the maximum pooling layer processing, It is Residual module processing of the layer;
[0016] The features processed by the encoder enter the bottleneck layer through maximum pooling. The bottleneck layer consists of 5 residual modules:
[0017] The features processed by the bottleneck layer enter the decoder. The decoder corresponds to the encoder one by one and uses skip connection to add and fuse the feature map of the corresponding encoder level with the upsampled feature map. The decoding process of each level in the decoder is:
[0018] ;
[0019] in, Indicates the Decoded features of the layer; represents the skip connection transformation of the encoded features, is the upsampling operation, It is Residual block processing of the layer;
[0020] The decoded features enter the multi-scale module for processing. After processing, the multi-scale module outputs the features to obtain the output:
[0021] ;
[0022] in, is the decoded feature after decoding, is the processing operation performed on the output, is the final output result;
[0023] Will output the result Out Divide and get a set of heat maps , two segmentation maps 、 , and normalize the heatmap to the interval [0, 1], where is the Sigmoid function:
[0024] ;
[0025] Get the heatmap regression loss function through the heatmap , the segmentation loss function is obtained by segmentation graph , multi-task loss function The weighted sum of the two is:
[0026] .
[0027] Furthermore, when the channel attention module processes image features, it includes the following steps:
[0028] Perform global average pooling on the image features, compress each channel into a real number, and obtain the global statistical information of the channel:
[0029] ;
[0030] in, is the first feature of the input image channels, m is the feature map height index, is the total height of the feature map, is the width index of the feature map, is the total width of the feature map, is the result of global average pooling;
[0031] Perform global maximum pooling on the image features to compress all spatial positions in each channel into a single maximum value, thereby capturing the most active response in the channel and obtaining the processed result :
[0032] ;
[0033] in, maxIt is the global maximum pooling process;
[0034] Perform feature splicing to obtain channel descriptors :
[0035] ;
[0036] Among them, [] is the splicing operation, It is the result after splicing;
[0037] Learn the weight coefficients of each channel through nonlinear transformation :
[0038] ;
[0039] in, is the Sigmoid function, is the ReLU function, 、 are the weight matrices of the first fully connected layer and the second fully connected layer respectively;
[0040] Finally, the learned weight coefficients are applied to the original features to complete the recalibration operation and obtain the output features. :
[0041] ;
[0042] in, It is a channel The weight coefficient of It is a recalibration operation.
[0043] Furthermore, when the decoded features enter the multi-scale module for processing:
[0044] First, three convolutions of different sizes are processed in parallel:
[0045] ;
[0046] ;
[0047] ;
[0048] in, represents a 1x1 convolution kernel, represents a 3x3 convolution kernel, represents a 5x5 convolution kernel, 、 、 Is the original feature The results after being processed by three convolution kernels respectively;
[0049] The three features are then concatenated and a 1×1 convolution is used to allow features of different scales to interact, thereby performing feature fusion:
[0050] ;
[0051] ;
[0052] in, It is splicing processing. This is the result of splicing. It is a 1×1 convolution process. is the output result after processing.
[0053] Furthermore, the heatmap regression loss function Responsible for training the heatmap regressor to accurately determine important locations such as wall connection points and opening endpoints:
[0054] ;
[0055] in, Represents the sample number of the heatmap regression loss function, Represents the actual point location information; It is the location information of the point predicted by the model; is the first uncertainty parameter; Indicates logarithmic operation;
[0056] Segmentation loss function Used for segmentation tasks to divide different areas in the floor plan:
[0057] ;
[0058] in, Represents the sample number of the segmentation loss function, is the real segmentation label, representing the category to which each region actually belongs. is the segmentation result predicted by the model, is the second uncertainty parameter, yes activation function, 、 They are room samples and icon samples respectively.
[0059] Furthermore, the Real-ESRGAN network adopts a GAN architecture, including a generator and a discriminator. An adaptive denoising module is introduced after the generator. The module dynamically adjusts the processing intensity according to the noise distribution, then gradually increases the resolution to the target size through upsampling operations. Finally, the number of channels is adjusted to meet the feature processing requirements at different stages.
[0060] Furthermore, when training the image super-resolution model based on the Real-ESRGAN network, the following steps are adopted:
[0061] Obtain the DF2K dataset, perform high-order degradation processing on the dataset to obtain a degraded low-resolution dataset, and construct the degraded low-resolution dataset and the original high-resolution dataset into a training set;
[0062] The degraded image in the low-resolution dataset is fed into the generator for feature extraction. The obtained feature maps are then subjected to upsampling and nearest interpolation operations, and finally a high-quality high-resolution image is generated.
[0063] The original high-resolution image in the original dataset and the generated high-resolution image are input into the discriminator. The discriminator calculates the error between the two and generates the loss function parameters for updating the generator.
[0064] The generator and discriminator are trained alternately based on the loss function. This process continues until the accuracy of the high-resolution image generated by the generator reaches a threshold. At this time, the training is stopped and a trained image super-resolution model is obtained.
[0065] Furthermore, the method for performing data enhancement processing on the original building plan data is:
[0066] S11: cropping the original building plan data to adjust the size to 256x256;
[0067] S12: performing a rotation operation on the cropped plan view data;
[0068] S13: Performing color dithering processing on the rotated plan view data.
[0069] The beneficial technical effects of the present invention are:
[0070] 1. This paper designs a channel attention module and a multi-scale module to enhance key features, allowing the network to focus on key information while also improving the model's ability to express objects of different sizes and solving the problem of object scale changes;
[0071] 2. The present invention proposes to combine image super-resolution reconstruction technology with plane map recognition technology. Before recognition, the image super-resolution reconstruction technology is used to process the low-resolution image to improve the resolution, thereby achieving better recognition.
[0072] 3. The method of the present invention performs better than the existing benchmark model in the plane map recognition task, indicating that the method proposed in the present invention can effectively improve the accuracy of the plane map recognition task. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] Figure 1 Flowchart of a method for indoor building outline recognition based on image super-resolution reconstruction processing in an embodiment of the present invention;
[0074] Figure 2 Schematic diagram of a channel attention module of an indoor building outline recognition method based on image super-resolution reconstruction processing in an embodiment of the present invention;
[0075] Figure 3 Schematic diagram of a multi-scale module of a method for recognizing indoor building outlines based on image super-resolution reconstruction processing in an embodiment of the present invention;
[0076] Figure 4 Schematic diagram of an image super-resolution model based on a Real-ESRGAN network for an indoor building outline recognition method based on image super-resolution reconstruction processing in an embodiment of the present invention;
[0077] Figure 5 Schematic diagram of an adaptive noise reduction module of an indoor building outline recognition method based on image super-resolution reconstruction processing in an embodiment of the present invention. DETAILED DESCRIPTION
[0078] The following is a further clear and complete description of the indoor building outline recognition method based on image super-resolution reconstruction processing provided by the present invention in conjunction with the accompanying drawings:
[0079] like Figure 1 As shown, the method for recognizing indoor building outlines based on image super-resolution reconstruction processing provided by this embodiment includes the following steps:
[0080] S1. Obtain an original building plan data set, perform data enhancement processing on the original building plan data, and obtain an expanded building plan data set;
[0081] S2. Improve the hourglass network. The hourglass network has an encoder-decoder structure. A channel attention module is introduced into the residual module of the hourglass network. After the decoder upsampling process, a multi-scale module is introduced. The improved hourglass network is trained using the expanded building floor plan dataset and optimized using the Adam optimizer. The training objective is to minimize the multi-task loss function. Finally, a multi-task floor plan recognition model is obtained.
[0082] S3. Build and train an image super-resolution model based on the Real-ESRGAN network. The trained image super-resolution model can perform super-resolution reconstruction on planar images.
[0083] S4. When recognizing a building floor plan, first perform a resolution test on the building floor plan to be recognized. If the resolution is higher than 600×600, no super-resolution reconstruction is performed. If the resolution is lower than 600×600, super-resolution reconstruction is performed. The floor plan recognition model is then used to recognize the processed building floor plan to obtain a recognition result.
[0084] Specifically, the method for training the improved hourglass network using the expanded building floor plan dataset is as follows:
[0085] First, the convolutional neural network (CNN) is used to extract features from the expanded building floor plan. Specifically, the expanded floor plan data is converted into a 64-channel feature map through a convolution layer with a convolution kernel size of 7x7, a stride of 2, and a padding of 3. The features are then extracted through batch normalization and ReLU activation function. :
[0086] ;
[0087] in, It is a convolutional layer with a kernel size of 7x7, a stride of 2, and a padding of 3. It is batch normalization processing, is the ReLU activation function;
[0088] After extraction, the features are processed by the encoder. The encoder consists of a four-layer network. Each layer includes a maximum pooling layer and multiple residual modules to gradually reduce the spatial dimension of the features. The encoding process of each layer can be expressed by the following formula:
[0089] ;
[0090] in, Indicates the The encoding features of the layer network, It is the maximum pooling layer processing, It is Residual modules of layers;
[0091] The features processed by the encoder enter the bottleneck layer through maximum pooling. In this area, the spatial dimension of the feature map is minimized and the number of channels is maximized, thereby capturing the semantic information in the image. The bottleneck layer is composed of five residual modules, which further enhance the expressive power of the features and allow the network to learn more complex feature representations.
[0092] The features processed by the bottleneck layer enter the decoder. The decoder layers correspond to the encoder layers one-to-one. Each layer doubles the spatial dimension of the feature map through the transposed convolution operation. In order to better restore the image details, the decoder uses a skip connection to add and fuse the feature map of the corresponding encoder layer with the upsampled feature map, making full use of the spatial detail information extracted from different layers in the encoder and alleviating the information loss problem that may occur during the upsampling process. The decoding process of each layer in the decoder is as follows:
[0093] ;
[0094] in, Indicates the Decoded features of the layer; represents the skip connection transformation of the encoded features, is the upsampling operation, It is Residual block processing of the layer;
[0095] The decoded features enter the multi-scale module for processing. After processing, the multi-scale module outputs the features to obtain the output:
[0096] ;
[0097] in, is the decoded feature after decoding, is the processing operation performed on the output, is the final output result.
[0098] Will output the result Out Divide and get a set of heat maps , two segmentation maps 、 , the heat map is used to accurately locate the key points such as wall connection points and opening endpoints, and one of the two segmentation maps is used to segment the background, room and wall, and the other is used to segment the openings of windows and doors; and the heat map is normalized to the [0, 1] interval through the Sigmoid function, where is the Sigmoid function:
[0099] ;
[0100] Get the heatmap regression loss function through the heatmap , the segmentation loss function is obtained by segmentation graph , multi-task loss function The weighted sum of the two is:
[0101] .
[0102] The heatmap regression loss function is responsible for training the heatmap regressor to accurately determine important locations such as wall connection points and opening endpoints. Its calculation formula is:
[0103] ;
[0104] in, Represents the sample number of the heatmap regression loss function, It represents the actual location information of the points, that is, the exact location of these points in reality; Indicates logarithmic operation; is the location information of the point predicted by the model, It is the first uncertainty parameter, which will be continuously adjusted and learned during the model training process; The weight of this part and Inversely proportional to the square of The smaller the value of , the greater the weight of this item, and the model will pay more attention to the prediction accuracy of this point; It is a regularization term. Its function is to avoid unreasonable prediction results of the model and make the model training more stable. By adding 1 and taking the logarithm, it is ensured that this term is always positive, which better plays a constraint role.
[0105] The segmentation loss function is used for segmentation tasks, that is, to accurately divide different areas in the floor plan, such as rooms, icons (doors, windows), etc.:
[0106] ;
[0107] in, Represents the sample number of the segmentation loss function, is the real segmentation label, representing the category to which each region actually belongs. is the segmentation result predicted by the model, yes activation function, 、 They are room samples and icon samples respectively; in this formula, the first half is the cross entropy loss, which is mainly used to measure the gap between the model prediction results and the actual labels. The model will continuously optimize and adjust according to the results to make the prediction results closer to the actual situation. The second half is It is a regularization term, which also serves to stabilize model training.
[0108] The channel attention module in this embodiment is as follows Figure 2 As shown in Figure 2, it is used to adaptively adjust the importance weight of each channel. When the channel attention module processes image features, it includes the following steps:
[0109] Perform global average pooling on the image features, compress each channel into a real number, and obtain the global statistical information of the channel:
[0110] ;
[0111] in, is the first feature of the input image channels, m is the feature map height index, is the total height of the feature map, is the width index of the feature map, is the total width of the feature map, is the result of global average pooling;
[0112] Perform global maximum pooling on image features to compress all spatial positions in each channel into a single maximum value, thereby capturing the most active response in that channel. Unlike global average pooling, which focuses on overall statistical information, maximum pooling focuses more on significant areas in the feature map. The process is as follows:
[0113] ;
[0114] in, max It is the global maximum pooling process, is the result after processing;
[0115] Perform feature splicing to obtain channel descriptors :
[0116] ;
[0117] Among them, [] is the splicing operation, It is the result after splicing;
[0118] Learn the weight coefficients of each channel through nonlinear transformation :
[0119] ;
[0120] in, is the Sigmoid function, is the ReLU function, 、 They are the weight matrices of fully connected layer 1 and fully connected layer 2 respectively;
[0121] Finally, the learned weight coefficients are applied to the original features to complete the recalibration operation and obtain the output features. :
[0122] ;
[0123] in, It is a channel The weight coefficient of It is a recalibration operation;
[0124] In this way, the channel attention module realizes adaptive feature learning of the channel dimension, enhances important features, and suppresses unimportant features. By applying the channel attention module to the residual module, the residual information is channel-weighted, enabling the network to focus on key features more effectively.
[0125] The multi-scale module in this embodiment is as follows Figure 3 As shown, when the decoded features enter the multi-scale module for processing:
[0126] First, three different sizes of convolution are processed in parallel. The receptive fields of convolution kernels of different sizes are significantly different: the receptive field of the 1x1 convolution kernel is the smallest and can only focus on a single pixel and its extremely local information; the receptive field of the 3x3 convolution kernel is relatively large and can capture a wider range of local features; the receptive field of the 5x5 convolution kernel is the largest and can obtain a wider range of context information. Through these three types of convolution, features are extracted to focus on targets of different sizes. The formula is as follows:
[0127] ;
[0128] ;
[0129] ;
[0130] in, represents a 1x1 convolution kernel, represents a 3x3 convolution kernel, represents a 5x5 convolution kernel, 、 、 Is the original feature The results after being processed by three convolution kernels respectively;
[0131] The three features are then concatenated and a 1×1 convolution is used to allow features of different scales to interact in order to perform feature fusion. The formula is as follows:
[0132] ;
[0133] ;
[0134] in, It is splicing processing. This is the result of splicing. It is a 1×1 convolution process. is the output result after processing;
[0135] Through the above operations, the modules capture features of different scales in the image respectively. After fusion, the diversity and richness of the features are increased, and the model's ability to express objects of different sizes is improved. Compared with the feature fusion in the decoder upsampling process, the multi-scale module focuses on the complementarity of different receptive fields at the same resolution to solve the problem of target scale change, while the feature fusion in the upsampling process focuses on the complementarity between features at different resolutions to solve the problem of resolution loss. The two fusions each capture different and complementary information, significantly improving the network's ability to understand complex scenes.
[0136] like Figure 4 As shown in the figure, the Real-ESRGAN network adopts a GAN architecture, including a generator and a discriminator. The main body of the generator consists of 23 RRDB (Residual in Residual Dense Block) modules, which are connected sequentially. The input of each RRDB module is the output of the previous RRDB module. After the output of the last RRDB module, it is residually connected with the initial input. An adaptive denoising module is introduced after the generator. The module dynamically adjusts the processing intensity according to the noise distribution, and then gradually increases the resolution to the target size through upsampling operations. Finally, the number of channels is adjusted to meet the feature processing requirements of different stages.
[0137] Specifically, the adaptive denoising module first generates a noise map through a noise estimator. The noise estimator consists of a 4-layer convolutional network structure. The noise estimator is used to analyze the input image and output a single-channel noise distribution map in the range [0, 1], which represents the noise intensity at each pixel position. After obtaining the noise map, the adaptive denoising module processes the features through three key components, such as Figure 5 As shown: First, the single-channel noise map is converted into a multi-channel spatial attention map through the spatial attention mechanism. The processing intensity required for each spatial position is determined according to the noise map, and the spatial attention weight is obtained:
[0138] ;
[0139] in, is the noise map, It is 3×3 convolution layer 1; is the 3×3 convolutional layer 2, yes activation function, The spatial attention weight is obtained, and then the spatial attention map is feature modulated. The spatial attention weight plus 1 is multiplied by the original feature. The feature strength is dynamically adjusted according to the noise distribution, allowing the network to adaptively enhance or suppress the features of different regions, as shown below:
[0140] ;
[0141] in, 、 They are convolutional layer 3 and convolutional layer 4 respectively. It is the result after feature modulation;
[0142] Finally, the channel attention mechanism is used to automatically learn and adjust the importance of different feature channels, and the residual connection is combined to retain the original information, and the result is output:
[0143] ;
[0144] ;
[0145] in, F is the image feature to be processed, 、 They are convolution layer 5 and convolution layer 6 respectively, AvgPool is average pooling, is the Sigmoid activation function, CA ( ) is the weight obtained after the channel attention mechanism is processed, AD ( x,N ) is the output result.
[0146] This multi-level adaptive design enables the network to perform differentiated processing based on the noise level in different areas, effectively removing noise while preserving image detail information.
[0147] The discriminator adopts a U-Net structure with spectral normalization, and outputs a feature map. Each pixel corresponds to a value representing the true probability. This can have a more accurate gradient for local texture, better handle complex inputs, and make more detailed judgments on local details of the image. In this embodiment, the discriminator adopts a symmetrical U-shaped structure, which includes a downsampling path, an upsampling path, and a jump connection connecting the two.
[0148] In detail, in this embodiment, when training the image super-resolution model based on the Real-ESRGAN network, the following steps are adopted:
[0149] Obtain the DF2K dataset, perform high-order degradation processing on the dataset to obtain a degraded low-resolution dataset, and construct the degraded low-resolution dataset and the original high-resolution dataset into a training set;
[0150] The degraded image in the low-resolution dataset is fed into the generator for feature extraction. The obtained feature maps are then subjected to upsampling and nearest interpolation operations, and finally a high-quality high-resolution image is generated.
[0151] The original high-resolution image in the original dataset and the generated high-resolution image are input into the discriminator. The discriminator calculates the error between the two and generates the loss function parameters for updating the generator.
[0152] The generator and discriminator are trained alternately based on the loss function. This process continues until the accuracy of the high-resolution image generated by the generator reaches a threshold. At this point, the training is stopped and a trained image super-resolution model is obtained.
[0153] It should be noted that the method for performing data enhancement processing on the original building plan data is:
[0154] S11: Crop the original building plan data to adjust the size to 256x256, and randomly select the coordinates of the upper left corner The cropping area is ;
[0155] S12: Rotate the training image to allow the model to be exposed to the image features of the plane image at different angles during the learning process; the specific operation is achieved through the matrix, and the image matrix is , the label matrix is , assuming a 90-degree clockwise rotation is performed, the transposition must be completed first, and then the columns must be flipped to obtain the resulting image. , similarly label results It can also be calculated like this;
[0156] S13: Given that actual floor plans may exhibit different color features due to factors such as lighting conditions, drawing styles, and scanning effects, it is possible to consider randomly adjusting the colors of the training images so that the model can learn floor plan features under multiple colors, prevent the model from over-relying on specific color pattern recognition, and improve the model's adaptability to color changes;
[0157] Specifically, the rotated planar image data is subjected to color dithering, and the image matrix is , randomly generate a Numbers in the range ,in is the parameter of brightness change, Add 1 to get the parameter , the image after brightness adjustment The calculation formula is as follows:
[0158] ;
[0159] Where, 0 It represents a completely black image;
[0160] Then convert the image to grayscale ,in 、 、 These are the red, green, and blue channels of the image:
[0161] ;
[0162] Calculate the mean of the grayscale image :
[0163] ;
[0164] in, is the pixel index in the image; is the total number of image pixels;
[0165] Randomly generate a Numbers in the range ,in is the parameter of contrast change, Add 1 to get the parameter , then according to Generate grayscale mean image (and Same type), contrast-adjusted image The calculation formula is:
[0166] ;
[0167] Randomly generate a Numbers in the range ,in is the parameter of saturation change, and Add 1 to get the parameter , the image after saturation adjustment The calculation formula is:
[0168] ;
[0169] As an example, in this embodiment, a verification experiment was carried out on the CubiCasa5K dataset; the experiment included a floor plan recognition task, and its prediction results were evaluated by five indicators widely used in floor plan recognition, namely, accuracy, overall accuracy, average accuracy, intersection-over-union (IoU), and average IoU. In the calculation of the evaluation indicators, this embodiment uses the relationship between the predicted value and the actual value, where TP is a true positive example, indicating that the area is correctly identified in the prediction result, and these areas also belong to the corresponding category in the true label; FP is a false positive example, indicating that the model predicts that an area is a certain category, but the area is not this category in the true label, that is, some areas that do not belong to a certain category are mistakenly identified as this category; FN is a false negative example, indicating that a certain category area exists in the true label, but the model does not recognize it; TN is a true negative example, indicating that the area is correctly identified by the model as non-rooms and icons.
[0170] This embodiment uses the accuracy rate ( Accuracy , Acc ) is used to measure the proportion of pixels correctly classified by the model. The accuracy is calculated as:
[0171] ;
[0172] The accuracy rate intuitively shows the accuracy of the model in identifying specific elements, and the overall accuracy rate ( Overall Accuracy , oAcc ) is calculated in the same way as the accuracy, but is calculated for all category samples in the object.
[0173] At the same time, this embodiment also chooses to use the average accuracy to reflect the average classification performance of the model across all categories. Because the recognition difficulty varies between different categories, only looking at the overall accuracy may mask the performance differences of the model in certain categories. The average accuracy mAcc The calculation formula is:
[0174] ;
[0175] in, represents the number of categories, Representation category The average accuracy rate can be used to understand the overall classification effect of the model on rooms and icons, which helps to discover the differences in the model's room and icon classification effects, and thus optimize the model in a targeted manner.
[0176] Intersection-over-Union Ratio ( ) is a key indicator for evaluating model performance in semantic segmentation tasks. It is used to measure the degree of overlap between the target area predicted by the model and the actual target area. Its formula is:
[0177] ;
[0178] The closer it is to 1, the higher the overlap between the predicted area and the true area, and the better the segmentation effect of the model on the target; on the contrary, if The closer it is to 0, the lower the overlap between the two and the worse the segmentation effect.
[0179] Average Intersection-Union Ratio ( ) is the average value of the intersection-over-union ratio of all categories, reflecting the segmentation performance of all categories. The calculation method is: first calculate the predicted area and the real area of each category separately Value, and then average to get .
[0180] The results of the method of this embodiment on the Cubicasa5K test set are compared with the results of other methods on the Cubicasa5K test set. The specific results are shown in Table 1:
[0181] Table 1
[0182]
[0183] UNet and DFPR (Deep Floor Plan Recognition Using a Multi-Task Network with Room-Boundary-Guided Attention) are both existing floor plan recognition methods. Table 1 shows the performance of existing methods and the method of this embodiment on the Cubicasa5K test set. The results show that the indoor building outline recognition method of this embodiment based on image super-resolution reconstruction processing has higher accuracy and intersection-over-union ratio, achieving better recognition results.
[0184] Thus far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art may make equivalent changes or substitutions to the relevant technical features, and the technical solutions after such changes or substitutions will fall within the scope of protection of the present invention.
Claims
1. A method for recognizing indoor building outlines based on image super-resolution reconstruction processing, characterized in that: The method comprises the following steps: Obtain an original building floor plan dataset, perform data enhancement processing on the original building floor plan data, and obtain an expanded building floor plan dataset; The hourglass network is improved. The hourglass network has an encoder-decoder structure. A channel attention module is introduced into the residual module of the hourglass network. After the decoder upsampling process, a multi-scale module is introduced. The improved hourglass network is trained using an expanded building floor plan dataset and optimized using the Adam optimizer. The training objective is to minimize the multi-task loss function. Finally, a multi-task floor plan recognition model is obtained. Build and train an image super-resolution model based on the Real-ESRGAN network. The trained image super-resolution model can perform super-resolution reconstruction on planar images. When recognizing building floor plans, the resolution of the building floor plan to be recognized is first detected. If the resolution is higher than 600×600, no super-resolution reconstruction is performed. If the resolution is lower than 600×600, super-resolution reconstruction is performed. The floor plan recognition model is then used to recognize the processed building floor plan to obtain the recognition result. The method for training the improved hourglass network using the expanded building floor plan dataset is as follows: First, the convolutional neural network (CNN) is used to extract features from the expanded building plan. The extracted features enter the encoder for processing. The encoder includes a four-layer network. Each layer of the network includes a maximum pooling layer and multiple residual modules. The encoding process of each layer of the network can be expressed by the following formula: ; in, Indicates the The encoding features of the layer network, It is the maximum pooling layer processing, It is Residual module processing of the layer; The features processed by the encoder enter the bottleneck layer through maximum pooling. The bottleneck layer consists of 5 residual modules: The features processed by the bottleneck layer enter the decoder. The decoder corresponds to the encoder one by one and uses skip connection to add and fuse the feature map of the corresponding encoder level with the upsampled feature map. The decoding process of each level in the decoder is: ; in, Indicates the Decoded features of the layer; represents the skip connection transformation of the encoded features, is the upsampling operation, It is Residual block processing of the layer; The decoded features enter the multi-scale module for processing. After processing, the multi-scale module outputs the features to obtain the output: ; in, is the decoded feature after decoding, is the processing operation performed on the output, is the final output result; Will output the result Out Divide and get a set of heat maps , two segmentation maps 、 , and normalize the heatmap to the interval [0, 1], where is the Sigmoid function: ; Get the heatmap regression loss function through the heatmap , the segmentation loss function is obtained by segmentation graph , multi-task loss function The weighted sum of the two is: ; Heatmap regression loss function Responsible for training the heatmap regressor to determine the locations of wall connection points and opening endpoints: ; in, Represents the sample number of the heatmap regression loss function, Represents the actual point location information; It is the location information of the point predicted by the model; is the first uncertainty parameter; Indicates logarithmic operation; Segmentation loss function Used for segmentation tasks to divide different areas in the floor plan: ; in, Represents the sample number of the segmentation loss function, is the real segmentation label, representing the category to which each region actually belongs. is the segmentation result predicted by the model, is the second uncertainty parameter, yes activation function, 、 They are room samples and icon samples respectively.
2. The method for indoor building outline recognition based on image super-resolution reconstruction processing according to claim 1 is characterized in that: When the channel attention module processes image features, it includes the following steps: Perform global average pooling on the image features, compress each channel into a real number, and obtain the global statistical information of the channel: ; in, is the first feature of the input image channels, m is the feature map height index, is the total height of the feature map, is the width index of the feature map, is the total width of the feature map, is the result of global average pooling; Perform global maximum pooling on the image features to compress all spatial positions in each channel into a single maximum value, thereby capturing the most active response in the channel and obtaining the processed result : ; in, max It is the global maximum pooling process; Perform feature splicing to obtain channel descriptors : ; Among them, [] is the splicing operation, It is the result after splicing; Learn the weight coefficients of each channel through nonlinear transformation : ; in, is the Sigmoid function, is the ReLU function, 、 are the weight matrices of the first fully connected layer and the second fully connected layer respectively; Finally, the learned weight coefficients are applied to the original features to complete the recalibration operation and obtain the output features. : ; in, It is a channel The weight coefficient of It is a recalibration operation.
3. The method for indoor building outline recognition based on image super-resolution reconstruction processing according to claim 1 is characterized in that: When the decoded features enter the multi-scale module for processing: First, three convolutions of different sizes are processed in parallel: ; ; ; in, represents a 1x1 convolution kernel, represents a 3x3 convolution kernel, represents a 5x5 convolution kernel, 、 、 Is the original feature The results after being processed by three convolution kernels respectively; The three features are then concatenated and a 1×1 convolution is used to allow features of different scales to interact, thereby performing feature fusion: ; ; in, It is splicing processing. This is the result of splicing. It is a 1×1 convolution process. is the output result after processing.
4. The method for indoor building outline recognition based on image super-resolution reconstruction processing according to claim 1 is characterized in that: The Real-ESRGAN network adopts a GAN architecture, including a generator and a discriminator. An adaptive denoising module is introduced after the generator. The module dynamically adjusts the processing intensity according to the noise distribution, then gradually increases the resolution to the target size through upsampling operations. Finally, the number of channels is adjusted to meet the feature processing requirements at different stages.
5. The method for indoor building outline recognition based on image super-resolution reconstruction processing according to claim 4 is characterized in that: When training an image super-resolution model based on the Real-ESRGAN network, the following steps are used: Obtain the DF2K dataset and perform high-order degradation processing on the dataset to obtain a degraded low-resolution dataset. The degraded low-resolution dataset and the original high-resolution dataset constitute the training set; The degraded image in the low-resolution dataset is fed into the generator for feature extraction. The obtained feature maps are then subjected to upsampling and nearest interpolation operations, and finally a high-quality high-resolution image is generated. The original high-resolution image in the original dataset and the generated high-resolution image are input into the discriminator. The discriminator calculates the error between the two and generates the loss function parameters for updating the generator. The generator and discriminator are trained alternately based on the loss function. This process continues until the accuracy of the high-resolution image generated by the generator reaches a threshold. At this time, the training is stopped and a trained image super-resolution model is obtained.
6. The method for indoor building outline recognition based on image super-resolution reconstruction processing according to claim 1, characterized in that: The method for data enhancement processing of the original building floor plan data is: S11: cropping the original building plan data to adjust the size to 256x256; S12: performing a rotation operation on the cropped plan view data; S13: Performing color dithering processing on the rotated plan view data.
Citation Information
Patent Citations
Face image super-resolution method based on facial structure prior fusion network
CN119624772A
Improved tissue ultrasound image segmentation lightweight system and method
CN119648716A