Facade Image Analysis Method and System Based on Completion and Refinement Network
Through the facade image analysis method based on the complete refinement network, the problem of inaccurate facade image analysis under occlusion or pollution in the prior art is solved, and a more accurate and robust facade image analysis effect is achieved.
Patent Information
- Application Number
- CN202410559658.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-08
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2044-05-08
AI Technical Summary
In the face of occlusion or pollution, the existing facade analysis method is difficult to accurately analyze the facade image, resulting in identification errors or inaccurate feature capture and degradation of classification performance.
The elevation image analysis method based on the complementary refinement network is adopted, and the original elevation image is preprocessed through the search and cropping method. The image analysis model including the initial encoder, feature completion module and feature refinement module is constructed. The feature completion module is used to complete the abnormal areas in the initial image features, and the feature refinement module is used to perform refinement processing to output the analysis results.
It improves the analytical accuracy of the facade images with occlusion or pollution, enhances the robustness of the model's anti-facial images and the richness of feature representation, and improves the classification performance and the accuracy of the analytical results.
Smart Images

Figure CN118397276B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of facade parsing, and specifically relates to a facade image parsing method and system based on a completion and refinement network. Background Art
[0002] With the rapid development of deep learning technology, facade parsing technology has made significant progress in the field of computer vision. Facade parsing refers to semantic segmentation of building facade images to accurately identify and separate different components, such as windows, doors, balconies, etc. Facade parsing has a wide range of applications in multiple fields, including urban development and construction, 3D reconstruction, urban road condition analysis, and autonomous driving.
[0003] Since the shapes and positional arrangements of most facade image elements are regular, existing facade parsing methods usually apply the K-means algorithm for color clustering, construct a computational graph to analyze the clustering results, and then perform semantic segmentation on the facade image. However, the current research goal of facade parsing has shifted to more complex scenarios, including facade situations suffering from occlusion or other interference and pollution. In these cases, facade images may encounter situations of local shape loss, global layout pattern damage, and feature confusion between categories. Existing facade parsing methods lack the ability to analyze and describe the high-level semantic information of facades, resulting in a decline in the parsing performance of existing facade parsing methods, inaccurate parsing of facade images with occlusion or pollution, which may cause the algorithm to misidentify objects in the image or fail to accurately capture the key features of the objects, leading to a decline in classification performance. Summary of the Invention
[0004] The present invention provides a facade image parsing method and system based on a completion and refinement network to solve the problem of inaccurate parsing of facade images with occlusion or pollution.
[0005] In a first aspect, the present invention provides a facade image parsing method based on a completion and refinement network, and the method includes the following steps:
[0006] Preprocess the original facade image to be parsed into a local facade image through a search and cropping method;
[0007] Construct an image parsing model based on a convolutional neural network, and the image parsing model includes an initial encoder, a feature completion module, and a feature refinement module;
[0008] Pre-train the initial encoder with the local facade image to obtain a trained target encoder;
[0009] Input the original facade image into the target encoder for downsampling to obtain an initial image feature;
[0010] Inputting the initial image features into the feature completion module, and completing the abnormal regions in the initial image features through the feature completion module to obtain basic image features, wherein the abnormal regions include information loss regions and information interference regions;
[0011] The obtained basic image features are refined by the feature refinement module, and the parsing result of the original facade image is output.
[0012] Optionally, the step of inputting the initial image features into the feature completion module, and completing the abnormal regions in the initial image features by the feature completion module to obtain basic image features comprises the following steps:
[0013] Inputting the initial image features into the feature completion module;
[0014] Using the average pooling layer in the feature completion module to compress the initial image features horizontally and vertically respectively, to obtain horizontal compression features and vertical compression features;
[0015] Inputting the horizontal compression feature and the vertical compression feature into the feature completion layer in the feature completion module, modeling the horizontal compression feature and the vertical compression feature through the feature completion layer, and completing the abnormal areas in the horizontal compression feature and the vertical compression feature respectively;
[0016] The horizontal compression feature and the vertical compression feature are fused into a basic image feature.
[0017] Optionally, the feature completion module further includes a classifier, and before inputting the initial image features into the feature completion module and completing the abnormal areas in the initial image features by the feature completion module to obtain the basic image features, the following steps are also included:
[0018] Acquire the target facade image where the window area is not blocked;
[0019] Compressing the target facade image in the horizontal direction and the vertical direction to obtain the actual layout in the horizontal direction and the vertical direction;
[0020] Based on the layout situation and by using matrix multiplication, a fuzzy completion image of the window area is obtained;
[0021] The fuzzy completion graph is used to complete the training of the feature completion module through the classifier.
[0022] Optionally, the loss function of the feature completion module training process is:
[0023]
[0024] Where: L FC represents the loss function of the training process of the feature completion module, N represents the total number of pixels of the target facade image, i represents the index of the pixels in the target facade image, j represents the index of the label categories in the target facade image, and y ij represents the true semantic label, represents the blurred completion map label obtained by converting the true semantic label, and P c-ij represents the predicted appearance parsing probability value obtained by the feature completion module.
[0025] Optionally, the step of refining the obtained basic image features by the feature refinement module and outputting the parsing result of the original facade image includes the following steps:
[0026] Input the basic image features into the feature refinement module;
[0027] Process the basic image features through the first convolutional layer in the feature refinement module to reduce the number of channels of the basic image features;
[0028] Capture the global spatial position features of the basic image features in the horizontal and vertical directions through the second convolutional layer in the feature refinement module;
[0029] Extract the local spatial position features of the basic image features through the third convolutional layer in the feature refinement module;
[0030] Capture the refined shape features of the basic image features through the fourth convolutional layer in the feature refinement module;
[0031] Fuse the basic image features, the global spatial position features, the local spatial position features, and the refined shape features through the fifth convolutional layer in the feature refinement module, and output the refined image features as the parsing result of the original facade image.
[0032] In a second aspect, the present invention also provides a facade image parsing system based on a completion and refinement network, and the system includes:
[0033] A preprocessing subsystem for preprocessing the original facade image to be parsed into a local facade image by a search and cropping method;
[0034] A model construction subsystem for constructing an image parsing model based on a convolutional neural network, and the image parsing model includes an initial encoder, a feature completion module, and a feature refinement module;
[0035] A pre-training subsystem for pre-training the initial encoder with the local facade image to obtain a trained target encoder;
[0036] The downsampling subsystem is used to input the original facade image into the target encoder for downsampling to obtain initial image features;
[0037] The completion subsystem is used to input the initial image features into the feature completion module, and the feature completion module completes the abnormal regions in the initial image features to obtain basic image features. The abnormal regions include information loss regions and information interference regions;
[0038] The refinement subsystem is used to refine the obtained basic image features through the feature refinement module and output the analysis result of the original facade image.
[0039] Optionally, the completion subsystem includes:
[0040] An image input unit for inputting the initial image features into the feature completion module;
[0041] A feature compression unit for horizontally and vertically compressing the initial image features respectively by using the average pooling layer in the feature completion module to obtain horizontally compressed features and vertically compressed features;
[0042] A feature completion unit for inputting both the horizontally compressed features and the vertically compressed features into the feature completion layer in the feature completion module, modeling the horizontally compressed features and the vertically compressed features through the feature completion layer, and respectively completing the abnormal regions in the horizontally compressed features and the vertically compressed features;
[0043] A feature fusion unit for fusing the horizontally compressed features and the vertically compressed features into basic image features.
[0044] Optionally, the feature completion module further includes a classifier, and the system further includes a feature completion training subsystem, and the feature completion training subsystem includes:
[0045] A training image acquisition unit for acquiring a target facade image with an unobstructed window area;
[0046] A training image compression unit for compressing the target facade image horizontally and vertically to obtain the layout facts in the horizontal and vertical directions;
[0047] A blur completion unit for obtaining a blurred completion map of the window area based on the layout facts and using matrix multiplication;
[0048] A feature completion training unit for training the feature completion module by using the blurred completion map and through the classifier.
[0049] Optionally, the loss function in the training process of the feature completion module is as follows:
[0050]
[0051] In the formula: L FC represents the loss function in the training process of the feature completion module, N represents the total number of pixels in the target facade image, i represents the index of the pixel in the target facade image, j represents the index of the label category in the target facade image, and y ij represents the true semantic label, represents the blurred completion map label obtained by converting the true semantic label, and P c-ij represents the predicted appearance parsing probability value obtained by the feature completion module.
[0052] Optionally, the refinement subsystem includes:
[0053] A feature input unit for inputting the basic image feature into the feature refinement module;
[0054] A first refinement unit for processing the basic image feature through the first convolutional layer in the feature refinement module to reduce the number of channels of the basic image feature;
[0055] A second refinement unit for capturing the global spatial position features of the basic image feature in the horizontal and vertical directions through the second convolutional layer in the feature refinement module;
[0056] A third refinement unit for extracting the local spatial position features of the basic image feature through the third convolutional layer in the feature refinement module;
[0057] A fourth refinement unit for capturing the refined shape features of the basic image feature through the fourth convolutional layer in the feature refinement module;
[0058] A fifth refinement unit for fusing the basic image feature, the global spatial position feature, the local spatial position feature, and the refined shape feature through the fifth convolutional layer in the feature refinement module, and outputting the refined image feature as the parsing result of the original facade image.
[0059] The beneficial effects of the present invention are:
[0060] During the main intervention process of the present invention, local cropped images are obtained through random cropping, and using these local images to pre-train the backbone can provide better initialization, allowing the parsing model to fully learn the local patterns of the facade, so as to perform more robustly in the face of various complex occlusion situations. In addition, through the completion module and the refinement module, the model can pay more comprehensive attention to different features, thereby being able to obtain richer and more refined feature representations, further improving the prediction accuracy of most categories, and ultimately improving the accuracy of facade image parsing. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Figure 1 It is a schematic flowchart of a facade image parsing method based on a completion and refinement network in one implementation manner of the present application.
[0062] Figure 2 It is a facade image with local loss and global loss due to occlusion in one implementation manner of the present application.
[0063] Figure 3 It is a schematic flowchart of the image preprocessing process in one implementation manner of the present application.
[0064] Figure 4 It is a schematic diagram of the problem to be solved by the feature completion module in one implementation manner of the present application.
[0065] Figure 5 It is a schematic diagram of the layout of the window area in one implementation manner of the present application.
[0066] Figure 6 It is a schematic diagram of the blurred completion map of the window area in one implementation manner of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0067] Next, the technical solutions in the embodiments of the present application will be clearly described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present application.
[0068] The terms "first", "second", etc. in the description and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first", "second", etc. are usually of the same category, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the description and claims means at least one of the connected objects, and the character " / ", generally represents an "or" relationship between the associated objects before and after.
[0069] Figure 1 FIG. is a schematic flowchart of a facade image parsing method based on a completion and refinement network in an embodiment. It should be understood that although Figure 1 each step in the flowchart of is shown sequentially according to the indication of the arrow, these steps are not necessarily executed in the order indicated by the arrow. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, Figure 1 at least a part of the steps in may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.
[0070] The facade image parsing method based on a completion and refinement network disclosed in the present invention refers to analyzing a building facade image to identify and classify different elements and features in the image, such as windows, doors, balconies, walls, etc. The purpose is to convert complex image data into a more understandable and usable information form. The meaning of image parsing is to deeply analyze the content in the image to identify specific objects and features in the image. This process usually includes steps such as image preprocessing, feature extraction, feature selection, and classification. The goal of parsing is to convert the visual information in the image into structured data that can be used for further analysis and decision-making.
[0071] In the present invention, the facade image parsing method based on the completion and refinement network can provide richer and more refined feature representations, meaning that it can capture more subtle differences and complex patterns in the image. These details and patterns are very important for distinguishing similar object categories, especially in complex scenes. Therefore, when learning can be based on richer information, the subsequent prediction accuracy for different categories will naturally increase. On the other hand, richer and more refined feature representations can directly affect the accuracy of facade image parsing. This is because the ability to accurately identify and distinguish different objects and features in the image will largely affect the accuracy of parsing. When more information can be obtained and better understood, the parsing result will naturally be more accurate and reliable. This not only improves the accuracy of single object recognition, but also improves the performance of the entire image parsing process, making the final parsing result more in line with the actual situation.
[0072] As Figure 1 shown, a specific facade image parsing method based on the completion and refinement network disclosed by the present invention specifically includes the following steps:
[0073] S101. Preprocess the original facade image to be parsed into a local facade image by a search and cropping method.
[0074] Among them, the original facade image usually contains multiple categories and may face large-area occlusion, as Figure 2 shown. The deep network needs to fully consider the differences and connections between different categories. However, in the case of insufficient training data, there may be many similar feature categories between some features, resulting in inaccurate encoding of all category features and potential feature confusion. In addition, when the facade image is largely occluded, the complete information in the original image is contaminated, and the shape and size of the contaminated area will have different effects on the final prediction. When the global layout is severely contaminated, clear local layout information can provide valuable prior knowledge. Based on this, referring to Figure 3 , decomposing the original image by a search and cropping method helps to prevent unfriendly contaminant information from appearing in network training, and at the same time extracts clear or slightly contaminated local areas from the original facade image. Therefore, this step can generate a large number of images with local information without an additional dataset.
[0075] The search and cropping method can include a manual search and cropping method and an automatic search and cropping method. Specifically, the manual search and cropping method is an artificial method for identifying different local layout patterns in elevation images. Manual search and cropping can be divided into several types, including horizontal ground layout pattern, horizontal window and balcony layout pattern, horizontal sky plane layout pattern, vertical window and balcony layout pattern, and vertical elevation overall layout pattern. The disassembled local images obtained by the manual search and cropping method usually only retain a small number of categories and a small amount of interference information. The advantage of the manual search and cropping method is that it can consciously evaluate the quality of the image, so as to determine the number of local images to be disassembled as required. In addition, the quality of the local images is relatively high, and the meaning of the local model is also clear.
[0076] Automatic random search and cropping is a technique that allows a computer to decide whether to crop randomly generated parts of the original image according to an evaluation criterion. A pre-trained evaluation network is used to score a region, and if the score exceeds a certain threshold, the region will be cropped accordingly. The evaluation metrics can be selected from total accuracy, class average accuracy, or mIoU. In an implementation, the total accuracy is used as the score for each randomly cropped local image. When the score exceeds the threshold (i.e., 0.8), the region is considered valid and saved for main intervention training. For the evaluation network, an extended fully convolutional neural network can be selected.
[0077] The automatic random cropping method allows the data in the dataset to score each other. The advantage of the automatic random cropping method is that it can determine which regions have less interference and obtain effective (slightly occluded) local disassembled images without introducing additional data. In addition, since the automatic random cropping method is based on randomness, the patterns decomposed by this method are diverse, rather than limited to several types as obtained by manual search and cropping.
[0078] S102. Construct an image parsing model based on a convolutional neural network.
[0079] Among them, the image parsing model includes an initial encoder, a feature completion module, a feature refinement module, and a decoder corresponding to the encoder.
[0080] S103. Pre-train the initial encoder with local elevation images to obtain a trained target encoder.
[0081] Among them, the initial encoder is usually an extended fully convolutional neural network. Pre-training the initial encoder with local elevation images can obtain the initial weights of the encoder. Configuring the initial encoder based on the initial weights can obtain the trained target encoder.
[0082] S104. Input the original elevation image into the target encoder for downsampling to obtain initial image features.
[0083] S105. Input the initial image features into the feature completion module, and complete the abnormal regions in the initial image features through the feature completion module to obtain the basic image features.
[0084] Among them, the abnormal regions include the information loss regions and the information interference regions. Compared with the clean elevation images, the elevation images with abnormal regions tend to lose and interfere with information. This effect also extends to the feature maps. Mainly due to the locality of the convolution operation, when there is lost or interfered information in the convolution window, this will lead to difficulties in learning the convolution. Refer to Figure 4 , such as Figure 4 shown, only some correct information in the picture can be observed, and the unknown pixels (i.e., the lost or occluded pixels) need to be inferred. Therefore, the abnormal regions in the initial image features can be completed through the feature completion module to obtain the basic image features.
[0085] S106. Refine the obtained basic image features through the feature refinement module and output the parsing result of the original elevation image.
[0086] Among them, in the elevation image, each category shows unique shape features and distribution patterns. For example, windows, balconies, and doors usually have a rectangular appearance, but their spatial distributions in the image show different rules. Specifically, windows and balconies show strong spatial regularity, while doors lack obvious spatial regularity. On the other hand, walls, shops, and the sky are usually concentrated in specific regions of the image, while chimneys have various shapes and appear at random positions throughout the image. Therefore, by refining the obtained basic image features through the feature refinement module, the parsing result of the original elevation image can be finally output.
[0087] In one implementation, step S105 specifically includes the following steps:
[0088] Input the initial image features into the feature completion module;
[0089] Use the average pooling layer in the feature completion module to horizontally compress and vertically compress the initial image features respectively to obtain the horizontally compressed features and the vertically compressed features;
[0090] Input both the horizontally compressed features and the vertically compressed features into the feature completion layer in the feature completion module. Through the feature completion layer, model the horizontally compressed features and the vertically compressed features, and complete the abnormal regions in the horizontally compressed features and the vertically compressed features respectively;
[0091] Fuse the horizontally compressed features and the vertically compressed features into the basic image features.
[0092] In this embodiment, the initial image features are first input into the feature completion module, and then the average pooling layer in the feature completion module is used to compress the image horizontally and vertically, so that the image clearly reveals the overall layout pattern in two directions, and obtains horizontal compression features and vertical compression features. Then, both the horizontal compression features and the vertical compression features are input into the feature completion layer in the feature completion module, and the horizontal compression features and the vertical compression features are modeled by the feature completion layer, and the missing information is filled in the two directions. Finally, matrix multiplication is used to fuse the pattern layout in these two directions and repair the lost or contaminated areas of the original image, and finally fuse them into the basic image features.
[0093] The above process is achieved through average pooling and one-dimensional convolution. Since the components of the facade image are mostly rectangular and have a global layout pattern (such as windows, balconies, roofs, etc.), this module is very effective for predicting and repairing contaminated facade images. In addition, compared with the ordinary convolution structure, the feature completion module has a simpler calculation method and is more space-saving. In another embodiment, the spatial attention mechanism can be further applied as another branch to assign weights to each position. Specifically, a spatial attention mechanism is configured in the feature completion module, and the initial image features are synchronously input into the spatial attention mechanism, which enables the feature completion module to understand which positions need to be completed using additional information. Finally, the output of the spatial attention mechanism and the results of the two branches of the basic image features are multiplied to obtain the final complete information.
[0094] In one embodiment, the feature completion module further includes a classifier, and before step S105, the following steps are also included:
[0095] Acquire the target facade image where the window area is not blocked;
[0096] Compress the target facade image in the horizontal and vertical directions to obtain the actual layout in the horizontal and vertical directions;
[0097] Based on the layout reality and by using matrix multiplication, a fuzzy completion image of the window area is obtained;
[0098] The feature completion module is trained using the fuzzy completion graph and the classifier.
[0099] In this embodiment, in the elevation parsing task, analyzing window elements is crucial. Window elements represent the basic layout information of the entire elevation and can provide key features for understanding the elevation structure. Different from unstructured categories, when a certain area of the elevation is occluded, it is usually possible to estimate which positions in the occluded area may contain windows based on known information. To make the feature completion module pay more attention to the overall layout pattern of windows in the image, it is necessary to perform fuzzy completion training on the feature completion module in advance. Specifically, first, obtain the target elevation image with the window area unoccluded, and then compress the target elevation image horizontally and vertically to obtain the layout ground truth in the horizontal and vertical directions, as Figure 5 shown. Next, based on the layout ground truth in the horizontal and vertical directions and using matrix multiplication, obtain the fuzzy completion map of the window area, as Figure 6 shown. The fuzzy completion map exhibits global regularity, which means that it mainly focuses on the regular layout of windows compared to the layout ground truth, including their shapes and alignments in the horizontal and vertical directions. Therefore, the fuzzy completion map with strong regularity is more suitable for the convolutional mode of the feature completion module. Finally, guide the training of the feature completion module through the classifier and using the fuzzy completion map of the window. Specifically, when there are large occlusions or contaminations in the elevation image, the fuzzy completion map can guide the feature completion module to repair missing components according to the characteristics of the windows. At the same time, the entire parsing model works in a progressive manner, first performing fuzzy completion and then refining the results.
[0100] In this embodiment, the loss function in the training process of the feature completion module is:
[0101]
[0102] where: L FC represents the loss function in the training process of the feature completion module, N represents the total number of pixels in the target elevation image, i represents the index of the pixel in the target elevation image, j represents the index of the label category in the target elevation image, y ij represents the true semantic label, represents the fuzzy completion map label obtained by converting the true semantic label, and P c-ij represents the predicted appearance parsing probability value obtained by the feature completion module.
[0103] In one of the embodiments, during the model training process of the image parsing model, the multi-class cross-entropy function is used to guide the learning and training of the image parsing model. The expression formula of the multi-class cross-entropy function is as follows:
[0104]
[0105] where: L CRO represents the multi-class cross-entropy function, C represents the total number of labels, and Pt-ij Represents the appearance parsing probability of the final prediction obtained by the image parsing model.
[0106] In summary, the total loss of the image parsing model consists of two parts: the fuzzy completion loss for guiding the feature completion module and the multi-class cross-entropy loss for guiding the entire network model. Therefore, the final total loss L of the image parsing model total The expression formula is as follows:
[0107] L total = L FC + L CRO
[0108] In one implementation, step S106 specifically includes the following steps:
[0109] Input the basic image features into the feature refinement module;
[0110] Process the basic image features through the first convolutional layer in the feature refinement module to reduce the number of channels of the basic image features;
[0111] Capture the global spatial position features of the basic image features in the horizontal and vertical directions through the second convolutional layer in the feature refinement module;
[0112] Extract the local spatial position features of the basic image features through the third convolutional layer in the feature refinement module;
[0113] Capture the refined shape features of the basic image features through the fourth convolutional layer in the feature refinement module;
[0114] Fuse the basic image features, global spatial position features, local spatial position features, and refined shape features through the fifth convolutional layer in the feature refinement module, and output the refined image features as the parsing result of the original facade image.
[0115] In this implementation, in the facade image, each category exhibits unique shape features and distribution patterns. For example, windows, balconies, and doors usually have a rectangular appearance, but their spatial distribution in the image shows different rules. Specifically, windows and balconies show strong spatial regularity, while doors lack obvious spatial regularity. On the other hand, walls, shops, and the sky are usually concentrated in specific areas of the image, while chimneys have various shapes and appear at random positions throughout the image. Therefore, it is crucial to comprehensively extract the shape features and spatial position features of different types of elements using multiple convolutional patterns. Based on this, the feature refinement module combination adopted in this implementation has multiple convolutions with different sizes and dilation ratios, and this module has a small overhead in terms of space consumption and also has the characteristics of plug-and-play.
[0116] Specifically, in this embodiment, the module first uses a first convolutional layer (1X1 convolution) to reduce the number of channels to 1 / 4 of the original number to process the input feature map. This can greatly save space resources in subsequent multiple convolutional processes. Next, a second convolutional layer (a combination of 11×1 and 1×11 one-dimensional dilated convolutions with a dilation ratio of 3) is used to capture global spatial position features in the horizontal and vertical directions. At the same time, a third convolutional layer (a combination of 7×1 and 1×7 one-dimensional dilated convolutions with a dilation ratio of 2) is used to supplement local spatial position features that may not have been extracted. In addition, a fourth convolutional layer (5×5 square convolution) is used to capture refined shape features. Finally, the above three features are combined with the original basic image features through a fifth convolutional layer (1×1 convolution), and the refined features are output as the parsing result of the original facade image.
[0117] The present invention also discloses a facade image parsing system based on a completion and refinement network, the system comprising:
[0118] A preprocessing subsystem for preprocessing the original facade image to be parsed into a local facade image by a search and cropping method;
[0119] A model construction subsystem for constructing an image parsing model based on a convolutional neural network, the image parsing model including an initial encoder, a feature completion module, and a feature refinement module;
[0120] A pre-training subsystem for pre-training the initial encoder using the local facade image to obtain a trained target encoder;
[0121] A downsampling subsystem for inputting the original facade image into the target encoder for downsampling to obtain initial image features;
[0122] A completion subsystem for inputting the initial image features into the feature completion module to complete the abnormal regions in the initial image features through the feature completion module to obtain basic image features, the abnormal regions including information loss regions and information interference regions;
[0123] A refinement subsystem for refining the obtained basic image features through the feature refinement module and outputting the parsing result of the original facade image.
[0124] In one of the embodiments, the completion subsystem includes:
[0125] An image input unit for inputting the initial image features into the feature completion module;
[0126] A feature compression unit for horizontally and vertically compressing the initial image features respectively using the average pooling layer in the feature completion module to obtain horizontally compressed features and vertically compressed features;
[0127] A feature completion unit, configured to input both the horizontally compressed feature and the vertically compressed feature into a feature completion layer in a feature completion module, model the horizontally compressed feature and the vertically compressed feature through the feature completion layer, and respectively complete abnormal regions in the horizontally compressed feature and the vertically compressed feature;
[0128] A feature fusion unit, configured to fuse the horizontally compressed feature and the vertically compressed feature into a basic image feature.
[0129] In one implementation, the feature completion module further includes a classifier, and the system further includes a feature completion training subsystem, where the feature completion training subsystem includes:
[0130] A training image acquisition unit, configured to acquire a target facade image with an unobstructed window area;
[0131] A training image compression unit, configured to compress the target facade image in the horizontal direction and the vertical direction to obtain the layout ground truth in the horizontal direction and the vertical direction;
[0132] A blur completion unit, configured to obtain a blurred completion map of the window area based on the layout ground truth and using matrix multiplication;
[0133] A feature completion training unit, configured to complete the training of the feature completion module by using the blurred completion map and through the classifier.
[0134] In one implementation, the loss function in the training process of the feature completion module is:
[0135]
[0136] In the formula: L FC represents the loss function in the training process of the feature completion module, N represents the total number of pixels in the target facade image, i represents the index of the pixel in the target facade image, j represents the index of the label category in the target facade image, y ij represents the true semantic label, represents the blurred completion map label obtained by converting the true semantic label, P c-ij represents the predicted appearance parsing probability value obtained by the feature completion module.
[0137] In one implementation, the refinement subsystem includes:
[0138] A feature input unit, configured to input the basic image feature into a feature refinement module;
[0139] A first refinement unit, configured to process the basic image feature through a first convolutional layer in the feature refinement module to reduce the number of channels of the basic image feature;
[0140] A second refinement unit, configured to capture the global spatial location features of the basic image features in the horizontal and vertical directions through a second convolutional layer in the feature refinement module;
[0141] A third refinement unit, configured to extract the local spatial location features of the basic image features through a third convolutional layer in the feature refinement module;
[0142] A fourth refinement unit, configured to capture the refined shape features of the basic image features through a fourth convolutional layer in the feature refinement module;
[0143] A fifth refinement unit, configured to fuse the basic image features, global spatial location features, local spatial location features, and refined shape features through a fifth convolutional layer in the feature refinement module, and output the refined image features as the parsing result of the original elevation image.
[0144] Those of ordinary skill in the art should understand that the discussion of any of the above embodiments is only exemplary, and is not intended to imply that the scope of protection of this application is limited to these examples; under the idea of this application, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of one or more embodiments in the present application as above, which are not provided in detail for the sake of brevity.
[0145] One or more embodiments of this application are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of this application. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of one or more embodiments of this application shall be included within the scope of protection of this application.
Claims
1. A facade image parsing method based on a completion and refinement network, characterized in that: The steps include: The original facade image to be analyzed is preprocessed into a local facade image by searching and cropping method; Building an image parsing model based on a convolutional neural network, the image parsing model includes an initial encoder, a feature completion module and a feature refinement module; Pre-training the initial encoder using the local facade image to obtain a trained target encoder; Inputting the original facade image into the target encoder for downsampling to obtain initial image features; Inputting the initial image features into the feature completion module, and completing the abnormal regions in the initial image features through the feature completion module to obtain basic image features, wherein the abnormal regions include information loss regions and information interference regions; The obtained basic image features are refined by the feature refinement module, and the parsing result of the original facade image is output.
2. The facade image parsing method based on the completion and refinement network according to claim 1 is characterized in that: The step of inputting the initial image features into the feature completion module and completing the abnormal area in the initial image features by the feature completion module to obtain the basic image features comprises the following steps: Inputting the initial image features into the feature completion module; Using the average pooling layer in the feature completion module to compress the initial image features horizontally and vertically respectively, to obtain horizontal compression features and vertical compression features; Inputting the horizontal compression feature and the vertical compression feature into the feature completion layer in the feature completion module, modeling the horizontal compression feature and the vertical compression feature through the feature completion layer, and completing the abnormal areas in the horizontal compression feature and the vertical compression feature respectively; The horizontal compression feature and the vertical compression feature are fused into a basic image feature.
3. The facade image parsing method based on the completion and refinement network according to claim 2 is characterized in that: The feature completion module further includes a classifier, and before inputting the initial image features into the feature completion module and completing the abnormal areas in the initial image features by the feature completion module to obtain the basic image features, the following steps are also included: Acquire the target facade image where the window area is not blocked; Compressing the target facade image in the horizontal direction and the vertical direction to obtain the actual layout in the horizontal direction and the vertical direction; Based on the layout situation and by using matrix multiplication, a fuzzy completion image of the window area is obtained; The feature completion module is trained by using the fuzzy completion graph and the classifier.
4. The facade image parsing method based on the completion and refinement network according to claim 3 is characterized in that: The loss function of the feature completion module training process is: Where: L FC represents the loss function of the feature completion module training process, N represents the total number of pixels in the target facade image, i represents the index of the pixel in the target facade image, j represents the index of the label category in the target facade image, and y ij represents the true semantic label, represents the fuzzy completion image label obtained by converting the true semantic label, P c-ij Represents the predicted appearance resolution probability value obtained by the feature completion module.
5. The facade image parsing method based on the completion and refinement network according to claim 1 is characterized in that: The step of refining the obtained basic image features by the feature refinement module and outputting the parsing result of the original facade image comprises the following steps: Inputting the basic image features into the feature refinement module; Processing the basic image features through the first convolutional layer in the feature refinement module to reduce the number of channels of the basic image features; Capturing global spatial position features of the basic image features in horizontal and vertical directions through a second convolutional layer in the feature refinement module; Extracting local spatial position features of the basic image features through the third convolutional layer in the feature refinement module; Capturing the refined shape features of the basic image features through the fourth convolutional layer in the feature refinement module; The basic image features, the global spatial position features, the local spatial position features and the refined shape features are fused through the fifth convolutional layer in the feature refinement module, and the refined image features are output as the parsing result of the original facade image.
6. A facade image parsing system based on a completion and refinement network, characterized in that: The system comprises: A preprocessing subsystem, used for preprocessing the original facade image to be parsed into a local facade image by searching and clipping method; A model building subsystem, used to build an image parsing model based on a convolutional neural network, wherein the image parsing model includes an initial encoder, a feature completion module, and a feature refinement module; A pre-training subsystem, configured to pre-train the initial encoder using the local facade image to obtain a trained target encoder; A downsampling subsystem, used for inputting the original facade image into the target encoder for downsampling to obtain initial image features; A completion subsystem, used for inputting the initial image features into the feature completion module, and completing the abnormal areas in the initial image features through the feature completion module to obtain basic image features, wherein the abnormal areas include information loss areas and information interference areas; The refinement subsystem is used to refine the basic image features obtained through the feature refinement module and output the analysis result of the original facade image.
7. The facade image parsing system based on the completion and refinement network according to claim 6 is characterized in that: The completion subsystem includes: An image input unit, used for inputting the initial image features into the feature completion module; A feature compression unit, used to compress the initial image features horizontally and vertically respectively by using the average pooling layer in the feature completion module to obtain horizontally compressed features and vertically compressed features; A feature completion unit, used for inputting the horizontal compression feature and the vertical compression feature into the feature completion layer in the feature completion module, modeling the horizontal compression feature and the vertical compression feature through the feature completion layer, and completing abnormal areas in the horizontal compression feature and the vertical compression feature respectively; The feature fusion unit is used to fuse the horizontal compression feature and the vertical compression feature into a basic image feature.
8. The facade image parsing system based on the completion and refinement network according to claim 7 is characterized in that: The feature completion module further includes a classifier, and the system further includes a feature completion training subsystem, and the feature completion training subsystem includes: A training image acquisition unit, used for acquiring a target facade image where the window area is not blocked; A training image compression unit, used for compressing the target facade image in the horizontal direction and the vertical direction to obtain the actual layout in the horizontal direction and the vertical direction; A fuzzy completion unit, used for obtaining a fuzzy completion image of the window area based on the layout situation and by using matrix multiplication; A feature completion training unit is used to complete the training of the feature completion module by using the fuzzy completion graph and the classifier.
9. The facade image parsing system based on the completion and refinement network according to claim 8, characterized in that: The loss function of the feature completion module training process is: Where: L FC represents the loss function of the feature completion module training process, N represents the total number of pixels in the target facade image, i represents the index of the pixel in the target facade image, j represents the index of the label category in the target facade image, and y ij represents the true semantic label, represents the fuzzy completion image label obtained by converting the true semantic label, P c-ij Represents the predicted appearance resolution probability value obtained by the feature completion module.
10. The facade image parsing system based on the completion and refinement network according to claim 6, characterized in that: The refinement subsystem includes: A feature input unit, used for inputting the basic image features into the feature refinement module; A first refinement unit, configured to process the basic image feature through a first convolutional layer in the feature refinement module to reduce the number of channels of the basic image feature; A second refinement unit, configured to capture global spatial position features of the basic image features in the horizontal direction and the vertical direction through a second convolutional layer in the feature refinement module; A third refinement unit, configured to extract local spatial position features of the basic image features through a third convolutional layer in the feature refinement module; a fourth refinement unit, configured to capture the refined shape features of the basic image features through a fourth convolutional layer in the feature refinement module; The fifth refinement unit is used to fuse the basic image features, the global spatial position features, the local spatial position features and the refined shape features through the fifth convolutional layer in the feature refinement module, and output the refined image features as the parsing result of the original facade image.
Citation Information
Patent Citations
Building facade texture automatic generation method and system
CN114117614A
Full convolution refining network gland image segmentation method based on information completion
CN114862747A