High-accuracy house type image inner contour segmentation method, system and equipment

By combining DETR and Mask2Former models, high accuracy segmentation of building structures in floor plans is achieved, which solves the problem of time-consuming, labor-intensive and inaccurate traditional manual annotation method, and improves the segmentation accuracy and the integrity of the results.

CN120182301APending Publication Date: 2025-06-20SOUTH CHINA AGRICULTURAL UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510106476.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

Traditional floor plan generation and labeling methods rely on manual labor, which is time-consuming and labor-intensive, and is prone to inaccurate drawings due to human error, making it difficult to ensure unified division and labeling standards.

Method used

The combination of efficient object detection capabilities based on DETR model and accurate segmentation capabilities of Mask2Former model is adopted to achieve high accuracy segmentation of building structures in floor plans.

Benefits of technology

It improves the accuracy of floor plan segmentation, reduces false detection and missed detection in structural overlap areas, and ensures the integrity and accuracy of segmentation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182301A_ABST
    Figure CN120182301A_ABST
Patent Text Reader

Abstract

The invention discloses a high-accuracy house type image inner contour segmentation method, system and device. The method comprises the following steps: carrying out normalization processing (such as data enhancement and noise optimization) on data; detecting a target area based on an improved DETR model, extracting multi-layer features through an optimized encoder and decoder, and generating a bounding box and confidence and category thereof; combining the bounding box with the original image, randomly expanding the bounding box to extract a house area, and reducing watermark and advertisement interference in the image; and on the basis of an improved Mask2Former model, an optimized Transform decoder and an attention mechanism are utilized to generate a segmentation mask and a classification result. According to the method, DETR and Mask2Former models are combined, all functional rooms of the house type image are quickly and accurately segmented, and the blank of a large-scale house type image processing method is filled.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of deep learning, and particularly relates to a method, system and device for segmenting the inner contour of a house floor plan with high accuracy. Background Art

[0002] The house floor plan is an important information carrier in the fields of architectural design, space planning, and house renovation, showing the layout, structure, and functional partition of each indoor space. The generation of high-precision house floor plans is of great significance for architectural design, property management, and real estate analysis. However, traditional methods for generating and annotating house floor plans often rely on manual drawing, measurement, and annotation, which are not only time-consuming and laborious but also prone to inaccuracies due to human errors. In addition, the details and structural types (such as walls, doors, windows, kitchens, bathrooms, etc.) included in different house floor plans are also relatively complex, and it is difficult to ensure unified segmentation and annotation standards through manual operations. Therefore, developing an automated and intelligent house floor plan segmentation and generation technology that can efficiently and accurately process various building structures in the drawings is an important direction for current technological development.

[0003] Currently, object detection and instance segmentation technologies based on deep learning are widely used in various image segmentation tasks. Object detection models and instance segmentation models have been proven to be able to accurately identify and segment different categories of objects in complex image scenes. However, in the application of house floor plans, due to the characteristics of the drawings, the compactness of building structures, and the similarity of lines, traditional models often have difficulty achieving accurate segmentation. Therefore, the house floor plan segmentation scheme based on object detection and instance segmentation models has become a key means to achieve automated house floor plan generation. Summary of the Invention

[0004] The main purpose of the present invention is to overcome the deficiencies of the prior art and provide a method, system, and device for segmenting the inner contour of a house floor plan with high accuracy. The present invention realizes the preliminary detection of building structures in the house floor plan based on the DETR model, automatically generates the bounding box of each structure, and further refines the segmentation in combination with the Mask2Former model to achieve accurate instance segmentation of structures such as walls, doors, windows, and kitchens. This method combines the efficient object detection ability of DETR and the accurate segmentation ability of Mask2Former, not only improving the segmentation accuracy of the house floor plan but also effectively reducing the missed detection and misdetection of structures in complex drawings.

[0005] To achieve the above object, the present invention adopts the following technical solutions:

[0006] In the first aspect, the present invention provides a method for segmenting the inner contour of a house floor plan with high accuracy, including the following steps:

[0007] Obtain the house type image, annotate the key structures in the house type image, perform data augmentation on the obtained house type image, and preprocess the data-augmented house type image;

[0008] Input the preprocessed house type image into the multi-layer feature fusion DETR model, which includes a backbone network, an improved encoder, and an improved decoder; the processing of the house type image by the multi-layer feature fusion DETR model is as follows: the backbone network extracts low-level image features with more boundary information through a convolutional neural network, and the improved encoder uses a multi-scale feature processing mechanism to perform feature enhancement processing on the extracted low-level image features with more boundary information, and further performs cross-scale feature integration on the enhanced features in the improved encoder to achieve the interaction of features within the scale and the fusion of features between scales; input the fused features into the improved decoder to output the bounding box coordinates, confidence, and category of each house type map;

[0009] Combine the bounding boxes generated by the multi-layer feature fusion DETR model with the original image, randomly expand the bounding boxes, and extract the house type map area to eliminate the interference of watermarks or advertising stickers in the original image on the segmentation result;

[0010] Fill the detection results of the multi-layer feature fusion DETR model to the original image size as the input of the multi-scale enhanced Mask2Former segmentation model. The multi-scale enhanced Mask2Former segmentation model extracts features of different scales through a multi-scale feature pyramid to more comprehensively capture targets of different sizes; based on a preset multi-scale feature depth fusion and optimized feature integration strategy, input the features of different scales into the optimized Transformer decoder, combine the design of the mask and the query, pay attention to the local area through cross-attention, and at the same time use the self-attention mechanism to further strengthen the feature representation, so as to improve the segmentation ability of complex structures, and finally output the binary mask and segmentation result of each structure;

[0011] Overlay the segmentation results of all structures to generate a complete segmentation map.

[0012] As a preferred technical solution, the obtaining of the house type image, the annotation of the key structures in the house type image, the data augmentation of the obtained house type image, and the preprocessing of the data-augmented house type image are specifically as follows:

[0013] Receive the house type image transmitted by the user terminal, and manually or automatically annotate the key structures in the house type image;

[0014] Perform data augmentation operations on the annotated house type image, including large-scale jitter, random flipping, image sharpening, and adding gravel noise, to enrich data diversity and improve the robustness of the model;

[0015] All enhanced house type images are scaled to a unified size within the set pixel range to ensure accurate annotation of categories, balanced category distribution, and consistent image size in the dataset, thereby forming a preprocessed house type map dataset.

[0016] As a preferred technical solution, the processing of the house type image by the multi-layer feature fusion DETR model specifically includes the following steps:

[0017] The preprocessed house type map image is input into the multi-layer feature fusion DETR model. The multi-layer feature fusion DETR model extracts low-level image features through a convolutional neural network. The low-level image features include edge information, texture features, and basic shape descriptions of the image, laying a foundation for the extraction of more advanced features;

[0018] The extracted low-level image features are input into an improved encoder. The improved encoder uses a multi-scale feature processing mechanism to enhance the low-level image features. By encoding feature layers of different scales, enhanced image features are generated, thereby comprehensively capturing the spatial relationships and structural details of the targets in the house type map;

[0019] The enhanced image features are further subjected to cross-scale feature integration in the encoder. Through a specially designed scale fusion interaction module, interaction of features within a scale and fusion of features between scales are achieved, ensuring that the extracted image features are optimized at both the global and local levels;

[0020] After generating multi-layer features, the multi-layer feature fusion DETR model introduces a refined feature selection strategy. The feature selection strategy enables the multi-layer feature fusion DETR model to screen out high-quality features from the encoded features and suppress redundant or low-confidence features;

[0021] The selected high-quality features are fed into an improved decoder. The improved decoder adopts a hierarchical structure and gradually refines the target features through multi-level decoding operations. The improved decoder adds an auxiliary prediction mechanism, which can provide additional constraints through auxiliary tasks and further improve the accuracy and stability of the decoding output;

[0022] The improved decoder outputs the detection results of each target, including the bounding box coordinates, category, and confidence of the target.

[0023] As a preferred technical solution, the bounding box generated by the multi-layer feature fusion DETR model is combined with the original image, and the bounding box is randomly expanded to extract the house type map area, avoiding interference from watermarks or advertisement stickers in the original image on the segmentation result. Specifically:

[0024] The bounding boxes generated by the multi-layer feature fusion DETR model are combined with the original image. By overlaying the bounding boxes on the original image, the spatial position information and structural features in the image are retained;

[0025] On the basis of the bounding box combination, a random expansion operation is performed on it. The purpose of the random expansion is to increase the area coverage on the basis of the original size of the bounding box to cope with the possible boundary offset or incomplete definition of the target area during the detection process by the model.

[0026] As a preferred technical solution, the processing of the detection results by the multi-scale enhanced Mask2Former segmentation model is specifically as follows:

[0027] The floor plan area extracted by the multi-layer feature fusion DETR model is used as the input of the multi-scale enhanced Mask2Former segmentation model;

[0028] The multi-scale enhanced Mask2Former segmentation model extracts multi-level features from the input data at different resolutions through the multi-scale feature pyramid module;

[0029] The extracted multi-scale features are input into the attention-focusing optimized Transformer decoder for processing. By introducing a more efficient multi-scale feature depth fusion and optimized feature integration strategy, the depth fusion and optimization of multi-scale features are achieved;

[0030] Mask2Former combines the mask and query mechanism to focus on specific local regions of the image through cross-attention; the mask is used to limit the calculation range of cross-attention within a specific target area to avoid interference from redundant information; at the same time, the model uses the self-attention mechanism to further strengthen the learned feature representation to ensure that both the global and local features of the target can be fully modeled, improving the segmentation ability for complex floor plan structures;

[0031] The improved segmentation model finally outputs the binary mask of each target structure and the corresponding classification result.

[0032] As a preferred technical solution, the specific multi-scale feature depth fusion and optimized feature integration strategy is as follows:

[0033] First, an adaptive weighting module is used to preprocess features at different scales. Weights are automatically assigned according to the importance of features in spatial position and semantic information. The weights of small-scale features are increased at the details of the target, and the weights of large-scale features are reasonably weighted in the background and overall structure areas. For small targets such as doors and windows, the boundary weights of small-scale features are highlighted.

[0034] Next, an attention-guided feature fusion unit is introduced. Calculate the feature correlation scores, and based on this, strengthen the fusion of highly correlated regions and weaken the low-correlated regions. For example, for wall structures, enhance the coherent fusion of their features at different scales.

[0035] Finally, adopt a progressive fusion method to gradually fuse from large-scale features to small-scale features, and calibrate and optimize with a small convolutional neural network after each fusion to ensure the accuracy of the fused feature space and semantics. For example, when fusing staircase structures, gradually fuse and fine-tune to clearly present their contours.

[0036] As a preferred technical solution, the superposition of the segmentation results of all structures to generate a complete segmentation map is specifically as follows:

[0037] Perform a superposition operation on the segmentation results of all structures to construct a complete segmentation map. During the generation process, use different colors to label each category to visually and intuitively view the segmentation results, so that various structures can be distinguished and recognized in the map for subsequent observation and evaluation of the segmentation situation;

[0038] With the help of image synthesis technical means, perform a synthesis process on the obtained segmentation results and the original house type plan, so that the segmentation results can be accurately superimposed on the original house type plan to obtain an accurate house type plan segmentation result.

[0039] In a second aspect, the present invention provides a high-accuracy inner contour segmentation system for house type plans, which applies the high-accuracy inner contour segmentation method for house type plans, including a data preprocessing module, a house type plan target detection module, a house type plan extraction module, a multi-scale feature segmentation module, and a post-processing module;

[0040] The data preprocessing module is used to obtain a house type image, label the key structures in the house type image, perform data augmentation on the obtained house type image, and perform preprocessing on the data-augmented house type image;

[0041] The house type plan target detection module is used to input the preprocessed house type image into a multi-layer feature fusion DETR model. The multi-layer feature fusion DETR model includes a backbone network, an improved encoder, and an improved decoder. The processing of the house type image by the multi-layer feature fusion DETR model is specifically as follows: The backbone network extracts low-level image features with more boundary information through a convolutional neural network. The improved encoder uses a multi-scale feature processing mechanism to perform feature enhancement processing on the extracted low-level image features with more boundary information, and further performs cross-scale feature integration on the enhanced features in the improved encoder to achieve the interaction of features within a scale and the fusion of features between scales; Input the fused features into the improved decoder to output the bounding box coordinates, confidence, and category of each house type plan;

[0042] The floor plan extraction module is used to combine the bounding boxes generated by the multi-layer feature fusion DETR model with the original image, randomly expand the bounding boxes, and extract the floor plan area to eliminate the interference of watermarks or advertising stickers in the original image on the segmentation result;

[0043] The multi-scale feature segmentation module is used to fill the detection result of the improved DETR model to the original image size as the input of the multi-scale enhanced Mask2Former segmentation model. The multi-scale enhanced Mask2Former segmentation model extracts features of different scales through a multi-scale feature pyramid to more comprehensively capture targets of various sizes; based on a preset multi-scale feature depth fusion and optimized feature integration strategy, the features of different scales are input into the optimized Transformer decoder, combined with the design of masks and queries, and local regions are focused through cross-attention, while the self-attention mechanism is used to further strengthen the feature representation, thereby improving the segmentation ability for complex structures, and finally outputting the binary mask and segmentation result of each structure;

[0044] The post-processing module is used to superimpose the segmentation results of all structures to generate a complete segmentation map.

[0045] In a third aspect, the present invention provides an electronic device, which includes:

[0046] At least one processor; and,

[0047] A memory communicatively connected to the at least one processor; wherein,

[0048] The memory stores computer program instructions executable by the at least one processor, and the computer program instructions are executed by the at least one processor so that the at least one processor can execute the high-accuracy floor plan inner contour segmentation method described above.

[0049] In a fourth aspect, the present invention provides a computer-readable storage medium storing a program, and when the program is executed by a processor, the high-accuracy floor plan inner contour segmentation method described above is implemented.

[0050] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0051] (1) Based on the combination of the DETR detection model and the Mask2Former segmentation model, the present invention realizes the efficient detection and accurate segmentation of building structures (such as walls, doors and windows, kitchens, bathrooms, etc.) in the floor plan. Compared with traditional segmentation methods, this method effectively improves the segmentation accuracy, reduces the misdetection and missed detection phenomena in the overlapping areas of structures, and ensures the integrity and accuracy of the segmentation result.

[0052] (2) The present invention only requires one detection and segmentation process, significantly reducing the preprocessing requirements for drawings and simplifying the operation process. Through optimization steps such as multi-scale feature extraction and boundary clipping, efficient segmentation of building structures at different scales in floor plans is achieved, which can adapt to floor plans of different styles and complexities, improving the applicability and robustness of the method.

[0053] (3) The present invention adopts an improved Transformer architecture, significantly improving the accuracy of the segmentation model. At the same time, the error propagation is effectively reduced through the attention mechanism, achieving a high-precision segmentation effect.

[0054] (4) The number of model parameters in the present invention is optimized, reducing the computational cost while ensuring the accuracy. Combined with an efficient inference algorithm, this method has strong real-time processing capabilities and is suitable for large-scale floor plan processing scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0056] Figure 1 It is a flowchart of the method for segmenting the inner contour of a floor plan with high accuracy according to an embodiment of the present invention;

[0057] Figure 2 It is an original floor plan example according to an embodiment of the present invention.

[0058] Figure 3 It is an effect diagram of stone noise data augmentation according to an embodiment of the present invention.

[0059] Figure 4 It is an effect diagram of the detection frame according to an embodiment of the present invention.

[0060] Figure 5 It is a segmentation result diagram according to an embodiment of the present invention.

[0061] Figure 6 It is a block diagram of the system for segmenting the inner contour of a floor plan with high accuracy according to an embodiment of the present invention.

[0062] Figure 7 It is a structural diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0063] To enable those skilled in the art to better understand the solution of this application, the technical solution in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of this application.

[0064] In this application, the mention of "embodiment" means that the specific features, structures or characteristics described in connection with the embodiment can be included in at least one embodiment of this application. The phrase appears in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described in this application can be combined with other embodiments.

[0065] As Figure 1 shown, the method for segmenting the inner contour of a floor plan with high accuracy in this embodiment includes the following steps:

[0066] (1) The user terminal transmits a floor plan image, annotates the key structures in the floor plan image, and performs data augmentation on the obtained floor plan image, including large-scale jitter, random flipping, image sharpening, and gravel noise. To unify the dataset specifications, all images are adjusted in shape proportionally to between 640 and 2048 to obtain a preprocessed floor plan image with accurate annotation for each category, balanced category distribution, and unified size.

[0067] Further, the step (1) is specifically:

[0068] (1-1) Receive the floor plan image transmitted by the user terminal, manually or automatically annotate the key structures in the image to ensure that the annotations for each category are clear and complete, laying a foundation for subsequent processing, as Figure 2 shown;

[0069] (1-2) Perform data augmentation operations on the annotated floor plan image, including large-scale jitter, random flipping, image sharpening, and adding gravel noise, etc., to enrich data diversity and improve the robustness of the model, as Figure 3 shown;

[0070] (1-3) Adjust all images to a unified size within the range of 640 to 2048 pixels proportionally to ensure accurate annotation of categories, balanced category distribution, and consistent image size in the dataset, thereby forming a preprocessed floor plan dataset.

[0071] (2) Input the preprocessed floor plan image into the multi-layer feature fusion DETR model. First, the model extracts low-level image features with more boundary information through a convolutional neural network, and further processes the extracted multi-layer features using an enhanced encoder to generate a high-quality image feature representation. On this basis, a more refined feature selection strategy is introduced in this embodiment to improve the utilization efficiency of key features. Finally, the improved decoder module outputs the bounding box coordinates, confidence, and category of each floor plan through an auxiliary prediction mechanism.

[0072] Further, the DETR detection model is used to obtain the target structure bounding box in step (2), specifically:

[0073] (2-1) The preprocessed floor plan image is input into the multi-layer feature fusion DETR model, and low-level image features are first extracted through the convolutional neural network of the model (usually a backbone network composed of ResNet or a similar architecture). These low-level features mainly contain the edge information, texture features, and basic shape descriptions of the image, laying a foundation for the extraction of more advanced features in the follow-up;

[0074] (2-2) The extracted low-level features are input into an improved feature encoder, and the encoder enhances the features using a multi-scale feature processing mechanism. By encoding feature layers at different scales, the model can generate higher-quality and richer feature representations, thereby comprehensively capturing the spatial relationships and structural details of the targets in the floor plan. In particular, the feature expression ability for targets of different sizes is significantly improved;

[0075] (2-3) The enhanced features are further integrated across scales in the encoder. Through a specially designed module, the model realizes the interaction of features within the scale and the fusion of features between scales, ensuring that the extracted image features are optimized at both the global and local levels. This integration strategy improves the synergy between features and enhances the model's perception ability for complex scenes;

[0076] (2-4) After generating multi-layer features, the model introduces a refined feature selection strategy. Through a specific algorithm, the model can screen out the most critical information from the encoded features and suppress redundant or low-confidence features. This feature selection mechanism improves the model's attention to key regions and ensures that only high-value image features are used in the subsequent decoding process;

[0077] (2-5) The screened high-quality features are fed into an improved decoder module. The decoder adopts a hierarchical structure and gradually refines the target features through multi-level decoding operations. On this basis, the decoder adds an auxiliary prediction mechanism, which can provide additional constraints through auxiliary tasks and further improve the accuracy and stability of the decoding output;

[0078] (2-6) Finally, the improved decoder module outputs the detection results of each target, including the bounding box coordinates, category, and confidence of the target. In this way, the model can accurately locate and classify various key structures in the floor plan, providing high-quality basic information for subsequent tasks (such as instance segmentation or structural analysis), such as Figure 4 shown.

[0079] (3) Combine the bounding boxes generated by the multi-layer feature fusion DETR model with the original image, randomly expand the bounding boxes, and accurately extract the floor plan area, effectively avoiding the interference of watermarks or advertisement stickers in the original image on the segmentation result.

[0080] Further, the specific steps of step (3) are as follows:

[0081] (3-1) Combine the bounding boxes generated by the multi-layer feature fusion DETR model with the original image. By overlaying the bounding boxes on the original image, the model can effectively retain the spatial position information and structural features in the image, laying a foundation for subsequent area extraction. This step ensures that the extracted target area can closely surround the floor plan area to be segmented actually, avoiding missing key information;

[0082] (3-2) On the basis of the bounding box combination, perform a random expansion operation on it. The purpose of the random expansion is to appropriately increase the area coverage on the basis of the original size of the bounding box to address issues such as possible boundary offsets or incomplete definition of the target area during the detection process by the model. By expanding the range of the bounding box, the model can more accurately extract the complete area containing the floor plan, reducing the missing of targets caused by improper cropping;

[0083] (3-3) By combining the bounding box expansion and the area extraction method, the interference of irrelevant content (such as watermarks, advertisement stickers, etc.) in the original image can be effectively avoided. These irrelevant information usually exists at the edge or background of the floor plan area and may have a negative impact on the segmentation result of the model. By accurately extracting the independent floor plan area, the segmentation model can focus more on the target structure in subsequent processing, improving the accuracy and reliability of the result.

[0084] (4) Use the floor plan extracted by the multi-layer feature fusion DETR model as the input of the multi-scale enhanced Mask2Former segmentation model. First, the model extracts features at different scales through a multi-scale feature pyramid to more comprehensively capture objects of various sizes; subsequently, a more effective multi-scale feature deep fusion and optimized feature integration strategy is introduced, and features at different scales are input into the optimized Transformer decoder. Combining the design of masks and queries, the model can focus on local regions through cross-attention and further strengthen the feature representation using the self-attention mechanism, thereby improving the segmentation ability for complex structures and finally outputting binary masks and classification results for each structure.

[0085] Further, the Mask2Former model is used for segmenting the floor plan in step (4), specifically:

[0086] (4-1) Use the floor plan area extracted by the multi-layer feature fusion DETR model as the input of the multi-scale enhanced Mask2Former segmentation model. This input data contains accurately extracted target area information, providing a high-quality starting point for the segmentation model and ensuring that the segmentation focuses on the detailed processing of key structures in the floor plan;

[0087] (4-2) Mask2Former extracts multi-level features from the input data at different resolutions through a multi-scale feature pyramid module. The multi-scale structure of the feature pyramid can capture the detailed information of objects with different sizes in space, paying attention to the fine features of small-sized structures and also being able to comprehensively represent the global information of large-scale regions, thereby improving the model's perception ability for diverse objects;

[0088] (4-3) The extracted multi-scale features are input into the attention-focusing optimized Transformer decoder for processing. By introducing a more efficient multi-scale feature deep fusion and optimized feature integration strategy, the model realizes the deep fusion and optimization of multi-scale features. This integration strategy can enhance the cooperation between features and effectively improve the adaptability and processing ability of the segmentation model for complex target structures;

[0089] (4-4) Mask2Former combines the mask and query mechanisms to focus on local regions of the image through cross-attention. The mask is used to limit the calculation range of cross-attention within specific target regions to avoid interference from redundant information. At the same time, the model uses the self-attention mechanism to further strengthen the learned feature representation, ensuring that both the global and local features of the target can be fully modeled and improving the segmentation ability for complex floor plan structures;

[0090] (4-5) The improved segmentation model finally outputs the binary mask of each target structure and the corresponding classification result. The binary mask clearly identifies the boundary of each target area, while the classification result provides specific category information of the structure. This refined segmentation output provides accurate data support for subsequent visualization and analysis, and can significantly improve the overall performance of floor plan segmentation.

[0091] (5) Overlay the segmentation results (masks) of all structures to generate a complete segmentation map, and use different colors to label each category for easy visual inspection of the results. Through image synthesis technology, overlay the segmentation results on the original floor plan to form a semi-transparent effect for comparative analysis of the segmentation accuracy.

[0092] Furthermore, the specific steps of step (5) are as follows:

[0093] (5-1) Perform an overlay operation on the segmentation results (masks) of all structures to construct a complete segmentation map. During the generation process, use different colors to label each category, which helps to clearly and intuitively view the segmentation results from a visual level, distinguish and identify various structures in the map, and facilitate subsequent observation and evaluation of the segmentation situation;

[0094] (5-2) By means of image synthesis technology, perform a synthesis process on the obtained segmentation results and the original floor plan, so that the segmentation results can be accurately overlaid on the original floor plan to obtain accurate floor plan segmentation results, as Figure 5 shown.

[0095] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously.

[0096] Based on the same idea as the high-accuracy floor plan inner contour segmentation method in the above embodiments, the present invention also provides a high-accuracy floor plan inner contour segmentation system, which can be used to execute the above high-accuracy floor plan inner contour segmentation method. For the sake of illustration, in the structural schematic diagram of the high-accuracy floor plan inner contour segmentation system embodiment, only the parts related to the embodiments of the present invention are shown. Those skilled in the art can understand that the illustrated structure does not constitute a limitation on the device, and may include more or fewer components than those illustrated, or combine certain components, or have different component arrangements.

[0097] Please refer to Figure 6, in another embodiment of the present application, a floor plan inner contour segmentation system 100 with high accuracy is provided. The system includes a data preprocessing module 101, a floor plan target detection module 102, a floor plan extraction module 103, a multi-scale feature segmentation module 104, and a post-processing module 105;

[0098] The data preprocessing module 101 is used to obtain a floor plan image, annotate the key structures in the floor plan image, perform data augmentation on the obtained floor plan image, and preprocess the floor plan image after data augmentation;

[0099] The floor plan target detection module 102 is used to input the preprocessed floor plan image into a multi-layer feature fusion DETR model. The multi-layer feature fusion DETR model includes a backbone network, an improved encoder, and an improved decoder. The processing of the floor plan image by the multi-layer feature fusion DETR model is specifically as follows: The backbone network extracts low-level image features with more boundary information through a convolutional neural network. The improved encoder uses a multi-scale feature processing mechanism to perform feature enhancement processing on the extracted low-level image features with more boundary information, and further performs cross-scale feature integration on the enhanced features in the improved encoder to achieve interaction of features within the scale and fusion of features between scales; The fused features are input into the improved decoder to output the bounding box coordinates, confidence, and category of each floor plan;

[0100] The floor plan extraction module 103 is used to combine the bounding box generated by the multi-layer feature fusion DETR model with the original image, randomly expand the bounding box, and extract the floor plan area to eliminate the interference of watermarks or advertisement stickers in the original image on the segmentation result;

[0101] The multi-scale feature segmentation module 104 is used to fill the detection result of the improved DETR model to the size of the original image as the input of the multi-scale enhanced Mask2Former segmentation model. The multi-scale enhanced Mask2Former segmentation model extracts features of different scales through a multi-scale feature pyramid to more comprehensively capture targets of different sizes; Based on a preset multi-scale feature depth fusion and optimized feature integration strategy, the features of different scales are input into the optimized Transformer decoder, combined with the design of mask and query, and cross-attention is used to focus on local areas, while the self-attention mechanism is used to further strengthen the feature representation, thereby improving the segmentation ability for complex structures, and finally outputting the binary mask and segmentation result of each structure;

[0102] The post-processing module 105 is used to superimpose the segmentation results of all structures to generate a complete segmentation map.

[0103] It should be noted that the high-accuracy floor plan inner contour segmentation system of the present invention corresponds one-to-one with the high-accuracy floor plan inner contour segmentation method of the present invention. The technical features and beneficial effects described in the embodiments of the above high-accuracy floor plan inner contour segmentation method are applicable to the embodiments of the high-accuracy floor plan inner contour segmentation. For specific content, reference can be made to the description in the method embodiments of the present invention, which will not be elaborated here. This is hereby declared.

[0104] In addition, in the implementation manner of the high-accuracy floor plan inner contour segmentation system in the above embodiments, the logical division of each program module is only an example. In actual applications, according to needs, for example, considering the configuration requirements of the corresponding hardware or the convenience of software implementation, the above functions can be assigned to different program modules to complete, that is, the internal structure of the high-accuracy floor plan inner contour segmentation system is divided into different program modules to complete all or part of the functions described above.

[0105] Please refer to Figure 7 , in an embodiment, an electronic device for implementing a high-accuracy floor plan inner contour segmentation method is provided. The electronic device 200 may include a first processor 201, a first memory 202, and a bus, and may also include a computer program stored in the first memory 202 and executable on the first processor 201, such as a high-accuracy floor plan inner contour segmentation program 203.

[0106] Among them, the first memory 202 includes at least one type of readable storage medium. The readable storage medium includes flash memory, mobile hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the first memory 202 may be an internal storage unit of the electronic device 200, such as the mobile hard disk of the electronic device 200. In other embodiments, the first memory 202 may also be an external storage device of the electronic device 200, such as a plug-in mobile hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the electronic device 200. Further, the first memory 202 may also include both the internal storage unit and the external storage device of the electronic device 200. The first memory 202 can be used not only to store application software installed in the electronic device 200 and various types of data, such as the code of the high-accuracy floor plan inner contour segmentation program 203, but also to temporarily store data that has been output or will be output.

[0107] In some embodiments, the first processor 201 may be composed of an integrated circuit. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple packaged integrated circuits with the same or different functions, including a combination of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The first processor 201 is the control core (Control Unit) of the electronic device, connecting various components of the entire electronic device through various interfaces and circuits, and by running or executing programs or modules stored in the first memory 202, and calling data stored in the first memory 202, to execute various functions of the electronic device 200 and process data.

[0108] Figure 7 Only the electronic device with components is shown. Those skilled in the art can understand that Figure 7 The shown structure does not constitute a limitation on the electronic device 200, and may include fewer or more components than shown, or combine certain components, or have a different component arrangement.

[0109] The high-accuracy floor plan inner contour segmentation program 203 stored in the first memory 202 of the electronic device 200 is a combination of multiple instructions. When running in the first processor 201, it can achieve:

[0110] Obtain a floor plan image, annotate the key structures in the floor plan image, perform data augmentation on the obtained floor plan image, and preprocess the data-augmented floor plan image;

[0111] Input the preprocessed floor plan image into a multi-layer feature fusion DETR model. The multi-layer feature fusion DETR model includes a backbone network, an improved encoder, and an improved decoder. The processing of the floor plan image by the multi-layer feature fusion DETR model is as follows: The backbone network extracts low-level image features with more boundary information through a convolutional neural network. The improved encoder uses a multi-scale feature processing mechanism to perform feature enhancement processing on the extracted low-level image features with more boundary information, and further performs cross-scale feature integration on the enhanced features in the improved encoder to achieve intra-scale feature interaction and inter-scale feature fusion; Input the fused features into the improved decoder to output the bounding box coordinates, confidence, and category of each floor plan;

[0112] Combine the bounding boxes generated by the multi-layer feature fusion DETR model with the original image, randomly expand the bounding boxes, and extract the floor plan area to eliminate the interference of watermarks or advertisement stickers in the original image on the segmentation result;

[0113] The detection results of the improved DETR model are filled to the size of the original image and used as the input of the multi-scale enhanced Mask2Former segmentation model. The multi-scale enhanced Mask2Former segmentation model extracts features of different scales through a multi-scale feature pyramid to more comprehensively capture targets of various sizes. Based on the preset multi-scale feature depth fusion and optimized feature integration strategy, features of different scales are input into the optimized Transformer decoder. Combining the designs of masks and queries, it focuses on local regions through cross-attention and further strengthens feature representation using the self-attention mechanism, thereby enhancing the segmentation ability for complex structures. Finally, binary masks and segmentation results of each structure are output.

[0114] The segmentation results of all structures are superimposed to generate a complete segmentation map.

[0115] Furthermore, if the modules / units integrated in the electronic device 200 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory).

[0116] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it may include the processes of the embodiments of the above methods. Among them, any reference to memory, storage, database, or other media used in the various embodiments provided in the present application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0117] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0118] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.

Claims

1. A high-accuracy method for segmenting the inner contour of a house plan, characterized in that: The steps include: Acquire apartment images, annotate key structures in the apartment images, perform data enhancement on the acquired apartment images, and preprocess the apartment images after data enhancement; The preprocessed house type image is input into the multi-layer feature fusion DETR model, which includes a backbone network, an improved encoder and an improved decoder; The multi-layer feature fusion DETR model processes apartment images as follows: the backbone network extracts low-level image features with more boundary information through a convolutional neural network, and the improved encoder uses a multi-scale feature processing mechanism to perform feature enhancement on the extracted low-level image features with more boundary information. The enhanced features are further integrated across scales in the improved encoder to achieve the interaction of intra-scale features and the fusion of inter-scale features; the fused features are input into the improved decoder to output the bounding box coordinates, confidence, and category of each apartment map; Combine the bounding box generated by the multi-layer feature fusion DETR model with the original image, randomly expand the bounding box, and extract the floor plan area to eliminate the interference of watermarks or advertising stickers in the original image on the segmentation result; The detection results of the multi-layer feature fusion DETR model are filled with the original image size as the input of the multi-scale enhanced Mask2Former segmentation model. The multi-scale enhanced Mask2Former segmentation model extracts features of different scales through a multi-scale feature pyramid to more comprehensively capture targets of different sizes; based on the preset multi-scale feature deep fusion and optimized feature integration strategy, features of different scales are input into the optimized Transformer decoder, combined with the design of mask and query, cross-attention is used to focus on local areas, and the self-attention mechanism is used to further strengthen the feature representation, thereby improving the segmentation ability of complex structures, and finally outputting the binary mask and segmentation results of each structure; The segmentation results of all structures are superimposed to generate a complete segmentation map.

2. The method for segmenting the inner contour of a house plan with high accuracy according to claim 1, characterized in that: The method of obtaining a house type image, annotating key structures in the house type image, performing data enhancement on the obtained house type image, and preprocessing the house type image after data enhancement is specifically as follows: Receive the apartment image transmitted by the user end, and manually or automatically mark the key structures in the apartment image; Perform data augmentation operations on the annotated apartment images, including large-scale jitter, random flipping, image sharpening, and adding rock noise to enrich data diversity and improve the robustness of the model; All enhanced floor plan images are proportionally adjusted to a uniform size within a set pixel range to ensure accurate labeling of categories in the dataset, balanced category distribution, and consistent image size, thus forming a preprocessed floor plan dataset.

3. The method for segmenting the inner contour of a house plan with high accuracy according to claim 1, characterized in that: The processing of apartment images by the multi-layer feature fusion DETR model includes the following steps: The preprocessed floor plan image is input into the multi-layer feature fusion DETR model. The multi-layer feature fusion DETR model extracts low-level image features through a convolutional neural network. The low-level image features include edge information, texture features, and basic shape descriptions of the image, laying the foundation for the subsequent extraction of higher-level features. The extracted low-level image features are input into the improved encoder, which uses a multi-scale feature processing mechanism to enhance the low-level image features. By encoding feature layers of different scales, enhanced image features are generated, thereby fully capturing the spatial relationship and structural details of the objects in the floor plan. The enhanced image features are further integrated across scales in the encoder. Through a specially designed scale fusion interaction module, the interaction of intra-scale features and the fusion of inter-scale features are achieved, ensuring that the extracted image features are optimized both globally and locally. After the multi-layer feature generation, the multi-layer feature fusion DETR model introduces a refined feature selection strategy, which enables the multi-layer feature fusion DETR model to filter out high-quality features from the encoded features and suppress redundant or low-confidence features; The filtered high-quality features are sent to the improved decoder, which adopts a hierarchical structure and gradually refines the target features through multi-level decoding operations. The improved decoder adds an auxiliary prediction mechanism, which can provide additional constraints through auxiliary tasks to further improve the accuracy and stability of the decoding output; The improved decoder outputs the detection results of each object, including the bounding box coordinates, category and confidence of the object.

4. The method for segmenting the inner contour of a house plan with high accuracy according to claim 1, characterized in that: The bounding box generated by the multi-layer feature fusion DETR model is combined with the original image, the bounding box is randomly expanded, the floor plan area is extracted, and the interference of watermarks or advertising stickers in the original image on the segmentation result is avoided. Specifically: The bounding box generated by the multi-layer feature fusion DETR model is combined with the original image to preserve the spatial location information and structural features in the image by superimposing the bounding box on the original image. Based on the combination of the bounding box, a random expansion operation is performed on it. The purpose of random expansion is to increase the area coverage based on the original size of the bounding box to cope with the problem of boundary offset or incomplete definition of the target area that may occur in the model during the detection process.

5. The method for segmenting the inner contour of a house plan with high accuracy according to claim 1, characterized in that: The multi-scale enhanced Mask2Former segmentation model processes the detection results as follows: The floor plan area extracted by the multi-layer feature fusion DETR model is used as the input of the multi-scale enhanced Mask2Former segmentation model; The multi-scale enhanced Mask2Former segmentation model extracts multi-level features at different resolutions of the input data through a multi-scale feature pyramid module; The extracted multi-scale features are input into the attention-focused optimized Transformer decoder for processing. By introducing a more efficient multi-scale feature deep fusion and optimized feature integration strategy, the deep fusion and optimization of multi-scale features are achieved. Mask2Former combines mask and query mechanisms to achieve focus on local areas of the image through cross-attention. Mask is used to limit the calculation range of cross-attention to a specific target area to avoid interference from redundant information. At the same time, the model uses the self-attention mechanism to further strengthen the learned feature representation, ensuring that both the global and local features of the target can be fully modeled, thereby improving the segmentation capability of complex apartment structures. The improved segmentation model finally outputs a binary mask of each target structure and the corresponding classification result.

6. The method for segmenting the inner contour of a house plan with high accuracy according to claim 5, characterized in that: The feature integration strategy of multi-scale feature deep fusion and optimization is specifically as follows: Firstly, the adaptive weighting module is used to preprocess the features of different scales. The weights are automatically assigned according to the importance of the features in spatial position and semantic information. The weights of small-scale features are increased in the target details, while the weights of large-scale features are reasonably weighted in the background and overall structure areas. Next, an attention-guided feature fusion unit is introduced to calculate the feature correlation score, thereby strengthening the fusion of high-correlation areas and weakening low-correlation areas, such as wall structures, to enhance the coherence fusion of features at different scales; Finally, a progressive fusion method is adopted to gradually fuse large-scale features to small-scale features, and a small convolutional neural network is used for calibration and optimization after each fusion to ensure the accuracy of the fused feature space and semantics.

7. The method for segmenting the inner contour of a house plan with high accuracy according to claim 1, characterized in that: The segmentation results of all structures are superimposed to generate a complete segmentation map, specifically: The segmentation results of all structures are superimposed to construct a complete segmentation map. During the generation process, each category is marked with different colors to clearly and intuitively view the segmentation results from a visual level, so that various structures can be distinguished and identified in the map for subsequent observation and evaluation of the segmentation situation; With the help of image synthesis technology, the obtained segmentation results are synthesized with the original floor plan, so that the segmentation results can be accurately superimposed on the original floor plan to obtain accurate floor plan segmentation results.

8. A high-accuracy house plan inner contour segmentation system, characterized by: A high-accuracy method for segmenting inner contours of a floor plan applied to any one of claims 1-7, comprising a data preprocessing module, a floor plan target detection module, a floor plan extraction module, a multi-scale feature segmentation module and a post-processing module; The data preprocessing module is used to obtain the apartment image, mark the key structures in the apartment image, perform data enhancement on the obtained apartment image, and preprocess the apartment image after data enhancement; The floor plan target detection module is used to input the preprocessed floor plan image into a multi-layer feature fusion DETR model, and the multi-layer feature fusion DETR model includes a backbone network, an improved encoder and an improved decoder; The multi-layer feature fusion DETR model processes apartment images as follows: the backbone network extracts low-level image features with more boundary information through a convolutional neural network, and the improved encoder uses a multi-scale feature processing mechanism to perform feature enhancement on the extracted low-level image features with more boundary information. The enhanced features are further integrated across scales in the improved encoder to achieve the interaction of intra-scale features and the fusion of inter-scale features; the fused features are input into the improved decoder to output the bounding box coordinates, confidence, and category of each apartment map; The floor plan extraction module is used to combine the bounding box generated by the multi-layer feature fusion DETR model with the original image, randomly expand the bounding box, and extract the floor plan area to eliminate the interference of watermarks or advertising stickers in the original image on the segmentation result; The multi-scale feature segmentation module is used to fill the detection results of the improved DETR model with the original image size as the input of the multi-scale enhanced Mask2Former segmentation model. The multi-scale enhanced Mask2Former segmentation model extracts features of different scales through a multi-scale feature pyramid to more comprehensively capture targets of different sizes; based on the preset multi-scale feature deep fusion and optimized feature integration strategy, the features of different scales are input into the optimized Transformer decoder, and the design of mask and query is combined to focus on the local area through cross attention, and the self-attention mechanism is used to further strengthen the feature representation, thereby improving the segmentation ability of complex structures, and finally outputting the binary mask and segmentation result of each structure; The post-processing module is used to superimpose the segmentation results of all structures to generate a complete segmentation map.

9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores computer program instructions that can be executed by the at least one processor, and the computer program instructions are executed by the at least one processor so that the at least one processor can execute the high-accuracy method for segmenting the inner contour of a floor plan as described in any one of claims 1-7.

10. A computer-readable storage medium storing a program, characterized in that: When the program is executed by a processor, the high-accuracy method for segmenting the inner contour of a floor plan described in any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Field unstructured road real-time segmentation system and method for agricultural machinery autonomous navigation

    CN120558196A