Real-time house type image outer contour detection and segmentation method, system and equipment

Through the improved DETR and SAM models, combined with sparse attention mechanism and feature pyramid network, the problem of traditional methods being difficult to accurately extract the outer contour of complex floor plans is solved, and efficient and accurate detection and segmentation of floor plans are achieved.

CN120182300APending Publication Date: 2025-06-20SOUTH CHINA AGRICULTURAL UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510105854.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

Traditional floor plan processing methods are inefficient and it is difficult to accurately extract the outer contours of complex floor plan, especially in the presence of fuzzy boundaries or diverse styles.

Method used

The improved DETR object detection model and SAM segmentation model are adopted. The DETR model is used for real-time detection through the Transformer architecture to identify the main structural areas of the floor plan. The SAM model is subjected to high-precision segmentation processing, combining the sparse attention mechanism and feature pyramid network to enhance the detection and segmentation capabilities of targets at different scales.

Benefits of technology

It realizes accurate detection and segmentation of the outer contours of complex floor plans, improves processing efficiency and segmentation quality, and is suitable for large-scale data or real-time image segmentation application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182300A_ABST
    Figure CN120182300A_ABST
Patent Text Reader

Abstract

The invention discloses a real-time house type image outer contour detection and segmentation method, system and equipment, and the method comprises the steps: employing a two-stage processing flow, firstly carrying out the rapid detection of an input image, namely a house type image region, through employing an improved DETR target detection model, and positioning the rectangular position of a house; using the rectangular position as a prompt input of a subsequent segmentation model so as to provide an accurate target area for a subsequent segmentation step; and carrying out high-fineness segmentation processing on the extracted outer contour region by using an improved SAM segmentation model. According to the method, the rapid target detection capability of the DETR model and the high-quality segmentation performance of the SAM model are combined, the method is suitable for a large-scale building house type drawing outer contour segmentation task, can be widely applied to the fields of intelligent house property analysis, building design assistance, house type drawing automatic archiving and the like, shows higher robustness and precision in the task, and can be widely applied to large-scale building house type drawing outer contour segmentation. And the automation level of building design and analysis can be obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision, and particularly relates to a real-time method, system and device for detecting and segmenting the outer contour of a house floor plan. Background Art

[0002] In recent years, with the rapid development and digital transformation of the construction industry, the demand for automatic processing of house floor plans has increased significantly. As an important data source in architectural design, real estate marketing and intelligent analysis, house floor plans usually contain complex information such as house layouts, wall structures, door and window distributions, etc. Traditionally, the processing and analysis of house floor plans mostly rely on manual operations, but this method is inefficient, costly, and extremely vulnerable to human factors, making it difficult to meet the current demand for large-scale data processing. Therefore, how to use advanced computer vision and deep learning technologies to achieve automatic detection and segmentation of house floor plans has become a research hotspot. In the field of computer vision, object detection and image segmentation technologies are the core of automatic house floor plan processing. Traditional image processing methods, such as techniques based on edge detection, contour extraction and threshold segmentation, although can achieve preliminary segmentation in simple images, perform poorly on interference factors such as fine structures, background noise and watermarks in complex house floor plans. Especially in the case of fuzzy boundaries or diverse styles, traditional methods often have difficulty in accurately extracting the outer contour of the house floor plan. Therefore, in recent years, deep learning-based algorithms have been introduced to utilize the feature extraction and semantic understanding capabilities of neural networks to significantly improve the accuracy of detection and segmentation. Summary of the Invention

[0003] The main objective of the present invention is to overcome the disadvantages and deficiencies of the prior art, and provide a real-time method, system and device for detecting and segmenting the outer contour of a house floor plan. The present invention uses an improved DETR object detection model to detect house floor plan objects. DETR uses a real-time detection model based on the Transformer architecture to detect the main areas of the house floor plan. The DETR model can effectively identify the main structural areas in the image and locate the boundary range of the entire house floor plan through object detection boxes. After the detection of the target area is completed, an improved SAM segmentation model is used to perform high-precision segmentation processing on the extracted outer contour area.

[0004] To achieve the above objective, the present invention adopts the following technical solutions:

[0005] In the first aspect, the present invention provides a real-time method for detecting and segmenting the outer contour of a house floor plan, including the following steps:

[0006] Perform preliminary image preprocessing on the input house floor plan image;

[0007] Use the improved DETR object detection model to detect the object of the house floor plan, locate the boundary range of the entire house floor plan through the object detection frame, and obtain the detection result of the house floor plan area; the improved DETR object detection model uses the lightweight backbone network ConvNeXt, and this lightweight backbone network ConvNeXt adopts a dynamic attention mechanism, uses a sparse attention calculation method to sparsify the attention matrix, so as to reduce the computational complexity and the number of parameters. At the same time, a feature pyramid network is integrated on the original DETR model to extract multi-scale features to enhance the detection ability for objects of different scales;

[0008] On the basis of obtaining the detection result of the house floor plan area in the object detection stage, use the improved SAM segmentation model to finely segment the outer contour to obtain the segmentation result of the house floor plan area; the improved SAM segmentation model optimizes the segmentation head on the basis of SAM, introduces a refinement network, and at the same time uses a higher-resolution feature map and a stronger boundary optimization strategy. At the same time, the segmentation loss of SAM is also improved; the improved segmentation loss is the sum of the binary cross-entropy loss BCE Loss and the Dice Loss;

[0009] On the basis of obtaining the segmentation result of the house floor plan area in the object segmentation stage, automatically annotate and optimize the processing result based on a preset continuous learning module. After each segmentation, perform secondary segmentation on the output result and the original image, and finally output a high-quality segmented image;

[0010] Post-process the output high-quality segmented image, and finally obtain the final output result of the house floor plan.

[0011] As a preferred technical solution, the preliminary image preprocessing of the input house floor plan image is specifically as follows:

[0012] Perform grayscale processing and edge enhancement on the input house floor plan image to highlight the important areas in the image;

[0013] Apply an image filtering method to remove slight noise and perform normalization processing to make the pixel value distribution uniform.

[0014] As a preferred technical solution, the sparse attention calculation method is used to sparsify the attention matrix, and the following calculation formula is adopted:

[0015]

[0016] Among them, SparseAttention(Q, K, V) represents a sparse attention mechanism used to reduce the computational complexity of the fully connected layer in the traditional attention mechanism; Q represents the query matrix, indicating the information at the current time step or position; K represents the key matrix, indicating the information of all time steps or positions; V represents the value matrix, indicating the actual information associated with the key; d k represents the dimension of the key, the feature dimension of the matrix K, which is used for scaling processing to prevent the dot product value from being too large; M is the sparsity mask matrix; φ(Qi) and ψ(Kj) respectively represent a certain non-linear transformation of the query vector and the key vector, α ij represents the query vector Q i and the key vector K j of the attention weight; Softmax normalizes the weight.

[0017] As a preferred technical solution, the feature pyramid network performs multi-scale features using the following calculation formula:

[0018]

[0019] Among them, F l-1 represents the feature map from the l-1 layer; F bottom represents the bottom feature of the current layer l generated after the convolution operation Conv; Conv() is the convolution operation used to extract features; Φ() represents the feature fusion function used to perform weighted combination of features from different sources; F topl represents the top feature of the current layer l, which can be obtained from a higher layer or upsampled features; α represents the fusion weight coefficient, controlling the contribution ratio of the top feature and the bottom feature; Ψ() represents a constraint function; F fused represents the fused feature obtained through the feature fusion function Φ; Conv(F fused ) represents performing a convolution operation on the fused feature to further extract deep information or generate the final feature.

[0020] As a preferred technical solution, in the improved SAM segmentation model, the specific improvements are as follows:

[0021] A convolution layer sensitive to boundaries and an attention mechanism are added to the segmentation head to strengthen the ability to capture edge information;

[0022] The refinement network gradually eliminates the roughness in the segmentation through an iterative optimization mechanism, improving the continuity and clarity of the edges;

[0023] By introducing a multi-scale feature fusion strategy, more details can be captured at different scales, so that key texture information can be effectively retained in the segmentation. The feature fusion is performed using the following formula:

[0024]

[0025] Among them, F fused represents the final output feature fused from L feature layers; F l represents the feature map of the l-th layer; L represents the number of layers participating in the fusion; β l is the weight coefficient of the feature of the l-th layer, which determines the contribution of the feature of this layer in the fusion process; β l represents the normalized weight calculated by using the Softmax function to ensure that the sum of the weights of all layers is 1; W l represents the learnable parameter of the feature of the l-th layer, which is used to represent the importance of this layer;

[0026] For the boundary optimization of complex scenarios, combined with the gradient-based edge detection algorithm and morphological post-processing operations, the segmentation accuracy of the model for edges is further improved.

[0027] As a preferred technical solution, the continuous learning module compares the segmentation result generated by the model with the original image, and performs secondary optimization on the segmented region through an automatic annotation and feedback mechanism to ensure the accuracy of the segmentation result; after each segmentation, the preliminary result will be used to further segment the original image, focusing on processing complex boundary regions or parts with missing details.

[0028] As a preferred technical solution, the post-processing of the output high-quality segmentation image is specifically as follows:

[0029] Crop the segmentation image to remove the edge watermark and text marking area.

[0030] In a second aspect, the present invention provides a real-time floor plan outer contour detection and segmentation system, which is applied to the real-time floor plan outer contour detection and segmentation method, and includes an input image preprocessing module, a target detection model module, a segmentation model module, a post-processing and continuous learning module, and a segmentation result output module;

[0031] The input image preprocessing module is used to perform preliminary image preprocessing on the input floor plan image;

[0032] The target detection model module is used to detect the target of the house floor plan using an improved DETR target detection model, locate the boundary range of the entire house floor plan through the target detection frame, and obtain the detection result of the house floor plan area; the improved DETR target detection model adopts a lightweight backbone network ConvNeXt, and this lightweight backbone network ConvNeXt adopts a dynamic attention mechanism, uses a sparse attention calculation method to sparsify the attention matrix, so as to reduce the computational complexity and the number of parameters. At the same time, a feature pyramid network is integrated on the original DETR model to extract multi-scale features to enhance the detection ability for targets of different scales;

[0033] The segmentation model module is used to, on the basis of obtaining the detection result of the house floor plan area in the target detection stage, use an improved SAM segmentation model to finely segment the outer contour to obtain the segmentation result of the house floor plan area; the improved SAM segmentation model optimizes the segmentation head on the basis of SAM, introduces a refinement network, and at the same time uses a higher-resolution feature map and a stronger boundary optimization strategy. At the same time, the segmentation loss of SAM is also improved; the improved segmentation loss is the sum of the binary cross-entropy loss BCE Loss and the Dice Loss;

[0034] The post-processing and continuous learning module is used to, on the basis of obtaining the segmentation result of the house floor plan area in the target segmentation stage, automatically annotate and optimize the processing result based on a preset continuous learning module. After each segmentation, the output result and the original image are subjected to secondary segmentation to finally output a high-quality segmented image;

[0035] The segmentation result output module is used to post-process the output high-quality segmented image to finally obtain the final output result of the house floor plan.

[0036] In a third aspect, the present invention provides an electronic device, and the electronic device includes:

[0037] At least one processor; and,

[0038] A memory communicatively connected to the at least one processor; wherein,

[0039] The memory stores computer program instructions executable by the at least one processor, and the computer program instructions are executed by the at least one processor so that the at least one processor can execute the real-time outer contour detection and segmentation method of the house floor plan.

[0040] In a fourth aspect, the present invention provides a computer-readable storage medium storing a program, and when the program is executed by a processor, the real-time outer contour detection and segmentation method of the house floor plan is implemented.

[0041] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0042] 1. The present invention adopts a technical solution that combines an improved DETR model and an improved SAM model, which can more accurately detect and segment the boundaries of floor plan drawings. Compared with traditional segmentation methods, the improved SAM processes details and complex edges more precisely, especially suitable for floor plan drawings with complex structures, ensuring clear and smooth edges.

[0043] 2. The DETR model of the present invention can achieve end-to-end object detection, reducing the complex processes of multi-step preprocessing and postprocessing in traditional methods, greatly shortening the processing time, and meeting the real-time requirements; the SAM model optimizes the inference speed on the premise of ensuring high-precision segmentation. While maintaining the segmentation quality, it improves the processing efficiency of the model, suitable for large-scale data or real-time image segmentation application scenarios.

[0044] 3. The present invention introduces a continuous learning mechanism. By automatically annotating and iteratively updating the training set of the model, the model has dynamic adaptability. As the dataset becomes richer, the model can better adapt to diverse floor plan styles and structures, enhancing the generalization ability; due to the excellent performance of the model under complex boundaries, overlapping regions, and noise interference, the present invention is applicable to various application scenarios such as architectural design assistance, real estate analysis, and image archiving, with broad industrial application potential.

[0045] 4. The present invention realizes an automated process for floor plan detection and segmentation. Compared with traditional methods that rely on manual annotation and segmentation, it significantly reduces the labor cost and human error; the continuous learning module can automatically generate annotation data and iteratively optimize the model during the process of generating the training set, further reducing the need for manual participation and improving the processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0047] Figure 1 is a flowchart of the method for detecting and segmenting the outer contour of a floor plan in an embodiment of the present invention;

[0048] Figure 2 is a model diagram of the improved DETR in an embodiment of the present invention;

[0049] Figure 3 is a model diagram of the improved SAM in an embodiment of the present invention;

[0050] Figure 4 is the original floor plan of the embodiment of the present invention;

[0051] Figure 5 is the result diagram after being detected by the target detection model in the embodiment of the present invention;

[0052] Figure 6 is the result diagram after being detected by the segmentation model in the embodiment of the present invention;

[0053] Figure 7 is the structural schematic diagram of the real-time floor plan outline detection and segmentation system in the embodiment of the present invention;

[0054] Figure 8 is the structural schematic diagram of the electronic device in the embodiment of the present invention. Detailed implementation manners

[0055] In order to enable those skilled in the art of the present technology to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.

[0056] In the present application, referring to "embodiment" means that the specific features, structures or characteristics described in combination with the embodiment may be included in at least one embodiment of the present application. The phrase appears in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art understand explicitly and implicitly that the embodiments described in the present application may be combined with other embodiments.

[0057] Please refer to Figure 1 , in an embodiment of the present application, a real-time floor plan outline detection and segmentation method is provided, including the following steps:

[0058] (1) Perform preliminary image preprocessing on the input floor plan image. Considering that the input image may contain irrelevant background regions (such as page edges, watermarks, marks or advertisements of pictures, etc.), and the requirements of the model for the size of the input picture. For the input original image, please refer to Figure 4 .

[0059] Furthermore, step (1) is specifically as follows:

[0060] First, we perform grayscale processing and edge enhancement on the image to highlight the important regions in the image. Subsequently, image filtering methods (such as Gaussian filtering, median filtering, etc.) are applied to remove minor noise, and further normalization is carried out to make the pixel value distribution uniform, which helps to improve the accuracy of subsequent detection and segmentation. In order to meet the requirements of the input of the DETR object detection model, the pixels of the image are processed.

[0061] (2) Use the improved DETR object detection model to detect the floor plan objects. Please refer to Figure 2 . First, the backbone is composed of ResNet to extract the basic features of the input image. Then, the data passes through an efficient hybrid encoder. This encoder consists of Adaptive Internal Feature Interaction (AIFI) and Cross-scale Feature Fusion (CCFM), which respectively fuse multi-scale features, extract cross-layer information and further process the fused features, and perform weighted optimization by combining cross-channel information. Then, the target queries are screened through the IoU-aware Query Selection perception mechanism to improve the detection efficiency. Finally, through the Decoder&Head module, the Transformer decoder module is used to process the queries and features.

[0062] Given that there are some irrelevant elements (such as watermarks, markings, etc.) in the input image, the detection results of DETR may be incomplete. To solve this problem, the model will set reasonable bounding boxes for the detection regions and perform multi-scale detection to ensure that the complete floor plan region is included as much as possible. For the candidate regions detected multiple times, in order to retain more edges and details, a strategy based on non-maximum suppression is adopted to screen the best detection results, and a tolerance boundary is set for the edge part to retain the floor plan information to the greatest extent. For the output results of the detection model, please refer to Figure 5 .

[0063] DETR uses a real-time detection model based on the Transformer architecture, but the Transformer architecture causes the model training time to be relatively long. Therefore, in this embodiment, some improvements are made on the basis of DETR, and the lightweight backbone network ConvNeXt is adopted. This network uses a dynamic attention mechanism, and its core is to sparsify the attention matrix. The calculation formula is as follows:

[0064]

[0065] Among them, SparseAttention(Q, K, V) represents a sparse attention mechanism used to reduce the computational complexity of the fully connected layer in traditional attention mechanisms (such as Transformer). Q represents the query matrix, which represents the information at the current time step or position. K represents the key matrix, which represents all time step or position information. V represents the value matrix, which represents the actual information associated with the key. d k represents the dimension of the key, the feature dimension of the matrix K, and is used for scaling processing to prevent the dot product value from being too large. M is the sparsity mask matrix. φ(Qi) and ψ(Kj) respectively represent a certain non-linear transformation of the query vector and the key vector, α ij represents the query vector Q i and the key vector K j of the attention weight. Softmax normalizes the weights.

[0066] Compared with DETR, it reduces the computational complexity and the number of parameters, and further improves the inference speed. Further, in order to improve the detection accuracy of the model for targets, in this embodiment, a feature pyramid network is integrated on DETR to extract multi-scale features and enhance the detection ability for targets of different scales. DETR detects the main regions of the floor plan. The model can effectively identify the main structural regions in the image and locate the boundary range of the entire floor plan through the target detection box.

[0067] How its feature pyramid network extracts multi-scale features is shown in the following formula:

[0068]

[0069] Among them, Fl-1 represents the feature map from the previous layer (layer l-1). F bottom represents the bottom feature of the current layer l generated after the convolution operation (Conv). Conv() is the convolution operation used to extract features. Φ() represents the feature fusion function used to perform weighted combination of features from different sources. F topl represents the top feature from the current layer l, which may be obtained from a higher layer or an upsampled feature. α represents the fusion weight coefficient, which controls the contribution ratio of the top feature and the bottom feature. Ψ() represents a constraint function. F fused represents the fused feature obtained through the feature fusion function Φ. Conv(F fused ) represents performing a convolution operation on the fused feature to further extract deep information or generate the final feature.

[0070] (3) Please refer to Figure 3, in this embodiment, the segmentation head of SAM is deeply improved. The improved SAM model consists of the original basic module SAM, that is, an image encoder and a mask decoder. Then, it consists of the improved modules HQ-OutputToken, global fusion module, and error correction module. Finally, the SAM mask is fused with the features of HQ-Features to finally generate a high-quality mask of HQ-SAM. Through these improved modules, convolutional layers sensitive to boundaries and attention mechanisms are added to strengthen the model's ability to capture edge information, enabling it to more precisely segment complex contours. Based on the original segmentation results, this embodiment designs and introduces a refinement network to further enhance the boundary details of the segmentation results. This network gradually eliminates the roughness in the segmentation through an iterative optimization mechanism, improving the continuity and clarity of the edges. To meet the requirements of high-precision tasks, this embodiment adjusts the model processing flow and adds support for high-resolution feature maps. By introducing a multi-scale feature fusion strategy, the model can capture more details at different scales, effectively retaining key texture information in the segmentation.

[0071] This embodiment also improves the segmentation loss of SAM to adapt to higher-quality boundary optimization or small-object segmentation. The improved loss function is as shown in formula (6):

[0072]

[0073] L dc (Dice Loss) represents the Dice loss, which is usually used to measure the overlap between the predicted segmentation and the ground truth segmentation. L budr represents the boundary loss, which focuses on the boundary region of the segmentation mask. λ dice is the weight coefficient of the Dice loss, used to adjust the importance of L dc in the total loss. λ boundary is the weight coefficient of the boundary loss, used to adjust the importance of L budr in the total loss. L sg represents the total segmentation loss.

[0074] (4) Based on obtaining the segmentation results of the house floor plan area in the target segmentation stage, the results after segmentation are further used for iterative optimization of the model. For the results of the segmentation model, please refer to Figure 6 . This system includes a continuous learning module that automatically annotates and optimizes the processing results. After each segmentation, the output results and the original image are segmented again. The high-quality segmentation image finally output by the model and the original image are stored in the database as further training data for the model, forming an adaptive training dataset.

[0075] Specifically, after obtaining the detection results of the house floor plan area in the target detection stage, this embodiment further introduces an improved Segment Anything Model (SAM) to perform refined segmentation on the outer contour. However, the original SAM model exhibits certain limitations when processing high-resolution images. Especially in cases such as dealing with complex boundaries and occlusion scenarios, its segmentation quality may be poor, specifically manifested as blurred boundaries or lost details.

[0076] This embodiment deeply improves the segmentation head of SAM, adding convolutional layers and attention mechanisms sensitive to boundaries to strengthen the model's ability to capture edge information, enabling it to more accurately segment complex contours. Based on the original segmentation results, a refinement network is designed and introduced to further enhance the boundary details of the segmentation results. This network gradually eliminates the roughness in the segmentation through an iterative optimization mechanism, improving the continuity and clarity of the edges.

[0077] To meet the requirements of high-precision tasks, we adjusted the model processing flow, adding support for high-resolution feature maps. By introducing a multi-scale feature fusion strategy, the model can capture more details at different scales, thereby effectively retaining key texture information in the segmentation. How it performs feature fusion is shown in formulas (7)(8).

[0078]

[0079] Among them, F fused represents the final output feature fused from L feature layers. F l represents the feature map of the l-th layer. L represents the number of layers participating in the fusion. β l is the weight coefficient of the l-th layer feature, determining the contribution of this layer feature in the fusion process. β l represents the normalized weight calculated using the Softmax function, ensuring that the sum of the weights of all layers is 1. W l represents the learnable parameter of the l-th layer feature, used to represent the importance of this layer.

[0080] For the boundary optimization in complex scenarios, we combined a gradient-based edge detection algorithm and morphological post-processing operations to further improve the model's segmentation accuracy for edges. This strategy effectively reduces the generation of pseudo-boundaries and accurately locates the target contour in complex backgrounds.

[0081] (5) Post-process the segmentation results output by the SAM model, that is, crop the image, remove irrelevant areas such as edge watermarks and text marks, and finally obtain the final output result of the house floor plan.

[0082] To enable the model to be continuously iteratively updated in the future, based on the original sample replay method of continual learning, a small amount of previous data is randomly saved, and this batch of mixed data is used to train the current model, which can alleviate the problem of insufficient training data and reduce the catastrophic forgetting that occurs during the iterative optimization of the model; based on the continual learning method, the idea of this method is to first store a part of the previous data, and at the same time make the direction of parameter update of the model similar to the update directions of all previous tasks (i.e., the included angle is an acute angle) when learning new tasks, and use episodic memory to minimize negative backward transfer. Specifically:

[0083] Save the data of the previously trained tasks, and define the memory of task k as M k ,on M k The loss function is:

[0084]

[0085] However, minimizing formula (2) will lead to overfitting of M k Therefore, transform the minimization of formula (9) into the solution formula (10):

[0086] minimize ι (f θ (x,k),y) (10)

[0087] where is the model after learning the previous task. It can be found that as long as it is ensured that the loss for the previous task does not increase after each parameter update, so in this embodiment, the angle between the loss gradient vectors of the previous tasks is calculated to judge the increase of the loss and the update direction under the previous tasks. Formula (11) can be rewritten as formula (12):

[0088]

[0089] If their included angle is an acute angle, then when learning the current task, the performance of task k will not increase. But it may not be the case that all included angles are acute angles. In this situation, this embodiment can project the gradient g onto the nearest gradient Thus, this embodiment obtains the following optimization objective:

[0090]

[0091] Based on the feedback iteration of continuous learning, when designing the docking of the detection model and the subsequent algorithm in this embodiment, a human-computer interaction function is added. The user can judge the marked images generated by the model according to the actual application. The model will select the modified partial images from the database as the data set for iterative update, and continuously iterate and update the model at the manually selected time interval to improve the discrimination accuracy of the entire system.

[0092] For the task of floor plan outer contour detection and segmentation, this application adopts a multi-module combination strategy. Among them, the improved DETR model is used for the floor plan detection model. The main advantage of this model is its high detection real-time performance, which can support the real-time processing requirements of floor plans, and at the same time ensure the accuracy of detection results to meet the requirements of the subsequent contour segmentation module of the floor plan.

[0093] To further improve the accuracy of detection and segmentation, the present invention introduces specific judgment conditions during the detection process, and specifically processes the image elements that may interfere with the segmentation accuracy. For example, for interference information such as noise, non-building edges, and watermarks in the floor plan, they are excluded through special judgment conditions, so as to ensure that the detected outer contour is clear and complete. Finally, through the floor plan contour detection and segmentation model, the effective boundary of the floor plan is accurately identified, irrelevant areas are removed, and high-quality contour data is provided for subsequent building analysis and drawing processing.

[0094] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously.

[0095] Based on the same idea as the real-time floor plan outer contour detection and segmentation method in the above embodiment, the present invention also provides a real-time floor plan outer contour detection and segmentation system, which can be used to execute the above real-time floor plan outer contour detection and segmentation method. For the sake of convenience of description, in the structural schematic diagram of the real-time floor plan outer contour detection and segmentation system embodiment, only the parts related to the embodiment of the present invention are shown. Those skilled in the art can understand that the illustrated structure does not constitute a limitation on the device, and may include more or fewer components than those illustrated, or combine certain components, or arrange different components.

[0096] Please refer to Figure 7 , in another embodiment of the present application, a real-time floor plan outer contour detection and segmentation system 100 is provided. The system includes an input image preprocessing module 101, a target detection model module 102, a segmentation model module 103, a post-processing and continuous learning module 104, and a segmentation result output module 105;

[0097] The input image preprocessing module 101 is used to perform preliminary image preprocessing on the input floor plan image;

[0098] The target detection model module 102 is used to detect the floor plan targets using an improved DETR target detection model, locate the boundary range of the entire floor plan through the target detection box, and obtain the floor plan area detection result; the improved DETR target detection model uses a lightweight backbone network ConvNeXt, and this lightweight backbone network ConvNeXt adopts a dynamic attention mechanism, uses a sparse attention calculation method to sparsify the attention matrix, so as to reduce the computational complexity and the number of parameters. At the same time, a feature pyramid network is integrated on the original DETR model to extract multi-scale features to enhance the detection ability for targets of different scales;

[0099] The segmentation model module 103 is used to perform fine segmentation on the outer contour using an improved SAM segmentation model on the basis of obtaining the floor plan area detection result in the target detection stage, and obtain the floor plan area segmentation result; the improved SAM segmentation model optimizes the segmentation head on the basis of SAM, introduces a refinement network, and at the same time uses a higher-resolution feature map and a stronger boundary optimization strategy. At the same time, the segmentation loss of SAM is also improved;

[0100] The post-processing and continuous learning module 104 is used to automatically annotate and optimize the processing result based on a preset continuous learning module on the basis of obtaining the floor plan area segmentation result in the target segmentation stage. After each segmentation, the output result and the original image are subjected to secondary segmentation to finally output a high-quality segmented image;

[0101] The segmentation result output module 105 is used to perform post-processing on the output high-quality segmented image, and finally obtain the final output result of the floor plan.

[0102] It should be noted that the real-time floor plan outer contour detection and segmentation system of the present invention corresponds one-to-one with the real-time floor plan outer contour detection and segmentation method of the present invention. The technical features and beneficial effects described in the embodiments of the above real-time floor plan outer contour detection and segmentation method are applicable to the embodiments of the real-time floor plan outer contour detection and segmentation. For specific content, reference can be made to the description in the method embodiments of the present invention, which will not be repeated here. This is hereby declared.

[0103] In addition, in the implementation of the real-time floor plan outline detection and segmentation system in the above embodiments, the logical division of each program module is only an example. In practical applications, according to needs, for example, considering the configuration requirements of the corresponding hardware or the convenience of software implementation, the above functions can be assigned to different program modules to complete, that is, the internal structure of the real-time floor plan outline detection and segmentation system is divided into different program modules to complete all or part of the functions described above.

[0104] Please refer to Figure 8 , in one embodiment, an electronic device for implementing a real-time floor plan outline detection and segmentation method is provided. The electronic device 200 may include a first processor 201, a first memory 202, and a bus, and may also include a computer program stored in the first memory 202 and executable on the first processor 201, such as a real-time floor plan outline detection and segmentation program 203.

[0105] Among them, the first memory 202 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, mobile hard disk, multimedia card, card-type memory (such as: SD or DX memory, etc.), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the first memory 202 may be an internal storage unit of the electronic device 200, such as the mobile hard disk of the electronic device 200. In other embodiments, the first memory 202 may also be an external storage device of the electronic device 200, such as a plug-in mobile hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the electronic device 200. Further, the first memory 202 may also include both the internal storage unit and the external storage device of the electronic device 200. The first memory 202 can be used not only to store application software installed on the electronic device 200 and various types of data, such as the code of the real-time floor plan outline detection and segmentation program 203, but also to temporarily store data that has been output or will be output.

[0106] In some embodiments, the first processor 201 may be composed of integrated circuits. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple packaged integrated circuits with the same or different functions, including one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and combinations of various control chips, etc. The first processor 201 is the control core (Control Unit) of the electronic device, connecting various components of the entire electronic device through various interfaces and circuits, and executing various functions of the electronic device 200 and processing data by running or executing programs or modules stored in the first memory 202 and calling data stored in the first memory 202.

[0107] Figure 8 Only the electronic device with components is shown. Those skilled in the art can understand that Figure 8 the shown structure does not constitute a limitation on the electronic device 200, and it may include fewer or more components than shown, or combine certain components, or have different component arrangements.

[0108] The real-time floor plan outline detection and segmentation program 203 stored in the first memory 202 of the electronic device 200 is a combination of multiple instructions. When running in the first processor 201, it can achieve:

[0109] Perform preliminary image preprocessing on the input floor plan image;

[0110] Use an improved DETR object detection model to detect floor plan objects, locate the boundary range of the entire floor plan through the object detection box, and obtain the floor plan area detection result; the improved DETR object detection model uses a lightweight backbone network ConvNeXt. This lightweight backbone network ConvNeXt adopts a dynamic attention mechanism, sparsifies the attention matrix using a sparse attention calculation method to reduce the computational complexity and the number of parameters. At the same time, a feature pyramid network is integrated on the original DETR model to extract multi-scale features to enhance the detection ability for objects of different scales;

[0111] Based on the floor plan area detection result obtained in the object detection stage, use an improved SAM segmentation model to finely segment the outer contour and obtain the floor plan area segmentation result; the improved SAM segmentation model optimizes the segmentation head on the basis of SAM, introduces a refinement network, and at the same time uses a higher-resolution feature map and a stronger boundary optimization strategy. At the same time, the segmentation loss of SAM is also improved;

[0112] Based on the result of the house type map area segmentation obtained in the target segmentation stage, the processing result is automatically labeled and optimized based on a preset continuous learning module. After each segmentation, the output result and the original image are subjected to secondary segmentation to finally output a high-quality segmented image;

[0113] The high-quality segmented image output is post-processed to finally obtain the final output result of the house type map.

[0114] Furthermore, if the modules / units integrated in the electronic device 200 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory).

[0115] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it may include the processes of the embodiments of the above methods. Among them, any reference to memory, storage, database or other media used in the various embodiments provided in the present application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0116] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0117] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.

Claims

1. A real-time floor plan outline detection and segmentation method, characterized in that: The steps include: Perform preliminary image preprocessing on the input floor plan image; An improved DETR target detection model is used to detect targets in a floor plan, and the boundary range of the entire floor plan is located by a target detection frame to obtain a detection result of the floor plan area; the improved DETR target detection model adopts a lightweight backbone network ConvNeXt, and the lightweight backbone network ConvNeXt adopts a dynamic attention mechanism and uses a sparse attention calculation method to sparse the attention matrix to reduce the complexity and parameter amount of the calculation, and at the same time integrates a feature pyramid network on the original DETR model to extract multi-scale features to enhance the detection capability of targets of different scales; On the basis of the floor plan area detection results obtained in the target detection stage, the improved SAM segmentation model is used to finely segment the outer contour to obtain the floor plan area segmentation result; the improved SAM segmentation model optimizes the segmentation head on the basis of SAM, introduces a refinement network, uses a higher resolution feature map and a stronger boundary optimization strategy, and also improves the segmentation loss of SAM; the improved segmentation loss is the sum of the two-class cross entropy loss BCE Loss and DiceLoss loss; Based on the segmentation results of the floor plan area obtained in the target segmentation stage, the processing results are automatically annotated and optimized based on the preset continuous learning module. After each segmentation, the output result and the original image are segmented twice to finally output a high-quality segmented image; The output high-quality segmented image is post-processed to obtain the final output result of the floor plan.

2. The real-time floor plan contour detection and segmentation method according to claim 1, characterized in that: The preliminary image preprocessing of the input floor plan image is specifically as follows: Grayscale and edge enhancement are performed on the input floor plan image to highlight the important areas in the image; Image filtering methods are applied to remove slight noise and normalization is performed to make the pixel values ​​evenly distributed.

3. The real-time floor plan outline detection and segmentation method according to claim 1, characterized in that: The sparse attention calculation method is used to sparse the attention matrix, and the following calculation formula is used: Among them, SparseAttention (Q, K, V) represents a sparse attention mechanism, which is used to reduce the computational complexity of full connection in the traditional attention mechanism; Q represents the query matrix, which represents the information of the current time step or position; K represents the key matrix, which represents the information of all time steps or positions; V represents the value matrix, which represents the actual information associated with the key; d k represents the dimension of the key, the characteristic dimension of the matrix K, which is used for scaling to prevent the dot product value from being too large; M is the sparsity mask matrix; φ(Qi) and ψ(Kj) represent some nonlinear transformations of the query vector and the key vector respectively, and α ij Denotes the query vector Q i and the key vector K j Attention weights; Softmax normalizes the weights.

4. The real-time floor plan outline detection and segmentation method according to claim 1, characterized in that: The feature pyramid network uses the following calculation formula for multi-scale features: Among them, F l-1 represents the feature map from layer l-1; F bottom represents the bottom feature of the current layer l generated after the convolution operation Conv; Conv() is a convolution operation used to extract features; Φ() represents a feature fusion function used to weightedly combine features from different sources; F topl represents the top feature from the current layer l, which can be obtained from higher layers or upsampled features; α represents the fusion weight coefficient, which controls the contribution ratio of the top feature and the bottom feature; Ψ() represents a constraint function; F fused Represents the fusion feature obtained by the feature fusion function Φ; Conv(F fused ) indicates that a convolution operation is performed on the fused features to further extract deep information or generate the final features.

5. The real-time floor plan outline detection and segmentation method according to claim 1, characterized in that: In the improved SAM segmentation model, the specific improvements are as follows: Add boundary-sensitive convolutional layers and attention mechanisms to the segmentation head to enhance the ability to capture edge information; The refinement network gradually eliminates the roughness in the segmentation and improves the continuity and clarity of the edges through an iterative optimization mechanism; By introducing a multi-scale feature fusion strategy, more details can be captured at different scales, thereby effectively retaining key texture information in segmentation. The following formula is used for feature fusion: Among them, F fused Represents the final output feature obtained by fusion of L feature layers; F l represents the feature map of the lth layer; L represents the number of layers involved in the fusion; β l The weight coefficient of the lth layer feature determines the contribution of the layer feature in the fusion process; β l W represents the normalized weight calculated by the Softmax function, ensuring that the sum of the weights of all layers is 1; l The learnable parameters representing the features of the lth layer are used to indicate the importance of this layer; Boundary optimization for complex scenes, combined with a gradient-based edge detection algorithm and morphological post-processing operations, further improves the model's edge segmentation accuracy.

6. The real-time floor plan outline detection and segmentation method according to claim 1, characterized in that: The continuous learning module compares the segmentation results generated by the model with the original image, and performs secondary optimization on the segmented area through automatic labeling and feedback mechanism to ensure the accuracy of the segmentation results; after each segmentation, the original image is further segmented using the preliminary results, focusing on processing areas with complex boundaries or parts with lost details.

7. The real-time floor plan outline detection and segmentation method according to claim 1, characterized in that: The post-processing of the output high-quality segmented image is specifically as follows: Crop the segmented image to remove edge watermarks and text marking areas.

8. Real-time floor plan outline detection and segmentation system, characterized by: A real-time floor plan outer contour detection and segmentation method applied to any one of claims 1-7, comprising an input image preprocessing module, a target detection model module, a segmentation model module, a post-processing and continuous learning module, and a segmentation result output module; The input image preprocessing module is used to perform preliminary image preprocessing on the input floor plan image; The target detection model module is used to detect the target of the floor plan using the improved DETR target detection model, locate the boundary range of the entire floor plan through the target detection frame, and obtain the detection result of the floor plan area; the improved DETR target detection model adopts a lightweight backbone network ConvNeXt, the lightweight backbone network ConvNeXt adopts a dynamic attention mechanism, and uses a sparse attention calculation method to sparse the attention matrix to reduce the complexity and parameter amount of the calculation, and integrates a feature pyramid network on the original DETR model to extract multi-scale features to enhance the detection capability of targets of different scales; The segmentation model module is used to obtain the floor plan area detection result in the target detection stage, and then use the improved SAM segmentation model to perform fine segmentation on the outer contour to obtain the floor plan area segmentation result; the improved SAM segmentation model optimizes the segmentation head on the basis of SAM, introduces a refinement network, uses a higher resolution feature map and a stronger boundary optimization strategy, and improves the segmentation loss of SAM; the improved segmentation loss is the sum of the two-class cross entropy loss BCE Loss and Dice Loss; The post-processing and continuous learning module is used to automatically annotate and optimize the processing results based on the preset continuous learning module on the basis of the segmentation results of the floor plan area obtained in the target segmentation stage, and after each segmentation, perform secondary segmentation on the output result and the original image to finally output a high-quality segmented image; The segmentation result output module is used to post-process the output high-quality segmentation image to finally obtain the final output result of the floor plan.

9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores computer program instructions that can be executed by the at least one processor, and the computer program instructions are executed by the at least one processor so that the at least one processor can execute the real-time floor plan outer contour detection and segmentation method as described in any one of claims 1-7.

10. A computer-readable storage medium storing a program, characterized in that: When the program is executed by a processor, the real-time floor plan outer contour detection and segmentation method described in any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Intelligent monitoring system and method for growth vigor of digital seedling-raising seedling tray based on SAM and improved RF-DETR

    CN121788778A