Space identification method and system based on YOLO model, terminal and medium
By combining the YOLOv8-seg model and the CGAL library, high-precision spatial recognition and automated annotation of architectural drawings are achieved, solving the problems of low efficiency and poor adaptability in existing technologies, and realizing high-precision mapping and automated processing from CAD drawings to architectural floor plans.
Patent Information
- Application Number
- CN202511084673.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-04
- Publication Date
- 2025-11-21
AI Technical Summary
Existing technologies have low spatial recognition efficiency and poor generalization ability in architectural drawings, making it difficult to adapt to complex layouts. Furthermore, it is difficult to reverse-map the recognition results back to the world coordinates of the DWG drawing and achieve semantic naming.
The YOLOv8-seg model is used for end-to-end training. Combined with the Python environment and CGAL library, it realizes spatial segmentation and coordinate transformation of CAD drawings, and outputs high-precision spatial region recognition and automatic annotation.
It achieves high-precision spatial recognition and classification from CAD drawings to architectural floor plans, has strong generalization ability and engineering practicality, supports scenarios with complex lines and inconsistent proportions, and has the ability to automate the entire process.
Smart Images

Figure CN120997866A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a spatial recognition method, system, terminal, and medium based on the YOLO model. Background Technology
[0002] In modern architectural design and intelligent operation and maintenance systems, the rapid and accurate identification of functional spaces such as rooms in CAD drawings is of great significance. Traditional methods rely on manual annotation or rule-based PNG image processing algorithms, which suffer from low efficiency, poor generalization ability, and difficulty in adapting to complex layouts.
[0003] In recent years, deep learning, especially the YOLO series models, has made significant progress in object detection and spatial segmentation. However, existing technologies are mostly focused on object recognition in general scenarios and have not yet effectively addressed the special challenges of complex spatial structures, dense lines, and inconsistent scales in architectural drawings. Furthermore, mapping the recognition results back to the world coordinates of the DWG drawing and achieving semantic naming remains a challenge for the industry.
[0004] Therefore, existing technologies still have shortcomings. Summary of the Invention
[0005] The technical problem to be solved by this invention is to provide a spatial recognition method, system, terminal, and medium based on the YOLO model, addressing the aforementioned deficiencies of the prior art. The technical solution adopted by this invention is as follows:
[0006] In a first aspect, the present invention provides a spatial recognition method based on the YOLO model, wherein the method includes:
[0007] Obtain CAD drawings and perform spatial segmentation and coordinate transformation on the CAD drawings to obtain PNG images corresponding to specified spatial regions in the CAD drawings and label files corresponding to the PNG images. Based on the PNG images and the label files, determine the training dataset. The label files are used to reflect the contours and categories corresponding to the segmented spatial regions.
[0008] Set up a Python runtime environment, install the YOLOv8-seg model and its related dependency libraries, and perform end-to-end training on the training dataset based on the training script of the YOLOv8-seg model to obtain the YOLO spatial segmentation model.
[0009] The architectural floor plan is spatially segmented based on the YOLO spatial segmentation model, the segmentation results are output, and the spatial areas are automatically labeled based on the segmentation results.
[0010] In one implementation, the step of spatially segmenting and transforming the CAD drawing to obtain a PNG image corresponding to a specified spatial region in the CAD drawing and a tag file corresponding to the PNG image includes:
[0011] Use polylines to draw the boundary of each spatial region in the CAD drawing, and record the coordinates of the polyline vertex of the specified spatial region;
[0012] Export a PNG image of a specified spatial region;
[0013] The coordinates of the vertices of the multi-line segment are transformed to obtain the tag file.
[0014] In one implementation, the step of performing coordinate transformation on the coordinates of the multi-segment vertices to obtain the tag file includes:
[0015] Obtain the world coordinate system of the print area of the CAD drawing;
[0016] Calculate the scaling factor based on the size of the PNG image and the size of the printing area of the CAD drawing;
[0017] Based on the scaling factor and the world coordinate system, a coordinate transformation matrix is determined, and the coordinate transformation of the vertex coordinates of the multi-segment is performed based on the coordinate transformation matrix to obtain the pixel coordinates of the PNG image;
[0018] After normalizing the pixel coordinates of the PNG image, the corresponding categories of the specified spatial regions are written into a YOLO format TXT file to obtain the tag file.
[0019] In one implementation, determining the training dataset based on the PNG image and the tag file includes:
[0020] The PNG images and the tag files are stored in a preset target structure to construct a training dataset that conforms to the YOLO format.
[0021] In one implementation, the step of spatially segmenting the architectural floor plan based on the YOLO spatial segmentation model and outputting the segmentation result includes:
[0022] Convert the YOLO spatial partitioning model into ONNX format;
[0023] Export the PNG image of the architectural floor plan and determine the world coordinate system of the architectural floor plan;
[0024] The PNG image of the architectural floor plan is preprocessed and inference is performed on the PNG image to obtain the bounding boxes of each spatial area in the architectural floor plan and their corresponding categories;
[0025] Based on the coordinate transformation matrix, the bounding boxes of each spatial region are mapped to the world coordinate system of the architectural plan drawing;
[0026] The multi-line segments extracted from the architectural floor plan are matched with the bounding boxes of each spatial region obtained through inference, and the segmentation result is output. The segmentation result includes the outline and category of each spatial region in the architectural floor plan.
[0027] In one implementation, the step of spatially segmenting the architectural floor plan based on the YOLO spatial segmentation model and outputting the segmentation result further includes:
[0028] Extract the geometric features of each spatial region and smooth the contours using the Snap Rounding algorithm from the CGAL library to reduce jagged edges and fragmented lines.
[0029] In one implementation, the automated annotation of spatial regions based on the segmentation results includes:
[0030] The OCR algorithm is used to recognize the text in architectural floor plans, and the BERT text classification model is used to perform semantic analysis on the text and extract semantic features.
[0031] The semantic features are matched with the categories of each spatial region, and the text with the closest semantics is selected from the semantic features as the final name of the corresponding spatial region in the architectural floor plan, thereby realizing the automatic labeling of spatial regions.
[0032] Secondly, embodiments of the present invention also provide a spatial recognition system based on the YOLO model, wherein the system is used to implement the steps of the spatial recognition method based on the YOLO model described in any of the above solutions, and the system includes:
[0033] The training dataset preparation module is used to acquire CAD drawings, perform spatial segmentation and coordinate transformation on the CAD drawings, obtain PNG images corresponding to specified spatial regions in the CAD drawings and label files corresponding to the PNG images, and determine the training dataset based on the PNG images and the label files. The label files are used to reflect the contours and categories corresponding to the segmented spatial regions.
[0034] The model training module is used to set up a Python runtime environment, install the YOLOv8-seg model and its related dependency libraries, and perform end-to-end training on the training dataset based on the training script of the YOLOv8-seg model to obtain the YOLO spatial segmentation model.
[0035] The model application module is used to perform spatial segmentation on architectural floor plans based on the YOLO spatial segmentation model, output the segmentation results, and automatically label the spatial areas based on the segmentation results.
[0036] Thirdly, embodiments of the present invention also provide a terminal, wherein the terminal includes a memory, a processor, and a spatial recognition program based on the YOLO model stored in the memory and executable on the processor. When the processor executes the spatial recognition program based on the YOLO model, it implements the steps of the spatial recognition method based on the YOLO model of any of the above-mentioned schemes.
[0037] Fourthly, embodiments of the present invention also provide a computer-readable storage medium, wherein the computer-readable storage medium stores a spatial recognition program based on the YOLO model, and the spatial recognition program based on the YOLO model implements the steps of the spatial recognition method based on the YOLO model as described in any of the above schemes on the computer-readable storage medium.
[0038] Beneficial Effects: Compared with existing technologies, this invention provides a spatial recognition method based on the YOLO model. First, this invention acquires CAD drawings and performs spatial segmentation and coordinate transformation on the CAD drawings to obtain PNG images corresponding to specified spatial regions in the CAD drawings and corresponding label files. Based on the PNG images and label files, a training dataset is determined. The label files reflect the contours and categories corresponding to the segmented spatial regions. Then, a Python runtime environment is set up, the YOLOv8-seg model and its related dependencies are installed, and end-to-end training is performed on the training dataset based on the training script of the YOLOv8-seg model to obtain a YOLO spatial segmentation model. Finally, the architectural floor plan is spatially segmented based on the YOLO spatial segmentation model, the segmentation results are output, and the spatial regions are automatically labeled based on the segmentation results. This invention uses the YOLO spatial segmentation model to achieve high-precision spatial recognition and classification with strong generalization ability. For special scenarios such as complex spatial structures, dense lines, and inconsistent proportions in architectural drawings, this invention can realize fully automated processing from spatial area recognition, contour correction, coordinate mapping, and automatic annotation of the original CAD drawings, and has high precision, strong adaptability, and engineering practicality. Attached Figure Description
[0039] Figure 1 A flowchart illustrating a preferred embodiment of the spatial recognition method based on the YOLO model provided in this invention.
[0040] Figure 2 This is a schematic diagram of the architecture of a spatial recognition system based on the YOLO model provided in an embodiment of the present invention.
[0041] Figure 3 A schematic diagram of the terminal provided in an embodiment of the present invention. Detailed Implementation
[0042] To make the objectives, technical solutions, and effects of this invention clearer and more explicit, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0043] The flowchart shown in the attached diagram is for illustrative purposes only and does not necessarily include all content, operations, or steps, nor does it require execution in the described order. For example, some operations or steps can be broken down, combined, or partially merged, so the actual execution order may change depending on the actual situation.
[0044] It should be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0045] It should be understood that, in order to clearly describe the technical solutions of the embodiments of the present invention, the terms "first" and "second" are used in the embodiments of the present invention to distinguish identical or similar items with essentially the same function and effect. For example, "first control information" and "second control information" are only used to distinguish different control information and do not limit their order.
[0046] Those skilled in the art will understand that the words "first" and "second" do not limit the quantity or the order of execution, and that the words "first" and "second" do not necessarily imply that they are different.
[0047] It should also be understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0048] To address the problems of existing technologies, this invention provides a spatial recognition method based on the YOLO model. This method, as described in this embodiment, automates the entire process from spatial region identification, contour correction, coordinate mapping, and automatic annotation of original CAD drawings, exhibiting high precision, strong adaptability, and engineering practicality. In specific applications, this embodiment first acquires CAD drawings and performs spatial segmentation and coordinate transformation on them, obtaining PNG images corresponding to specified spatial regions in the CAD drawings and corresponding tag files. Based on the PNG images and tag files, a training dataset is determined. The tag files reflect the contours and categories corresponding to the segmented spatial regions. Then, a Python runtime environment is set up, the YOLOv8-seg model and its related dependencies are installed, and end-to-end training is performed on the training dataset based on the training script of the YOLOv8-seg model to obtain a YOLO spatial segmentation model. Finally, the architectural floor plan is spatially segmented based on the YOLO spatial segmentation model, the segmentation results are output, and automatic annotation of spatial regions is performed based on the segmentation results. This invention uses the YOLO spatial segmentation model to achieve high-precision spatial recognition and classification with strong generalization ability.
[0049] The spatial recognition method based on the YOLO model in this embodiment can be applied to terminals, including intelligent product terminals such as computers, smart TVs, and mobile phones. In this embodiment, as shown... Figure 1 As shown, the spatial recognition method based on the YOLO model includes the following steps:
[0050] Step S100: Obtain CAD drawings and perform spatial segmentation and coordinate transformation on the CAD drawings to obtain PNG images corresponding to specified spatial regions in the CAD drawings and label files corresponding to the PNG images. Based on the PNG images and the label files, determine the training dataset. The label files are used to reflect the contours and categories corresponding to the segmented spatial regions.
[0051] First, the CAD drawing is opened using software. The CAD drawing is in DWG format. Various closed spatial regions (such as bedrooms, kitchens, bathrooms, stairwells, elevator shafts, living rooms, etc.) in the CAD drawing are spatially segmented and mapped, resulting in PNG images corresponding to the specified spatial regions in the CAD drawing, and corresponding label files for these PNG images. The label files reflect the outline and category of the segmented spatial regions. In this embodiment, the label files can be JSON label files in LabelMe format or TXT format label files. Finally, based on the PNG images and the label files, the training dataset is determined.
[0052] In one implementation, the spatial segmentation and coordinate transformation in this embodiment include the following steps:
[0053] Step S101: Use polysegments to draw the boundary of each spatial region in the CAD drawing, and record the coordinates of the polysegment vertices of the specified spatial region;
[0054] Step S102: Export the PNG image of the specified spatial region;
[0055] Step S103: Perform coordinate transformation on the coordinates of the vertex of the multi-line segment and obtain the tag file.
[0056] Specifically, this embodiment uses closed polylines to draw the boundaries of each spatial region in the CAD drawing. This embodiment can develop a C# plugin, which operates based on CAD vector graphics, possessing high precision and supporting object snapping, making the annotation process more accurate and convenient, conforming to the operating habits and usage needs of architectural engineers. Based on this C# plugin, a PNG image of the specified spatial region can be automatically exported, and the vertex coordinates of the polylines used to draw the specified spatial region can be determined based on the PNG image. Furthermore, the aforementioned C# plugin can also establish the world coordinate system of the printed area of the CAD drawing for subsequent coordinate transformation. In this embodiment, the coordinate transformation is used to convert between the pixel coordinates of the PNG image and the world coordinate system of the CAD drawing. Specifically, in the world coordinate system of the CAD drawing, the origin is located at the lower left corner of the view, the X-axis extends to the right, and the Y-axis extends upwards; while in the pixel coordinate system of the PNG image, the origin is usually located at the upper left corner of the image, the X-axis extends to the right, and the Y-axis extends downwards. To achieve the conversion between the two coordinate systems, this embodiment first obtains the dimensions of the PNG image and the printed area of the CAD drawing, which can be determined by the minimum bounding rectangle of the display area. Then, a scaling factor is calculated based on the dimensions of the PNG image and the printed area of the CAD drawing. Furthermore, a coordinate transformation matrix is determined based on the scaling factor and the world coordinate system. The coordinate transformation matrix is then used to transform the coordinates of the polyline vertex to obtain the pixel coordinates of the PNG image. Specifically, in determining the scaling factor, this embodiment calculates the ratio scale = width / imgWidth based on the width (imgWidth) of the PNG image and the actual width (width) of the printed area of the CAD drawing, where scale is the scaling factor. During coordinate transformation, the Y-coordinate of the PNG image is first inverted (i.e., multiplied by -1) to align its direction with the Y-axis of the world coordinate system of the CAD drawing. Then, the X and Y components are multiplied by the scaling factor scale for scaling. Finally, a translation operation is used to move the origin of the PNG image to (minX, maxY), completing the mapping from the pixel coordinates of the PNG image to the world coordinates (i.e., dwg coordinates) of the CAD drawing.
[0057] In this embodiment, the specific coordinate transformation matrix can be constructed in the following way:
[0058]
[0059]
[0060] After converting the vertex coordinates of the polyline segments in the CAD drawing into pixel coordinates of the PNG image using the aforementioned coordinate transformation matrix, the pixel coordinates of the PNG image are normalized and then written into a YOLO format TXT file along with the category of the specified spatial region to obtain the tag file. Each line of data in the tag file represents a spatial segmentation contour, and its data format is as follows: <class-index> <x1> <y1> <x2> <y2> ... <xn> <yn>,in <class-index>Indicates the target category index. <xn>and <yn>respectively, are the normalized n-th vertex coordinate values. The label file generated in this embodiment indicates that the data has been converted into the format required by the YOLO instance segmentation, and can be used for subsequent model training and inference. When generating a spatial segmentation contour based on the YOLO model, the coordinate conversion matrix is also used to convert back to the world coordinate system of the CAD drawing, so as to realize accurate positioning and application of the recognition result in the original CAD drawing.
[0061] Further, the PNG image and its corresponding label file can be divided into a training set and a validation set in this embodiment. The PNG image and the label file are stored in the corresponding directory structure, and a training data set conforming to the YOLO format is constructed. The structure of the training data set includes two main folders, images and labels, which respectively contain the image and label file of the training set (train) and the validation set (val), facilitating subsequent model training and performance evaluation.
[0062] In addition, in other implementations, the embodiment can also use a pre-annotation combined with a manual correction mode when constructing the training data. Specifically, the new drawing is first pre-segmented using the trained model, and the initial contour is generated. Then, manual fine-tuning is performed, which can reduce more than 70% of manual operations. In addition, the CAD plug-in can be optimized to support the annotation of circular arcs and spline curves, and automatically discrete the curves into high-precision polygons to avoid errors caused by manual drawing of multiple line segments. In addition, existing coordinate conversion assumes that there is no nonlinear deformation in the CAD drawing, but in actual engineering, the drawing may be stretched or tilted (such as scanning distortion). The embodiment can introduce a perspective transformation matrix or a thin plate spline interpolation (TPS) to dynamically correct the nonlinear deformation by marking four or more reference points (such as axis net intersection points) in the CAD drawing, so that the coordinate conversion error is reduced from ±5 pixels to within ±1 pixel.
[0063] Step S200, build a Python running environment, install YOLOv8-seg model and its related dependent libraries, and perform end-to-end training on the training data set based on the training script of the YOLOv8-seg model to obtain a YOLO spatial segmentation model.
[0064] The embodiment first configures a Python running environment, and installs a YOLOv8-seg model and related dependent libraries, including core frameworks such as OpenCV, PyTorch and ONNX, which provides complete technical support for subsequent model training, inference and deployment. The embodiment writes a data configuration file in yaml format, which contains the training set path, validation set path and category list. Then, the training script provided by YOLOv8 is used to perform end-to-end training on the constructed training data set. To improve training efficiency and accelerate model convergence, a GPU can be selected for acceleration, thereby significantly improving the overall training speed and model recognition performance.
[0065] In addition, in other implementations, the embodiment can also use a model with a larger receptive field (such as YOLOv9-seg, Mask R-CNN+FPN), or add deformable convolution (DCN) to the backbone to enhance the ability to capture irregular contours. Alternatively, an attention mechanism (such as CBAM) can also be introduced to focus the model on key boundary features such as walls, doors and windows, and reduce the interference of non-key lines (such as size annotations). In addition, contrastive learning can also be introduced to constrain the similarity of different augmented samples in the same space during training, improving the adaptability of the model to changes in drawing style. The embodiment can also use a multi-modal fusion training method, taking the vector features in the CAD drawing (such as the length, angle and layer ID of the wall line) as additional input, combining them with image features through a cross-modal fusion module (such as attention fusion), and letting the model learn prior knowledge such as "wall layer lines are more likely to be room boundaries", which can improve the segmentation accuracy by 5%-10%.
[0066] Step S300, performing spatial segmentation on the architectural plan drawing based on the YOLO spatial segmentation model, outputting a segmentation result, and performing automatic labeling of spatial regions based on the segmentation result.
[0067] The embodiment converts the trained YOLO instance segmentation model into an ONNX format to support cross-platform deployment and efficient inference. The YOLO spatial segmentation model is used to perform spatial segmentation on the architectural plan drawing, and a segmentation result is output.
[0068] In one implementation, the embodiment includes the following steps when segmenting the architectural plan:
[0069] Step S301, converting the YOLO spatial segmentation model into an ONNX format;
[0070] Step S302, exporting a PNG image of the architectural plan drawing and determining a world coordinate system of the architectural plan drawing;
[0071] Step S303, the PNG image of the building plan paper is preprocessed, and the PNG image is inferred to obtain the bounding box of each space region in the building plan paper and the corresponding category;
[0072] Step S304, based on the coordinate conversion matrix, the bounding box of each space region is mapped to the world coordinate system of the building plan paper;
[0073] Step S305, the multi-line segment in the building plan paper is matched with the bounding box of each space region obtained by inference, and the segmentation result is output, the segmentation result includes the contour of each space region in the building plan paper and the category thereof.
[0074] Specifically, the embodiment utilizes CAD secondary development interface or ODA technology to export the building plan paper as a PNG image, and synchronously records the world coordinate system of the building plan paper in the exported area, which is used for subsequent coordinate mapping. At this time, the building plan paper is also a CAD paper, i.e., a dwg format paper. Then, the C++ OpenCV library is used to preprocess the PNG image of the building plan paper. Then, the ONNX Runtime library is called to load the YOLO space segmentation model in ONNX format to infer the PNG image, to obtain the bounding box of each space region in the building plan paper and the corresponding category. The embodiment extracts the area, aspect ratio and shape complexity of each space region, and uses the Snap Rounding algorithm of the CGAL library to smooth the contour, so as to reduce the jagged and fragmented line segments, effectively eliminate the jagged phenomenon caused by image scaling, although the contour after topological correction has little difference in vision from the original inference result, but the actual contour has significantly reduced the fragmented line segments, which greatly reduces the calculation complexity for subsequent line segment merging and boundary optimization process, and improves the overall processing efficiency.
[0075] In addition, in other implementations, the embodiment can also introduce a topological repair module, for example, using the arrangements class of the CGAL library to detect self-intersecting line segments and cut, and using the minimum cost closing algorithm for the notch to ensure that the contour is strictly closed, so as to better segment the space region in the subsequent step. In addition, the embodiment can also add geometric constraint correction for regular building structures, such as rectangles or L shapes, cluster the angles of the contour vertices (such as 90°, 180°), correct the angles deviating from the threshold (such as ±5°) to the standard value, and use the minimum circumscribed rectangle or convex hull to assist in correcting irregular contours, so that they are more in line with the building design specifications.
[0076] Furthermore, this embodiment maps the bounding boxes of each spatial region to the world coordinate system of the architectural floor plan based on the coordinate transformation matrix. Then, multi-line segments are extracted from the architectural floor plan. These multi-line segments are decomposed into independent straight line segments, which are then matched with the bounding boxes of each spatial region obtained through inference, outputting a segmentation result. The segmentation result includes the outline and category of each spatial region in the architectural floor plan. During matching, this embodiment searches for the edge line in the architectural floor plan that is closest in distance and direction to the bounding boxes of each spatial region obtained through inference. When a match is successful, the matched line segments are merged into a closed polygonal outline as the outline of the spatial region. This allows the outlines of each spatial region to be segmented from the architectural floor plan and their corresponding categories matched, resulting in a segmentation result. This embodiment can also perform outline correction to further improve the consistency and accuracy of the recognition result with the geometry of the original drawing.
[0077] Furthermore, this embodiment can also use OCR algorithms to recognize text in architectural floor plans, such as "balcony," "living room," "bedroom," and "bathroom." In practical applications, edge detection can be used to remove interfering lines within the text area before text recognition. Next, the BERT text classification model is used to perform semantic analysis on the text and extract semantic features. Then, the semantic features are matched with the categories of each spatial area, and the text with the closest semantics is selected as the final name of the corresponding spatial area in the architectural floor plan, realizing automated labeling of spatial areas. For example, the trained BERT text classification model is used to match and label the final room names, outputting the final spatial recognition result. In addition, since text in architectural floor plans may be far from the corresponding spatial area (e.g., "Bedroom 1" is labeled at the door), this embodiment can also use spatial association rules to assist in matching the corresponding text. For example, there is a certain correlation between the distance and orientation of the text and the spatial area. Combined with the semantic similarity calculation of the BERT text classification model, the corresponding text can be identified, reducing mismatches of multiple texts in the same area or text misalignment.
[0078] Furthermore, this embodiment can also visually annotate the segmented spatial regions in PNG images using different colors and output JSON structured data format, including information such as layout name, drawing frame name, drawing frame ID, area, and outer contour coordinates. This invention supports data interaction interfaces with BIM platforms or other building software.
[0079] In summary, the present invention has the following advantages compared with the prior art:
[0080] 1. Implement a complete closed-loop processing flow from CAD drawings in DWG format to image recognition and then to world coordinate mapping;
[0081] 2. The YOLO instance segmentation model is used to achieve high-precision spatial recognition and classification with strong generalization ability;
[0082] 3. Introduce a coordinate transformation mechanism to ensure that the recognition results can be accurately reproduced in CAD drawings;
[0083] 4. Combining the CGAL algorithm with CAD edge matching effectively corrects contour jaggedness and deviation;
[0084] 5. Automate spatial naming using the BERT text classification model to enhance the system's semantic understanding capabilities;
[0085] 6. Strong compatibility, supports multiple CAD drawing styles, suitable for automated analysis of actual engineering projects;
[0086] 7. It has good scalability and is easy to integrate into platforms such as Building Information Modeling (BIM), indoor navigation, and intelligent design.
[0087] Based on the above embodiments, the present invention also provides a spatial recognition system based on the YOLO model. The system in this embodiment is used to implement the steps of the above method embodiments. Specifically, as... Figure 2 As shown, the system in this embodiment includes: a training dataset preparation module 10, a model training module 20, and a model application module 30. Specifically, the training dataset preparation module 10 is used to acquire CAD drawings, perform spatial segmentation and coordinate transformation on the CAD drawings, obtain PNG images corresponding to specified spatial regions in the CAD drawings and label files corresponding to the PNG images, and determine the training dataset based on the PNG images and the label files. The label files are used to reflect the contours and categories corresponding to the segmented spatial regions. The model training module 20 is used to set up a Python runtime environment, install the YOLOv8-seg model and its related dependent libraries, and perform end-to-end training on the training dataset based on the training script of the YOLOv8-seg model to obtain a YOLO spatial segmentation model. The model application module 30 is used to perform spatial segmentation on architectural floor plans based on the YOLO spatial segmentation model, output the segmentation results, and perform automatic annotation of spatial regions based on the segmentation results.
[0088] The working principle of each module in the YOLO-based spatial recognition system of this embodiment is the same as that of each step in the above method embodiment, and will not be repeated here.
[0089] The modules in the YOLO-based spatial recognition system described above can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of the terminal in hardware form or independent of it, or stored in the terminal's memory in software form, so that the processor can call and execute the corresponding operations of each module.
[0090] Based on the above embodiments, the present invention also provides a terminal, the principle block diagram of which can be as follows: Figure 3 As shown. The terminal may include one or more processors 100 ( Figure 3 (Only one is shown in the image), memory 101, and computer program 102 stored in memory 101 and executable on one or more processors 100. For example, a spatial recognition program based on the YOLO model. When one or more processors 100 execute computer program 102, they can implement the various steps in the embodiment of the spatial recognition method based on the YOLO model. Alternatively, when one or more processors 100 execute computer program 102, they can implement the functions of various modules / units in the embodiment of the spatial recognition system based on the YOLO model, which is not limited here.
[0091] In one embodiment, the processor 100 may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0092] In one embodiment, memory 101 may be an internal storage unit of an electronic device, such as a hard drive or RAM. Memory 101 may also be an external storage device of the electronic device, such as a plug-in hard drive, smart media card (SMC), secure digital (SD) card, flash card, etc. Furthermore, memory 101 may include both internal and external storage units. Memory 101 is used to store computer programs and other programs and data required by the terminal. Memory 101 can also be used to temporarily store data that has been output or will be output.
[0093] Those skilled in the art will understand that Figure 3 The block diagram shown is merely a partial structural diagram related to the present invention and does not constitute a limitation on the terminal to which the present invention is applied. A specific terminal may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0094] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the methods described above. Any references to memory, storage, operational databases, or other media used in the embodiments provided by this invention can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual operating data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAM bus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0095] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.< / yn> < / xn> < / yn> < / xn> < / y2> < / x2> < / y1> < / x1> < / class-index>
Claims
1. A spatial recognition method based on the YOLO model, characterized in that, The method includes: Obtain CAD drawings and perform spatial segmentation and coordinate transformation on the CAD drawings to obtain PNG images corresponding to specified spatial regions in the CAD drawings and label files corresponding to the PNG images. Based on the PNG images and the label files, determine the training dataset. The label files are used to reflect the contours and categories corresponding to the segmented spatial regions. Set up a Python runtime environment, install the YOLOv8-seg model and its related dependency libraries, and perform end-to-end training on the training dataset based on the training script of the YOLOv8-seg model to obtain the YOLO spatial segmentation model. The architectural floor plan is spatially segmented based on the YOLO spatial segmentation model, the segmentation results are output, and the spatial areas are automatically labeled based on the segmentation results.
2. The spatial recognition method based on the YOLO model according to claim 1, characterized in that, The step of performing spatial segmentation and coordinate transformation on the CAD drawing to obtain a PNG image corresponding to a specified spatial region in the CAD drawing and a tag file corresponding to the PNG image includes: Use polylines to draw the boundary of each spatial region in the CAD drawing, and record the coordinates of the polyline vertex of the specified spatial region; Export a PNG image of a specified spatial region; The coordinates of the vertices of the multi-line segment are transformed to obtain the tag file.
3. The spatial recognition method based on the YOLO model according to claim 2, characterized in that, The process of performing coordinate transformation on the vertex coordinates of the multi-segment network to obtain the tag file includes: Obtain the world coordinate system of the print area of the CAD drawing; Calculate the scaling factor based on the size of the PNG image and the size of the printing area of the CAD drawing; Based on the scaling factor and the world coordinate system, a coordinate transformation matrix is determined, and the coordinate transformation is performed on the coordinates of the vertex coordinates of the multi-segment to obtain the pixel coordinates of the PNG image. After normalizing the pixel coordinates of the PNG image, the corresponding categories of the specified spatial regions are written into a YOLO format TXT file to obtain the tag file.
4. The spatial recognition method based on the YOLO model according to claim 3, characterized in that, The step of determining the training dataset based on the PNG image and the tag file includes: The PNG images and the tag files are stored in a preset target structure to construct a training dataset that conforms to the YOLO format.
5. The spatial recognition method based on the YOLO model according to claim 4, characterized in that, The process of spatially segmenting architectural floor plans based on the YOLO spatial segmentation model and outputting the segmentation results includes: Convert the YOLO spatial partitioning model into ONNX format; Export the PNG image of the architectural floor plan and determine the world coordinate system of the architectural floor plan; The PNG image of the architectural floor plan is preprocessed and inference is performed on the PNG image to obtain the bounding boxes of each spatial area in the architectural floor plan and their corresponding categories; Based on the coordinate transformation matrix, the bounding boxes of each spatial region are mapped to the world coordinate system of the architectural plan drawing; The multi-line segments extracted from the architectural floor plan are matched with the bounding boxes of each spatial region obtained through inference, and the segmentation result is output. The segmentation result includes the outline and category of each spatial region in the architectural floor plan.
6. The spatial recognition method based on the YOLO model according to claim 5, characterized in that, Based on the YOLO spatial segmentation model, the architectural floor plan is spatially segmented, and the segmentation results are output, including: Extract the geometric features of each spatial region and smooth the contours using the Snap Rounding algorithm from the CGAL library to reduce jagged edges and fragmented lines.
7. The spatial recognition method based on the YOLO model according to claim 6, characterized in that, The automated annotation of spatial regions based on the segmentation results includes: The OCR algorithm is used to recognize the text in architectural floor plans, and the BERT text classification model is used to perform semantic analysis on the text and extract semantic features. The semantic features are matched with the categories of each spatial region, and the text with the closest semantics is selected from the semantic features as the final name of the corresponding spatial region in the architectural floor plan, thereby realizing the automatic labeling of spatial regions.
8. A spatial recognition system based on the YOLO model, characterized in that, The system is used to implement the steps of the spatial recognition method based on the YOLO model according to any one of claims 1-7, the system comprising: The training dataset preparation module is used to acquire CAD drawings, perform spatial segmentation and coordinate transformation on the CAD drawings, obtain PNG images corresponding to specified spatial regions in the CAD drawings and label files corresponding to the PNG images, and determine the training dataset based on the PNG images and the label files. The label files are used to reflect the contours and categories corresponding to the segmented spatial regions. The model training module is used to set up a Python runtime environment, install the YOLOv8-seg model and its related dependency libraries, and perform end-to-end training on the training dataset based on the training script of the YOLOv8-seg model to obtain the YOLO spatial segmentation model. The model application module is used to perform spatial segmentation on architectural floor plans based on the YOLO spatial segmentation model, output the segmentation results, and automatically label the spatial areas based on the segmentation results.
9. A terminal, characterized in that, The terminal includes a memory, a processor, and a spatial recognition program based on the YOLO model stored in the memory and executable on the processor. When the processor executes the spatial recognition program based on the YOLO model, it implements the steps of the spatial recognition method based on the YOLO model as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a spatial recognition program based on the YOLO model, and the spatial recognition program based on the YOLO model implements the steps of the spatial recognition method based on the YOLO model as described in any one of claims 1-7 on the computer-readable storage medium.
Citation Information
Patent Citations
Space recognition method and device for CAD drawing, electronic equipment and storage medium
CN111008597A
Residential subspace structure identification method and equipment based on CAD (Computer Aided Design) drawing
CN119229468A
Engineering drawing label identification method and system based on multi-modal information extraction
CN119964171A
Cited By
AI-based picture frame generation and intelligent arrangement method and system, terminal and medium
CN121259261A
A method, system, terminal, and medium for AI-based frame generation and intelligent layout.
CN121259261B
CAD electrical scheme identification method and system
CN121616894A
Rock core photograph automatic correction and cutting method based on YOLOv8-seg and perspective transformation
CN121707887A
PCB schematic diagram picture labeling and data set generating method and system
CN121766253A