Automatic building engineering drawing frame generation system and method based on LLM and YOLO, terminal and medium
By combining LLM and YOLO models, the intelligent system automatically recognizes the content of drawings and generates drawing frames, solving the problem of low efficiency in manually inserting drawing frames and improving the efficiency and quality of architectural engineering drawing design.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN CAPOL INT & ASSOC CO LTD
- Filing Date
- 2025-12-04
- Publication Date
- 2026-04-28
AI Technical Summary
In existing technologies, the process of inserting drawing frames into architectural engineering drawings relies on manual judgment, which leads to low efficiency, easy errors, and high labor costs, making it difficult to meet complex design requirements.
An intelligent system based on LLM and YOLO is adopted. It understands user needs through a large language model, recognizes drawing content by combining YOLO model, and automatically generates and inserts appropriate title frames, thus realizing automatic title frame generation.
It improves the efficiency and quality of drawing design, reduces manual intervention, ensures the accuracy and consistency of drawing frame insertion, and is suitable for single and batch drawing processing.
Smart Images

Figure CN121935997A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent processing technology for architectural engineering drawings, and in particular to an automatic generation system, method, terminal, and medium for architectural engineering drawing frames based on LLM and YOLO. Background Technology
[0002] In the architectural design industry, after the drawings are completed, the subsequent steps of inserting the drawing frame, filling in the title block information, and printing the drawings are crucial. Currently, designers mainly rely on manually judging the drawing area to estimate the approximate range and insert a drawing frame that matches the drawing size and scale. However, this process is full of variables. If the inserted drawing frame is inappropriate, such as being too large or too small, the inserted drawing frame must be deleted and a drawing frame of the appropriate size must be manually inserted again, or the drawing frame scale must be manually modified. If the insertion position is off, the insertion point of the drawing frame must also be manually adjusted until the desired effect is achieved.
[0003] As construction projects become increasingly complex and the number of design drawings increases dramatically, the drawbacks of the traditional method of inserting drawing frames become apparent. It is inefficient, consuming a lot of time in the process of repeatedly adjusting the drawing frames; moreover, manual operation is prone to errors, affecting the quality of the drawings; at the same time, the entire process is highly dependent on manual labor, resulting in high labor costs.
[0004] Therefore, existing technologies still have shortcomings. Summary of the Invention
[0005] The technical problem to be solved by this invention is to provide an automatic generation system, method, terminal, and medium for architectural engineering drawing frames based on LLM and YOLO, addressing the aforementioned deficiencies of the prior art. The technical solution adopted by this invention is as follows: In a first aspect, the present invention provides a method for automatically generating architectural engineering drawing frames based on LLM and YOLO, the method comprising: The intelligent agent receives natural language input from the user, understands the natural language based on the configured large language model, obtains the demand information, and performs task planning based on the demand information to obtain the task planning result. The drafting tool is driven to acquire architectural drawings input into the drafting tool, and the content recognition of the architectural drawings is performed based on the trained YOLO model to obtain the inference result of the drawing frame range; Based on the task planning results and the inference results of the drawing frame range, a target drawing frame is automatically generated, inserted into the architectural engineering drawings, and saved.
[0006] In one implementation, the method further includes: Develop MCP services and MCP interfaces, wherein the MCP interface is used to establish a connection between the large language model and the mapping tool.
[0007] In one implementation, task planning is performed based on the required information to obtain the task planning result, including: The requirement information is broken down to determine the drawing processing tasks in the requirement information. The drawing processing tasks include single drawing processing tasks and batch drawing processing tasks. Obtain available MCP interfaces from the MCP service, and based on the MCP interfaces, obtain the specified tag type from natural language and determine the frame category; Based on the specified label type, the frame category, and the drawing processing task, the task planning result is obtained.
[0008] In one implementation, task planning based on the required information to obtain task planning results further includes: If the specified tag type cannot be obtained from natural language, a predefined tag type will be obtained.
[0009] In one implementation, the architectural drawings are subjected to content recognition based on a trained YOLO model to obtain a bounding box inference result, including: The architectural drawings are preprocessed by parsing and recognizing the preprocessed drawings based on a trained YOLO model to obtain geometric data and characteristic data of the primitives. The preprocessing includes: drawing repair, drawing cleanup, extrinsic parameter binding, and format conversion. The geometric data of the primitives includes the insertion point of the block reference, the bounding box, the drawing scaling ratio, the vertex coordinates of the sub-primitives, the coordinates of the annotation text, and the coordinate transformation matrix. The characteristic data includes the block name, handle, attributes, and layer. The trained YOLO model outputs bounding box range inference results based on primitive geometric data and feature data.
[0010] In one implementation, based on the task planning result and the bounding box range inference result, a target bounding box is automatically generated, including: Based on the task planning results and the inference results of the map frame range, determine the map frame size and map frame scale; The target frame is generated based on the frame size and frame scale.
[0011] In one implementation, determining the frame size and frame scale based on the task planning result and the frame range inference result includes: Based on the task planning results, all candidate frames are traversed, and the candidate frames are matched with the frame range reasoning results in ascending order of size. If the width and height of a candidate frame are both greater than the width and height in the frame range reasoning result, then obtain the frame size and frame scale of the candidate frame. If all candidate frames fail to match the frame range inference result, the frame insertion is determined to have failed.
[0012] Secondly, embodiments of the present invention also provide an automatic generation system for architectural engineering drawing frames based on LLM and YOLO, wherein the system is used to implement the steps of the automatic generation method for architectural engineering drawing frames based on LLM and YOLO as described in any of the above solutions, and the system includes: The language understanding and task planning module is used by the intelligent agent to receive natural language input from the user, understand the natural language based on the configured large language model, obtain the requirement information, and perform task planning based on the requirement information to obtain the task planning result. The drawing recognition and reasoning module is used to drive the drafting tool, acquire the architectural engineering drawings input into the drafting tool, and perform content recognition on the architectural engineering drawings based on the trained YOLO model to obtain the drawing frame range reasoning result. The drawing frame generation and saving module is used to automatically generate a target drawing frame based on the task planning results and the drawing frame range reasoning results, insert the generated target drawing frame into the architectural engineering drawings and save it.
[0013] Thirdly, embodiments of the present invention also provide a terminal, wherein the terminal includes a memory, a processor, and an automatic generation program for architectural engineering drawing frames based on LLM and YOLO stored in the memory and executable on the processor. When the processor executes the automatic generation program for architectural engineering drawing frames based on LLM and YOLO, it implements the steps of the automatic generation method for architectural engineering drawing frames based on LLM and YOLO of any of the above-mentioned solutions.
[0014] Fourthly, embodiments of the present invention also provide a computer-readable storage medium, wherein the computer-readable storage medium stores an automatic generation program for architectural engineering drawing frames based on LLM and YOLO, the automatic generation program for architectural engineering drawing frames based on LLM and YOLO implementing the steps of the automatic generation method for architectural engineering drawing frames based on LLM and YOLO as described in any of the above schemes on the computer-readable storage medium.
[0015] Beneficial Effects: Compared with existing technologies, this invention provides an automatic generation method for architectural engineering drawing frames based on LLM and YOLO. First, the intelligent agent receives natural language input from the user, understands the natural language based on a configured large language model to obtain requirement information, and performs task planning based on this requirement information to obtain the task planning result. Then, it drives a drafting tool to acquire the architectural engineering drawings input into the tool, and performs content recognition on the architectural engineering drawings based on a trained YOLO model to obtain a frame range inference result. Finally, based on the task planning result and the frame range inference result, a target frame is automatically generated, inserted into the architectural engineering drawing, and saved. This invention leverages LLM to understand the user's natural language input, utilizes a YOLO model to recognize architectural engineering drawings to determine the appropriate frame range, and then uses a drafting tool to automatically insert the frame and save it, achieving automatic frame generation and improving design efficiency and quality. Attached Figure Description
[0016] Figure 1 This is a flowchart illustrating a preferred embodiment of the method for automatically generating architectural drawing frames based on LLM and YOLO, provided in this invention.
[0017] Figure 2 This is a technical principle framework diagram of the automatic generation method for architectural engineering drawing frames based on LLM and YOLO provided in the embodiments of the present invention.
[0018] Figure 3 This is a business process diagram of the intelligent agent in the automatic generation method of architectural engineering drawing frame based on LLM and YOLO provided in the embodiments of the present invention.
[0019] Figure 4 This is a flowchart illustrating the preprocessing of architectural engineering drawings in the automatic generation method of architectural engineering drawing frames based on LLM and YOLO provided in the embodiments of the present invention.
[0020] Figure 5 This is a schematic diagram of the automatic generation system for architectural engineering drawing frames based on LLM and YOLO, provided in an embodiment of the present invention.
[0021] Figure 6 A schematic diagram of a terminal provided in an embodiment of the present invention. Detailed Implementation
[0022] To make the objectives, technical solutions, and effects of this invention clearer and more explicit, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0023] The flowchart shown in the attached diagram is for illustrative purposes only and does not necessarily include all content, operations, or steps, nor does it require execution in the described order. For example, some operations or steps can be broken down, combined, or partially merged, so the actual execution order may change depending on the actual situation.
[0024] It should be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0025] It should be understood that, in order to clearly describe the technical solutions of the embodiments of the present invention, the terms "first" and "second" are used in the embodiments of the present invention to distinguish identical or similar items with essentially the same function and effect. For example, the first control information and the second control information are only used to distinguish different control information and do not limit their order.
[0026] Those skilled in the art will understand that the words "first" and "second" do not limit the quantity or the order of execution, and that the words "first" and "second" do not necessarily imply that they are different.
[0027] It should also be understood that the term “and / or” as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0028] Currently, artificial intelligence is developing rapidly, with the emergence of large language models (LLM) and computer vision technologies (such as the YOLO model). Against this backdrop, exploring how to effectively apply artificial intelligence to the architectural design industry to achieve automatic generation of drawing frames and improve design efficiency and quality has become a research direction of great practical significance. To address the problems of existing technologies, this embodiment provides a method for automatically generating architectural engineering drawing frames based on LLM and YOLO. Based on this method, this embodiment can use LLM to understand the natural language input by the user, use the YOLO model to identify the architectural engineering drawings to determine the appropriate drawing frame range, and then automatically insert the drawing frame into the drawing tool and save it, thus achieving automatic drawing frame generation and improving design efficiency and quality. In specific applications, firstly, the intelligent agent receives the natural language input by the user, understands the natural language based on the configured large language model to obtain requirement information, and performs task planning based on the requirement information to obtain the task planning result. Then, it drives the drawing tool to obtain the architectural engineering drawings input into the drawing tool, and performs content recognition on the architectural engineering drawings based on the trained YOLO model to obtain the drawing frame range inference result. Finally, based on the task planning results and the inference results of the drawing frame range, a target drawing frame is automatically generated, inserted into the architectural engineering drawings, and saved.
[0029] The automatic generation method for architectural engineering drawing frames based on LLM and YOLO in this embodiment can be applied to a terminal, which can be an intelligent terminal product such as a computer. Specifically, as shown in the example... Figure 1 As shown in the figure, the automatic generation method for architectural engineering drawing frames based on LLM and YOLO in this embodiment includes the following steps: Step S100: The intelligent agent receives natural language input from the user, understands the natural language based on the configured large language model, obtains the demand information, and performs task planning based on the demand information to obtain the task planning result.
[0030] This embodiment first selects a suitable large language model and then deploys and configures it. Specifically, this embodiment does not employ post-training or fine-tuning of open-source models, as this incurs significant development costs for most enterprise application departments and presents considerable difficulties in preparing the data corpus. Therefore, a mainstream, general-purpose large language model is chosen and integrated via API. The selection of the large language model considers stability, performance, and cost to ensure its application effectiveness and benefits. The selection and evaluation criteria are as follows: 1. Language understanding: The large language model accurately recognizes the semantics and captures the intent of architectural instructions, ensuring that the task instructions are accurately received and avoiding erroneous output due to misunderstandings; 2. Language Generation: Evaluate whether the generated task results are professional, clear, and in line with industry standards, ensuring that the interactive output can effectively guide or complete practical construction tasks; 3. Controllability: This reflects the degree to which the large language model follows task requirements such as the rules of frame overlay, whether it generates results according to the set conditions, and whether it can controllably complete the task; 4. Multi-turn dialogue: This examines the model's ability to maintain contextual continuity and task coherence during multi-turn interactions; 5. Large Language Model Performance: The time taken from task input to the first utterance of a large language model reflects the efficiency of the large language model in understanding and reasoning. The context length of a large language model reflects the effectiveness of complex task applications during task execution. 6. Token consumption: The number of inputs and outputs of a large language model, quantifying the cost of model application.
[0031] Based on the above selection and evaluation criteria, this embodiment initially selected Qwen-max, Doubao-1.5-pro-32k, Ernie-4.5-turbo-128K, and Deepseek-v3 as candidate large language models. Then, each of these large language models was used to perform specific tasks, using natural language: "Please apply a drawing frame to the current drawing file, with the title block version as default and the drawing frame type as standard drawing frame. Save and close the file after applying the drawing frame to each file." Multiple rounds (no less than 10 times) of testing were conducted, and the test results are shown in Table 1 below. Table 1 compares the performance of the various language models.
[0032] Table 1
[0033] As shown in Table 1, Qwen-max exhibits the best stability and a good balance between performance and cost; therefore, Qwen-max can be selected as the large language model in this embodiment. The method in this embodiment can be implemented based on an agent mounted on the user terminal, specifically an insert frame agent. Combined with... Figure 2As shown, the agent on the client can receive natural language input from the user and understand it using a large language model. When deploying the large language model, this embodiment can choose an online large model service from a vendor, or directly use the API provided by the cloud vendor to which the large language model belongs. During the deployment phase, when selecting a client, this embodiment needs to consider that the client can quickly deploy large models from various vendors, provide agent creation and configuration functions, and support the local MCP protocol to drive local drawing tools (such as AutoCAD). This invention chooses the open-source CherryStudio as the system client. CherryStudio is a professional-grade cross-platform AI deployment client tool specifically designed for AI users. When configuring the large model service, taking Qwen-max service as an example, the API key is obtained, and the model service in CherryStudio is configured to set the API address. When configuring the large language model, this embodiment configures the response API key and API address according to the cloud vendor to which the large language model belongs. In order to give full play to the performance of the large language model and make the generated results closer to the user's needs, considering the stability required for the execution of the scenario task, the model temperature is set to 1, the Top-P is set to 1, the number of contexts is set to 10 or more (the maximum number of tokens is canceled by default and is determined by the limitations of the large language model itself).
[0034] Furthermore, this embodiment also develops an MCP service and an MCP interface. The MCP service is a model context protocol designed to standardize communication between large language models and external data sources and tools. The MCP interface is used to establish a connection between the large language model and the plotting tool. The implementation of the MCP interface mainly includes defining the MCP interface name, interface description, parameter name, parameter type, parameter description, return value, and return status description. To complete the plot frame insertion operation, several MCP interfaces are defined, as shown in Table 2.
[0035] Table 2
[0036] Regarding the development of the MCP service program, this embodiment uses Python script format to define the service entry point and interface call method. Combined with... Figure 2As shown, the agent can start the MC service through parameters, and then obtain the MCP interface from the MCP service for the large language model to call. When the large language model calls the MCP interface, it passes in specific parameters. The MCP service receives the call request from the large language model, parses the call parameters, and then establishes a connection with the drawing tool (AutoCAD) through port communication, passing the parameters to the drawing tool. After receiving the call parameters, the drawing tool starts to execute the task and finally returns the execution result to the MCP service. The MCP service then returns the execution result to the large language model.
[0037] For the design of intelligent agents' business scenarios, this embodiment first collects user usage scenarios, such as... Figure 3 As shown, Figure 3 This is a business process diagram for the intelligent agent. Users typically input commands like "Please insert title blocks into drawing xxx," or specify a task like "Please insert title blocks into all drawings under the 'D:\\test' folder," using natural language to outline the available title block types and title block categories, and specify default options. Alternatively, the intelligent agent can prompt the user for input, including confirming the title block type, title block type, and save path. The agent then executes the automatic title block insertion task. Upon completion, it prompts the user to save and close the drawing. After each title block insertion task, it returns the execution status, such as the number of successful insertions, the number of failed insertions, or the reason for failure. If it's a batch execution, the results are summarized and displayed at the end.
[0038] The design of prompts for the intelligent agent helps it operate more effectively as expected by the user. When designing prompts, the agent's role is first defined, such as "Your role is a CAD drawing frame generation engineer. Your goal is to efficiently parse user commands, strictly follow the drawing frame generation steps, and complete the automatic insertion process of drawing frames." Then, the interface call method is defined, such as "If the user directly specifies the drawing path, the drawing frame insertion interface is called directly" or "If the user specifies a folder, all dwg files in the folder are retrieved first, and then the drawing frame insertion interface is iterated through." Next, the interface parameter retrieval rules are defined, such as the "drawing frame type" parameter, which includes the following: "Standard Label," "Custom Label 1," and "Custom Label 2," with "Standard Label" as the default. If the user explicitly specifies one, it is used; otherwise, these options are listed for the user to choose from. Furthermore, the agent's prompts also set interface call rules, such as whether to continue the operation or exit if a drawing insertion fails during batch operations. You can also control the large language model through prompts to determine how the task results are presented, such as "summarize the results in a table".
[0039] When configuring the MCP tool on the intelligent agent, after writing the MCP interface and MCP service in the previous steps, the service needs to be started. In Cherry Studio, specify the MCP service name, such as "Intelligent Agent Drawing Frame Insertion," select "Standard Input / Output" for the service type, enter "python" as the command, specify the Python script path and timeout for the call parameters, and finally save the configuration. After saving, you can see the previously written MCP interface in Cherry Studio. For batch operations, such as inserting drawing frames into all drawings in a folder, you need to execute the "Get all dwg files in the specified folder" operation. This operation can utilize the existing file service, which provides MCP interfaces for common operations on local files, such as reading files, writing files, and traversing files.
[0040] In practical applications, users can input "Help me apply drawing frames to D:\\test\\a1.dwg" or "Help me apply drawing frames to all drawings in the D:\\test folder" into the agent input box. They can also add descriptions such as "The title tag type is a standard title tag," "The drawing frame category is a standard drawing frame," and "Automatically save and close after insertion." Upon receiving the user's natural language input, the agent in the client begins to invoke the large language model. The large language model first understands the user's natural language, obtains the requirement information, and performs task planning based on this information, obtaining the task planning result. Specifically, the requirement information is first broken down to determine the drawing processing tasks, which include single drawing processing tasks and batch drawing processing tasks. Then, an available MCP interface is obtained from the MCP service, and the specified title tag type is obtained from the natural language based on the MCP interface, and the drawing frame category is determined. At this point, the large language model analyzes the natural language to see if the user has specified a tag type. If not, it retrieves the predefined tag type from the prompts. If the user has explicitly specified a tag type, such as "tag type is standard tag," the large language model also analyzes the prompts to see if the user's specified tag type matches the predefined tag type. Next, the agent performs task planning based on the specified tag type, the frame category, and the drawing processing task, obtains the task planning result, and presents the result to the user.
[0041] Step S200: Drive the drafting tool, obtain the architectural engineering drawings input into the drafting tool, and perform content recognition on the architectural engineering drawings based on the trained YOLO model to obtain the inference result of the drawing frame range.
[0042] This embodiment selects a drafting tool, such as AutoCAD, and inputs architectural drawings into it. Next, content recognition is performed on the architectural drawings based on a trained YOLO model. To improve the success rate and speed of drawing parsing and recognition, this embodiment preprocesses the input architectural drawings and combines... Figure 4 As shown, preprocessing includes drawing repair, drawing cleanup, extrinsic parameter binding, and format conversion. Drawing cleanup aims to remove irrelevant elements and layers from architectural drawings, improving parsing speed and accuracy. For drawings with extrinsic parameters, extrinsic parameter binding is required. Since architectural drawings may be created using other tools, they may contain custom elements that may not be recognized, necessitating format conversion.
[0043] After the architectural drawings are preprocessed, a trained YOLO model is used to parse and identify the preprocessed drawings, obtaining primitive geometric data and characteristic data. The primitive geometric data includes the insertion point of the block reference, the bounding box, the drawing scaling ratio, the vertex coordinates of the sub-primitives, the coordinates of the annotation text, and the coordinate transformation matrix; the characteristic data includes the block name, handle, attributes, and layer. The trained YOLO model can then output the bounding box range inference result based on the primitive geometric data and characteristic data.
[0044] The training process of the YOLO model in this embodiment includes: 1. Data preparation: Collect over 1,000 DWG drawings covering more than 10 categories, including catalogs, descriptions, and floor plans. After manually removing the original drawing frames, export standardized PNG images and label the drawing frame range and type according to multiple dimensions such as content modality, profession, and number of drawing frames.
[0045] 2. Model Training: Configure the YOLO dataset YAML file, use GPU acceleration to train for 300-500 rounds, set the resolution to 1024×1024, monitor metrics such as box_loss (boundary box regression loss) and cls_loss (classification loss) to ensure convergence and generalization, so that the YOLO model learns the correspondence between the exported PNG image and the bounding box range.
[0046] 3. Model Validation: The trained YOLO model is exported after manual verification using a validation set. The trained YOLO model can then export PNG images based on input architectural drawings, output the bounding box inference results, and, combined with a coordinate transformation matrix, achieve accurate conversion from pixel coordinates of the PNG image to world coordinates of the architectural drawing. In practical applications, in the world coordinate system of CAD software, the origin is located at the lower left corner of the view, with the X-axis extending to the right and the Y-axis extending upwards. However, in the pixel coordinate system of a PNG image, the origin is typically located at the upper left corner of the image, with the X-axis extending to the right and the Y-axis extending downwards. To achieve conversion between the two coordinate systems, it is first necessary to determine the minimum bounding rectangle of the exported area of the architectural drawing (i.e., the CAD drawing) and obtain the world coordinates (minX, maxY) of its upper left corner, which will serve as the reference origin for subsequent coordinate transformations. Next, a scaling factor, `scale`, is set, its value being equal to the ratio of the width of the exported target image (`imgWidth`) to the actual width of the exported area of the CAD drawing (`width`), i.e., `scale = width / imgWidth`. During coordinate transformation, the Y-coordinates of the PNG image points are first inverted (i.e., multiplied by -1) to align their direction with the Y-axis in the architectural drawing. Then, the X and Y components are scaled by multiplying by `scale`. Finally, a translation operation is used to move the origin of the coordinate system to (`minX`, `maxY`), completing the mapping from the pixel coordinates of the PNG image to the world coordinates of the architectural drawing, and obtaining the coordinate transformation matrix. Therefore, this embodiment maps the frame range inference result to the world coordinate system of the architectural drawing based on the aforementioned coordinate transformation matrix.
[0047] Step S300: Based on the task planning results and the inference results of the drawing frame range, automatically generate the target drawing frame, insert the generated target drawing frame into the architectural engineering drawing and save it.
[0048] After the agent has planned the task as described above, it starts to drive the MCP tool to execute. First, it calls the tool_insert_frame interface, passes in the task planning result, and then waits for the tool to insert the execution frame. If the execution is completed normally, it will return the result and status to the user. If the tool encounters an error, it will return the error status and error message to the user. In addition, the agent will return the token consumption for each dialogue.
[0049] In practical applications, the agent determines the frame size and scale based on the task planning results and the frame range inference results. Then, it generates the target frame based on the frame size and scale. Specifically, the agent can iterate through all candidate frames based on the task planning results and match the candidate frames with the frame range inference results inferred from the YOLO model in ascending order of size. If the width and height of a candidate frame are both greater than the width and height in the frame range inference results, the frame size and scale of that candidate frame are obtained. Otherwise, larger candidate frames are used to match the frame inference results. If all candidate frames fail to match the frame range inference results, the frame insertion is determined to have failed. After the target frame is inserted, the agent saves and closes the architectural drawings according to user needs. The processing results are directly displayed on the CAD main interface. The agent returns the status of each MCP interface call, such as call parameters, returned content, whether there are errors, and finally summarizes the number of successes and failures.
[0050] Therefore, this application can leverage LLM to understand user-input natural language, utilize YOLO models to identify architectural engineering drawings to determine the appropriate drawing frame range, and then automatically insert and save the drawing frame using a plotting tool, thereby achieving automatic drawing frame generation and improving design efficiency and quality. Compared with existing technologies, this invention has at least the following advantages: 1. For a single drawing, the intelligent agent can automatically identify the drawing area, automatically match the appropriate drawing size, and calculate the insertion point of the drawing frame without manual judgment.
[0051] 2. For batch drawings, the intelligent agent can automatically perform batch insertion of drawing frames after the task is planned and the parameters are set, without the need for manual intervention.
[0052] 3. Strong applicability: Drawing recognition based on the YOLO model can be applied to new drawing frame standards as long as data annotation, model training, and model verification are performed, without the need to rewrite code and redevelop.
[0053] 4. Good scalability: Based on the scenario of automatically inserting title blocks, subsequent related operations can be extended, such as automatic title block updates, automatic printing to PDF format, and automatic uploading of drawings.
[0054] Based on the above embodiments, the present invention also provides an automatic generation system for architectural engineering drawing frames based on LLM and YOLO, which is used to implement the steps in the above method embodiments. Figure 5As shown in the diagram, the system in this embodiment includes: a language understanding and task planning module 10, a drawing recognition and reasoning module 20, and a drawing frame generation and saving module 30. Specifically, the language understanding and task planning module 10 is used to acquire natural language input by the user, understand the natural language based on a configured large language model to obtain requirement information, and perform task planning based on the requirement information to obtain a task planning result. The drawing recognition and reasoning module 20 is used to drive a drafting tool, acquire architectural engineering drawings input into the drafting tool, and perform content recognition on the architectural engineering drawings based on a trained YOLO model to obtain a drawing frame range reasoning result. The drawing frame generation and saving module 30 is used to automatically generate a target drawing frame based on the task planning result and the drawing frame range reasoning result, insert the generated target drawing frame into the architectural engineering drawing, and save it.
[0055] The automatic generation method for architectural engineering drawing frames based on LLM and YOLO in this embodiment is the same as the principle of each terminal and module in the above system embodiment, and will not be repeated here.
[0056] Based on the above embodiments, the present invention also provides a terminal, the principle block diagram of which can be as follows: Figure 6 As shown. The terminal may include one or more processors 100 ( Figure 6 (Only one is shown in the image), memory 101, and computer program 102 stored in memory 101 and executable on one or more processors 100. For example, an automatic generation program for architectural drawing frames based on LLM and YOLO. When one or more processors 100 execute computer program 102, they can implement the various steps in the embodiment of the method for automatically generating architectural drawing frames based on LLM and YOLO. Alternatively, when one or more processors 100 execute computer program 102, they can implement the functions of various modules / units in the embodiment of the system for automatically generating architectural drawing frames based on LLM and YOLO, which is not limited here.
[0057] In one embodiment, the processor 100 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0058] In one embodiment, memory 101 may be an internal storage unit of an electronic device, such as a hard drive or RAM. Memory 101 may also be an external storage device of the electronic device, such as a plug-in hard drive, smart media card (SMC), secure digital card (SD), flash card, etc. Furthermore, memory 101 may include both internal and external storage units. Memory 101 is used to store computer programs and other programs and data required by the terminal. Memory 101 can also be used to temporarily store data that has been output or will be output.
[0059] Those skilled in the art will understand that Figure 6 The block diagram shown is merely a partial structural diagram related to the present invention and does not constitute a limitation on the terminal to which the present invention is applied. A specific terminal may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0060] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, operational databases, or other media used in the embodiments provided by this invention can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual operating data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0061] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for automatically generating architectural engineering drawing frames based on LLM and YOLO, characterized in that, The method includes: The intelligent agent receives natural language input from the user, understands the natural language based on the configured large language model, obtains the demand information, and performs task planning based on the demand information to obtain the task planning result. The drafting tool is driven to acquire architectural drawings input into the drafting tool, and the content recognition of the architectural drawings is performed based on the trained YOLO model to obtain the inference result of the drawing frame range; Based on the task planning results and the inference results of the drawing frame range, a target drawing frame is automatically generated, inserted into the architectural engineering drawings, and saved.
2. The method for automatically generating architectural engineering drawing frames based on LLM and YOLO according to claim 1, characterized in that, The method further includes: Develop MCP services and MCP interfaces, wherein the MCP interface is used to establish a connection between the large language model and the mapping tool.
3. The method for automatically generating architectural engineering drawing frames based on LLM and YOLO according to claim 2, characterized in that, Based on the aforementioned requirements information, task planning is performed to obtain the task planning results, including: The requirement information is broken down to determine the drawing processing tasks in the requirement information. The drawing processing tasks include single drawing processing tasks and batch drawing processing tasks. Obtain available MCP interfaces from the MCP service, and based on the MCP interfaces, obtain the specified tag type from natural language and determine the frame category; Based on the specified label type, the frame category, and the drawing processing task, the task planning result is obtained.
4. The method for automatically generating architectural engineering drawing frames based on LLM and YOLO according to claim 3, characterized in that, Based on the aforementioned requirements information, task planning is performed to obtain the task planning results, which also includes: If the specified tag type cannot be obtained from natural language, a predefined tag type will be obtained.
5. The method for automatically generating architectural engineering drawing frames based on LLM and YOLO according to claim 4, characterized in that, Based on the trained YOLO model, content recognition is performed on the architectural drawings to obtain the inference results of the drawing frame range, including: The architectural drawings are preprocessed by parsing and recognizing the preprocessed drawings based on a trained YOLO model to obtain geometric data and characteristic data of the primitives. The preprocessing includes: drawing repair, drawing cleanup, extrinsic parameter binding, and format conversion. The geometric data of the primitives includes the insertion point of the block reference, the bounding box, the drawing scaling ratio, the vertex coordinates of the sub-primitives, the coordinates of the annotation text, and the coordinate transformation matrix. The characteristic data includes the block name, handle, attributes, and layer. The trained YOLO model outputs bounding box range inference results based on primitive geometric data and feature data.
6. The method for automatically generating architectural engineering drawing frames based on LLM and YOLO according to claim 5, characterized in that, Based on the task planning results and the bounding box range inference results, a target bounding box is automatically generated, including: Based on the task planning results and the inference results of the map frame range, determine the map frame size and map frame scale; The target frame is generated based on the frame size and frame scale.
7. The method for automatically generating architectural engineering drawing frames based on LLM and YOLO according to claim 6, characterized in that, Based on the task planning results and the inference results of the map frame range, the map frame size and map frame scale are determined, including: Based on the task planning results, all candidate frames are traversed, and the candidate frames are matched with the frame range reasoning results in ascending order of size. If the width and height of a candidate frame are both greater than the width and height in the frame range reasoning result, then obtain the frame size and frame scale of the candidate frame. If all candidate frames fail to match the frame range inference result, the frame insertion is determined to have failed.
8. An automatic generation system for architectural engineering drawing frames based on LLM and YOLO, characterized in that, The system is used to implement the steps of the automatic generation method for architectural engineering drawing frames based on LLM and YOLO as described in any one of claims 1-7, and the system includes: The language understanding and task planning module is used by the intelligent agent to receive natural language input from the user, understand the natural language based on the configured large language model, obtain the requirement information, and perform task planning based on the requirement information to obtain the task planning result. The drawing recognition and reasoning module is used to drive the drafting tool, acquire the architectural engineering drawings input into the drafting tool, and perform content recognition on the architectural engineering drawings based on the trained YOLO model to obtain the drawing frame range reasoning result; The drawing frame generation and saving module is used to automatically generate a target drawing frame based on the task planning results and the drawing frame range reasoning results, insert the generated target drawing frame into the architectural engineering drawings and save it.
9. A terminal, characterized in that, The terminal includes a memory, a processor, and an automatic generation program for architectural engineering drawing frames based on LLM and YOLO, which is stored in the memory and can run on the processor. When the processor executes the automatic generation program for architectural engineering drawing frames based on LLM and YOLO, it implements the steps of the automatic generation method for architectural engineering drawing frames based on LLM and YOLO as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores an automatic generation program for architectural engineering drawing frames based on LLM and YOLO, and the automatic generation program for architectural engineering drawing frames based on LLM and YOLO implements the steps of the automatic generation method for architectural engineering drawing frames based on LLM and YOLO as described in any one of claims 1-7 on the computer-readable storage medium.