Cross-format intelligent conversion method and system for engineering drawings
By introducing a multimodal fusion structure identification engine and a difference compensation network into the drawing conversion system, the problems of structural distortion, semantic confusion and visual inconsistency in drawing format conversion are solved, and efficient, accurate conversion and reliable application of drawings are achieved.
Patent Information
- Application Number
- CN202510518450.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-04-24
AI Technical Summary
The lack of structural differences perception and modeling capabilities, difficulty in decoupling multimodal information, and insufficient support for personalized styles of users in the prior art, resulting in the problems of structural distortion, semantic confusion and inconsistent visual expression in the drawings during the format conversion process.
Through an intelligent cross-format conversion method of engineering drawings, the multimodal fusion structure identification engine and the difference compensation network are used to perform element extraction, format adaptation optimization and semantic consistency verification to establish a second format engineering drawing.
The reduction of drawing structure, semantic recognition accuracy and visual style matching are improved, and the problem of efficient conversion, accurate recognition and reliable application of drawings in complex engineering contexts is solved.
Smart Images

Figure CN120047960A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of intelligent conversion of drawing formats, and in particular to a method and system for intelligent conversion of engineering drawings across formats. Background Art
[0002] As the complexity of engineering projects in industries such as construction, machinery, and electricity continues to increase, engineering drawings face increasingly complex format conversion and graphic element recognition requirements in collaborative design, version transfer, and system integration. Especially in scenarios where multiple design tools and multiple parties collaborate, drawings of different formats have significant differences in structural expression, layer semantics, graphic element encoding, and visual style, which poses challenges to the unified management and automatic recognition of engineering data. In order to improve work efficiency and data consistency, more and more systems are trying to use artificial intelligence and multimodal recognition technology to perform structured analysis and intelligent conversion of drawings. However, there are still many key defects in the existing technology.
[0003] At present, most of the existing drawing format conversion methods rely on static rule matching and template mapping, lacking the ability to perceive and model the differences between drawing structures, which leads to structural distortion or information omissions when converting between different formats. Secondly, multimodal fusion processing has problems such as high data coupling and difficulty in decoupling. Especially when processing mixed information such as images, vectors, and text, it is often impossible to accurately distinguish the semantics of the graphics elements, affecting the usability and accurate restoration of the drawings. Thirdly, existing methods generally ignore the personalized needs of users in the visual expression style of graphics elements. The converted drawings may not be consistent with the user's aesthetic preferences or industry standards, reducing the user experience. In addition, most systems have not yet established a task-oriented attention mechanism or semantic verification mechanism, and cannot dynamically optimize the recognition path according to the task scenario. They also lack the ability to feedback results based on semantic consistency verification, making it difficult for the system to adapt to the precise conversion needs in complex engineering contexts.
[0004] In summary, the existing technology has problems such as structural distortion, semantic confusion and inconsistent visual expression during the format conversion process due to the lack of structural difference perception and modeling capabilities, difficulty in decoupling multimodal information and insufficient support for user personalized style adaptation, which further affects the efficient conversion, accurate identification and reliable application of drawings in complex engineering contexts. Summary of the invention
[0005] The purpose of this application is to provide a method and system for intelligent cross-format conversion of engineering drawings, so as to solve the problems in the prior art that drawings are prone to structural distortion, semantic confusion and inconsistent visual expression during format conversion due to the lack of structural difference perception and modeling capabilities, difficulty in decoupling multimodal information and insufficient support for user personalized style adaptation, which further affects the technical problems of efficient conversion, accurate identification and reliable application of drawings in complex engineering contexts.
[0006] In view of the above problems, the present application provides a method and system for intelligent cross-format conversion of engineering drawings.
[0007] In a first aspect, the present application provides a method for cross-format intelligent conversion of engineering drawings, which is implemented through a cross-format intelligent conversion system for engineering drawings, including: after loading a first format engineering drawing, activating a multimodal fusion structure recognition engine and obtaining auxiliary input information from a user; performing feature extraction based on the auxiliary input information, and establishing recognition attention constraints, wherein the recognition attention constraints include type attention constraints and task attention constraints; after initializing the multimodal fusion structure recognition engine using the recognition attention constraints, performing multimodal content extraction of the first format engineering drawing, and establishing a primitive extraction result, wherein the primitive extraction result is bound to a semantic identifier; parsing the auxiliary input information, and obtaining a second format, wherein the second format is a conversion target format; performing multi-dimensional format difference modeling on the first format and the second format, and generating a difference compensation network; using the difference compensation network to perform format adaptation optimization of the primitive extraction result, and using the bound semantic identifier for optimization verification, to establish a second format engineering drawing.
[0008] Preferably, the method for intelligent cross-format conversion of engineering drawings also includes: activating a structural difference perception encoder in a difference compensation network, performing structural mapping processing on the primitive extraction result, using spatial topological relationship features, semantic identifiers, and layer attributes to perform structural vectorization encoding to obtain a primitive structure representation; activating a format-driven difference mapper in the difference compensation network, using the format-driven difference mapper to perform primitive type mapping, layer semantic reconstruction, and expression conversion adaptation of the primitive structure representation to generate an adaptation optimization result.
[0009] Preferably, the method for intelligent cross-format conversion of engineering drawings also includes: calling the adaptation optimization result, performing multimodal content extraction under identification focus constraints on the adaptation optimization result, and generating a verification extraction result; performing semantic consistency verification using the verification semantic identifier of the verification extraction result and the bound semantic identifier, and generating a semantic verification score; and completing optimization verification according to the semantic verification score.
[0010] Preferably, the method for intelligent cross-format conversion of engineering drawings also includes: calling a type-associated vocabulary library according to the type attention constraint, and setting the correlation coefficient within the type-associated vocabulary library; establishing a task-oriented attention activation mechanism according to the task attention constraint, the identification attention vector of the task-oriented attention activation mechanism includes a regional attention vector and a priority identification vector; and using the type-associated vocabulary library and the task-oriented attention activation mechanism to complete the initialization of the multimodal fusion structure recognition engine.
[0011] Preferably, the method for intelligent cross-format conversion of engineering drawings also includes: after using the multimodal fusion structure recognition engine to decouple the multimodal data of the first format engineering drawings, establishing a shared representation space, the multimodal data in the shared representation space is data fused into a unified vector representation; introducing a task-oriented attention activation mechanism to identify attention constraints to guide the activation of feature channels to perform traversal extraction in the shared representation space, and calling the type association vocabulary for semantic recognition to complete multimodal content extraction.
[0012] Preferably, the method for intelligent cross-format conversion of engineering drawings also includes: obtaining a calibrated zoom ratio of an engineering drawing in a first format; performing a zoom display verification on the engineering drawing in a second format according to the calibrated zoom ratio, and generating a zoom display verification result; if the zoom display verification result is a verification failure result, reporting a display abnormality.
[0013] Preferably, the method for intelligent cross-format conversion of engineering drawings also includes: determining whether the user selects a display adaptation optimization instruction; if the user selects a display adaptation optimization instruction, using the zoom display verification result to locate anomalies; using the anomaly location to establish display optimization feedback, and using the display optimization feedback to optimize the engineering drawings in the second format.
[0014] Preferably, the method for intelligent cross-format conversion of engineering drawings also includes: reading the user's graphic element visual expression style database; using the graphic element visual expression style database to perform color feature, line width feature, and attention font feature detail conversion of the adaptation optimization results; and establishing a second format engineering drawing based on the detail conversion results.
[0015] Preferably, the method for intelligent cross-format conversion of engineering drawings also includes: determining whether the engineering drawings in the first format are joint task drawings; if the engineering drawings in the first format are joint task drawings, establishing a temporary shared encoder during the conversion of the engineering drawings in the first format; and using the temporary shared encoder to perform task collaborative optimization of all joint task drawings.
[0016] In a second aspect, the present application also provides a cross-format intelligent conversion system for engineering drawings, which is used to execute a cross-format intelligent conversion method for engineering drawings as described in the first aspect, including: an engine activation module, which is used to activate a multimodal fusion structure recognition engine after loading a first format engineering drawing, and obtain auxiliary input information from a user; a feature extraction module, which is used to perform feature extraction according to the auxiliary input information, and establish recognition attention constraints, wherein the recognition attention constraints include type attention constraints and task attention constraints; a primitive extraction module, which is used to initialize the multimodal fusion structure recognition engine using the recognition attention constraints, perform multimodal content extraction of the first format engineering drawing, and establish a primitive extraction result, wherein the primitive extraction result is bound to a semantic identifier; an auxiliary parsing module, which is used to parse the auxiliary input information and obtain a second format, wherein the second format is a conversion target format; a difference modeling module, which is used to perform multi-dimensional format difference modeling on the first format and the second format, and generate a difference compensation network; a format optimization module, which is used to perform format adaptation optimization of the primitive extraction result using the difference compensation network, and perform optimization verification using the bound semantic identifier to establish a second format engineering drawing.
[0017] The technical solution provided in this application has at least the following technical effects or advantages: by realizing the technical goal of intelligent drawing format conversion based on difference perception modeling and multimodal fusion recognition, the technical effect of improving the drawing structure restoration, semantic recognition accuracy and visual style matching is achieved.
[0018] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented according to the contents of the specification, and in order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are specifically cited below. It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become easy to understand through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the present application or the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings in the following description are only exemplary, and for ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0020] Figure 1 A schematic diagram of a process for intelligently converting engineering drawings across formats for this application;
[0021] Figure 2 This is a schematic diagram of the structure of a cross-format intelligent conversion system for engineering drawings in this application.
[0022] Explanation of the reference numerals: engine activation module 11 , feature extraction module 12 , primitive extraction module 13 , auxiliary parsing module 14 , difference modeling module 15 , format optimization module 16 . DETAILED DESCRIPTION
[0023] This application provides a method and system for intelligent cross-format conversion of engineering drawings, which solves the existing problems in the prior art, such as the lack of structural difference perception and modeling capabilities, difficulty in decoupling multimodal information, and insufficient support for user personalized style adaptation, which leads to structural distortion, semantic confusion, and inconsistent visual expression in the process of format conversion of drawings, further affecting the efficient conversion, accurate recognition, and reliable application of drawings in complex engineering contexts. By achieving the technical goal of intelligent drawing format conversion based on difference perception modeling and multimodal fusion recognition, the technical effect of improving the degree of structural restoration of drawings, the accuracy of semantic recognition, and the degree of matching of visual styles is achieved.
[0024] Below, the technical solutions in the present application will be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all of the embodiments of the present application. It should be understood that the present application is not limited to the example embodiments described herein. Based on the embodiments of the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present application. It should also be noted that, for the convenience of description, only the parts related to the present application are shown in the accompanying drawings, rather than all of them.
[0025] For example, please refer to the attached Figure 1 The present application provides a method for cross-format intelligent conversion of engineering drawings, which is applied to a cross-format intelligent conversion system for engineering drawings, and specifically includes the following steps:
[0026] S1: After loading the first format engineering drawing, activate the multimodal fusion structure recognition engine and obtain the user's auxiliary input information.
[0027] Specifically, after loading the first-format engineering drawing, an original drawing file provided by the user is received and read. The drawing may be in a standard format such as DWG, DXF or PDF. The first format refers to the data encoding method currently used by the engineering drawing, including graphic information, annotation text, layer structure, etc. Loading includes opening the file, parsing the elements in the drawing, and establishing a preliminary data representation structure to provide an input basis for subsequent intelligent recognition.
[0028] Next, the multimodal fusion structure recognition engine is a comprehensive recognition system that combines image processing, natural language understanding, and spatial structure analysis. Multimodal representation can simultaneously process multiple types of data such as graphic lines, text annotations, legend symbols, layer information, etc., integrate multimodal data into a unified computing framework, extract different component units from the drawings and their spatial and logical relationships, such as identifying walls, pipelines, text descriptions, and their relative positional relationships to obtain a multimodal fusion structure recognition engine. Activate the multimodal fusion structure recognition engine, obtain the user's auxiliary input information, and coordinate the processing of information from different modes so that the structural meaning of the entire drawing can be accurately understood. The user's auxiliary input information is provided by the user through interface interaction during the system operation process, such as specifying the target drawing format, setting the layers to be retained or ignored, annotation preferences, scaling parameters, output style templates, etc. Auxiliary input information may include information descriptions of drawings. For example, if the first format of engineering drawings is described as architectural drawings, rectangles will be identified as architectural-related content. It may also include user attention information, such as parts that need to be fully displayed, or parts that need to be partially enlarged. This information is then used for targeted optimization of the recognition engine, which can determine recognition priorities, filter out irrelevant content, and improve the personalization and accuracy of conversion results.
[0029] S2: Extract features according to the auxiliary input information and establish recognition attention constraints, where the recognition attention constraints include type attention constraints and task attention constraints.
[0030] Specifically, after receiving the auxiliary information provided by the user, the key content of the drawing information is actively extracted from the drawing. The auxiliary input information may include key areas marked by the user, layers of interest, specified target formats, preset style templates, etc. The length, color, position of the primitives, font style, etc. of the lines are extracted by scanning, recognizing and encoding the drawing content for subsequent analysis and processing. After the feature extraction is completed, the recognition attention constraints are established based on the features. The recognition attention constraints are a set of rules or guiding mechanisms used to guide the recognition engine to prioritize the information when processing the drawing content, so as to focus on the areas that the user cares about and reduce unnecessary information interference, thereby improving the recognition efficiency and accuracy.
[0031] Recognition attention constraints are further divided into type attention constraints and recognition attention constraints. Type attention constraints refer to attention to specific content in the drawing. For example, if the user only wants to identify switches, sockets and lines in the electrical drawing, the recognition will be guided to focus on switches and sockets, and not to identify irrelevant building walls or furniture elements. Task attention constraints include prioritizing the content in the drawing, thereby allocating more resources and computing power when extracting content.
[0032] S3: After initializing the multimodal fusion structure recognition engine using the recognition attention constraint, perform multimodal content extraction of the engineering drawing in the first format and establish a primitive extraction result, wherein the primitive extraction result is bound with a semantic identifier.
[0033] Specifically, before the recognition task begins, the recognition engine is activated and configured using recognition attention constraints for a set of recognition target ranges and priorities defined by user input or preset rules. Then, multimodal content extraction of the first-format engineering drawing is performed to identify and analyze the various data contents in the original drawing, extract graphic primitives (such as line segments, circles, text annotations, etc.), text information (such as room names, material descriptions), and data such as layers and layout structures, and establish primitive extraction results. Primitives are the basic components of drawings, including lines, arcs, polygons, symbols, text blocks, etc. The primitive extraction results are bound to semantic identifiers, and corresponding semantic information will be attached to each primitive while extracting it. Semantic identifiers are descriptions of the function or meaning of primitives. For example, a line segment is not just a geometric figure, but may represent a pipe, wire, or wall boundary. Semantic identifiers are generated through a preset type vocabulary or machine learning model, making subsequent format conversion and information extraction more intelligent and efficient.
[0034] S4: Parse the auxiliary input information to obtain a second format, where the second format is a conversion target format.
[0035] Specifically, the additional guidance data provided by the user is analyzed during the processing. The auxiliary input information is usually manually input by the user, or it can come from a preset template, interactive operation, or historical usage habits. The parsing process includes semantic understanding of the auxiliary input information, format recognition, and matching with the current drawing structure to ensure that the subsequent conversion direction is clear and in line with expectations, and then obtain the second format and determine the target format to convert the drawing from the original format.
[0036] S5: Perform multi-dimensional format difference modeling on the first format and the second format to generate a difference compensation network.
[0037] Specifically, multi-dimensional format difference modeling is performed on the first format and the second format, that is, a comprehensive modeling analysis is performed on the differences between the original engineering drawing format and the target output format at the structural level and the semantic level, and then not only the visual presentation of the drawing is considered, but also the organizational structure of the graphic elements, the expression of data fields, the layer hierarchy, the attribute nesting rules, the unit system, the coordinate system reference and other technical differences are generated to generate a difference compensation network, which can solve the technical problems of incompatibility or misalignment between different formats. Among them, the difference compensation network can be a combination of a rule engine and a neural network, or it can be a structure optimized by a graph neural network or an attention mechanism.
[0038] S6: Utilize the difference compensation network to perform format adaptation optimization of the primitive extraction result, and utilize the bound semantic identifier to perform optimization verification to establish a second format engineering drawing.
[0039] Specifically, the neural network used for format conversion is called to adjust the structure and style of the extracted primitives. The difference compensation network is a structure based on deep learning that can identify the differences in expression rules between engineering drawings in the first format and the second format, and adaptively process the primitives accordingly. The result of primitive extraction is the graphic elements identified from the original drawing in the previous step, which may include lines, graphics, symbols, and text, etc. The process of format adaptation optimization is to make these graphic elements display correctly in the target format and meet the semantic and visual specifications, such as converting the left-aligned annotation in the source drawing to center-aligned in the target drawing to meet new software specifications or industry standards.
[0040] Next, optimization verification is performed using bound semantic identifiers. Semantic identifiers refer to the text or symbol information added to each graphic element during the primitive extraction process to describe its specific meaning. Bound semantic identifiers ensure that the primitive does not lose its original meaning during the format conversion process. Optimization verification is to compare the primitive structure after difference compensation with the original semantics to ensure that the primitives in the new drawing are not only correct in position and style, but also retain their functional meaning.
[0041] Finally, the second-format engineering drawing is created. The second-format engineering drawing is the conversion target, which is a drawing version suitable for different platforms, software or standards, contains all the elements that have been adapted to the format, and meets user expectations in terms of visual style and semantic accuracy.
[0042] Furthermore, the present application also includes: activating the structural difference perception encoder in the difference compensation network, performing structural mapping processing on the primitive extraction result, using spatial topological relationship features, semantic identifiers, and layer attributes to perform structural vectorization encoding to obtain the primitive structure representation; activating the format-driven difference mapper in the difference compensation network, using the format-driven difference mapper to perform primitive type mapping, layer semantic reconstruction, and expression conversion adaptation of the primitive structure representation to generate an adaptation optimization result.
[0043] Specifically, activating the structural difference perception encoder in the difference compensation network means calling a coding module with structural understanding ability in the difference compensation stage to identify and process the structural differences between the input primitives, that is, the structural difference perception encoder is used to perceive and analyze the relative position, connection mode and arrangement mode of the primitives to obtain the spatial topological relationship characteristics. For example, in an engineering drawing, the walls, doors and windows of a room may be represented by different primitive combinations in different formats. Semantic identification is a descriptive label bound in the process of primitive extraction, which is used to explain the meaning of a certain primitive, such as a polygon is a wall and a line is a cable. Layer attributes include the layer name, layer color, transparency, printing control, etc. of the primitive, which determines the context and presentation level of the primitive. Structural mapping processing is to structure the relationship between primitives into a graph, use nodes and edges to represent the primitives and their connections, and convert the representation of the graph into a numerical vector through structural vectorization encoding, that is, the primitive structure representation, which provides an input basis for subsequent format adaptation.
[0044] Next, activating the format-driven difference mapper in the difference compensation network means starting a submodule that specifically performs structural conversion for format differences. The format-driven difference mapper is responsible for a series of intelligent conversions of the above-mentioned primitive structural representations, including primitive type mapping, that is, mapping the primitive type defined in one format to the corresponding or equivalent type in another format. Layer semantic reconstruction is to adjust the layer organization structure to adapt to the different requirements of the target format for layer classification or hierarchical structure, ensuring the semantic consistency of the primitives. Expression conversion adaptation is to deal with the differences in the presentation of primitives at the visual and data levels, such as from vector to dot matrix, from continuous line segment expression to closed graphic representation, etc. The result finally generated by processing is the adaptation optimization result, which lays a stable and usable foundation for the final output of subsequent format conversion. Table 1 is a record of the most recent adaptation optimization results.
[0045]
[0046] Furthermore, the present application also includes: calling the adaptation optimization result, performing multimodal content extraction under identification focus constraints on the adaptation optimization result, and generating a verification extraction result; performing semantic consistency verification using the verification semantic identifier and the bound semantic identifier of the verification extraction result, and generating a semantic verification score; and completing optimization verification according to the semantic verification score.
[0047] Specifically, calling the adaptation optimization result means taking the content that has completed primitive type mapping, layer semantic reconstruction and expression conversion as new input, entering the processing flow of the current step, and generating a verification extraction result by performing multimodal content extraction under the recognition attention constraint, that is, extracting information again from the optimized drawing to verify its accuracy. Recognition attention constraints are used to limit the key focus areas, focus types or priority levels when extracting information. Multimodal content extraction refers to the comprehensive extraction of various types of data such as graphics, text, symbols, structures, etc. from drawings.
[0048] Next, the extracted primitive semantic labels are compared with the original semantic labels established previously. The verification semantic identification is based on the semantic information automatically identified in the current optimized data, while the bound semantic identification is the semantic description manually or model-identified attached to the original primitive in the initial extraction stage. The purpose of semantic consistency verification is to determine the degree of match between the two in terms of understanding, classification and attribution, that is, to verify whether the converted drawings remain consistent at the semantic level, and finally generate a semantic verification score. For example, if the similarity reaches 90 points or more, the conversion can be considered accurate, and if it is less than 60 points, it needs to be re-optimized.
[0049] Finally, the optimization check is completed according to the semantic verification score, and the optimization process is reviewed and confirmed based on the generated score results. If the score is high enough, it means that the converted drawing content is highly semantically restored to the original design and can be directly output; if the score is low, it is necessary to return to the previous step to readjust the format mapping, layer semantics or content extraction strategy to ensure that the output drawings are not only formatted correctly, but also semantically complete, to avoid design errors or construction deviations caused by misunderstandings in engineering applications.
[0050] Furthermore, the present application also includes: calling a type-associated vocabulary library according to the type attention constraint, and setting a correlation coefficient within the type-associated vocabulary library; establishing a task-oriented attention activation mechanism according to the task attention constraint, and the recognition attention vector of the task-oriented attention activation mechanism includes a regional attention vector and a priority recognition vector; and using the type-associated vocabulary library and the task-oriented attention activation mechanism to complete the initialization of the multimodal fusion structure recognition engine.
[0051] Specifically, according to the type of object of interest set by the user, such as doors and windows, electrical equipment, pipes, etc., the key words, graphic symbols and semantic labels related to these objects are called out from the built-in type-related vocabulary. The type-related vocabulary is a structured knowledge base that contains the names, abbreviations, shape features and semantic information of common graphic elements in various engineering drawings. For example, if the type of interest set by the user is the fire protection system, then fire-related terms such as sprinkler heads, fire hydrants, alarms, etc. are automatically extracted for subsequent recognition and matching of drawing content. The relevance coefficient within the type-related vocabulary is set to assign a matching degree score with the current recognition task to different terms in the vocabulary, which is used to affect the processing priority of the terms during recognition, so that terms with high relevance will be given priority.
[0052] Next, a task-oriented attention activation mechanism is established based on the task attention constraints, thereby enhancing the recognition engine's ability to focus on key areas and high-priority content. The task attention constraint is the recognition task content set by the user, which will guide the task-oriented attention activation mechanism to focus on analyzing the specified part of the drawing. The task-oriented attention activation mechanism is a mechanism that simulates human observation habits to give more computing resources to more important information in the drawing. In the task-oriented attention activation mechanism, the recognition attention vector is composed of a regional attention vector and a priority recognition vector. The regional attention vector is a spatial vector that represents the user's attention area and is used to locate the focus area in the drawing; while the priority recognition vector represents the importance level of the primitive or content. High-priority content will be processed first to improve the accuracy and efficiency of recognition.
[0053] Finally, the type-related vocabulary and task-oriented attention activation mechanism will be used to initialize the multimodal fusion structure recognition engine. Before the recognition task officially begins, the parameters, model weights, and processing procedures within the recognition engine are configured and adjusted based on the obtained vocabulary, priority, attention distribution, and other information.
[0054] Furthermore, the present application also includes: after using the multimodal fusion structure recognition engine to decouple the multimodal data of the first format engineering drawing, a shared representation space is established, and the multimodal data in the shared representation space is data fused into a unified vector representation; a task-oriented attention activation mechanism is introduced to identify attention constraints to guide the activation of feature channels to perform traversal extraction in the shared representation space, and call the type association vocabulary for semantic recognition to complete multimodal content extraction.
[0055] Specifically, when processing the input engineering drawings, the multimodal fusion structure recognition engine is used to classify and process different forms of information in the drawings. Multimodality refers to the various data forms that exist in the drawings, such as image information (such as primitive shapes), text information (such as annotations and instructions), and spatial structure information (such as layers and relative positions). Different types of data are extracted independently for subsequent fusion and analysis. Next, a shared representation space is established to convert the decoupled image, text, and structural information into the same numerical expression, namely the unified vector representation. The unified vector representation encodes different types of data into vectors with a unified structure, which can then be collaboratively calculated and processed in the same algorithm model.
[0056] The task-oriented attention activation mechanism refers to a mechanism that determines the more important channels of data according to the needs of the recognition task when processing multimodal data. Feature channels refer to the channels of various feature mappings after the drawings are processed by the neural network, such as color channels, edge channels, text channels, etc. When the task-oriented attention activation mechanism is activated, traversal extraction will be performed in the shared representation space, that is, a comprehensive scan of the entire vector space to find high-response areas related to the target of attention. At the same time, the type-associated vocabulary is called for semantic recognition. The type-associated vocabulary provides keywords, graphic primitive symbols, and semantic labels related to the recognition task. By comparing the extracted vectors with the entries in the vocabulary, the actual meaning of each graphic primitive or element is determined, thereby completing the extraction of multimodal content.
[0057] Furthermore, the present application also includes: obtaining a calibrated zoom ratio of an engineering drawing in a first format; performing a zoom display check on the engineering drawing in a second format according to the calibrated zoom ratio, and generating a zoom display check result; if the zoom display check result is a check failure result, reporting a display abnormality.
[0058] Specifically, the calibration scaling factor or calibration conversion ratio of the current drawing in the first format is identified and extracted, which is used to map the drawing from the original design size to the size range of screen display or print output. The calibration scaling ratio may come from the engineering software setting or the user's manual input. For example, if the display width of an architectural drawing with an original size of 1000 mm is set to 500 mm, the calibration scaling ratio is 0.5.
[0059] Subsequently, the drawing content in the second format is proportionally adjusted according to the calibrated scaling ratio in the first format, and the adjusted drawing is proofread in the view to check whether its display size, scale, and content layout are consistent with the original drawing, and generate a scaling display verification result.
[0060] If the scaling display verification result is a verification failure, it means that there are errors or deviations in the display between the second format drawing after proportional adjustment and the original drawing, such as text misalignment, element deformation, or boundaries outside the expected display range. A display abnormality is reported to remind the user that there may be information loss or proportional distortion problems in the drawing during the conversion process, and manual intervention or readjustment of drawing parameters are required to ensure accuracy.
[0061] Furthermore, the present application also includes: determining whether the user selects a display adaptation optimization instruction; if the user selects a display adaptation optimization instruction, using the zoom display verification result to locate anomalies; using the anomaly location to establish display optimization feedback, and using the display optimization feedback to optimize the second format engineering drawings.
[0062] Specifically, when executing the drawing conversion and verification process, it is detected whether the user has actively enabled the relevant display optimization function. The display adaptation optimization command is an option that the user can check or operate in the interface to control whether to start the subsequent process of automatically adjusting the drawing display, thereby improving the consistency of drawing display between different formats.
[0063] Next, if the user selects the display adaptation optimization command, the scaling differences or display anomalies identified during the verification process will be used as a reference to accurately find the areas with display deviations in the drawings. Abnormal location mainly relies on the error data, element offset information or boundary exceeding records contained in the scaling display verification results. For example, element offset exceeding 10 mm, text overlap, layer information loss, etc. can all be used as abnormal points.
[0064] Subsequently, all located problem areas are integrated into a feedback mechanism. The feedback information includes abnormal position, offset, affected layer, etc., which are used to guide the subsequent drawing correction action, that is, the second format engineering drawing is optimized based on the display optimization feedback. The optimization content may include element position adjustment, layer rearrangement, font size correction, etc., so that the final display effect is highly consistent with the first format drawing.
[0065] Furthermore, the present application also includes: reading the user's graphic primitive visual expression style database; using the graphic primitive visual expression style database to perform color feature, line width feature, and attention font feature detail conversion of the adaptation optimization results; and establishing a second format engineering drawing based on the detail conversion results.
[0066] Specifically, data resources containing the visual style of drawings are extracted from user presets or historical records. The graphic element visual expression style database stores the color matching, line thickness, font type and other contents commonly used by users in drawing drawings. It is a collection of user personalized settings to ensure that the converted drawings conform to user habits.
[0067] Next, the adaptation optimization results are converted into details of color features, line width features, and attention font features using the graphic element visual expression style database. The adaptation optimization result refers to the drawing data obtained after structural reconstruction and semantic adjustment. Although the content is consistent with the original drawing, its visual style may not have been adjusted to the style that the user is accustomed to. Therefore, according to the information provided by the database, the color attributes, line thickness, and fonts used for annotations of the graphic elements are finely adjusted in turn to keep their visual style consistent with the user's original drawing. Then, the second format engineering drawing is established based on the completed detail conversion results. All graphic elements that have undergone detail conversion are reorganized and arranged to ensure that each graphic element visually conforms to the preset style, and finally a new drawing that meets the user's visual habits is generated.
[0068] Furthermore, the present application also includes: determining whether the engineering drawings in the first format are joint task drawings; if the engineering drawings in the first format are joint task drawings, establishing a temporary shared encoder during the conversion of the engineering drawings in the first format; and using the temporary shared encoder to perform task collaborative optimization of all joint task drawings.
[0069] Specifically, before the conversion work begins, the input engineering drawing format is identified to analyze whether it contains multiple design tasks or content drawn collaboratively by multiple submodules. Joint task drawings refer to drawings completed by multiple disciplines (such as structure, electrical, HVAC, etc.) or multiple design stages. They may integrate the design information of multiple subsystems and require unified coordination and management. Therefore, by identifying the complexity and multi-task characteristics of the drawings, we can prepare for the subsequent conversion.
[0070] Then, if the recognition result shows that the drawing belongs to a joint task drawing, a temporary shared encoder will be established during the drawing conversion process. The temporary shared encoder is an intermediate processing module that is used to uniformly encode and standardize the information of multiple different tasks or sub-drawings, that is, to establish a temporary translation center to convert the "language" of each sub-task into a common format, so that all information can circulate in the same semantic space to avoid conflicts or losses during the conversion process.
[0071] Then, a temporary shared encoder is used to perform task collaborative optimization on all joint task drawings. Task collaborative optimization not only converts the content of each sub-drawing, but also pays attention to the mutual relationship and dependency between each sub-drawing, adjusts and optimizes the consistency of information expression between tasks, the uniformity of drawing formats, and the logical relationship between graphics. For example, the position of a load-bearing beam in a structural drawing may also affect the wiring path in the electrical drawing. Collaborative optimization can avoid contradictions and conflicts in these cross-designs.
[0072] To sum up, the method for intelligent cross-format conversion of engineering drawings provided in this application has the following technical effects: by realizing the technical goal of intelligent drawing format conversion based on difference perception modeling and multimodal fusion recognition, the technical effect of improving the drawing structure restoration, semantic recognition accuracy and visual style matching is achieved.
[0073] Embodiment 2, based on the same inventive concept as the method for cross-format intelligent conversion of engineering drawings in the above embodiment, the present application also provides a system for cross-format intelligent conversion of engineering drawings, see the attached Figure 2 , including: an engine activation module 11, which is used to activate the multimodal fusion structure recognition engine after loading the first format engineering drawing, and obtain the user's auxiliary input information; a feature extraction module 12, which is used to perform feature extraction according to the auxiliary input information and establish recognition attention constraints, wherein the recognition attention constraints include type attention constraints and task attention constraints; a primitive extraction module 13, which is used to perform multimodal content extraction of the first format engineering drawing after initializing the multimodal fusion structure recognition engine using the recognition attention constraints, and establish primitive extraction results, wherein the primitive extraction results are bound to semantic identifiers; an auxiliary parsing module 14, which is used to parse the auxiliary input information and obtain a second format, wherein the second format is a conversion target format; a difference modeling module 15, which is used to perform multi-dimensional format difference modeling on the first format and the second format, and generate a difference compensation network; a format optimization module 16, which is used to perform format adaptation optimization of the primitive extraction results using the difference compensation network, and perform optimization verification using the bound semantic identifiers to establish a second format engineering drawing.
[0074] Furthermore, the cross-format intelligent conversion system for engineering drawings is also used to: activate the structural difference perception encoder in the difference compensation network, perform structural mapping on the primitive extraction results, use spatial topological relationship features, semantic identifiers, and layer attributes to perform structural vectorization encoding to obtain primitive structural representation; activate the format-driven difference mapper in the difference compensation network, use the format-driven difference mapper to perform primitive type mapping, layer semantic reconstruction, and expression conversion adaptation of the primitive structural representation, and generate adaptation optimization results.
[0075] Furthermore, the cross-format intelligent conversion system for engineering drawings is also used to: call the adaptation optimization result, perform multimodal content extraction under identification focus constraints on the adaptation optimization result, and generate a verification extraction result; perform semantic consistency verification using the verification semantic identifier and the bound semantic identifier of the verification extraction result to generate a semantic verification score; and complete optimization verification according to the semantic verification score.
[0076] Furthermore, the cross-format intelligent conversion system for engineering drawings is also used to: call the type association vocabulary according to the type attention constraint, and set the correlation coefficient within the type association vocabulary; establish a task-oriented attention activation mechanism according to the task attention constraint, and the recognition attention vector of the task-oriented attention activation mechanism includes a regional attention vector and a priority recognition vector; and use the type association vocabulary and the task-oriented attention activation mechanism to complete the initialization of the multimodal fusion structure recognition engine.
[0077] Furthermore, the cross-format intelligent conversion system for engineering drawings is also used to: establish a shared representation space after decoupling the multimodal data of the first format engineering drawings using the multimodal fusion structure recognition engine, wherein the multimodal data in the shared representation space is data fused into a unified vector representation; introduce a task-oriented attention activation mechanism to identify attention constraints to guide the activation of feature channels to perform traversal extraction in the shared representation space, and call the type-associated vocabulary for semantic recognition to complete multimodal content extraction.
[0078] Furthermore, the cross-format intelligent conversion system for engineering drawings is also used to: obtain a calibrated zoom ratio of an engineering drawing in a first format; perform zoom display verification on the engineering drawing in a second format according to the calibrated zoom ratio, and generate a zoom display verification result; if the zoom display verification result is a verification failure result, report a display abnormality.
[0079] Furthermore, the cross-format intelligent conversion system for engineering drawings is also used to: determine whether the user selects the display adaptation optimization instruction; if the user selects the display adaptation optimization instruction, use the zoom display verification result to locate the abnormality; use the abnormality location to establish display optimization feedback, and use the display optimization feedback to optimize the second format engineering drawings.
[0080] Furthermore, the cross-format intelligent conversion system for engineering drawings is also used to: read the user's graphic element visual expression style database; use the graphic element visual expression style database to perform color feature, line width feature, and attention font feature detail conversion of the adaptation optimization results; and establish a second format engineering drawing based on the detail conversion results.
[0081] Furthermore, the cross-format intelligent conversion system for engineering drawings is also used to: determine whether the engineering drawings in the first format are joint task drawings; if the engineering drawings in the first format are joint task drawings, establish a temporary shared encoder during the conversion of the engineering drawings in the first format; and use the temporary shared encoder to perform task collaborative optimization of all joint task drawings.
[0082] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The method for cross-format intelligent conversion of engineering drawings and the specific examples in the aforementioned embodiment 1 are also applicable to a system for cross-format intelligent conversion of engineering drawings in this embodiment. Through the aforementioned detailed description of the method for cross-format intelligent conversion of engineering drawings, those skilled in the art can clearly understand the system for cross-format intelligent conversion of engineering drawings in this embodiment, so for the sake of brevity of the specification, it will not be described in detail here.
[0083] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
[0084] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application belong to the scope of the present application and its equivalent technology, the present application is also intended to include these modifications and variations.
Claims
1. A method for intelligently converting engineering drawings across formats, characterized in that: include: After loading the engineering drawing in the first format, activating the multimodal fusion structure recognition engine and obtaining the auxiliary input information of the user; Perform feature extraction according to the auxiliary input information and establish recognition attention constraints, wherein the recognition attention constraints include type attention constraints and task attention constraints; After initializing the multimodal fusion structure recognition engine by using the recognition attention constraint, performing multimodal content extraction of the engineering drawing in the first format, and establishing a primitive extraction result, wherein the primitive extraction result is bound with a semantic identifier; Parsing the auxiliary input information to obtain a second format, where the second format is a conversion target format; Performing multi-dimensional format difference modeling on the first format and the second format to generate a difference compensation network; The difference compensation network is used to perform format adaptation optimization of the primitive extraction result, and the bound semantic identifier is used to perform optimization verification to establish a second format engineering drawing.
2. The method for intelligently converting engineering drawings across formats as claimed in claim 1, characterized in that: The method of using the difference compensation network to perform format adaptation optimization of the primitive extraction result includes: Activate the structural difference perception encoder in the difference compensation network, perform structural mapping processing on the primitive extraction result, use spatial topological relationship features, semantic identifiers, and layer attributes to perform structural vectorization encoding, and obtain primitive structural representation; The format-driven difference mapper in the difference compensation network is activated, and the format-driven difference mapper is used to perform primitive type mapping, layer semantic reconstruction and expression conversion adaptation of primitive structure representation to generate an adaptation optimization result.
3. The method for intelligently converting engineering drawings across formats as claimed in claim 2, characterized in that: The optimization verification using the bound semantic identifier includes: Calling the adaptation optimization result, performing multimodal content extraction under identification focus constraints on the adaptation optimization result, and generating a verification extraction result; Performing semantic consistency verification using the verification semantic identifier of the verification extraction result and the bound semantic identifier to generate a semantic verification score; The optimization verification is completed according to the semantic verification score.
4. The method for cross-format intelligent conversion of engineering drawings according to claim 1, characterized in that: The initializing the multimodal fusion structure recognition engine by using the recognition attention constraint includes: Invoking a type-related vocabulary according to the type-focused constraint, and setting a correlation coefficient within the type-related vocabulary; Establishing a task-oriented attention activation mechanism according to the task attention constraint, wherein the identification attention vector of the task-oriented attention activation mechanism includes a regional attention vector and a priority identification vector; The multimodal fusion structure recognition engine is initialized using the type-associated vocabulary and the task-oriented attention activation mechanism.
5. The method for intelligently converting engineering drawings across formats as claimed in claim 4, characterized in that: The performing of multimodal content extraction of the engineering drawing in the first format and establishing a primitive extraction result includes: After decoupling the multimodal data of the engineering drawing in the first format by using the multimodal fusion structure recognition engine, a shared representation space is established, wherein the multimodal data in the shared representation space is data fused into a unified vector representation; A task-oriented attention activation mechanism is introduced to identify attention constraints to guide the activation of feature channels to perform traversal extraction in the shared representation space, and call the type-associated vocabulary for semantic recognition to complete multimodal content extraction.
6. The method for intelligently converting engineering drawings across formats as claimed in claim 1, characterized in that: After the second format engineering drawing is established, the following steps are included: Obtaining a nominal scaling ratio of a first-format engineering drawing; Performing a zoom display check of the engineering drawing in the second format according to the calibrated zoom ratio, and generating a zoom display check result; If the zoom display verification result is a verification failure result, a display abnormality is reported.
7. The method for intelligently converting engineering drawings across formats as claimed in claim 6, characterized in that: If the zoom display verification result is a verification failure result, a display abnormality is reported, including: Determine whether the user chooses to display the adaptation optimization instruction; If the user selects the display adaptation optimization instruction, the scaling display verification result is used to locate the abnormality; The abnormal positioning is used to establish display optimization feedback, and the second-format engineering drawing is optimized with the display optimization feedback.
8. The method for cross-format intelligent conversion of engineering drawings according to claim 1, characterized in that: Before establishing the second format engineering drawing, the method further includes: Read the user's graphic element visual expression style database; Utilizing the graphic primitive visual expression style database to convert color features, line width features, and attention font features of the adaptation optimization results; Create a second format engineering drawing based on the detail conversion results.
9. The method for intelligently converting engineering drawings across formats as claimed in claim 1, characterized in that: The step of loading the engineering drawing in the first format includes: determining whether the engineering drawing in the first format is a joint task drawing; If the first-format engineering drawing is a joint task drawing, establishing a temporary shared encoder during the conversion of the first-format engineering drawing; The temporary shared encoder is used to perform task collaborative optimization of all joint task drawings.
10. An intelligent cross-format conversion system for engineering drawings, characterized in that: The steps for implementing the method for cross-format intelligent conversion of engineering drawings as described in any one of claims 1 to 9 include: The engine activation module is used to activate the multimodal fusion structure recognition engine after loading the engineering drawing in the first format, and obtain the auxiliary input information of the user; A feature extraction module, used to extract features according to the auxiliary input information and establish recognition attention constraints, wherein the recognition attention constraints include type attention constraints and task attention constraints; A primitive extraction module, configured to perform multimodal content extraction of engineering drawings in a first format after initializing the multimodal fusion structure recognition engine using the recognition attention constraint, and to establish a primitive extraction result, wherein the primitive extraction result is bound with a semantic identifier; An auxiliary parsing module, used for parsing the auxiliary input information to obtain a second format, where the second format is a conversion target format; A difference modeling module, used for performing multi-dimensional format difference modeling on the first format and the second format to generate a difference compensation network; The format optimization module is used to use the difference compensation network to perform format adaptation optimization of the primitive extraction result, and use the bound semantic identifier to perform optimization verification to establish a second format engineering drawing.
Citation Information
Patent Citations
Two-dimensional electronic technical drawing format conversion and vectorization interaction system
CN109636887A
Drawing difference identification method, device and system and storage medium
CN116798062A
Method and device for digital intelligent management and application of drawings based on SVG (Scalable Vector Graphics) standard
CN118295963A
Data format conversion method and system, storage medium and program product
CN119088759A
Drawing difference intelligent detection method and system
CN119094663A
Cited By
Intelligent drawing management method and device and storage medium
CN120766306A
A drawing intelligent management method and device and a storage medium
CN120766306B
DWG drawing intelligent translation and information verification integrated method and system
CN121413630A
Cross-platform CAD drawing intelligent collaborative management method and system and medium
CN122174295A
Cross-platform cad drawing intelligent collaborative management method and system and medium
CN122174295B