A two-dimensional engineering drawing manufacturability reasoning method based on a multi-modal large model
By using a multimodal large model to automatically analyze the manufacturing parameters and process intentions of engineering drawings, the problem of time-consuming manual review is solved, and efficient and reliable manufacturability assessment is achieved, thereby improving the collaborative efficiency between design and manufacturing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- WUXI XUELANG DIGITAL TECH CO LTD
- Filing Date
- 2026-04-28
- Publication Date
- 2026-05-29
AI Technical Summary
In existing technologies, manufacturability analysis of engineering drawings relies on manual review, which is time-consuming, labor-intensive, difficult to scale up, affects production efficiency, and makes it difficult to guarantee the consistency and accuracy of analysis results.
A multimodal large model-based approach is adopted. By acquiring raw engineering data and determining the view to be inferred, a pre-trained multimodal large model is used to perform semantic decoding of manufacturing parameters and semantic recognition of process intent, generating a manufacturability report and achieving automatic parsing and manufacturing feasibility assessment.
It significantly improves analysis efficiency and the consistency and reliability of evaluation results, reduces human intervention, and enhances the collaborative efficiency of design and manufacturing.
Smart Images

Figure CN122113680A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of drawing processing technology, and more specifically, to a method for manufacturability reasoning of two-dimensional engineering drawings based on a multimodal large model. Background Technology
[0002] In mechanical design and intelligent manufacturing, two-dimensional engineering drawings can encompass all dimensions of manufacturing constraints, including geometric definitions, dimensional tolerances, and material processes. As product complexity increases and R&D cycles shorten, design departments urgently need real-time feedback on Design for Manufacturability (DFM) before drawings are released to avoid processing failures, rework, scrap, and supply chain delays caused by unreasonable designs.
[0003] Currently, manufacturability analysis of engineering drawings mainly relies on manual review by experienced process engineers. However, this approach is not only time-consuming and labor-intensive, but also difficult to scale up in scenarios with a large number and high complexity of drawings, thus impacting manufacturing efficiency. Summary of the Invention
[0004] The purpose of this application is to address the shortcomings of the prior art by providing a method for manufacturability reasoning of two-dimensional engineering drawings based on a multimodal large model, so as to solve the problems of long time consumption, high labor costs, and impact on production and manufacturing efficiency in the prior art regarding manufacturability analysis of engineering drawings.
[0005] To achieve the above objectives, the technical solutions adopted in the embodiments of this application are as follows: In a first aspect, one embodiment of this application provides a method for manufacturability reasoning based on a multimodal large model of two-dimensional engineering drawings, the method comprising: Obtain the original engineering data and determine at least one view to be inferred corresponding to the original engineering data; Based on the first multimodal large model obtained through pre-training, the manufacturing parameter semantic decoding is performed on each of the views to be reasoned to obtain the manufacturing parameter set corresponding to each of the views to be reasoned to. The manufacturing parameter set includes information on multiple manufacturing parameters, and the information on each manufacturing parameter includes: parameter label, parameter value, parameter unit, and the bounding box corresponding to the parameter. Based on the pre-trained second multimodal large model, process intention semantic recognition is performed on each of the views to be reasoned, and material process context structure information corresponding to each of the views to be reasoned is obtained. The material process context structure information is used to characterize the manufacturing intention related to materials and processes. Based on the set of manufacturing parameters corresponding to each of the views to be inferred and the material process context information, manufacturability inference is performed to generate a manufacturability report corresponding to the original engineering data.
[0006] Secondly, another embodiment of this application provides a manufacturability reasoning device for two-dimensional engineering drawings based on a multimodal large model, the device comprising: An acquisition module is used to acquire raw engineering data and determine at least one view to be inferred corresponding to the raw engineering data; The manufacturing parameter semantic decoding module is used to perform manufacturing parameter semantic decoding on each of the views to be reasoned based on the first multimodal large model obtained through pre-training, to obtain the manufacturing parameter set corresponding to each of the views to be reasoned. The manufacturing parameter set includes information on multiple manufacturing parameters, and the information on each manufacturing parameter includes: parameter label, parameter value, parameter unit, and the bounding box corresponding to the parameter. The process intent semantic recognition module is used to perform process intent semantic recognition on each of the views to be reasoned based on the pre-trained second multimodal large model, and obtain the material process context structure information corresponding to each of the views to be reasoned. The material process context structure information is used to characterize the manufacturing intent related to materials and processes. The manufacturability reasoning module is used to perform manufacturability reasoning based on the set of manufacturing parameters corresponding to each of the views to be reasoned and the material process context structure information, and generate a manufacturability report corresponding to the original engineering data.
[0007] Thirdly, another embodiment of this application provides an electronic device, including: a processor, a storage medium, and a bus, wherein the storage medium stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the storage medium via the bus, and the processor executes the machine-readable instructions to perform the steps of any of the methods described in the first aspect above.
[0008] Fourthly, another embodiment of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, performs the steps of any of the methods described in the first aspect above.
[0009] The beneficial effects of this application are as follows: By identifying multiple views to be inferred corresponding to the original engineering data, and performing semantic decoding of manufacturing parameters and semantic recognition of process intent on each view based on multiple pre-trained multimodal large models, the set of manufacturing parameters and material process context information corresponding to each view to be inferred are obtained. Based on the set of manufacturing parameters and material process context information corresponding to each view to be inferred, manufacturability inference is performed, generating a manufacturability report. This enables automatic parsing, semantic understanding, and manufacturing feasibility assessment of the original engineering data. Simultaneously, it reduces human intervention in manufacturability inference, significantly improving analysis efficiency, consistency and reliability of assessment results, and effectively enhancing the collaborative efficiency of design and manufacturing. Attached Figure Description
[0010] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 A flowchart illustrating a two-dimensional engineering drawing manufacturability reasoning method based on a multimodal large model, provided in an embodiment of this application; Figure 2 This is a schematic diagram of a process for generating a manufacturability report corresponding to the original engineering data in the manufacturability reasoning method for two-dimensional engineering drawings based on a multimodal large model provided in this application embodiment; Figure 3 This is a schematic diagram of a process for generating a manufacturability report corresponding to the original engineering data in the manufacturability reasoning method for two-dimensional engineering drawings based on a multimodal large model provided in this application embodiment; Figure 4 This is a flowchart illustrating the process of determining at least one risk information, at least one optimization information corresponding to each risk information, and efficiency information corresponding to each optimization information in the two-dimensional engineering drawing manufacturability reasoning method based on a multimodal large model provided in this application embodiment. Figure 5 This is a flowchart illustrating the process of obtaining the set of manufacturing parameters corresponding to each view to be reasoned in the two-dimensional engineering drawing manufacturability reasoning method based on a multimodal large model provided in this application embodiment. Figure 6 A schematic diagram of a spatial anchoring model in a two-dimensional engineering drawing manufacturability reasoning method based on a multimodal large model provided in this application embodiment; Figure 7This is a flowchart illustrating the process of obtaining the candidate region descriptor set corresponding to each view to be reasoned in the two-dimensional engineering drawing manufacturability reasoning method based on a multimodal large model provided in this application embodiment; Figure 8 This is a flowchart illustrating the process of obtaining the material and process context information corresponding to each view to be reasoned in the two-dimensional engineering drawing manufacturability reasoning method based on a multimodal large model provided in this application embodiment. Figure 9 This is a flowchart illustrating the process of determining at least one text image block corresponding to each view to be reasoned and the type label of each text image block in the manufacturability reasoning method for two-dimensional engineering drawings based on a multimodal large model provided in this application embodiment. Figure 10 This is a flowchart illustrating the process of obtaining the material and process context information corresponding to each view to be reasoned in the two-dimensional engineering drawing manufacturability reasoning method based on a multimodal large model provided in this application embodiment. Figure 11 This is a flowchart illustrating the process of determining at least one view to be reasoned in the two-dimensional engineering drawing manufacturability reasoning method based on a multimodal large model provided in this application embodiment; Figure 12 A schematic diagram of a two-dimensional engineering drawing manufacturability reasoning device based on a multimodal large model provided in this application embodiment; Figure 13 This is a schematic diagram of the electronic device structure provided in an embodiment of this application. Detailed Implementation
[0012] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the accompanying drawings in this application are for illustrative and descriptive purposes only and are not intended to limit the scope of protection of this application. Furthermore, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate operations implemented according to some embodiments of this application. It should be understood that the operations in the flowcharts may not be implemented in sequence, and steps without logical contextual relationships may be reversed or implemented simultaneously. In addition, those skilled in the art, guided by the content of this application, may add one or more other operations to the flowcharts, or remove one or more operations from the flowcharts.
[0013] Furthermore, the described embodiments are merely some, not all, of the embodiments of this application. The components of the embodiments of this application described and illustrated herein can typically be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0014] It should be noted that the term "comprising" will be used in the embodiments of this application to indicate the presence of the features declared thereafter, but does not exclude the addition of other features.
[0015] Currently, manufacturability analysis of engineering drawings mainly relies on manual review by experienced process engineers. However, this approach is not only time-consuming and labor-intensive, but also difficult to scale up in scenarios with a large number and high complexity of drawings, impacting production efficiency. Furthermore, it suffers from other problems such as difficulty in reusing experience, slow response to rule updates, and challenges in ensuring the consistency and accuracy of analysis results.
[0016] This application proposes a manufacturability reasoning method for two-dimensional engineering drawings based on a multimodal large model, addressing the aforementioned problems. By identifying multiple views to be reasoned about corresponding to the original engineering data, and using pre-trained multimodal large models to perform semantic decoding of manufacturing parameters and semantic recognition of process intent for each view, the method obtains the set of manufacturing parameters and material and process context information corresponding to each view. Based on these information, manufacturability reasoning is performed, generating a manufacturability report. This method enables automatic parsing, semantic understanding, and manufacturing feasibility assessment of the original engineering data. Furthermore, it reduces manual intervention in manufacturability reasoning, improves analysis efficiency and the consistency and reliability of assessment results, and enhances design and manufacturing collaboration efficiency.
[0017] The following describes in detail the manufacturability reasoning method for two-dimensional engineering drawings based on a multimodal large model provided in this application, with reference to several embodiments.
[0018] Figure 1 A flowchart illustrating a method for manufacturability reasoning based on a multimodal large model for two-dimensional engineering drawings provided in this application is shown below. Figure 1 As shown, the executing entity of this method can be any electronic device with processing capabilities, and the method includes: S101. Obtain the original engineering data and determine at least one view to be inferred corresponding to the original engineering data.
[0019] Optionally, the original engineering data can be obtained. Original engineering data refers to the engineering drawing data that needs to be analyzed. Original engineering data can be various types of engineering drawing data, such as two-dimensional vector drawings, scanned images, or three-dimensional solid models.
[0020] Optionally, after obtaining the original engineering data, the original engineering data can be processed to obtain multiple views to be inferred corresponding to the original engineering data.
[0021] In one example, if the original engineering data is a two-dimensional vector drawing, then the original engineering data is subjected to view recognition and cropping to obtain at least one view to be inferred corresponding to the original engineering data.
[0022] In another example, if the original engineering data is an image-type scanned drawing, the original engineering drawing is segmented into a structured view to identify multiple intermediate views such as the main view, side view, and sectional view. Each intermediate view is then cropped and scaled according to the minimum bounding rectangle to obtain multiple candidate views. Each candidate view is then subjected to quality inspection. Based on the quality inspection results of each candidate view, at least one view to be inferred is obtained.
[0023] Specifically, the quality inspection of each candidate view includes: detecting the number of valid geometric elements in each candidate view; if the number of valid geometric elements in a candidate view is greater than a preset threshold, the quality inspection result of the candidate view is determined to be passed. If the quality inspection result of a candidate view is passed, the candidate view is used as a view to be inferred.
[0024] S102. Based on the first multimodal large model obtained through pre-training, perform semantic decoding of manufacturing parameters for each view to be inferred to obtain the set of manufacturing parameters corresponding to each view to be inferred.
[0025] Optionally, after obtaining each view to be reasoned, the manufacturing parameter semantic decoding process can be performed on each view to be reasoned based on the first multimodal large model pre-trained and the first prompt words pre-constructed, to decode the manufacturing parameter set corresponding to each view to be reasoned.
[0026] For example, each view to be reasoned and the pre-constructed first prompt word can be input into the first multimodal large model. The first multimodal large model performs explicit semantic recognition and decoding on the manufacturing parameters in each view to be reasoned to obtain the set of manufacturing parameters corresponding to each view to be reasoned.
[0027] Specifically, the first multimodal large model utilizes embedded engineering semantic prior knowledge and cross-modal context modeling capabilities to perform robust parsing during inference. Specifically, it dynamically associates all visual elements within the labeled area through a cross-attention mechanism. When a local character cannot be recognized due to interference, it relies on already recognized context elements (e.g., detected verticality symbols). The model infers reasonable values for missing or ambiguous parts from the benchmark A. Meanwhile, the output of the first multimodal large model is constrained by engineering semantic legality rules, automatically filtering out analytical results that do not conform to the drawing specifications (such as non-numerical tolerance values and isolated diameter symbols), thus achieving end-to-end logical self-consistency and automatic error correction.
[0028] The manufacturing parameter set includes information on multiple manufacturing parameters, each of which includes: parameter label, parameter value, parameter unit, and the corresponding bounding box.
[0029] For example, a parameter label refers to the functional role that the manufacturing parameter plays in the manufacturing semantic space, including: dimensions with diameter symbols, perpendicularity, and surface roughness (symbol and value). A parameter value refers to the parsed value of the parameter.
[0030] Among them, the first multimodal large model is a visual-language multimodal model finely tuned based on mechanical engineering drawing corpus. The first multimodal large model integrates pixel-level visual features and engineering semantic priors through a cross-attention mechanism to achieve end-to-end structured analysis from visual perception to manufacturing logic.
[0031] Specifically, during fine-tuning, the first multimodal large model learns the syntactic structural constraints of manufacturing annotations through a large number of perturbed synthetic drawing samples. Using structured manufacturing semantic objects as the supervised target, it internalizes the semantic coupling relationship between symbols, values, and references. This allows it to achieve logically consistent semantic parsing based on the overall context even when faced with line intersections, font distortions, or partial occlusions during inference, avoiding global semantic errors caused by misidentification of local characters. The structured manufacturing semantic objects can include geometric tolerance types, tolerance values, references to references, etc.
[0032] S103. Based on the pre-trained second multimodal large model, perform process intent semantic recognition on each view to be reasoned to obtain the material process context structure information corresponding to each view to be reasoned.
[0033] Optionally, after obtaining each view to be reasoned, based on the pre-trained second multimodal large model and the pre-constructed second prompt words, process intention semantic recognition is performed on each view to be reasoned to identify the material process context structure information corresponding to each view to be reasoned.
[0034] For example, each view to be reasoned and the pre-constructed second prompt word can be input into the second multimodal large model, and the second multimodal large model can perform implicit semantic recognition on the process intention in each view to be reasoned to obtain the manufacturing parameter set corresponding to each view to be reasoned.
[0035] The material process context information is used to characterize the manufacturing intent related to materials and processes. This information includes multiple sets of structured data, each containing standardized and mapped material grades, heat treatment process specifications, surface treatment technical requirements, and corresponding physical or functional performance constraints. This information characterizes the manufacturing intent in the original engineering data that is directly related to material selection and subsequent process execution.
[0036] The second multimodal large model can employ a cascaded architecture including a dual-path heterogeneous coding layer and an interactive transformer (Transformer) to decouple deep features. Specifically, the second multimodal large model includes a positional encoding module and an attention mechanism. The positional encoding module, designed for tabular data, uses positional feature vectors to capture the two-dimensional coordinate relationships of text within rows and columns, resolving semantic breaks caused by cross-row / cross-cell descriptions. The attention mechanism, through the Transformer architecture's self-attention mechanism, automatically establishes logical weights between core keywords (such as "material," "grade," and "heat treatment") and their subsequent attribute values in long text paragraphs.
[0037] By performing semantic recognition of the process intent of each view to be inferred, the material process context structure information corresponding to each view to be inferred can be obtained. This can address the difficulties in extracting material information in the original engineering data, such as discrete layout positions, non-standard description formats, and high mixing with massive amounts of technical requirement text, and achieve full-drawing dimension material semantic search.
[0038] S104. Based on the set of manufacturing parameters corresponding to each view to be inferred and the material process context information, perform manufacturability inference and generate a manufacturability report corresponding to the original engineering data.
[0039] Optionally, after obtaining the set of manufacturing parameters and material process context information corresponding to each view to be inferred, manufacturability inference can be performed based on the set of manufacturing parameters and material process context information corresponding to each view to be inferred, as well as the pre-trained heuristic algorithm and process decision model, thereby generating a manufacturability report corresponding to the original engineering data.
[0040] For example, heuristic algorithms and process decision models can perform manufacturability reasoning on the set of manufacturing parameters and material process context information from dimensions such as cost, efficiency and supply chain standardization, thereby generating a manufacturability report corresponding to the original engineering data.
[0041] In this embodiment, by identifying multiple views to be inferred corresponding to the original engineering data, and performing semantic decoding of manufacturing parameters and semantic recognition of process intent on each view based on multiple pre-trained multimodal large models, the set of manufacturing parameters and material process context information corresponding to each view to be inferred are obtained. Based on these information, manufacturability inference is performed, generating a manufacturability report. This enables automatic parsing, semantic understanding, and manufacturing feasibility assessment of the original engineering data. Furthermore, it reduces human intervention in manufacturability inference, significantly improving analysis efficiency, consistency and reliability of assessment results, and effectively enhancing the collaborative efficiency of design and manufacturing.
[0042] In one possible implementation, Figure 2 This application provides a flowchart illustrating the process of generating a manufacturability report corresponding to the original engineering data in a two-dimensional engineering drawing manufacturability reasoning method based on a multimodal large model. (Refer to...) Figure 2 As shown, in S104 above, manufacturability reasoning is performed based on the manufacturing parameter set corresponding to each view to be reasoned and the material process context information, generating a manufacturability report corresponding to the original engineering data, including: S201. Perform multi-dimensional manufacturability reasoning on the set of manufacturing parameters and the material process context information to determine at least one risk information corresponding to the original engineering data in multiple dimensions, at least one optimization information corresponding to each risk information, and efficiency information corresponding to each optimization information.
[0043] Optionally, multi-dimensional manufacturability reasoning can be performed on the set of manufacturing parameters and material process context information from dimensions such as cost, efficiency and supply chain standardization to determine at least one risk information corresponding to the original engineering data in multiple dimensions, at least one optimization information corresponding to each risk information, and efficiency information corresponding to each optimization information.
[0044] For example, the set of manufacturing parameters and the material process context information can be input into a pre-trained multimodal manufacturability reasoning model. The multimodal manufacturability reasoning model can then perform multi-dimensional manufacturability reasoning on the set of manufacturing parameters and the material process context information from dimensions such as cost, efficiency, and supply chain standardization to obtain at least one risk information corresponding to the original engineering data in multiple dimensions, at least one optimization information corresponding to each risk information, and efficiency information corresponding to each optimization information.
[0045] Risk information indicates potential problems that may lead to quality defects, cost overruns, efficiency losses, or production disruptions during the manufacturing process. Optimization information indicates specific, actionable improvement suggestions based on the risk information. Efficiency information refers to the quantifiable expected improvements in manufacturing efficiency, cost, or resource consumption after adopting optimization information.
[0046] S202. Based on each risk information, each optimization information corresponding to each risk information, and each efficiency information corresponding to each optimization information, generate a manufacturability report corresponding to the original engineering data.
[0047] Optionally, after obtaining each risk information, each optimization information corresponding to each risk information, and each efficiency information corresponding to each optimization information, the risk information, the optimization information corresponding to each risk information, and the efficiency information corresponding to each optimization information can be filled into a preset template to generate a manufacturability report corresponding to the original engineering data.
[0048] By using multi-dimensional manufacturability reasoning, at least one risk information, at least one optimization information corresponding to each risk information, and efficiency information corresponding to each optimization information can be identified under multiple dimensions. This can upgrade single-point checks in manufacturability reasoning to comprehensive checks, thereby covering the entire manufacturing chain and achieving accurate risk positioning and traceability. This generates manufacturability reports corresponding to the original engineering data, significantly improving analysis efficiency, consistency and reliability of evaluation results, and effectively improving the collaborative efficiency of design and manufacturing.
[0049] In one possible implementation, Figure 3 This application provides a flowchart illustrating the process of generating a manufacturability report corresponding to the original engineering data in a two-dimensional engineering drawing manufacturability reasoning method based on a multimodal large model. (Refer to...) Figure 3 As shown, in step S202 above, a manufacturability report corresponding to the original engineering data is generated based on each risk information, each optimization information corresponding to each risk information, and each efficiency information corresponding to each optimization information. This report includes: S301. Determine the risk weight and severity of each risk information under multiple dimensions, and determine the comprehensive risk index of the original engineering data based on each risk information and its risk weight and severity under multiple dimensions.
[0050] Optionally, based on preset risk weights, severity, and mapping relationships of risk information, the risk weights and severity of each risk information under multiple dimensions can be matched, and a comprehensive risk index of the original engineering data can be calculated based on each risk information and its risk weights and severity under multiple dimensions. This achieves comparability, weighting, and interpretability of multi-source heterogeneous risks, and significantly improves the robustness and engineering practicality of manufacturability reasoning.
[0051] Optionally, the risk weights and severity of each risk information under multiple dimensions can be predicted based on the pre-trained risk weight assessment model and severity assessment model, and the comprehensive risk index of the original engineering data can be calculated based on each risk information and its risk weights and severity under multiple dimensions.
[0052] The risk weight and severity are affected by the following factors: geometric constraints (such as aspect ratio, thin walls, and radius reachability), process capabilities (such as tolerance requirements, the matching degree between geometric tolerances and machine tool accuracy), and resource constraints (such as material inventory and special tooling gaps).
[0053] For example, after obtaining the risk weight and severity of each risk information, the product of the risk weight and severity of each risk information can be calculated to obtain the risk result of each risk information, and the sum of the risk results of each risk information can be calculated as a comprehensive risk index.
[0054] S302. Based on the comprehensive risk index, determine the production execution conclusion corresponding to the original engineering data.
[0055] Optionally, the comprehensive risk index can be mapped to the mapping relationship between the risk index and the production execution conclusion to determine the production execution conclusion corresponding to the original engineering data.
[0056] The production execution conclusions include: recommended production scheduling, prohibited production scheduling, or optimized production scheduling.
[0057] Specifically, recommended production scheduling refers to a perfect match between existing equipment, tools, and materials. Prohibited production scheduling refers to situations where physical interference exists or precision limits are exceeded. Optimized production scheduling refers to situations where potential risks have been identified.
[0058] S303. Fill the production execution conclusion, the optimization information corresponding to each risk information, and the efficiency information corresponding to each optimization information into the preset template to generate a manufacturability report corresponding to the original engineering data.
[0059] Optionally, the production execution conclusions, the optimization information corresponding to each risk information, and the efficiency information corresponding to each optimization information can be filled into a preset template to generate a manufacturability report corresponding to the original engineering data.
[0060] In one possible implementation, Figure 4 This application provides a flowchart illustrating the process of determining at least one risk information corresponding to the original engineering data in multiple dimensions, at least one optimization information corresponding to each risk information, and efficiency information corresponding to each optimization information in the manufacturability reasoning method for two-dimensional engineering drawings based on a multimodal large model. (Refer to...) Figure 4 As shown, in S201 above, multi-dimensional manufacturability reasoning is performed on the manufacturing parameter set and material process context information to determine at least one risk information corresponding to the original engineering data in multiple dimensions, at least one optimization information corresponding to each risk information, and efficiency information corresponding to each optimization information, including: S401. Based on the pre-trained third multimodal large model, perform cost manufacturability reasoning on the manufacturing parameter set and material process context structure information, determine at least one risk information corresponding to the original engineering data in the cost dimension and the optimization information corresponding to each risk information, and determine the efficiency information corresponding to the optimization information.
[0061] Optionally, the manufacturing parameter set, material process context structure information, and pre-constructed third prompt words are input into the pre-trained third multimodal large model. The pre-trained third multimodal large model performs cost manufacturability reasoning to determine at least one risk information corresponding to the original engineering data in the cost dimension, as well as the optimization information corresponding to each risk information.
[0062] Among them, the third multimodal large model was obtained by fine-tuning based on the engineering design experience library and the manufacturing resource library.
[0063] Specifically, cost manufacturability reasoning includes: accuracy redundancy manufacturability reasoning, process complexity redundancy manufacturability reasoning, and material specification redundancy manufacturability reasoning.
[0064] For example, precision redundancy manufacturability reasoning refers to determining whether tolerances or roughness far exceeding the standard machining grade are marked on non-mating surfaces. Non-mating surfaces refer to the outer contour decorative surfaces or non-contact support surfaces of a part.
[0065] For example, when performing semantic assembly analysis on a set of manufacturing parameters, if it is found that a certain surface has no assembly datum, no relative motion fit, and a roughness requirement of Ra0.8 (requiring fine grinding), while the conventional functional requirements only require Ra3.2 (milling), then the surface is determined to be redundant in terms of precision.
[0066] For example, manufacturability reasoning based on process complexity redundancy refers to determining whether geometric features that are extremely difficult to manufacture but are not core functional features have been designed.
[0067] For example: a completely flat bottom is required at the bottom of a blind hole, or a clear corner is required at a non-load-bearing corner. Specifically, by using the manufacturing parameter set and material process context information, the number of tool changes or special process costs caused by this feature are calculated. If its contribution to part performance is lower than a preset threshold, it is determined to be redundant in process complexity. Special process costs can be, for example, electrical discharge machining (EDM).
[0068] For example, material specification redundancy manufacturability reasoning refers to determining whether there are selected materials whose performance specifications far exceed their actual operating requirements.
[0069] For example, if ordinary carbon steel can meet the requirements, but the material process context information specifies an expensive special alloy, then the material specification is considered redundant.
[0070] For example, after identifying the existence of risk information, the third multimodal large model infers optimization information and calculates the cost difference before and after optimization to obtain the efficiency information corresponding to the optimization information.
[0071] For example, when the risk information is identified as excessive roughness of the non-assembly hole on the left, the generated optimization information is: "It is recommended to relax the roughness of the non-mating surface A from Ra0.8 to Ra3.2." The generated efficiency information is: "It is expected to reduce machining time by 20% and reduce tool depreciation costs by about 15%."
[0072] S402. Based on the pre-trained fourth multimodal large model, perform efficiency manufacturability reasoning on the manufacturing parameter set and material process context structure information, determine at least one risk information corresponding to the original engineering data in the efficiency dimension and the optimization information corresponding to each risk information, and determine the efficiency information corresponding to the optimization information.
[0073] Optionally, based on the pre-trained fourth multimodal large model, efficiency manufacturability reasoning is performed on the manufacturing parameter set and material process context structure information, the geometric spatial features of the drawings are transformed into manufacturing time cost, at least one risk information corresponding to the original engineering data in the efficiency dimension and the optimization information corresponding to each risk information are determined, and the efficiency information corresponding to the optimization information is calculated.
[0074] Among them, the fourth multimodal large model was obtained by fine-tuning based on the engineering design experience library and the manufacturing resource library.
[0075] Specifically, efficiency manufacturability reasoning includes: clamping scheme manufacturability reasoning and processing risk manufacturability reasoning.
[0076] For example, manufacturability reasoning of a clamping scheme refers to evaluating the efficiency of a clamping scheme by analyzing the geometric features of the part. Specifically, for each machining feature (such as a hole, slot, or plane) in the set of manufacturing parameters, a corresponding set of normal vectors is established based on its machining infeed direction. Then, a clustering algorithm is used to calculate the minimum number of coordinate system flips required to complete the machining of all features to identify which features can be grouped together and completed on the same machining surface.
[0077] For example, if the original design requires "six-sided clamping" due to a slight local tilt, the fourth multimodal large model will evaluate whether rotating the feature to be perpendicular to the main reference plane can reduce the number of clamping operations to "two-sided clamping". This will reduce the time spent on reversing clamping, secondary alignment and repeated positioning, thereby directly reducing the cumulative positioning error and improving the production line turnover rate.
[0078] For example, in the manufacturability reasoning of processing risks, for high-difficulty features such as deep cavities and narrow slits, the depth of the deep cavity features is extracted. With width Calculate its aspect ratio, when it is identified Exceeding a preset threshold (e.g.) When this feature is detected, the fourth multimodal large model determines that it will lead to unstable cutting forces and difficulty in chip removal, thus triggering an efficiency warning. Simultaneously, the tool library is automatically searched to assess whether switching to a high-cost long-shank tool or using electrical discharge machining (EDM) is necessary. If the feature allows, the output optimization information suggests adding a relief groove or increasing the inner diameter of the tool. Angles are used to support larger diameter tools to perform high-speed helical milling, thereby reducing the wasted time of corner clearing and rework.
[0079] For example, the efficiency before and after optimization is quantitatively compared using a preset efficiency model function to obtain efficiency information. Specifically, this can be achieved using the following formula:
[0080] in, For cutting time, For automatic tool change time, For manual or robotic arm clamping and alignment time, Here, n represents efficiency information, n represents the number of risk information items, and m represents the number of manual or robotic arm clamping actions required.
[0081] Output: The system compares and recommends solutions. The difference will be used as an evaluation indicator to output the "percentage reduction in working hours" in the decision-making report.
[0082] S403. Based on the pre-trained fifth multimodal large model, perform supply chain manufacturability reasoning on the manufacturing parameter set and material process context structure information, determine at least one risk information corresponding to the original engineering data in the supply chain dimension and the optimization information corresponding to each risk information, and determine the efficiency information corresponding to the optimization information.
[0083] Optionally, the manufacturing parameter set, material and process context information, and the pre-constructed fifth prompt word are input into the pre-trained fifth multimodal large model to perform supply chain manufacturability reasoning on the manufacturing parameter set and material and process context information. The extracted part features are then matched with the factory standard resource library in a multidimensional semantic manner to determine at least one risk information corresponding to the original engineering data in the supply chain dimension, as well as the optimization information corresponding to each risk information. The efficiency information corresponding to the optimization information is also determined, thereby achieving the optimization of supply chain costs.
[0084] Among them, the fifth multimodal large model was obtained by fine-tuning based on the engineering design experience library and the manufacturing resource library.
[0085] Specifically, supply chain manufacturability reasoning includes: in-library redirection reasoning for fastener and thread features and normalized spatial mapping reasoning for blank specifications.
[0086] For example, the in-library redirection reasoning for fasteners and thread features includes: retrieving commonly used tools and standard specification tables from the standard parts library based on the thread parameters in the manufacturing parameter set, and performing a difference determination, thereby determining risk information and corresponding optimization information based on the difference determination results.
[0087] For example, the differentiation determination can be as follows: if the search results show that the thread is a non-standard fine thread (i.e. not in the list), then a standard coarse thread with equivalent performance is found by using a preset neighborhood search algorithm.
[0088] For example, the change in connection strength after modifying the thread specification is calculated, and optimization information is output while ensuring functional safety. The cost of purchasing custom taps / gauges is compared with the cost of calling up standard inventory, and the difference is marked as the direct cost reduction amount. The increased production line uptime due to reduced tool change preparation time is estimated to obtain efficiency information.
[0089] For example, the normalized spatial mapping reasoning for blank specifications includes: calculating the maximum external dimensions of the part, deriving the minimum blank requirement, and matching the required dimensions with the stock standard sheet metal. If the part thickness is found to be slightly smaller than the stock standard thickness, optimization information is generated: suggesting fine-tuning non-functional surfaces or directly utilizing the standard thickness to eliminate single-sided / double-sided pre-processing planing operations. The reduced roughing time and material removal rate are calculated, and the optimized production time curve is output as efficiency information.
[0090] In one possible implementation, Figure 5 This application provides a flowchart illustrating the process of obtaining the set of manufacturing parameters corresponding to each view to be reasoned in the two-dimensional engineering drawing manufacturability reasoning method based on a multimodal large model, as shown in the embodiments of this application. Figure 5 As shown, in S102 above, based on the pre-trained first multimodal large model, the manufacturing parameter semantic decoding is performed on each view to be reasoned, to obtain the manufacturing parameter set corresponding to each view to be reasoned, including: S501. Input each view to be reasoned into the pre-trained annotation space anchoring model, and use the annotation space anchoring model to perform omnidirectional annotation space anchoring on each view to be reasoned, thereby obtaining a set of candidate region descriptors corresponding to each view to be reasoned.
[0091] Optionally, each view to be inferred is input into a pre-trained annotation space anchoring model, which performs omnidirectional annotation space anchoring on each view to be inferred, thereby achieving accurate selection of the omnidirectional annotation region and obtaining a set of candidate region descriptors corresponding to each view to be inferred.
[0092] The base model of the annotation space anchoring model can be a YOLOv12 model trained on a large-scale mechanical engineering drawing set. The annotation space anchoring model can be obtained by integrating multiple enhancement modules on the base model.
[0093] The candidate region descriptor set includes multiple candidate region descriptors, which include: view identifier, bounding box parameters, category labels corresponding to the bounding box parameters, and image cropping blocks corresponding to the bounding box parameters.
[0094] Specifically, the view identifier is used to indicate the view to which the candidate region descriptor belongs, and the bounding box parameters include: center point coordinates (x, y), rotation angle θ, box width w, and box height h.
[0095] Specifically, the category label corresponding to the bounding box parameter refers to the semantic type of the manufacturing element carried by the bounding box. For example, the category label corresponding to the bounding box parameter may include: dimension annotation, geometric tolerance frame, surface roughness symbol, etc.
[0096] Specifically, the image cropping block corresponding to the bounding box parameters refers to the positively aligned sub-image generated by performing an affine transformation and normalization on the original image corresponding to the bounding box according to the bounding box parameters.
[0097] Optionally, the candidate region descriptor may also include a confidence level.
[0098] S502. Input the set of candidate region descriptors corresponding to each view to be inferred into the first multimodal large model, and perform semantic decoding by the first multimodal large model to obtain the set of manufacturing parameters.
[0099] Optionally, the candidate region descriptor set corresponding to each view to be inferred is input into the first multimodal large model, and the first multimodal large model performs semantic decoding to obtain the manufacturing parameter set.
[0100] For example, the manufacturing parameter set obtained by semantic decoding of the first multimodal large model can be achieved by referring to the following mapping relationship:
[0101] in, This is the set of candidate region descriptors corresponding to each view to be inferred. These are the model parameters for the first multimodal large model. This represents the end-to-end decoding function of the first multimodal large model. To create a set of parameters.
[0102] By using a spatial anchoring model to perform omnidirectional spatial anchoring on each view to be inferred, a set of candidate region descriptors corresponding to each view is obtained. Semantic decoding is then performed using a first multimodal large model to obtain a set of manufacturing parameters. This allows for the accurate extraction of key manufacturing parameters from complex engineering graphic backgrounds through heterogeneous feature collaborative decoding. Furthermore, it can treat complex combinations containing geometric symbols, letters, and numbers as a single semantic unit for closed-loop parsing, effectively avoiding the risks of character breaks or garbled text caused by overlapping lines in single-character recognition.
[0103] In one possible implementation, Figure 6 This is a schematic diagram of a spatially anchored model used in the manufacturability reasoning method for two-dimensional engineering drawings based on a multimodal large model, as provided in an embodiment of this application. Figure 7 This application provides a flowchart illustrating the process of obtaining the candidate region descriptor set corresponding to each view to be reasoned in the manufacturability reasoning method for two-dimensional engineering drawings based on a multimodal large model, as shown in the embodiments of this application. Figure 6 as well as Figure 7 As shown, the annotation space anchoring model includes: a multi-scale feature fusion module, a rotation-sensitive detection head module, a coordinate refinement module, a geometric topology post-verification module, and a backbone network; in S501 above, each view to be reasoned is input into the pre-trained annotation space anchoring model, and the annotation space anchoring model performs omnidirectional annotation space anchoring on each view to be reasoned, obtaining a set of candidate region descriptors corresponding to each view to be reasoned, including: S701. Input the view to be reasoned into the multi-scale feature fusion module. The multi-scale feature fusion module perceives the symbols and long dimension lines in the view to be reasoned and obtains the first enhanced feature.
[0104] Optionally, the view to be reasoned is input into the multi-scale feature fusion module, which performs unified perception on the symbols and long lines in the view to be reasoned to obtain the first enhanced feature. The first enhanced feature strengthens the joint representation of the symbols and long lines, solving the problem of short lines being submerged by long lines or characters being isolated and difficult to locate.
[0105] Specifically, the multi-scale feature fusion module can be referred to as the multi-scale feature fusion neck, which is used to uniformly perceive cross-scale manufacturing annotation elements. The first enhanced feature can be represented as a structured tensor. One channel of this structured tensor corresponds to one manufacturing element.
[0106] For example, the multi-scale feature fusion module includes a micro-scale enhancement layer, a meso-scale integration layer, and a macro-scale continuity refinement layer.
[0107] S702. Input the view to be reasoned into the rotation-sensitive detection head module. The rotation-sensitive detection head module regresses the rotation direction in the view to be reasoned to obtain the second enhanced feature.
[0108] Optionally, the view to be inferred is input into the rotation-sensitive detection head module, which performs end-to-end regression on the rotation direction in the view to be inferred to obtain the second enhanced feature. Through the second enhanced feature, rotation-invariant geometric priors are provided, so that the model has the ability to generalize to any initial rotation angle (0°–360° continuous value) of the view to be inferred in subsequent processing, and can directly output the correct rotation bounding box without the need for preset correction steps.
[0109] The rotation-sensitive detection head module is used to semantically model the inherent geometric orientation of manufacturing elements. The second enhancement feature can be characterized using a direction prior tensor field. For example, the second enhancement feature may include three channels, each representing the rotation direction cosine, sine, and direction prediction confidence of each spatial location in the view to be reasoned.
[0110] S703. Input the view to be reasoned into the coordinate refinement module. The coordinate refinement module predicts the local offset in the view to be reasoned and obtains the third enhanced feature.
[0111] Optionally, the view to be inferred is input into the coordinate refinement module, which predicts the local feature offsets in the view to be inferred to obtain the third enhanced feature. The positioning accuracy is then improved to the sub-pixel level through the third enhanced feature.
[0112] The coordinate refinement module specifically includes a sub-pixel level coordinate refinement sub-network. The third enhancement feature can be characterized by a sub-pixel level geometric offset field, used to accurately align the coarse bounding box with the true geometric boundaries of the manufacturing elements.
[0113] S704. Input the view to be reasoned into the geometric topology post-verification module. The geometric topology post-verification module verifies the spatial relationships in the view to be reasoned and obtains the fourth enhanced feature.
[0114] Optionally, the view to be inferred is input into the geometric topology post-verification module, which performs consistency verification on the spatial relationships in the view to be inferred to obtain the fourth enhanced feature.
[0115] The geometric topology post-verification module is used to perform consistency verification based on the spatial dependencies between manufacturing elements, eliminating isolated false detections. These spatial dependencies include: the connectivity between dimension lines and contour edges, the center alignment of dimension text and dimension lines, the nested structure of symbols, values, and datums within the geometric tolerance frame, and the directional relationship between leader lines and controlled features.
[0116] The fourth enhancement feature is used to perform logical final review and credibility weighting on candidate region descriptors.
[0117] S705. Input the view to be reasoned, the first enhanced feature, the second enhanced feature, the third enhanced feature, and the fourth enhanced feature into the backbone network to obtain the candidate region descriptor set corresponding to the view to be reasoned.
[0118] Optionally, the view to be inferred, the first enhanced feature, the second enhanced feature, the third enhanced feature, and the fourth enhanced feature are input into the backbone network, and the backbone network performs inference to obtain a set of candidate region descriptors corresponding to the view to be inferred.
[0119] The backbone network is a YOLOv12 model trained on a large-scale mechanical engineering drawing set.
[0120] By generating first, second, third, and fourth enhanced features, multi-source priors can be generated. These priors are then processed in parallel by the backbone network to obtain a set of candidate region descriptors corresponding to the view to be inferred. This allows the backbone network to simultaneously complete four cognitive tasks—perception, modeling, refinement, and verification—in a single forward pass during inference. Consequently, geometric localization accuracy is significantly improved, pose robustness is enhanced, semantic parsing consistency is improved, and computational efficiency and deployment friendliness are enhanced.
[0121] In one possible implementation, Figure 8 This application provides a flowchart illustrating the process of obtaining the material and process context information corresponding to each view to be reasoned in the two-dimensional engineering drawing manufacturability reasoning method based on a multimodal large model, as shown in the embodiments of this application. Figure 8As shown, in S103 above, based on the pre-trained second multimodal large model, process intent semantic recognition is performed on each view to be reasoned, and the material process context structure information corresponding to each view to be reasoned is obtained, including: S801. Based on the pre-trained layout analysis model, perform layout analysis on each view to be inferred to obtain multiple semantic information corresponding to each view to be inferred.
[0122] Optionally, each view to be reasoned can be input into a pre-trained layout analysis model, which will then perform global semantic segmentation and geometric layout modeling on each view to be reasoned, thereby obtaining multiple semantic information corresponding to each view to be reasoned.
[0123] The semantic information includes text image blocks, type labels for the text image blocks, and the content of the text image blocks. The type labels indicate the function of the text image block in the view to be processed, such as: title bar, detail table, technical requirements, local annotation area, etc.
[0124] The layout analysis model is pre-built using a deep convolutional neural network.
[0125] By using a pre-trained layout analysis model, layout analysis is performed on each view to be inferred, and multiple semantic information in each view to be inferred is obtained. While completing the macro-region segmentation, pixel-level cleaning of material information within each semantic region is achieved, thereby transforming discrete material grades and technical descriptions that are interfered with by lines into clean and recognizable text blocks, providing accurate spatial indexes for subsequent deep feature extraction.
[0126] S802. Input the semantic information into the second multimodal large model, and the second multimodal large model performs context awareness and parsing to obtain at least one engineering string entity and the semantic constraints of each engineering string entity.
[0127] Optionally, the semantic information is input into the second multimodal large model, which performs context awareness and parsing to obtain at least one engineering string entity and the semantic constraints of each engineering string entity.
[0128] Among them, semantic constraints are used to indicate the process constraints or material constraints corresponding to the engineering string entity.
[0129] Specifically, the engineering string entity contains the original material text extracted from the technical requirements, text area, or title block, such as material names with abbreviations, industry slang, or specific writing conventions (e.g., "45#"). The semantic constraints of the engineering string entity contain the original process text extracted from the technical requirements, text area, or title block, such as heat treatment processes (e.g., "quenching and tempering") and hardness values (e.g., "HRC28-32").
[0130] For example, if the fifth record of the technical requirement is “5. Material: 45 steel, heat treatment to HRC28-32”, the second multimodal large model can accurately extract the engineering string entity as “45 steel” and the semantic constraint of the engineering string entity as “heat treatment” through semantic reference relationship.
[0131] S803. Standardize and map the semantic constraints of each engineering string entity to obtain the material process context structure information corresponding to each view to be reasoned.
[0132] Optionally, after obtaining each engineering string entity and its semantic constraints, knowledge alignment, standardization mapping, and structured output can be performed on each engineering string entity and its semantic constraints to obtain the material and process context structure information corresponding to each view to be reasoned.
[0133] By using a pre-trained layout analysis model, layout analysis is performed on each view to be reasoned, obtaining multiple semantic information corresponding to each view. The visual-to-text conversion can be completed in advance, reducing the processing burden of the large model. The semantic information is then input into a second multimodal large model, which performs context awareness and parsing to obtain at least one engineering string entity and the semantic constraints of each engineering string entity. This allows for accurate understanding of the semantic relationships between text blocks, extraction of engineering entities and their constraints, and standardized mapping of each engineering string entity and its semantic constraints. This yields the material and process context structure information corresponding to each view to be reasoned, enabling knowledge accumulation and reuse. Ultimately, this achieves high-precision, strong-structure, interpretable, and easy-to-maintain intelligent parsing and reasoning of engineering drawings.
[0134] In one possible implementation, Figure 9 This application provides a flowchart illustrating the process of determining at least one text image block corresponding to each view to be reasoned and the type label of each text image block in the manufacturability reasoning method for two-dimensional engineering drawings based on a multimodal large model. (Refer to...) Figure 9 As shown, in step S801 above, layout analysis is performed on each view to be reasoned based on the pre-trained layout analysis model, resulting in multiple semantic information in each view to be reasoned, including: S901. Based on the pre-trained layout analysis model, perform region detection and pixel scanning on each view to be inferred, and identify each semantic region, the bounding box corresponding to each semantic region, and the category label of each semantic region in each view to be inferred.
[0135] Optionally, each view to be reasoned is input into a pre-trained layout analysis model, which performs semantic region detection and pixel scanning on each view to be reasoned, and identifies each semantic region, the bounding box corresponding to each semantic region, and the category label of each semantic region in each view to be reasoned.
[0136] Among them, the layout analysis model can be implemented based on OCR.
[0137] Specifically, semantic regions include: Title Block, Bill of Materials (BOM), Technical Requirements text area, and local annotation area. Each semantic region corresponds to a bounding box and a category label. The category label is used to indicate the semantic entity type of the semantic region.
[0138] S902. Based on the preset edge detection operator, extract non-textual geometric features from each semantic region and generate a binary mask matrix corresponding to each semantic region.
[0139] Optionally, by using a preset edge detection operator and preset geometric feature filtering rules, geometric interference masking and mask generation are performed on each semantic region to extract non-textual geometric features and generate a binary mask matrix corresponding to each semantic region. The non-textual geometric features include table lines, guide lines, and profile lines, etc.
[0140] Specifically, for the current semantic region, if the current pixel is a non-text geometric pixel, the corresponding position in the binary mask matrix is assigned a value of 0; if the current pixel is not a non-text geometric pixel, the corresponding position in the binary mask matrix is assigned a value of 1.
[0141] S903. Generate clean region images corresponding to each semantic region based on the binary mask matrix corresponding to each semantic region and each semantic region.
[0142] Optionally, a bitwise AND operation is performed between the binary mask matrix corresponding to each semantic region and the original image corresponding to each semantic region to obtain the clean region image corresponding to each semantic region, thereby physically shielding geometric interference noise using the masking effect. The clean region image is the region image with geometric line interference removed.
[0143] S904. Based on the clean region image corresponding to each semantic region, determine the multiple text image blocks corresponding to each view to be inferred, the type label of each text image block, and the content of each text image block.
[0144] Optionally, after obtaining the clean region images corresponding to each semantic region, the text detection branch of the OCR model is used to detect the clean region images, locate the text position, output the bounding box of each text line, and crop the image within the bounding box according to the bounding box of each text line to obtain text image blocks.
[0145] Optionally, the category label of the semantic region to which the text image block belongs can be used as the type label for each text image block.
[0146] Optionally, the text recognition branch of the OCR model can be used to perform text recognition on the text image blocks to obtain the content of each text image block.
[0147] By using a pre-trained layout analysis model and edge detection operator, layout analysis is performed on each view to be inferred, and multiple semantic information in each view to be inferred is obtained. This physically shields line interference, significantly reduces the false detection rate of the text detection branch, and improves the accuracy of character recognition, thereby improving the accuracy of text recognition.
[0148] In one possible implementation, Figure 10 This application provides a flowchart illustrating the process of obtaining the material and process context information corresponding to each view to be reasoned in the two-dimensional engineering drawing manufacturability reasoning method based on a multimodal large model, as shown in the embodiments of this application. Figure 10 As shown, in S803 above, the semantic constraints of each engineering string entity are standardized and mapped to obtain the material process context structure information corresponding to each view to be reasoned, including: S1001. Based on each project string entity and its semantic constraints, construct at least one semantic vector corresponding to each project string entity.
[0149] Optionally, the engineering string entity is concatenated with the semantic constraint and then input into a pre-trained semantic encoding model to obtain a joint semantic vector, which serves as the semantic vector corresponding to the engineering string entity.
[0150] S1002. Based on each semantic vector, the preset material grade knowledge base, and the preset synonym weights, determine the standard material grade corresponding to each semantic vector.
[0151] Optionally, the cosine similarity between the semantic vector and the standard term vectors in the preset material grade knowledge base is calculated, and the weighted adjustment is performed in combination with the preset synonym weights. The standard grade with the highest similarity is selected as the matching result to obtain the standard material grade corresponding to the semantic vector.
[0152] Optionally, if the highest similarity is lower than a preset threshold, it is marked as "awaiting manual confirmation" and Top-K candidates are output.
[0153] For example, a normalization operator is used to compare the semantic vector with a pre-defined material grade knowledge base (Standard Material Library, abbreviated as...). Multidimensional feature matching is performed. Specifically, the normalization operator uses cosine similarity or edit distance algorithms to calculate the distance between the semantic vector and the standard terms in the material grade knowledge base within the semantic embedding space.
[0154] For example, when the semantic vector is the vector corresponding to {"45 steel", "45#", "Steel45"}, the normalization operator uses the synonym weight allocation of the knowledge graph to uniformly cluster them and redirect them to the standard grade ID, for example, it can be... .
[0155] S1003. Based on the standard material grade corresponding to each semantic vector, determine the material properties and process intentions corresponding to each engineering string entity, and use the material properties and process intentions corresponding to each engineering string entity as material process context structure information.
[0156] Optionally, material properties such as basic attributes, mechanical properties, and physical properties associated with standard grades can be retrieved from the material grade knowledge base and used as the material properties corresponding to each engineering string entity. Combined with the process description in the semantic constraints, the material properties are mapped to standard process codes (such as heat treatment status codes) through the process knowledge base and used as the process intents corresponding to each engineering string entity. The material properties and process intents are used as material process context structure information and output in a structured format such as JSON.
[0157] By constructing at least one semantic vector corresponding to each engineering string entity and its semantic constraints, unstructured text is transformed into computable mathematical objects. Then, using each semantic vector, a pre-defined material grade knowledge base, and pre-defined synonym weights, the standard material grade corresponding to each semantic vector is determined, establishing a reliable mapping from non-standard representations to standard entities. Based on the standard material grade corresponding to each semantic vector, the material properties and process intentions of each engineering string entity are determined. These material properties and process intentions are used as material process context information, ensuring the accuracy and robustness of the obtained material process context information. Furthermore, reusability is improved.
[0158] In one possible implementation, Figure 11 This application provides a flowchart illustrating the process of determining at least one view to be reasoned about in a two-dimensional engineering drawing manufacturability reasoning method based on a multimodal large model, as described in the embodiments of this application. Figure 11As shown, determining at least one view to be reasoned about corresponding to the original engineering data in S101 above includes: S1101. If the original engineering data is a three-dimensional solid model, perform multi-dimensional projection mapping processing on the original engineering data to obtain multiple intermediate two-dimensional views.
[0159] Optionally, if the original engineering data is a three-dimensional solid model, then the original engineering data is subjected to multi-dimensional projection mapping to obtain multiple intermediate two-dimensional views.
[0160] For example, based on the geometric features of the part (such as a rotating body, a plate, or a box) and the identification of the manufacturing datum surface, the direction that best reflects the machining and clamping direction and the distribution of the main features is determined as the direction of the main view, and an initial main view is obtained.
[0161] Specifically, point cloud sampling is performed on the surface of the 3D solid model to calculate the spatial distribution entropy of the hole / groove / thin-wall region. Based on the spatial distribution entropy of each region, the structural type label of the 3D solid model is determined. Based on the structural type label of the 3D solid model and the preset structure and main view mapping relationship, the main view direction is determined. The 3D solid model is then projected according to the main view direction to obtain the initial main view.
[0162] For example, after determining the orientation of the main view, feature recognition is performed on the 3D solid model. Based on the orientation of the main view and the feature recognition results of the 3D solid model, a cutting plane that can expose the internal structure and facilitate tool accessibility analysis is determined, and an initial sectional view is dynamically generated. The feature recognition results are used to indicate spatial features such as holes, cavities, thin walls, and deep cavities in the 3D solid model.
[0163] For example, for a 3D solid model, based on the initial front view and initial sectional view, other views such as the initial side view and initial top view are determined. The Product Manufacturing Information (PMI) embedded in the 3D model is automatically mapped to the corresponding positions in the corresponding views. For areas lacking PMI, necessary manufacturing annotations (such as default chamfers and unspecified tolerance grades) are inferred based on geometric features to obtain alternative views. The product manufacturing information includes dimensions, tolerances, surface roughness, and annotations.
[0164] For example, each candidate view is automatically arranged according to ISO or GB drafting standards to eliminate view overlap and reserve sufficient annotation space to ensure the accuracy of subsequent OCR and semantic parsing. At the same time, risk identification is performed on each candidate view, and potential high-risk manufacturing features are visually enhanced or isolated by layers to obtain each intermediate two-dimensional view. Potential high-risk manufacturing features may include: deep holes, narrow slots, thin walls, and minute features.
[0165] S1102. Based on the pre-trained image quality assessment model, perform quality assessment on each intermediate two-dimensional view to obtain the quality index corresponding to each intermediate two-dimensional view.
[0166] Optionally, each intermediate two-dimensional view is input into a pre-trained image quality assessment model, which then performs automated quality measurement on each intermediate two-dimensional view to obtain the corresponding quality index.
[0167] The quality metrics include spatial resolution, edge sharpness, grayscale contrast, and random noise level. Specifically, spatial resolution is estimated through frequency domain Fourier transform analysis or edge modulation transfer function (MTF); edge sharpness is quantized using Laplacian variance or Tenengrad gradient operator; grayscale contrast is calculated based on the RMS contrast of global or local histograms; and random noise level is obtained by calculating the standard deviation of pixel grayscale or the variance of wavelet coefficients in non-annotated areas of the drawing (such as blank backgrounds or uniformly filled areas).
[0168] S1103. Determine the enhancement strategy for each intermediate two-dimensional view based on the quality indicators corresponding to each intermediate two-dimensional view.
[0169] Optionally, an enhancement strategy corresponding to each intermediate two-dimensional view can be obtained by matching the quality index corresponding to each intermediate two-dimensional view and the preset mapping relationship between the quality index and the enhancement strategy.
[0170] The enhancement strategies include one or more of the following: adaptive denoising, histogram equalization, edge sharpening, and tilt correction based on Hough transform.
[0171] S1104. Based on the enhancement strategy corresponding to each intermediate two-dimensional view, perform image enhancement processing on each intermediate two-dimensional view to obtain at least one view to be inferred.
[0172] Optionally, after obtaining the enhancement strategy corresponding to each intermediate two-dimensional view, image enhancement processing is performed on each intermediate two-dimensional view according to the enhancement strategy corresponding to each intermediate two-dimensional view to obtain multiple views to be inferred, thereby eliminating geometric distortion and visual noise generated by physical scanning.
[0173] Optionally, pixel-level resolution unification can be performed on each view to be inferred. Based on the scale information, coordinate scale or known size features explicitly marked in the original engineering data, the precise mapping relationship between real physical units and image pixels can be automatically calculated and established, thereby converting the image coordinate system into a physically meaningful engineering coordinate system, providing a high-precision data benchmark for subsequent sub-pixel level geometric measurement and tolerance analysis of feature elements.
[0174] By performing multidimensional projection mapping on the original engineering data when it is a three-dimensional solid model, multiple intermediate two-dimensional views are obtained. The quality of each intermediate two-dimensional view is then evaluated. Based on the quality indicators corresponding to each intermediate two-dimensional view, image enhancement processing is performed on each intermediate two-dimensional view to obtain multiple views to be inferred. This can actively encode three-dimensional geometric information into a two-dimensional symbol space that conforms to the cognitive paradigm of human process engineers, thereby ensuring the reliability of the data for downstream large model inference and improving the accuracy of inference.
[0175] Based on the same inventive concept, this application also provides a two-dimensional engineering drawing manufacturability reasoning device based on a multimodal large model, which corresponds to the two-dimensional engineering drawing manufacturability reasoning method based on a multimodal large model. Since the principle of the device in this application is similar to the two-dimensional engineering drawing manufacturability reasoning method based on a multimodal large model described above, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.
[0176] Reference Figure 12 As shown, Figure 12 This is a schematic diagram of a manufacturability reasoning device for two-dimensional engineering drawings based on a multimodal large model, provided in an embodiment of this application. The device includes: an acquisition module 1201, a manufacturing parameter semantic decoding module 1202, a process intent semantic recognition module 1203, and a manufacturability reasoning module 1204; wherein... The acquisition module 1201 is used to acquire the original engineering data and determine at least one view to be inferred corresponding to the original engineering data; The manufacturing parameter semantic decoding module 1202 is used to perform manufacturing parameter semantic decoding on each view to be reasoned based on the first multimodal large model obtained through pre-training, to obtain the manufacturing parameter set corresponding to each view to be reasoned. The manufacturing parameter set includes information on multiple manufacturing parameters, and the information on each manufacturing parameter includes: parameter label, parameter value, parameter unit, and the bounding box corresponding to the parameter. The process intent semantic recognition module 1203 is used to perform process intent semantic recognition on each view to be reasoned based on the pre-trained second multimodal large model, and obtain the material process context structure information corresponding to each view to be reasoned. The material process context structure information is used to characterize the manufacturing intent related to materials and processes. The manufacturability reasoning module 1204 is used to perform manufacturability reasoning based on the set of manufacturing parameters corresponding to each view to be reasoned and the material process context structure information, and generate a manufacturability report corresponding to the original engineering data.
[0177] Optionally, the manufacturability reasoning module 1204 is specifically used for: Multi-dimensional manufacturability reasoning is performed on the set of manufacturing parameters and the material process context information to determine at least one risk information, at least one optimization information corresponding to each risk information, and efficiency information corresponding to each optimization information in the original engineering data under multiple dimensions. Based on each risk information, each corresponding optimization information, and each corresponding efficiency information, a manufacturability report corresponding to the original engineering data is generated.
[0178] Optionally, the manufacturability reasoning module 1204 is specifically used for: Determine the risk weight and severity of each risk information under multiple dimensions, and determine the comprehensive risk index of the original engineering data based on each risk information and its risk weight and severity under multiple dimensions; Based on the comprehensive risk index, the production execution conclusions corresponding to the original engineering data are determined. The production execution conclusions include: production scheduling recommendation, production scheduling prohibition, or optimized production scheduling. Fill the preset template with the production execution conclusions, the optimization information corresponding to each risk information, and the efficiency information corresponding to each optimization information to generate a manufacturability report corresponding to the original engineering data.
[0179] Optionally, the manufacturability reasoning module 1204 is specifically used for: Based on the pre-trained third multimodal large model, cost manufacturability reasoning is performed on the manufacturing parameter set and material process context structure information to determine at least one risk information corresponding to the original engineering data in the cost dimension, as well as the optimization information corresponding to each risk information, and the efficiency information corresponding to the optimization information. Based on the pre-trained fourth multimodal large model, efficiency manufacturability reasoning is performed on the manufacturing parameter set and material process context structure information to determine at least one risk information corresponding to the original engineering data in the efficiency dimension, as well as the optimization information corresponding to each risk information, and the efficiency information corresponding to the optimization information. Based on the pre-trained fifth multimodal large model, supply chain manufacturability reasoning is performed on the manufacturing parameter set and material process context structure information to determine at least one risk information corresponding to the original engineering data in the supply chain dimension, as well as the optimization information corresponding to each risk information, and the efficiency information corresponding to the optimization information.
[0180] Optionally, a parameter semantic decoding module 1202 is manufactured, specifically for: Each view to be reasoned is input into the pre-trained annotation space anchoring model. The annotation space anchoring model performs omnidirectional annotation space anchoring on each view to be reasoned, resulting in a set of candidate region descriptors for each view to be reasoned. The set of candidate region descriptors includes multiple candidate region descriptors, which include: view identifier, bounding box parameters, category labels corresponding to the bounding box parameters, and image cropping blocks corresponding to the bounding box parameters. The candidate region descriptor set corresponding to each view to be inferred is input into the first multimodal large model, and the first multimodal large model performs semantic decoding to obtain the manufacturing parameter set.
[0181] Optionally, the annotation space anchoring model includes: a multi-scale feature fusion module, a rotation-sensitive detection head module, a coordinate refinement module, a geometric topology post-verification module, and a backbone network; the manufacturing parameter semantic decoding module 1202 is specifically used for: The view to be reasoned is input into the multi-scale feature fusion module, which perceives the symbols and long-size lines in the view to be reasoned and obtains the first enhanced features. The view to be reasoned is input into the rotation-sensitive detection head module, which then regresses the rotation direction in the view to be reasoned to obtain the second enhanced feature. The view to be reasoned is input into the coordinate refinement module, which then predicts the local offsets in the view to be reasoned, thus obtaining the third enhanced feature. The view to be reasoned is input into the geometric topology post-verification module, which verifies the spatial relationships in the view to be reasoned, and obtains the fourth enhanced feature. The view to be inferred, the first enhanced feature, the second enhanced feature, the third enhanced feature, and the fourth enhanced feature are input into the backbone network to obtain the set of candidate region descriptors corresponding to the view to be inferred.
[0182] Optionally, the process intent semantic recognition module 1203 is specifically used for: Based on the pre-trained layout analysis model, layout analysis is performed on each view to be inferred, and multiple semantic information corresponding to each view to be inferred is obtained. The semantic information is input into the second multimodal large model, which performs context awareness and parsing to obtain at least one engineering string entity and the semantic constraints of each engineering string entity. Standardize and map the semantic constraints of each engineering string entity to obtain the material and process context structure information corresponding to each view to be reasoned.
[0183] Optionally, the process intent semantic recognition module 1203 is specifically used for: Based on the pre-trained layout analysis model, region detection and pixel scanning are performed on each view to be inferred, and the semantic regions, the bounding boxes corresponding to each semantic region, and the category labels of each semantic region in each view to be inferred are identified. Based on the preset edge detection operator, non-textual geometric features are extracted from each semantic region to generate a binary mask matrix corresponding to each semantic region. Based on the binary mask matrix corresponding to each semantic region and each semantic region, generate the clean region image corresponding to each semantic region; Based on the clean region image corresponding to each semantic region, determine the multiple text image blocks corresponding to each view to be inferred, the type label of each text image block, and the content of each text image block.
[0184] Optionally, the process intent semantic recognition module 1203 is specifically used for: Based on each project string entity and its semantic constraints, at least one semantic vector corresponding to each project string entity is constructed. Based on each semantic vector, the pre-set material grade knowledge base, and the pre-set synonym weights, the standard material grade corresponding to each semantic vector is determined; Based on the standard material grade corresponding to each semantic vector, determine the material properties and process intentions corresponding to each engineering string entity, and use the material properties and process intentions corresponding to each engineering string entity as material process context structure information.
[0185] The processing flow of each module in the device and the interaction flow between each module can be referred to the relevant descriptions in the above method embodiments, and will not be detailed here.
[0186] This application also provides an electronic device, such as... Figure 13 As shown, Figure 13 The schematic diagram of the electronic device structure provided in this application embodiment includes: a processor 1301 and a memory 1302, and optionally, a bus 1303. The memory 1302 stores machine-readable instructions executable by the processor 1301. When the electronic device is running, the processor 1301 and the memory 1302 communicate via the bus 1303. The processor 1301 executes the machine-readable instructions to perform the steps of the above-described method for manufacturability reasoning based on two-dimensional engineering drawings of a multimodal large model.
[0187] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the above-described method for manufacturability reasoning based on multimodal large models of two-dimensional engineering drawings.
[0188] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and devices described above can be referred to the corresponding processes in the method embodiments, and will not be repeated here. In the several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some communication interfaces; the indirect coupling or communication connection of devices or modules can be electrical, mechanical, or other forms.
[0189] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. If the functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes: USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, optical disks, and other media capable of storing program code.
[0190] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. A method for manufacturability reasoning of two-dimensional engineering drawings based on a multimodal large model, characterized in that, include: Obtain the original engineering data and determine at least one view to be inferred corresponding to the original engineering data; Based on the first multimodal large model obtained through pre-training, the manufacturing parameter semantic decoding is performed on each of the views to be reasoned to obtain the manufacturing parameter set corresponding to each of the views to be reasoned to. The manufacturing parameter set includes information on multiple manufacturing parameters, and the information on each manufacturing parameter includes: parameter label, parameter value, parameter unit, and the bounding box corresponding to the parameter. Based on the pre-trained second multimodal large model, process intention semantic recognition is performed on each of the views to be reasoned, and material process context structure information corresponding to each of the views to be reasoned is obtained. The material process context structure information is used to characterize the manufacturing intention related to materials and processes. Based on the set of manufacturing parameters corresponding to each of the views to be inferred and the material process context information, manufacturability inference is performed to generate a manufacturability report corresponding to the original engineering data.
2. The method according to claim 1, characterized in that, The process of performing manufacturability reasoning based on the manufacturing parameter set corresponding to each of the views to be reasoned and the material process context information, and generating a manufacturability report corresponding to the original engineering data, includes: Multi-dimensional manufacturability reasoning is performed on the set of manufacturing parameters and the material process context structure information to determine at least one risk information, at least one optimization information, and efficiency information corresponding to each of the original engineering data in multiple dimensions. Based on the risk information, the optimization information corresponding to each risk information, and the efficiency information corresponding to each optimization information, a manufacturability report corresponding to the original engineering data is generated.
3. The method according to claim 2, characterized in that, The step of generating a manufacturability report corresponding to the original engineering data based on each of the aforementioned risk information, each of the aforementioned optimization information, and each of the aforementioned efficiency information includes: Determine the risk weight and severity of each risk information under multiple dimensions, and determine the comprehensive risk index of the original engineering data based on each risk information and its risk weight and severity under multiple dimensions; Based on the comprehensive risk index, the production execution conclusion corresponding to the original engineering data is determined. The production execution conclusion includes: production scheduling recommendation, production scheduling prohibition, or optimized production scheduling. The production execution conclusions, the optimization information corresponding to each risk information, and the efficiency information corresponding to each optimization information are filled into a preset template to generate a manufacturability report corresponding to the original engineering data.
4. The method according to claim 2, characterized in that, The step of performing multi-dimensional manufacturability reasoning on the manufacturing parameter set and the material process context information to determine at least one risk information corresponding to the original engineering data in multiple dimensions, at least one optimization information corresponding to each risk information, and efficiency information corresponding to each optimization information includes: Based on the pre-trained third multimodal large model, cost manufacturability reasoning is performed on the manufacturing parameter set and the material process context structure information to determine at least one risk information corresponding to the original engineering data in the cost dimension and optimization information corresponding to each risk information, and efficiency information corresponding to the optimization information is determined. Based on the pre-trained fourth multimodal large model, efficiency manufacturability reasoning is performed on the manufacturing parameter set and the material process context structure information to determine at least one risk information corresponding to the original engineering data in the efficiency dimension and optimization information corresponding to each risk information, and efficiency information corresponding to the optimization information is determined. Based on the pre-trained fifth multimodal large model, supply chain manufacturability reasoning is performed on the manufacturing parameter set and the material process context structure information to determine at least one risk information corresponding to the original engineering data in the supply chain dimension, as well as optimization information corresponding to each risk information, and efficiency information corresponding to the optimization information.
5. The method according to claim 1, characterized in that, The first multimodal large model, based on pre-trained data, performs semantic decoding of manufacturing parameters on each of the views to be reasoned, obtaining a set of manufacturing parameters corresponding to each view to be reasoned, including: Each of the views to be reasoned is input into a pre-trained annotation space anchoring model. The annotation space anchoring model performs omnidirectional annotation space anchoring on each of the views to be reasoned, resulting in a set of candidate region descriptors for each view to be reasoned. The set of candidate region descriptors includes multiple candidate region descriptors, each of which includes: view identifier, bounding box parameter, category label corresponding to the bounding box parameter, and image cropping block corresponding to the bounding box parameter. The candidate region descriptor set corresponding to each of the views to be inferred is input into the first multimodal large model, and the first multimodal large model performs semantic decoding to obtain the manufacturing parameter set.
6. The method according to claim 5, characterized in that, The labeled spatial anchoring model includes: a multi-scale feature fusion module, a rotation-sensitive detection head module, a coordinate refinement module, a geometric topology post-verification module, and a backbone network; The step involves inputting each of the views to be reasoned into a pre-trained annotation space anchoring model, whereby the annotation space anchoring model performs omnidirectional annotation space anchoring on each of the views to be reasoned, resulting in a set of candidate region descriptors corresponding to each view to be reasoned, including: The view to be reasoned is input into the multi-scale feature fusion module, which perceives the symbols and long dimension lines in the view to be reasoned and obtains the first enhanced feature. The view to be reasoned is input into the rotation-sensitive detection head module, which then regresses the rotation direction in the view to be reasoned to obtain the second enhanced feature. The view to be reasoned is input into the coordinate refinement module, which predicts the local offsets in the view to be reasoned to obtain the third enhanced feature. The view to be reasoned is input into the geometric topology post-verification module, which verifies the spatial relationships in the view to be reasoned to obtain the fourth enhanced feature. The view to be reasoned, the first enhanced feature, the second enhanced feature, the third enhanced feature, and the fourth enhanced feature are input into the backbone network to obtain a set of candidate region descriptors corresponding to the view to be reasoned.
7. The method according to claim 1, characterized in that, The second multimodal large model, based on pre-trained data, performs process intent semantic recognition on each of the views to be reasoned, obtaining material process context structure information corresponding to each view, including: Based on the pre-trained layout analysis model, layout analysis is performed on each of the views to be inferred, and multiple semantic information corresponding to each of the views to be inferred is obtained. The semantic information is input into the second multimodal large model, which performs context awareness and parsing to obtain at least one engineering string entity and the semantic constraints of each engineering string entity. The semantic constraints of each of the engineering string entities are standardized and mapped to obtain the material process context structure information corresponding to each of the views to be reasoned.
8. The method according to claim 7, characterized in that, The step involves performing layout analysis on each of the views to be reasoned based on a pre-trained layout analysis model, obtaining multiple semantic information corresponding to each view to be reasoned, including: Based on the pre-trained layout analysis model, region detection and pixel scanning are performed on each view to be inferred, and the semantic regions, the bounding boxes corresponding to each semantic region, and the category labels of each semantic region in each view to be inferred are identified. Based on the preset edge detection operator, non-textual geometric features are extracted from each semantic region to generate a binary mask matrix corresponding to each semantic region. Based on the binary mask matrix corresponding to each semantic region and each semantic region, generate the clean region image corresponding to each semantic region; Based on the clean region image corresponding to each semantic region, determine the multiple text image blocks corresponding to each view to be inferred, the type label of each text image block, and the content of each text image block.
9. The method according to claim 7, characterized in that, The standardization mapping of each of the engineering string entities and their semantic constraints to obtain the material process context structure information corresponding to each of the views to be reasoned includes: Based on each of the project string entities and the semantic constraints of each of the project string entities, at least one semantic vector corresponding to each of the project string entities is constructed; Based on each semantic vector, a preset material grade knowledge base, and preset synonym weights, the standard material grade corresponding to each semantic vector is determined; Based on the standard material grade corresponding to each semantic vector, determine the material properties and process intentions corresponding to each engineering string entity, and use the material properties and process intentions corresponding to each engineering string entity as the material process context structure information.
10. An electronic device, characterized in that, include: The processor and memory, the memory storing machine-readable instructions executable by the processor, wherein when the electronic device is running, the processor executes the machine-readable instructions to perform the steps of the two-dimensional engineering drawing manufacturability reasoning method based on a multimodal large model as described in any one of claims 1 to 9.