Automatic modeling agent system based on MCP protocol and large model

Through the automatic modeling agent system based on MCP protocol and large models, the existing three-dimensional modeling tools are solved inefficient and complex design problems, and an efficient, automated and intelligent three-dimensional modeling process is realized.

CN120259557AActive Publication Date: 2025-07-04BEIJING URBAN CONSTRUCTION DESIGN & DEVELOPMENT GROUP CO LIMITED

Patent Information

Application Number
CN202510725151.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-07-04
Estimated Expiration
2045-06-03

AI Technical Summary

Technical Problem

Existing three-dimensional modeling tools rely on manual operational efficiency and high skill requirements. AI-based tools are difficult to deal with complex semantic requirements and high degree of freedom design, and the modeling process lacks dynamic intervention and processes that cannot be traced.

Method used

The automatic modeling agent system based on MCP protocol and large models is adopted, including a multimodal input interface module, a semantic understanding hub module, a MCP instruction compiler module and a modeling execution engine module, to realize the structured processing, semantic analysis and automated modeling of user intentions.

Benefits of technology

Three-dimensional modeling is completed without users understanding professional modeling terms, greatly improving the level of modeling automation and intelligence, and ensuring transparent and controllable modeling process and efficient execution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259557A_ABST
    Figure CN120259557A_ABST
Patent Text Reader

Abstract

The invention relates to an automatic modeling agent system based on an MCP protocol and a large model, and the system comprises a multi-modal input interface module which is used for receiving, analyzing and standardizing modeling instructions or intention descriptions inputted by a user through different media to form structured feature representations; the semantic understanding central module is used for performing semantic analysis on the structured feature representation by using a pre-trained large language model and outputting a user modeling intention; the MCP instruction compiler module is used for converting a user modeling intention into an MCP instruction sequence following an MCP protocol; and the modeling execution engine module is used for analyzing and automatically executing the MCP instruction sequence to complete modeling of the target three-dimensional model. According to the method, three-dimensional modeling can be completed without knowing professional modeling terms and concepts by a user, and the modeling automation and intelligence level is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of three-dimensional modeling technology, and particularly to an automatic modeling intelligent agent system based on the MCP protocol and large models. Background Art

[0002] Traditional 3D modeling tools, such as Blender and 3ds Max, mainly rely on manual operations to create and edit models. Although this method can provide highly customized results, its efficiency is relatively low, and it requires a high level of skills from the operator. In addition, most existing AI-based 3D modeling tools can only generate simple models according to fixed templates, and it is difficult to handle complex semantic requirements or achieve highly flexible designs.

[0003] In recent years, the development of large language models (such as GPT-4) has significantly improved the ability of multimodal semantic understanding. However, these advanced models still have deficiencies in deeply integrating with professional fields (such as 3D modeling), especially in how to transform high-level semantic understanding into specific modeling instructions.

[0004] The Model Context Protocol (MCP) defines the way to exchange context information between applications and AI models. This enables developers to connect various data sources, tools, and functions to AI models in a consistent manner. The goal of MCP is to create a general standard to make the development and integration of AI applications simpler and more unified. Through MCP, more efficient data interaction and model control can be achieved, providing a more flexible foundation for automated modeling.

[0005] The currently closest prior art is deep learning-based parametric modeling, whose technical architecture includes three core components. Semantic parsing module: uses a convolutional neural network (CNN) to extract features from the input text, captures local semantic features through convolutional operations, maps the extracted semantic features to a preset template library, and calls the corresponding modeling parameter template when the feature matching degree exceeds the threshold. If the matching fails, approximate parameters are generated through an algorithm. After analysis, it is found that this solution has multiple technical limitations: template dependence bottleneck, by combining standard parameters from the preset template library, resulting in insufficient innovation in the generated models and relying heavily on the content of the existing template library; lack of interactive control: the modeling process lacks a dynamic intervention mechanism, and users cannot correct parameters in real time during the generation process; the process cannot be traced: the system internally processes modeling instructions in an unstructured binary stream, resulting in a black box for the generation process. Summary of the Invention

[0006] To solve the above problems, the purpose of the embodiments of the present invention is to provide an automatic modeling intelligent agent system based on the MCP protocol and large models.

[0007] An automatic modeling intelligent agent system based on the MCP protocol and large models, comprising: A multimodal input interface module, configured to receive, parse, and standardize modeling instructions or intention descriptions input by users through different media to form a structured feature representation; A semantic understanding central module, configured to perform semantic analysis on the structured feature representation using a pre-trained large language model to output the user's modeling intention; An MCP instruction compiler module, configured to convert the user's modeling intention into an MCP instruction sequence that follows the MCP protocol; A modeling execution engine module, configured to parse and automatically execute the MCP instruction sequence to complete the modeling of the target 3D model.

[0008] Preferably, the multimodal input interface module includes: An input channel manager, configured to identify the type of user input data and input it into the corresponding pre-processor; the types of the input data include: text, image, and sketch; A text pre-processor, configured to pre-process text data to obtain structured text information; An image pre-processor, configured to pre-process image data to obtain structured image information; A sketch pre-processor, configured to pre-process image data to obtain an output structured sketch representation; A standardization fusion unit, configured to fuse the structured text information, structured image information, and structured sketch representation to form a structured feature representation.

[0009] Preferably, the semantic understanding central module includes: A multimodal fusion parser, configured to perform feature enhancement on the structured feature representation to form a comprehensive semantic representation vector or graph structure; A large model processing unit, configured to perform semantic analysis on the comprehensive semantic representation vector or graph structure using a pre-trained large language model to output the user's modeling intention.

[0010] Preferably, the MCP instruction compiler module includes: A semantic-to-instruction mapper, configured to map the user's modeling intention to corresponding MCP atomic instructions to form an MCP atomic instruction sequence; An instruction serialization and optimizer, configured to use an MCP instruction sequence intelligent optimization algorithm to sort the generated MCP atomic instruction sequence according to the logical dependency relationship to form an ordered instruction sequence.

[0011] Preferably, in the instruction serialization and optimizer, it includes: Calculate the dependency degree between two instructions to form a dependency graph; Traverse the nodes in the dependency graph, identify consecutive and mergable instruction types, and for the mergable instruction groups, calculate the total cost of separate execution and the execution cost of the merged instructions; Use the merge decision function to execute the eligible instructions in combination, and replace the original multiple nodes with a new merged node; According to the dependency relationship between instructions, divide the instructions without mutual dependencies into groups that can be executed in parallel. For each group of instructions that can be executed in parallel, calculate the serial execution cost and the parallel execution cost of all instructions in the group; When the cost of parallel execution is lower than the cost of serial execution, keep the group of instructions executing in parallel; otherwise, sort the group of instructions in descending order of execution cost for serial execution.

[0012] Preferably, the calculation formula for the dependency degree between two instructions is:

[0013] Among them, is the set of output objects of instruction i , is the set of input objects of instruction j , is the dependency type adjustment factor, represents the dependency degree between instruction i and instruction j .

[0014] Preferably, the calculation formula for the execution cost of an instruction is:

[0015] Among them: is the execution cost, is the weight of the i th instruction, is the basic cost of the i th instruction, is the parallel penalty factor, is the additional overhead of the j th parallel group.

[0016] Preferably, the merge decision function is:

[0017] Among them: is the cost of separately executing the first instruction, is the cost of separately executing the second instruction, is the execution cost after merging, is the merging threshold; when the value of the merging decision function is 1, the instructions are merged and executed, and at the same time, multiple original nodes are replaced by a new merged node.

[0018] Preferably, the modeling execution engine module includes: The MCP instruction parser is used to read and parse the MCP instruction sequence one by one, verify the legality of the instructions, and use the instructions that pass the verification as the parsed MCP instructions; The core modeling library interface is used to convert the parsed MCP instructions into calls to functions in the underlying modeling library of the modeling engine; The scene state manager is used to represent the state information of the three-dimensional scene being built in real time; The geometric calculation and constraint solver is used to optimize the initially built three-dimensional model by using the modeling accuracy optimization algorithm of constraint propagation; The real-time feedback interface is used to send the intermediate results or final model data during the modeling process to the user interface for real-time preview.

[0019] Preferably, in the geometric calculation and constraint solver, it includes: Obtain the initially built three-dimensional model and the geometric constraint list; the geometric constraints include: parallelism constraint, perpendicularity constraint, distance constraint, and symmetry constraint; Calculate the priority of each constraint in the geometric constraint list; among them, the priority calculation formula of the constraint is:

[0020] Among them, is the weight of the i-th constraint feature, is the i-th constraint feature function, is the relevance weight, is the constraint correlation degree, represents the priority level of constraint C in the entire modeling process; Continuously optimize the initially built three-dimensional model according to the priority of each constraint until the termination condition is met.

[0021] According to the specific embodiments provided by the present invention, the following technical effects of the present invention are disclosed: The present invention relates to an automatic modeling intelligent agent system based on the MCP protocol and a large model. Compared with the prior art, the present invention can complete three-dimensional modeling without the user's understanding of professional modeling terms and concepts, greatly improving the automation and intelligence level of modeling.

[0022] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following specific preferred embodiments are given, and in conjunction with the accompanying drawings, the detailed description is as follows. Description of the Drawings

[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0024] Figure 1 Schematic diagram of an automatic modeling intelligent agent system based on the MCP protocol and large model provided by the present invention; Figure 2 Schematic diagram of the multi-modal input interface module provided by the present invention; Figure 3 Schematic diagram of the semantic understanding central module provided by the present invention; Figure 4 Schematic diagram of the MCP instruction compilation module provided by the present invention; Figure 5 Schematic diagram of the modeling execution engine module provided by the present invention; Figure 6 Schematic diagram of the three-dimensional model output module provided by the present invention. Detailed implementation manners

[0025] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be construed as a limitation of the present invention.

[0026] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality" means two or more unless otherwise specifically defined.

[0027] In the present invention, unless otherwise clearly specified or limited, terms such as "installation", "connection", "linkage", "fixation" and the like shall be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be a direct connection or an indirect connection through an intermediate medium, and it may be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0028] Please refer to Figure 1 , an automatic modeling intelligent agent system based on the MCP protocol and large models, comprising: A multimodal input interface module, configured to receive, parse and standardize modeling instructions or intention descriptions input by a user through different media to form a structured feature representation; Please refer to Figure 2 , the multimodal input interface module is the perception front end of the system, responsible for receiving, parsing and standardizing modeling instructions or intention descriptions input by a user through different media (text, image, sketch). Its core goal is to convert heterogeneous and unstructured user inputs into unified, structured feature representations containing preliminary semantic information for downstream semantic understanding centers to process. Among them, the multimodal input interface module includes: An input channel manager: identify the type of input data (such as.txt,.jpg,.png, real-time stream of a graphics tablet, etc.), and route the data to the corresponding preprocessor.

[0029] A text preprocessor: receive natural language text input, perform sentence segmentation, word segmentation, part-of-speech tagging, named entity recognition (recognize object nouns, attribute words, spatial relation words, etc.), and dependency syntax analysis (preliminarily recognize the modification and domination relationships between words) on the input content. Output structured text information, a word sequence with annotations and preliminary semantic primitives.

[0030] An image preprocessor: receive a reference picture or photo in bitmap format. Perform image denoising, color correction, and size normalization on the input picture; use an object detection model to identify the main objects and their bounding boxes in the image; use an image segmentation model to extract object contours or masks; use a feature extraction network to extract global and local visual features. Output an object list (including category, position, mask), a set of visual feature vectors.

[0031] A sketch preprocessor: receive a hand-drawn sketch in vector format (such as SVG) or bitmap format.

[0032] Perform denoising, binarization, and skeleton extraction on the input content for the bitmap sketch; use the stroke segmentation algorithm to split continuous lines into meaningful stroke segments; use the geometric primitive recognition algorithm to recognize straight lines, circles, or more complex curve recognition methods based on learning, and abstract the strokes into geometric primitives; analyze the connection relationships and topological structures between the primitives. Output a structured sketch representation, a stroke sequence, a list of geometric primitives, and their connection relationships.

[0033] Standardized fusion unit: Receive the structured information from each preprocessor. Map the information of different modalities to a shared feature space to generate a unified internal data structure, such as a composite feature object containing information such as text semantic primitives, visual object lists, and sketch geometric structures. This object should contain metadata such as timestamps and input sources.

[0034] Through diverse preprocessors and standardization processes, the requirements for the user input format and professionalism are reduced, laying a foundation for accurately understanding the user's intention subsequently.

[0035] Semantic understanding central module, used to perform semantic analysis on the structured feature representation using a pre-trained large language model and output the user's modeling intention; Please refer to Figure 3 , the semantic understanding central module serves as the cognitive core of the system. This module is responsible for performing in-depth semantic analysis and reasoning on the standardized multi-modal features provided by the input interface module, accurately and completely understanding the user's modeling intention, and parsing it into a series of clear high-level semantic instructions or parametric descriptions oriented to the modeling task.

[0036] Furthermore, the semantic understanding central module includes: Multi-modal fusion parser: Receive the standardized composite feature object, and adopt mechanisms such as cross-attention in the Transformer architecture to deeply fuse text, image, and sketch features, capturing cross-modal correlation information (for example, the correlation between the text description "red cube" and the red area in the image and the square contour in the sketch). Use a graph neural network (GNN) to jointly model the text semantic primitives and the sketch topological structure.

[0037] Output: A highly fused comprehensive semantic representation vector or graph structure rich in context information.

[0038] Large model processing unit: Receive the fused semantic representation. This large model is a large language model (LLM) or visual language model (VLM) that has been fine-tuned following instructions in the modeling domain.

[0039] Intent Recognition: Determine whether the user's core instruction is to create a new object, modify an existing object, combine objects, or query the status, etc.

[0040] Entity and Attribute Extraction: Accurately identify the object to be operated on (such as "table", "chair"), the attributes of the object (such as "length, width, and height", "color is brown", "material is wood"), the spatial relationship between objects (such as "above...", "adjacent to", "aligned"), quantity (such as "three"), etc.

[0041] Constraint Resolution: Understand and process the constraint conditions in the user's description (such as "height does not exceed 1 meter", "parallel to the wall").

[0042] Ambiguity Resolution and Information Completion: When the user input is vague or insufficient in information, use the knowledge and reasoning ability of the large model to clarify (may need to interact with the user and feedback through the input interface module) or make reasonable completion based on common sense (such as if the number of chair legs is not specified, default to 4).

[0043] Finally, output a structured high-level semantic description, such as a JSON object or a similar data structure that contains a list of objects, a property dictionary for each object, a relationship graph between objects, and global constraint conditions.

[0044] The MCP Instruction Compiler Module is used to convert the user's modeled intent into a sequence of MCP instructions that follow the MCP protocol; Please refer to Figure 4 , the MCP Instruction Compiler Module serves as a bridge between semantic understanding and specific execution. This module is responsible for accurately translating the abstract, high-level semantic description output by the semantic understanding center into a series of low-level, executable atomic instruction sequences that follow a predefined modeling control protocol.

[0045] Semantic to Instruction Mapper: Maintain a mapping rule library to map high-level semantic concepts (such as "create a cuboid", attributes: length = 10, width = 5, height = 2) to one or more MCP atomic instructions: CREATE_PRIMITIVE(type = CUBOID, size = [10,5,2], id = obj1)) Handle the decomposition of complex operations, such as "place cylinder A at the center directly above cube B", decomposed into the following multiple instructions: GET_BOUNDING_BOX(id = B) CALCULATE_CENTER(box =...) GET_BOUNDING_BOX(id = A) CALCULATE_BOTTOM_CENTER(box=...) MCP instruction set definition library: Define all atomic operations of the MCP protocol, including but not limited to: Geometry creation: CREATE_PRIMITIVE (parameters: type, size, ID), CREATE_FROM_CURVE(parameters: curve data, ID) Transformation operations: TRANSLATE, ROTATE, SCALE (parameters: object ID, transformation parameters) Boolean operations: BOOLEAN_UNION, BOOLEAN_DIFFERENCE, BOOLEAN_INTERSECT (parameters: list of object IDs) Modifier application: APPLY_MODIFIER (parameters: object ID, modifier type, parameters) Materials and appearance: ASSIGN_MATERIAL (parameters: object ID, material ID / parameters), SET_COLOR (parameters: object ID, color value) Scene management: GROUP_OBJECTS, SET_PARENT (parameters: list of object IDs / parent-child ID) Each instruction has strict parameter definitions and expected behaviors.

[0046] Furthermore, the MCP instruction compiler module includes: Material database interface: According to the material information in the semantic description (such as "wooden", "metal brushed"), query the pre-set material database (this library contains material names, PBR parameters, texture map paths, etc.). Obtain the matching material ID or detailed parameters and fill them into relevant instructions such as ASSIGN_MATERIAL.

[0047] Instruction serialization and optimizer: Utilize the intelligent optimization algorithm for MCP instruction sequences to sort the generated atomic instructions according to logical dependencies (such as objects must be created before transformation) to form an ordered instruction sequence. Merge consecutive identical transformations, eliminate redundant operations, and adjust the instruction order according to the characteristics of the execution engine to improve efficiency. This module outputs the final, optimized MCP instruction stream (for example, a JSON array, where each element is an MCP instruction and its parameters). This module solves the problem of lack of standardization in the modeling process, ensures the accurate transmission of modeling intent to execution operations, and lays a foundation for automated execution.

[0048] The intelligent optimization algorithm for MCP instruction sequences can optimize the execution efficiency of instruction sequences, ensure the correctness of instruction sequences, reduce redundant operations, and improve the efficiency of parallel instruction execution.

[0049] The formula for calculating the instruction cost is: Where: is the total cost, is the weight of the i-th instruction, is the basic cost of the i-th instruction, is the parallel penalty factor, is the additional overhead of the j-th parallel group.

[0050] This formula is used to assign weights to each node when constructing a graph and to evaluate the costs of serial and parallel execution strategies during the final optimization.

[0051] (2)Calculation of instruction dependency Degree of dependency between two instructions: Where: is the set of output objects of instruction i, is the set of input objects of instruction j, is the dependency type adjustment factor.

[0052] This formula is used to obtain the tightness of the association between instructions and establish the dependency relationship between instructions.

[0053] (3)Instruction merging evaluation function Merging decision function: Where: 、 are the costs of executing the two instructions separately, is the execution cost after merging, is the merging threshold.

[0054] This formula is used to determine whether it is worthwhile to merge multiple simple operations (such as consecutive transformations) into a more complex but potentially more efficient operation to reduce the number of instructions and potential overhead.

[0055] (4)Parallelism optimization formula Calculation of the optimal parallelism: Where: is the serial execution time, is the execution time of a single instruction, is the parallel overhead factor, is the maximum parallelism supported by the hardware This formula is used to identify groups of instructions that can be executed concurrently based on dependencies, and determine whether to execute the instruction groups in parallel or serially ultimately, as well as the optimal number of instructions for parallel execution, according to the cost model and hardware limitations.

[0056] The algorithm flow is as follows: 1. Start: Input the original MCP instruction sequence instruction_sequence.

[0057] 2. Build a dependency graph: Traverse the input instruction_sequence.

[0058] Calculate node costs: For each instruction inst, calculate its execution cost cost according to its type, such as CREATE_PRIMITIVE, TRANSFORM, using the predefined instruction cost calculation formula (1), and create an InstructionNode.

[0059] Establish instruction dependencies: Through the input input_objects and output output_objects of the instructions. Use formula (2) to calculate the object overlap information to judge the dependency strength and form a dependency graph.

[0060] 3. Merge mergeable instructions: Traverse the nodes in the constructed dependency graph to identify consecutive and mergeable instruction types (such as consecutive TRANSFORM operations).

[0061] Evaluate the merge benefits: For the mergeable instruction groups, calculate the total cost of separate execution and the cost of the merged instructions. Use the logic of formula (3) to make a decision: If the cost reduction exceeds the threshold, then perform the merge and replace the original multiple nodes with a new merged node.

[0062] 4. Analyze the possibility of parallelism: According to the dependencies between instructions, divide the instructions without mutual dependencies into groups that can be executed in parallel.

[0063] 5. Calculate the execution costs of each group: For each group of instructions that can be executed in parallel, use formula (1) again to calculate the total serial cost and the estimated parallel execution cost of all instructions within the group.

[0064] 6. Optimize the sequence according to the cost and parallel strategy Compare the serial cost and the parallel cost. Apply formula (4) for decision analysis: If the cost of parallel execution is lower than the serial cost, then keep the group of instructions executed in parallel; otherwise, it may be necessary to sort the group of instructions in descending order of cost for serial execution, or limit the degree of parallelism.

[0065] Determine the final instruction execution order according to the optimization decision (maintaining parallelism, serializing, limiting parallelism).

[0066] 7. End: Output the optimized MCP instruction sequence.

[0067] This algorithm introduces graph-based dependency analysis to ensure that instruction optimization does not disrupt the execution order, uses a cost model for quantitative analysis to achieve the optimal execution strategy. It supports instruction merging and parallel optimization. Considering hardware characteristics, it dynamically adjusts the parallelism of instruction execution.

[0068] Please refer to Figure 5 , the modeling execution engine module, which is used to parse and automatically execute the MCP instruction sequence to complete the modeling of the target 3D model.

[0069] As the system execution core, the modeling execution engine module is responsible for receiving and precisely parsing the MCP instruction sequence, calling the underlying geometric modeling library functions, gradually constructing or modifying the model in the virtual 3D space, and maintaining the scene state.

[0070] MCP Instruction Parser: Read the MCP instruction stream one by one. Verify the legality of the instructions (whether the instruction name exists, whether the parameter types and quantities are correct). Parse the instruction parameters to extract the operation type and specific values.

[0071] Core Modeling Library Interface: This is an adaptation layer that converts the parsed MCP instructions into function calls to specific underlying modeling libraries. This core modeling library interface docks with the underlying modeling libraries of various modeling engines on the market (such as Blender Python API, Open CASCADE Technology, Unity Engine API, or self-developed geometric libraries, etc.).

[0072] For example, CREATE_PRIMITIVE(type = CUBOID, size = [10,5,2], id = obj1) may be translated to bpy.ops.mesh.primitive_cube_add(size = 1, scale=(10,5,2)) and name the returned object obj1.

[0073] Scene State Manager: Maintain an internal data structure (such as a scene graph Scene Graph) to represent the state of the current 3D scene in real time, including the positions, poses, geometries, materials, hierarchical relationships, etc. of all objects. When executing instructions, update the corresponding nodes and attributes in the scene graph. Provide query interfaces for the instruction compiler or semantic understanding module to obtain the current scene information (for constraint checking or relative operations).

[0074] Geometric calculation and constraint solver: Using an algorithm for optimizing modeling accuracy based on constraint propagation to provide underlying geometric operation support (such as calculating bounding boxes, distances, intersection points, etc.). Handling complex constraints involved in MCP instructions (such as alignment, equidistance).

[0075] The algorithm for optimizing modeling accuracy based on constraint propagation can improve modeling accuracy, keep the model under geometric constraints, optimize the model topology, and ensure modeling stability.

[0076] (1) Constraint priority calculation Constraint priority scoring Where: is the weight of different constraint features, is the constraint feature function, is the relevance weight, is the constraint correlation degree.

[0077] This formula is used to guide the order and focus of optimization to ensure that key constraints are processed first.

[0078] (2) Geometric constraint error calculation Parallelism constraint error Perpendicularity constraint error Where: is the first vector, representing a certain direction vector in the model, is the second vector, representing another direction vector in the model, is the dot product of the two vectors, is the vector 's modulus (length), is the vector 's modulus (length), is the parallelism error, and the smaller the value, the more parallel the two vectors are, is the perpendicularity error, and the smaller the value, the more perpendicular the two vectors are.

[0079] Distance constraint error Where: is the coordinate vector of the first point, is the coordinate vector of the second point, is the target distance value, is the actual distance between the two points, is the distance error, and the smaller the value, the closer the actual distance is to the target distance.

[0080] Symmetry constraint error Where: is the coordinate vector of the original point, is the coordinate vector of the mapping point of point p with respect to the symmetry plane, is the parameter of the symmetry plane (normal vector and distance), is the point the distance to the plane, is the point the distance to the plane, is the symmetry error, and the smaller the value, the better the symmetry.

[0081] The geometric error formula is the core calculation of each iteration, quantifying the deviation degree of the current model from each constraint.

[0082] (3)Constraint propagation weight calculation Propagation weight matrix: Where: is the topological distance between constraints i and j, is the propagation attenuation factor This formula expresses the mutual influence between constraints, enabling the adjustment of one constraint to be intelligently transmitted to the relevant parts.

[0083] (4)Model optimization objective function Where: is the error of the i-th constraint, is the regularization term, is the smoothness metric, 、 、 are the weight coefficients.

[0084] This formula defines the ultimate goal of the algorithm, which is a comprehensive manifestation of all constraint errors being satisfied, model regularization, and smoothness.

[0085] (5)Topological optimization quality metric Patch quality assessment: Where: is the patch area, is the patch side length.

[0086] Edge sharpening metric: Where: is the current dihedral angle, is the target dihedral angle.

[0087] The topological optimization quality metric formula is used to further improve the model quality after iterative optimization, ensuring that there are no cracks, bad faces, or unnatural sharp edges.

[0088] (6) Convergence judgment criterion: Iteration termination condition Where: is the error of the i-th constraint in the t-th iteration, is the convergence threshold, is the maximum number of iterations.

[0089] This formula is used to determine whether the iterative process has achieved a good enough result and stop the iteration.

[0090] The algorithm flow is as follows: 1. Start: Input the initial 3D model model and a list of geometric constraints constraints.

[0091] 2. Constraint layering: Traverse the constraints list.

[0092] For each constraint, calculate its priority Priority(C) using formula (1), which comprehensively considers factors such as constraint type, complexity of the objects involved, and correlation with other constraints.

[0093] Divide the constraints into different levels constraint_layers according to the calculated priority. Constraints with high priority will be preferentially satisfied and have a greater influence in subsequent optimizations.

[0094] 3. Build a constraint network: Based on the layering results, build an internal constraint network data structure. This data includes the constraints themselves and the association information between constraints. 4. Initialize the iteration: Set the iteration counter t = 0 and initialize the model state.

[0095] 5. Enter the iterative optimization loop: Start the main optimization loop with the condition that the iteration count t has not reached the maximum value and the model has not converged.

[0096] 6. Traverse the constraint network: In each iteration, traverse the constraints in the constraint network. The traversal order is affected by the priority.

[0097] 7. Apply a single constraint: Use equations (2), (3), (4), and (5) to calculate the constraint error under the current constraints. Adjusting one constraint may affect other associated constraints or model parts, and the influence strength can be calculated using weights similar to this formula to guide the direction and magnitude of model updates).

[0098] 8. Record the maximum error: Record the maximum error calculated for all constraints in this iteration.

[0099] 9. Check for convergence: After completing a round of constraint traversal or reaching the established evaluation point, use equation (10) to determine whether the convergence condition is met: Compare whether the maximum error or error change amount in the current iteration is less than the threshold.

[0100] 10. Update the model state: If not converged, calculate the update amount of the model parameters (such as vertex coordinates) based on all the constraint errors calculated in this iteration (and possible propagation effects). This update aims to minimize the overall optimization objective function equation (7), which is usually a weighted sum of all constraint errors, including regularization terms and smoothness terms. Update the model and increment the iteration count t = t + 1, and return to step 1 for the next iteration.

[0101] 11. Perform topology optimization: After the iteration loop ends (converged or reached the maximum number of iterations), enter the topology optimization phase. Perform a series of operations to improve the mesh quality and structure of the model, such as: boundary repair, singularity handling, patch optimization can use equation (8), and edge sharpening can use equation (9).

[0102] 12. End: Output the final optimized 3D model that meets the constraints and has a good topological structure.

[0103] This algorithm ensures that important constraints are satisfied first through multi-level and constraint hierarchical processing, introduces a constraint propagation mechanism to achieve global optimization, integrates topology optimization to improve model quality, and guarantees the convergence of the algorithm through an adaptive iteration strategy.

[0104] Error handling and rollback mechanism: If an error occurs when executing an instruction (such as a boolean operation failure, invalid parameters), record the error and attempt to recover or skip. Support transaction processing, allowing rollback to the previous stable state when a critical step fails.

[0105] Real-time feedback interface: Provide an API that allows sending intermediate results or final model data during the modeling process to the user interface for real-time preview. This module solves the problem of low modeling efficiency by automatically executing standard instructions to complete the modeling; solves the problem of unstable model quality because the execution results of standard instructions are more controllable.

[0106] Please refer to Figure 6 The 3D model output module, as the delivery end of the system, is responsible for exporting the internal scene data structure finally constructed by the modeling execution engine into one or more standard 3D file formats according to user requirements, and can perform necessary optimization and post-processing.

[0107] Format exporter set: It contains dedicated exporters for multiple target formats (such as obj, fbx, gltf / glb, stl, step, etc.). Each exporter is responsible for converting the internal scene graph data (vertices, faces, normals, UV coordinates, materials, skeletal animations, etc.) into the data structure of the target format specification.

[0108] Model data converter: Handles coordinate system conversion (such as from left-handed system to right-handed system). Handles unit conversion (such as converting from internal unit meters to centimeters). Handles the conversion and packaging of material and texture information (such as generating bin and png files for glTF).

[0109] Post-processing and optimization unit, providing optional post-processing functions: Mesh optimization: Reduces the number of faces (Decimation), triangulation, removes redundant vertices. Normal calculation and smoothing group processing. UV unwrapping and optimization. Texture baking. Users can configure optimization parameters (such as the target number of faces, whether to retain details, etc.).

[0110] Output configuration and file writer: Receives the output format, file path, and export options (such as whether to export animations, whether to embed textures) specified by the user. Invokes the corresponding format exporter and optimization unit. Writes the final data to the file system.

[0111] Provides a standardized output interface, facilitating users to use the model on different platforms and software; the optimization options help improve the performance and quality of the final model.

[0112] The present invention integrates text descriptions, design drawings, and hand-drawn sketches as diverse inputs, and uses the understanding ability of a pre-trained large model to perform in-depth semantic analysis on the fused information to output the user's modeling intention. The MCP instruction compiler translates the modeling intention information into a standardized and executable atomic instruction sequence following the MCP protocol. The modeling execution engine will parse and automatically execute this MCP instruction sequence, and call the underlying modeling functions to complete the 3D model construction, without the user having to understand a large number of professional modeling terms and concepts, greatly improving the level of modeling automation and standardization.

[0113] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of technical solutions of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims described above.

Claims

1. An automatic modeling intelligent agent system based on the MCP protocol and large models, characterized in that It includes: A multimodal input interface module for receiving, parsing, and standardizing the modeling instructions or intention descriptions input by the user through different media to form a structured feature representation; A semantic understanding central module for using a pre-trained large language model to perform semantic analysis on the structured feature representation and output the user's modeling intention; An MCP instruction compiler module for converting the user's modeling intention into an MCP instruction sequence that follows the MCP protocol; A modeling execution engine module for parsing and automatically executing the MCP instruction sequence to complete the modeling of the target 3D model.

2. The automatic modeling intelligent agent system based on the MCP protocol and large model according to claim 1, characterized in that, The multimodal input interface module includes: An input channel manager for identifying the type of user input data and inputting it into the corresponding pre-processor; the types of the input data include: text, image, and sketch; A text pre-processor for pre-processing text data to obtain structured text information; An image pre-processor for pre-processing image data to obtain structured image information; A sketch pre-processor for pre-processing image data to obtain an output structured sketch representation; A standardization fusion unit for fusing the structured text information, structured image information, and structured sketch representation to form a structured feature representation.

3. An automatic modeling intelligent agent system based on the MCP protocol and a large model according to claim 2, characterized in that, The semantic understanding central module includes: A multimodal fusion parser for performing feature enhancement on the structured feature representation to form a comprehensive semantic representation vector or graph structure; A large model processing unit for using a pre-trained large language model to perform semantic analysis on the comprehensive semantic representation vector or graph structure and output the user's modeling intention.

4. An automatic modeling intelligent agent system based on the MCP protocol and large models according to claim 3, characterized in that, The MCP instruction compiler module includes: A semantic-to-instruction mapper for mapping the user's modeling intention to the corresponding MCP atomic instructions to form an MCP atomic instruction sequence; An instruction serialization and optimizer for using the MCP instruction sequence intelligent optimization algorithm to sort the generated MCP atomic instruction sequence according to the logical dependency relationship to form an ordered instruction sequence.

5. The automatic modeling intelligent agent system based on the MCP protocol and large model according to claim 4, wherein In the instruction serialization and optimizer, it includes: Calculating the dependency degree between two instructions to form a dependency relationship graph; Traversing the nodes in the dependency relationship graph, identifying consecutive and mergable instruction types, and for the mergable instruction groups, calculating the total cost of separate execution and the execution cost of the merged instructions; Using the merge decision function to merge and execute the eligible instructions, and replacing the original multiple nodes with a new merged node at the same time; According to the dependency relationship between instructions, dividing the instructions without mutual dependencies into groups that can be executed in parallel, and for each group of instructions that can be executed in parallel, calculating the serial execution cost and parallel execution cost of all instructions in the group; When the parallel execution cost is lower than the serial execution cost, keep the group of instructions to be executed in parallel; otherwise, sort the group of instructions in descending order of execution cost for serial execution.

6. The automatic modeling intelligent agent system based on the MCP protocol and large model according to claim 5, characterized in that, The calculation formula for the dependency degree between two instructions is: ; Among them, is the output object set of the instruction i . is the input object set of the instruction j . is the dependency type adjustment factor indicating the degree of dependency i between the instruction j and the instruction 7. The automatic modeling intelligent agent system based on the MCP protocol and large model according to claim 6, wherein The calculation formula for the execution cost of an instruction is: ; Wherein: is the execution cost, is the weight of the i th instruction, is the basic cost of the i th instruction, is the parallel penalty factor, is the overhead of the j th parallel group.

8. The automatic modeling intelligent agent system based on the MCP protocol and the large model according to claim 7, characterized in that, The merge decision function is: ; Wherein: is the cost of separately executing the first instruction, is the cost of separately executing the second instruction, is the execution cost after merging, is the merging threshold; when the value of the merging decision function is 1, the instructions are executed merged, and at the same time, multiple original nodes are replaced by a new merged node.

9. The automatic modeling intelligent agent system based on the MCP protocol and the large model according to claim 8, characterized in that, The modeling execution engine module includes: An MCP instruction parser for reading and parsing the MCP instruction sequence one by one, and verifying the legality of the instructions and taking the instructions passed the verification as the parsed MCP instructions; Core modeling library interface, used to convert the parsed MCP instructions into calls to functions in the underlying modeling library of the modeling engine; Scene state manager, used to represent the state information of the three-dimensional scene currently being built in real time; Geometry calculation and constraint solver, used to optimize the initially built three-dimensional model using the modeling accuracy optimization algorithm of constraint propagation; Real-time feedback interface, used to send the intermediate results or final model data during the modeling process to the user interface for real-time preview.

10. The automatic modeling intelligent agent system based on the MCP protocol and large model according to claim 9, wherein In the geometry calculation and constraint solver, it includes: Obtain the initially built three-dimensional model and the list of geometric constraints; the geometric constraints include: parallelism constraint, perpendicularity constraint, distance constraint, and symmetry constraint; Calculate the priority of each constraint in the geometric constraint list; among them, the priority calculation formula of the constraint is: ; Among them, is the weight of the i-th constraint feature, is the i-th constraint feature function, is the relevance weight, is the constraint correlation degree, indicates the priority level of constraint C in the entire modeling process; Continuously optimize the initially built three-dimensional model according to the priority of each constraint until the termination condition is met.

Citation Information

Patent Citations

  • Modeling processing method and device based on natural language, equipment and storage medium

    CN118151908A

  • Semantic perception recommendation method combining large model and knowledge graph

    CN119003787A

  • Protocol method for realizing intelligent connection of PLC (Programmable Logic Controller) with AI (Artificial Intelligence) large model

    CN119484673A

  • Method and system for realizing automatic code generation by using multiple AI Agents

    CN119861916A

  • 3D Modeling Construction System and Method through Text-Based Commands

    KR102616381B1

Cited By

  • Equipment abnormal fault recovery method and device based on large model, equipment and medium

    CN120469848A

  • Configuration method and device of storage cluster, electronic equipment and storage medium

    CN120762781A

  • Business service method, system and device based on user request, and electronic equipment

    CN121078011A

  • Engineering intelligent design method and system based on LLM and MCP

    CN121189202A

  • Government affair intelligent writing method and system based on multi-Agent collaboration and long memory technology

    CN121327156A