Three-dimensional automatic modeling method and system based on AI large model technology
Through the three-dimensional automatic modeling method based on AI large-scale model technology, multi-modal data is processed, conflicts are analyzed and models are generated, and the problem of insufficient modeling efficiency and intelligence in the existing technology is solved, and efficient and reliable three-dimensional modeling is achieved.
Patent Information
- Application Number
- CN202510535354.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-04-27
AI Technical Summary
When processing multimodal data, it is difficult for the prior art to effectively identify semantic associations and conflicts between data, resulting in misalignment of model attributes or logical contradictions, and insufficient modeling efficiency and intelligence level.
The three-dimensional automatic modeling method based on AI big model technology is adopted. By obtaining the multimodal data input by the user, it converts it into a feature vector in a unified encoding format, and a dynamic arbitration mechanism is used to analyze conflicts. The initial three-dimensional model is generated by combining the large language model and neural radiation field, and real-time physical verification and dynamic correction are carried out.
Significantly improve the modeling efficiency and reliability in complex scenarios, ensure that the generated three-dimensional model meets the dual standards in geometric accuracy and semantic consistency, reduces the design iteration cycle, and reduces the design and communication costs.
Smart Images

Figure CN120088409A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of digital modeling, and relates to a three-dimensional automatic modeling method and system based on AI large model technology. Background Art
[0002] Existing methods mainly rely on manual intervention or multi-software collaborative processing, and there are core problems of insufficient modeling efficiency and intelligence level. Limited by the parsing ability of a single data source, when the input data contains multi-modal information such as images, texts, and videos, traditional algorithms are difficult to effectively identify the semantic associations and conflict relationships between data, resulting in model attribute misalignment or logical contradictions.
[0003] Traditional solutions mostly focus on single-modal data processing, and solve data conflicts through rule-based feature matching or static priority configuration. The typical modeling process sets data priorities manually or uses an independent physical simulation tool for offline verification after model generation. Although it can handle simple cases, it cannot cope with dynamically changing mixed input scenarios. The phased modeling tool causes the separation of geometric modeling and physical verification, and the lag of defect feedback prolongs the design iteration cycle.
[0004] Based on the above problems, traditional multi-modal interaction technology stays at the basic parameter adjustment level. Fuzzy semantic instructions are difficult to be accurately converted into model modification actions, and users need to go through multiple trials and errors to achieve the expected effect, greatly increasing the design and communication costs. Summary of the Invention
[0005] In order to solve the above problems, the present invention provides a three-dimensional automatic modeling method and system based on AI large model technology.
[0006] In the first aspect, the present invention provides a three-dimensional automatic modeling method based on AI large model technology, adopting the following technical solutions: The three-dimensional automatic modeling method based on AI large model technology includes the following steps: S1. Obtain multi-modal data input by the user, including at least two types of images, texts, and videos, and convert the multi-modal data into feature vectors in a unified encoding format; S2. Use a dynamic arbitration mechanism to perform conflict resolution on the feature vectors. When contradictions in the content of different modal data are detected, generate a conflict arbitration result based on the priority rules preset by the user or the default physical rationality rules; S3. According to the conflict arbitration result, generate an initial three-dimensional model through a jointly trained large language model and neural radiance field. The large language model is used to parse semantic constraint conditions, and the neural radiance field is used to reconstruct the three-dimensional volume according to the image space information; S4. Input the initial 3D model into the physical simulation engine for real-time physical verification, including balance detection, material load-bearing simulation, and dynamic collision testing, and output a corrected model with structural defect markings; S5. Receive dynamic correction instructions input by the user through natural language or gestures, and parse the correction instructions into model parameterization adjustment commands; S6. Based on the parameterization adjustment commands, perform incremental calculation and re-rendering on the associated areas in the model to generate the final 3D model, and update the verification results of the physical simulation engine.
[0007] In a further solution of the present invention, converting multi-modal data into feature vectors in a unified coding format includes the following steps: Input channel separation: allocate the input stream to the corresponding parsing module according to the type of multi-modal data; data dimension alignment: use a cross-modal embedding layer to convert features from different sources into hidden vectors of the same dimension; normalize and fuse the feature vectors after dimension alignment to eliminate the dimensional differences between different modal data.
[0008] In a further solution of the present invention, using a dynamic arbitration mechanism to resolve conflicts in feature vectors includes the following steps: Calculate the credibility weights of different modalities using a predefined priority rule library, and the priority rule library includes keyword trigger rules, device type preference rules, and scene default physical rules; According to the priority rule library, calculate the credibility weights between different modal data, and the credibility weights are positively correlated with data clarity and user instruction clarity; When the difference in credibility weights exceeds the set threshold, automatically mask the conflicting content in the low-weight modal data, otherwise generate a visual conflict prompt interface for the user to manually select.
[0009] In a further solution of the present invention, generating an initial 3D model through a jointly trained large language model and neural radiance field includes the following steps: Perform cross-modal attention alignment on text semantic vectors and image feature vectors to establish a weight mapping matrix between semantics and spatial features; Use a multiple constraint loss function to jointly optimize the decoding process, and the loss function includes geometric structure loss, semantic matching loss, and physical pre-verification loss.
[0010] In a further solution of the present invention, inputting the initial 3D model into the physical simulation engine for real-time physical verification includes the following steps: Discretize the initial 3D model into grid cells based on finite element analysis; Apply virtual mechanical loads to each grid cell, and the load type is automatically selected according to the scenario. Gravity and wind loads are applied in the indoor scenario, and torsion and impact loads are applied in the industrial part scenario; When the model deformation amount exceeds the safety threshold, a red highlight mark is generated at the corresponding grid position.
[0011] A further solution of the present invention is to parse the correction instruction into a model parameterization adjustment command, including the following steps: Map the natural language instruction to a parameter modification type through an intent recognition model, and the parameter modification type includes dimension adjustment, material replacement, and structural addition and deletion; Directly standardize the instruction containing exact numerical values into model parameters, and generate interpolation-style candidate solutions for instructions with fuzzy descriptions.
[0012] A further solution of the present invention is to perform incremental calculation on the associated area in the correction model and re-render to generate the final three-dimensional model, including the following steps: Define a local calculation area according to the influence range of the parameter adjustment command, calculate the association strength of adjacent voxels to determine the calculation boundary; set constraint equations for displacement gradient matching and strain energy conservation at the boundary to ensure the continuity of the old and new geometric structures.
[0013] A further solution of the present invention is to re-render to generate the final three-dimensional model, including the following steps: Adopt a multi-resolution rendering strategy, first present the calculation result with a low-precision grid, and asynchronously calculate the physical properties of the high-precision grid at the same time. Trigger surface parameterization reconstruction to complete the final rendering after the user confirms.
[0014] A further solution of the present invention is to update the verification result of the physical simulation engine, including the following steps: Only perform incremental finite element analysis on the calculation area and reuse the simulation data in the unmodified area; Dynamically adjust the weight parameters in the multiple constraint loss function according to the verification result to optimize the subsequent model generation process.
[0015] In a second aspect, the present invention provides a three-dimensional automatic modeling system based on AI large model technology, adopting the following technical solutions: A multi-modal data acquisition and conversion module, which is used to acquire at least two types of data among the images, texts, and videos input by the user and convert them into feature vectors in a unified coding format; A dynamic arbitration module, which is used to detect the content conflicts of different modal feature vectors and generate conflict arbitration results based on the priority rules preset by the user or the default physical rationality rules; A joint model generation module, according to the conflict arbitration result, parses the semantic constraint conditions through a jointly trained large language model, combines the neural radiance field to perform three-dimensional volume reconstruction on the image space information, and outputs an initial three-dimensional model; A physical simulation verification module, which is used to input the initial three-dimensional model into a physical simulation engine for real-time physical verification, including balance detection, material load-bearing simulation, and dynamic collision testing, and outputs a corrected model with structure defect marks; A dynamic correction instruction parsing module, which is used to receive dynamic correction instructions input by the user through natural language or gestures, and parse the correction instructions into model parametric adjustment commands; An incremental solving and rendering module, based on the parametric adjustment commands, performs incremental solving on the associated areas in the corrected model and re-renders to generate the final three-dimensional model, and updates the verification results of the physical simulation engine.
[0016] In summary, the present invention includes the following beneficial technical effects: 1. Significantly improve the modeling efficiency and reliability in complex scenarios. By using a dynamic arbitration mechanism to intelligently analyze potential conflicts in multi-modal input data, automatically apply preset priority rules or physical laws for content adjudication, effectively avoiding delays and errors in manual intervention decisions; when there are attribute contradictions between the uploaded images and text descriptions, generate arbitration results based on quantitative indicators such as data clarity and user instruction clarity, improving the fusion accuracy of cross-modal data; 2. Integrate large language models and neural radiance fields to achieve joint modeling of semantics and space, ensuring double compliance of the generated three-dimensional models in terms of geometric accuracy and semantic consistency. The real-time embedding of the physical simulation engine can identify structural defects in the initial modeling stage; 3. For the interactive optimization stage in response to user feedback, adopt an incremental solving and multi-resolution rendering strategy to achieve efficient response to local modifications and resource optimization; when the user adjusts model parameters through natural language instructions, only perform constraint solving on the associated areas instead of reconstructing the entire model. Combining low-precision grid real-time preview and asynchronous high-precision calculation, the dynamic correction time of complex models is reduced. Description of the Drawings
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. The drawings are used to provide a further understanding of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0018] Figure 1 Disclosed is a flow diagram of a three-dimensional automatic modeling method based on AI large model technology.
[0019] Figure 2 Disclosed are schematic diagrams of the comparison of the original feature distribution and the normalized feature distribution.
[0020] Figure 3A schematic diagram of structural deformation simulation under finite element analysis is disclosed.
[0021] Figure 4 A schematic diagram of a style transfer optimization curve is disclosed.
[0022] Figure 5 A structural schematic diagram of a three-dimensional automatic modeling system based on AI large model technology is disclosed. Detailed implementation manners
[0023] In order to make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.
[0024] The following is a preferred and detailed description of the present invention in conjunction with the attached Figures 1 - 5 drawings.
[0025] Referring to the attached Figure 1 drawings, the present invention provides a three-dimensional automatic modeling method based on AI large model technology, including the following steps: S1. Obtain multimodal data input by the user, including at least two types of images, texts, and videos, and convert the multimodal data into feature vectors in a unified coding format; S2. Use a dynamic arbitration mechanism to resolve conflicts in the feature vectors. When content contradictions in different modal data are detected, generate a conflict arbitration result based on the priority rules preset by the user or the default physical rationality rules; S3. According to the conflict arbitration result, generate an initial three-dimensional model through a jointly trained large language model and a neural radiance field. The large language model is used to parse semantic constraint conditions, and the neural radiance field is used to reconstruct the three-dimensional volume according to the image space information; S4. Input the initial three-dimensional model into a physical simulation engine for real-time physical verification, including balance detection, material load-bearing simulation, and dynamic collision testing, and output a corrected model with structural defect marks; S5. Receive dynamic correction instructions input by the user through natural language or gestures, and parse the correction instructions into model parameterization adjustment commands; S6. Based on the parameterization adjustment command, perform incremental calculation on the associated areas in the corrected model and re-render to generate the final three-dimensional model, and update the verification result of the physical simulation engine.
[0026] In one embodiment of the present invention, step S1 includes the following steps: Specifically, when a user uploads multimodal data through a terminal device, the initial processing of the multimodal data is completed through the following process; 1. Input channel separation: according to the type of multimodal data, the input stream is assigned to the corresponding parsing module. For example, image data enters the convolutional neural network to extract spatial features, text data enters the large language model to extract semantic keywords, and video data is decomposed frame by frame and then synchronously performs time series analysis and key frame extraction; 2. Data dimension alignment: Use a cross-modal embedding layer to convert features from different sources into latent vectors of the same dimension. Text features are compressed into 256-dimensional vectors through attention pooling, image features are reduced to 256 dimensions through global average pooling, and video features are output as 256-dimensional vectors after aggregating temporal information through a bidirectional long short-term memory network. 3. Refer to the attached Figure 2 , feature normalization fusion, linear transformation is performed on the above vectors to achieve dimensional unification by satisfying the following formula: ,in, Represents the normalized The dedimensionalization process is completed by using a feature vector, which provides a numerical consistency basis for subsequent conflict resolution. Indicates Input feature vectors represent raw data features from different modalities, such as convolutional features of images or semantic embedding vectors of text. Represents the mean of the data distribution, which is calculated by counting the feature vectors of all training data or the current processing batch. It represents the standard deviation of the characteristic data, reflects the degree of discreteness of the characteristic value, and is used to eliminate the influence of different dimensions on subsequent processing. It represents a very small constant to prevent division by zero errors, and tries to avoid division by zero errors when the standard deviation approaches zero.
[0027] Multimodal data, which includes input forms of more than two interactive media, such as a combination of users uploading room photos (images), verbalizing renovation requirements (text), and on-site panoramic videos (video) in an architectural design scenario; Unified encoding format, the process of converting heterogeneous data into the same mathematical expression, which realizes the conversion of unstructured data into vectorized representation through the embedding layer of the deep learning model, eliminating the interference of modal differences on subsequent processing; Feature vector, a high-dimensional numerical expression obtained through abstraction of a neural network, for example, the text description "round glass table" is converted into a 256-dimensional real-valued vector containing shape and material attributes.
[0028] For example, assume that the user provides a photo of a living room and a voice command "change the wall to light gray and keep the original wood floor", and the following process is performed: The RGB value of the wall color recognized in the image data is , and the texture of the floor material conforms to the characteristics of wood; after the voice command is converted into text through voice recognition, the keywords "wall", "light gray", and "retain the original wooden floor" are extracted; the weight given to the image feature vector , and the weight given to the text feature vector . After weighted fusion, a composite vector pointing to "light gray wall + original wooden floor" is generated.
[0029] In one embodiment of the present invention, step S2 includes the following steps: Establish a priority rule library, including keyword trigger rules for user voice commands, device type preference rules, and scene default physical rules; Specifically, the priority rule library represents a preset rule set for arbitrating multimodal data conflicts and includes the following core modules. Keyword trigger rules, which identify enhanced expressions of user commands through natural language processing. For example, when the voice command contains "must be retained" or "subject to the text", the priority of the corresponding modality is automatically increased; the keyword trigger rules are implemented jointly through a preset keyword list and a real-time intention recognition model. Device type preference rules, which assign weights according to the characteristics of the hardware where the data source is located. For example, a wide-angle video taken by a mobile phone is more suitable for spatial scale modeling than a text description, and the priority of an image taken by a professional single-lens reflex camera is higher than that of a sketch in terms of material details. Scene default physical rules, which embed strong constraint conditions of industry norms. For example, the position of load-bearing walls in a building scene cannot be modified, and the mechanical movement trajectory in an industrial scene must conform to mechanical principles.
[0030] Among them, natural language processing is a technology that converts the voice or text input by the user into structured semantic tags. The real-time intention recognition model is a deep learning model based on the attention mechanism, which is used to extract the core requirements and limiting conditions of the user command.
[0031] Exemplarily, when the user uploads a photo of a bedroom taken by a mobile phone and sends a voice command "change this bed to a round shape", the following process is executed: The keyword trigger rule detects that "change to" in the voice is a forced action command and sets the priority of the text modality to the highest. The device type preference rule recognizes that the photo is taken by a mobile phone and there may be perspective distortion, and gives priority to the shape of the bed described in the text. The scene default physical rule verifies whether the diameter of the round bed meets the remaining space in the bedroom. If it cannot be satisfied, a secondary confirmation is triggered.
[0032] Calculate the credibility weights of different modality data according to the priority rule library. The calculation of the credibility weights satisfies the following formula: , where represents the The credibility weight of a modality, where a larger value indicates a higher priority of the modality in conflict resolution. Represents the data clarity score, which is quantified by objective indicators such as image resolution, text statement integrity, and video frame rate. For example, The image clarity is , and the blurred image is . Represents the user instruction clarity score, which is calculated by analyzing the number of specific parameters included in the instruction. For example, the text description "a solid wood table that is 2 meters long and 1.5 meters wide" scores , and only the description "a large table" scores . Represents the scene adaptability score, which is set according to the matching degree between the data modality and the current modeling field. For example, in a mechanical design scene, the adaptability of engineering drawings is , and the adaptability of hand-drawn sketches is . , , Represents the dynamic adjustment coefficient, which is a weighted parameter that changes in real time according to the conflict scenario type and is determined by the real-time decision result of the priority rule library. For example, when the keyword trigger rule takes effect, the value is increased to to strengthen the impact of the user instruction.
[0033] When the difference in credibility weights exceeds the set threshold, the low-weight conflict content is automatically blocked; otherwise, a visual conflict prompt interface is generated; Specifically, the threshold is set using a dynamic double-interval decision method and satisfies the following conditions: The first threshold (absolute threshold), when the weight of a certain modality is higher than this value and exceeds twice the weight of other modalities, the low-weight conflict content is directly blocked; The second threshold (relative threshold), when the weights of all modalities are lower than the first threshold but the difference between the highest weight and the second-highest weight is greater than , arbitration is automatically executed; If the above conditions are not met, a conflict prompt interface is generated, and the conflict prompt interface displays the conflict area in a comparison view and provides three options: "Force Application", "Compromise Solution", and "Manual Adjustment".
[0034] Among them, the dynamic double-interval decision method represents a decision-making mechanism that adaptively selects an arbitration strategy based on the weight distribution. The comparison view is an interactive interface that superimposes different modality data and marks the conflicting areas with a semi-transparent mask and a highlighted box.
[0035] Exemplarily, the user inputs the following data to generate a clothing store display rack. Video shooting shows a "metal shelf with an L-shaped layout", and the text description is "a U-shaped wooden shelf is needed for customers to browse around conveniently". According to the calculation of the credibility weight, the video , the text , the weight difference between the credibility weights of the video and the text is , set the first threshold (absolute threshold) to , set the second threshold (relative threshold) to .
[0036] The weight difference is lower than the second threshold , generate a conflict prompt interface: the three-dimensional model of the L-shaped metal shelf of the video is displayed in the left area; the simulation effect of the U-shaped wooden shelf described in the text is displayed in the right area; the bottom toolbar provides three options: "Keep the video structure and use wooden material instead", "Rebuild the shelf according to the text description", and "Hybrid solution: U-shaped metal shelf".
[0037] In one embodiment of the present invention, step S3 includes the following steps: Perform cross-modal attention alignment between the text semantic output vector and the image feature vector; Specifically, during the joint training process, the text vector obtained by parsing the text semantics and the extracted image feature vector are subjected to cross-modal feature association, and a weight mapping matrix between the text vector and the image vector is established through the self-attention mechanism to assign the interaction weights of different modal features. In the feature dimension alignment stage, the multi-head attention mechanism is used to calculate the spatial matching degree of the text and image features, and a fused joint feature vector is generated. For cross-modal attention alignment, a semantic association relationship between different modal data is established through a weight assignment mechanism, and the attention score is used to measure the spatial matching priority of different modal features.
[0038] Exemplarily, input the text data "a modern-style chair with an arc-shaped waistline" and the corresponding multi-angle chair frame images. The text data is input into the large language model to extract semantic keywords, and semantic encodings of "arc, modern style, chair frame structure" are obtained. The image feature vector extracts the edge contour features from the side view of the chair. During the cross-modal attention alignment, the weight matrix of the text data and the image features is calculated, and it is found that the weight of the "arc" semantics and the curve part of the image features is the highest, and the weight of the "chair frame structure" semantics and the support frame of the image features is the second. Finally, a joint vector containing high-weight features is generated by fusion.
[0039] The neural network decoding layer sets a multiple constraint loss function, and the loss function includes a geometric structure loss, a semantic matching loss, and a physical pre-verification loss; Specifically, in the decoding stage of the large language model, a geometric structure loss function is defined to constrain the topological integrity of the generation model, a semantic matching loss is used to ensure the consistency between the model attributes and the input semantics, and a physical pre-verification loss is used to pre-optimize the local structures that are prone to physical failures. The three are combined with dynamic weights in the total loss, satisfying the following formula: , specifically, represents the semantic matching loss, which is the negative value of the cosine similarity between the text vector and the model attribute vector. represents the physical pre-verification loss, which is a penalty term generated based on the deformation of the finite element pre-analysis model. represents the total loss function. represents the geometric structure loss, specifically the distance between the generation model and the real three-dimensional shape, , and are the predicted point cloud and the real point cloud respectively. , , represent dynamically adjusted hyperparameters.
[0040] Among them, The distance is a metric for measuring the maximum and minimum distance between two point sets in three-dimensional space, used to evaluate the geometric error on the surface of the generation model. Finite element pre-analysis is a simulation method in which the model is discretized into mesh elements and the stress distribution is pre-calculated, used to identify high-stress areas in advance during the training phase.
[0041] Exemplarily, training to generate a three-dimensional model of a robotic arm part: Geometric structure loss, constraining the surface profile error between the generated part and the original design drawing to be less than 2 mm; Semantic matching loss, ensuring that the "hinged structure" attribute in the model matches the description of "rotatable joint" in the text input by more than 90%; Physical pre-verification loss, during the training phase, it is found that there is stress concentration at the connection, and the part thickness is adjusted in advance to reduce the stress peak by 40%.
[0042] In one embodiment of the present invention, step S4 includes the following steps: Based on finite element analysis, discretize the initial three-dimensional model into mesh elements; Specifically, perform discretization processing on the geometric surface and internal structure of the initial three-dimensional model, divide the model using tetrahedral or hexahedral mesh elements, ensure that the vertex coordinates of each unit are aligned with the feature points of the original model, and increase the mesh density in the stress concentration areas (connections, bending parts, etc.). The side length of the mesh element is set according to the accuracy requirements of the initial three-dimensional model, with a default side length of 5 mm for indoor scenes and reduced to 1 mm for industrial part scenes.
[0043] Among them, finite element analysis refers to a method of decomposing a continuous geometric body into a finite number of mesh elements, and realizing structural strength prediction by calculating the stress and strain of each element.
[0044] A mesh element represents a closed polyhedron structure composed of geometric vertices and edges, and is used to approximately express the mechanical behavior of complex shapes.
[0045] Exemplarily, when performing mesh division on a household kitchen island model, it is detected that the connection between the tabletop and the support column is a stress concentration area, and the mesh side length of this area is automatically compressed from 5 mm to 2 mm, while the adjacent areas maintain the original density.
[0046] A virtual mechanical load is applied to each mesh element. The virtual mechanical load is a set of equivalent external forces applied to the mesh element and is used to simulate the stress conditions under real working conditions. The type of virtual mechanical load is automatically selected according to the scenario: gravity and wind loads are applied in the indoor scene, and torsion and impact loads are applied in the industrial part scene; Specifically, for an indoor furniture model, a vertically downward gravity load is applied to each mesh element. The magnitude of the gravity load is calculated by the product of the volume of the network element, the material density, and the acceleration due to gravity, and satisfies the following formula, , where, represents the gravity load vector; represents the volume of the mesh element, and the volume of the tetrahedron or hexahedron is calculated from the vertex coordinates; represents the material density; represents the acceleration due to gravity; for an outdoor building model, a horizontal wind load is superimposed on each mesh element, and the wind pressure distribution is calculated using the fluid dynamics formula.
[0047] In the industrial part scene, a torque load is applied to parts such as gears, and satisfies the following formula, , where, represents the magnitude of the torque, represents the material torsional stiffness coefficient, represents the applied rotation angle; the impact load needs to be simulated by an instantaneous velocity change, for example, the momentum impact generated by hitting at speed.
[0048] Refer to Appendix Figure 3 , when the deformation of the initial three-dimensional model exceeds the safety threshold, a red highlight mark is generated at the corresponding mesh position; Specifically, through a finite element solver, the displacement of each mesh node under the action of the load is calculated, and the deformation is the maximum absolute value of the node displacement, and the deformation exceeds the preset safety threshold (the allowable deformation threshold of indoor furniture materials is of the length dimension, and that of steel parts is Trigger the alarm mechanism. In the 3D visualization interface, the over-limit elements are assigned a red highlight material, and the transparency of the material is adjusted according to the severity of the deformation. The greater the deformation, the deeper the red color. Generate a defect report, recording the coordinates of the over-limit elements, the deformation values, and the recommended corrective measures.
[0049] Among them, the safety threshold, which is the maximum allowable deformation of the material under safe working conditions, is preset according to the material mechanics performance table. The red highlight mark represents the visualization alarm identifier, which is used to guide the user to quickly locate the problem area.
[0050] Exemplarily, when testing the solid wood dining table model of pine material, it is detected that the deformation amount of the grid element in the middle of the table leg under the load is , and the threshold value of the pine material is set to . The deformation amount of the grid element in the middle of the table leg under the load exceeds the threshold value of the pine material, and the corresponding area is marked dark red, prompting "It is recommended to add a support beam or replace it with a high-density material".
[0051] In one embodiment of the present invention, step S5 includes the following steps: Map the natural language instruction to a parameter modification type through the intent recognition model. The types include size adjustment, material replacement, and structure addition and deletion; Specifically, mapping the natural language instruction to a parameter modification type through the intent recognition model includes the following steps: The user inputs an instruction in natural language, such as "Increase the width of the sofa by 20%" or "Replace the wall material with marble", receives the instruction through the speech recognition module or the text interface, and performs word segmentation and grammar analysis; Use the intent recognition model to parse the preprocessed text. The intent recognition model is trained based on a multi-layer architecture, including a semantic encoder and a classifier module. The encoder converts the input statement into a high-dimensional semantic vector, and the classifier outputs a probability distribution according to the preset modification type labels; When the highest probability label output by the classifier belongs to the preset parameter modification type (size adjustment, material replacement, structure addition and deletion), generate the corresponding modification type code. For multi-intent mixed statements (such as "Widen the sofa and replace the wood material"), the intent recognition model supports multiple label classification outputs.
[0052] Among them, the parameter modification type represents a set of predefined operation labels. Size adjustment involves changes in the geometric dimensions of the model, material replacement corresponds to modifications of the surface physical properties, and structure addition and deletion refer to topological addition and deletion of the model components. The intent recognition model represents a natural language processing model based on the architecture, which is fine-tuned through a large amount of dialogue data in the design field and is used to analyze the core operation intent of the user's instruction and extract quantifiable parameters.
[0053] Exemplarily, when the user inputs "lower the desk height to 75 cm", keywords "lower", "height", and "75 cm" are extracted. The semantic vector has a matching degree of 0.93 (the highest probability label) with the dimension adjustment type features, generating a modification type code: DIMENSION - ADJUST, and extracting the target parameter value of 75 cm.
[0054] If the instruction contains an explicit numerical description, it is directly converted into an absolute parameter value; if the instruction contains a fuzzy numerical description, a style transfer model is called to generate a set of candidate solutions. Specifically, for instructions with explicit numerical descriptions (such as "reduce by 20%" and "rotate by 45 degrees"), the numbers and units are extracted through regular expression matching and standardized into the unit system of the modeling software. For example, "reduce by 20%" is converted into a scaling factor of 0.8, and "rotate by 45 degrees" is converted into a radian value. For instructions lacking precise numerical descriptions (such as "a more retro look" and "increase the modern feel"), the style transfer model is activated. The adaptive instance normalization technique is adopted, and the target style is encoded as a direction vector in the feature space. The feature vector of the current model and the style vector are linearly interpolated to generate 5 intermediate results with different weight ratios for the user to interactively select.
[0055] Among them, the style transfer model refers to a visual conversion model based on the generative adversarial network (GAN), and its hidden layer activation value guiding matrix controls the stylization intensity of materials, colors, and textures. Adaptive instance normalization refers to a technique for dynamically adjusting the input data distribution, which decouples the style features and content features by statistically calculating the feature mean and variance.
[0056] Refer to Appendix Figure 4 , the calculation of the style interpolation equation satisfies the following formula. Among them, represents the output result feature vector, which is formed by superimposing the difference between the current model feature and the target style feature after weight correction, representing the new state representation of the model after style mixing. represents the current model feature vector, which is the three - dimensional model hidden feature representation extracted through a deep neural network, including the encoding information of geometric structure, material attributes, and light response, and its dimension is determined by the network architecture. represents the feature vector of the target style, which is generated by aggregating the features of a specific style (such as "retro" and "modern") in the pre - trained style library. represents the style mixing coefficient, whose value range is between 0 and 1, and determines the penetration intensity of the target style feature into the original model. When it is 0, the output completely retains the original model features. When it is 1, it is completely replaced with the target style.
[0057] Exemplarily, the user adjusts the industrial part model from the "ordinary steel" style (corresponding to ) to the "polished metal" style (corresponding to ), sets = 0.6, and follows the following process: Calculate the difference direction , and the vector includes the characteristic change trends such as enhanced metal reflection and reduced surface roughness. Apply 60% of the difference amount, and the specular reflection intensity of the original model is increased to 60% of the target value, while retaining 40% of the diffuse reflection attributes of the original material. The parts in the output result model show a semi-polished effect.
[0058] In one embodiment of the present invention, step S6 includes the following steps: Delimit the boundary of the local solution area according to the influence range of the parametric adjustment command; Specifically, when the user triggers the model modification instruction through natural language or the interaction interface, analyze the set of geometric voxels affected by the modification operation, and determine the solution boundary according to the correlation strength calculation formula. The correlation strength is positively correlated with the force conduction coefficient between adjacent voxels and the material connection method. After the boundary is delimited, lock the topological structure of the non-solution area; Among them, the boundary of the local solution area is a dynamic range that expands outward with the action point of the modification instruction as the center. The voxels closer to the action point are more sensitive to the adjustment instruction; the calculation of the correlation strength satisfies the following formula, , where represents the correlation strength between adjacent voxels, represents the difference in displacement amounts of adjacent voxels, represents the material elastic modulus adjustment coefficient, represents the voxel spacing.
[0059] Set a transition constraint equation at the boundary to ensure the continuity of the new and old geometric structures; Specifically, establish displacement and stress coordination conditions on the boundary surface between the local solution area and the non-solution area, and use a two-parameter constraint equation to ensure a smooth transition between the modified area and the original model, satisfying the following equation, Equation 1, the gradient matching condition, requires that the displacement gradient vector at the boundary remains consistent inside and outside the solution area, , represents the gradient vector of displacement, represents the internal displacement of the solution area, represents the external displacement of the original model; Equation 2, the energy conservation condition, restricts the change rate of the strain energy at the boundary not to exceed a preset threshold, , Represents the strain energy on the boundary, Represents the time of the strain energy change, Represents that the maximum energy change rate allowed by the material is the preset threshold, which is dynamically set according to the structural safety standard.
[0060] Among them, the displacement gradient vector is a vector index describing the direction and intensity of geometric deformation, and its dimension is consistent with the model space coordinate system; the strain energy is the elastic potential energy stored during the material deformation process, and is obtained by integrating the product of the principal stress and principal strain of all boundary elements.
[0061] Exemplarily, when the user shortens the length of the table leg in the furniture model from 1 meter to 0.8 meters, the above constraint equation is applied to the boundary where the table leg is connected to the tabletop to ensure that no abnormal deformation occurs in the tabletop support structure.
[0062] Adopt a multi-resolution rendering strategy, first present the solution results with a low-precision grid, and perform full-precision optimization after the user confirms; Specifically, generate a three-level grid hierarchy for the local solution area: the basic layer is a low-face-count grid discretized by an octree, the middle layer adopts a quadrilateral subdivision transition structure, and the detail layer retains the NURBS surface parameters of the original model; when rendering, the basic layer grid is preferentially displayed, and at the same time, the physical properties of the detail layer are calculated asynchronously in the background; after the user confirms the modification plan, perform surface parameterization reconstruction to complete full-precision rendering; Among them, octree discretization represents a data structure that recursively divides three-dimensional space into multiple cubic units, and each cubic unit corresponds to a grid vertex. NURBS surface parameters represent the mathematical expression form of non-uniform rational B-splines, which are used to accurately describe the shape of complex surfaces.
[0063] Exemplarily, when the user adjusts the diameter of the round hole of an industrial part, first display a low-precision ring composed of 24 triangular faces, and after the user confirms that the size is correct, then restore the real engineering precision model composed of 256 faces. Hierarchical rendering reduces real-time interaction latency, while retaining the high-fidelity details of the final output, and avoiding waste of computing resources during repeated adjustment processes.
[0064] See Appendix Figure 5 As shown, the present invention also proposes a three-dimensional automatic modeling system based on AI large model technology, including the following modules: A multi-modal data acquisition and conversion module, which is used to acquire at least two data types of images, texts, and videos input by the user and convert them into feature vectors in a unified coding format; A dynamic arbitration module, which is used to detect content conflicts of different modal feature vectors and generate conflict arbitration results based on the priority rules preset by the user or the default physical rationality rules; The joint model generation module, according to the conflict arbitration result, parses the semantic constraint conditions of the large language model trained jointly, combines the neural radiance field for three-dimensional volume reconstruction of the image spatial information, and outputs an initial three-dimensional model; The physical simulation verification module is used to input the initial three-dimensional model into the physical simulation engine for real-time physical verification, including balance detection, material load-bearing simulation, and dynamic collision testing, and outputs a corrected model with structural defect marks; The dynamic correction instruction parsing module is used to receive the dynamic correction instructions input by the user through natural language or gestures, and parse the correction instructions into model parameterized adjustment commands; The incremental solution and rendering module, based on the parameterized adjustment command, performs incremental solution and re-rendering on the associated areas in the corrected model to generate the final three-dimensional model, and updates the verification result of the physical simulation engine.
[0065] Each of the above-mentioned modules can be implemented in whole or in part by software, hardware, and their combination, supporting the hardware form to be embedded in or independent of the processor in the computer device, and at the same time also supporting the software form to be stored in the memory of the computer device, facilitating the processor to call and execute the operations corresponding to each of the above-mentioned modules.
[0066] It should be noted that the human body information (including but not limited to human body device information and personal information, etc.) and data (including but not limited to data for analysis, stored data, and displayed data, etc.) involved in the present invention are all information and data authorized by the human body or fully authorized by all parties. The collection, use, and processing of relevant data require relevant legal standards.
[0067] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.
Claims
1. A three-dimensional automatic modeling method based on AI large model technology, characterized in that: The following steps are involved: S1. Obtain multimodal data input by a user, including at least two types of images, texts, and videos, and convert the multimodal data into a feature vector in a unified coding format; S2. Analyze the conflicts of feature vectors using a dynamic arbitration mechanism. When conflicts are detected in different modal data, generate conflict arbitration results based on user-preset priority rules or default physical rationality rules. S3. Based on the conflict arbitration result, an initial 3D model is generated by jointly training a large language model and a neural radiation field. The large language model is used to parse semantic constraints, and the neural radiation field is used to reconstruct the 3D volume based on the image spatial information. S4, inputting the initial three-dimensional model into a physical simulation engine for real-time physical verification, including balance detection, material load-bearing simulation, and dynamic collision testing, and outputting a revised model containing structural defect markers; S5, receiving a dynamic correction instruction input by a user through natural language or gesture, and parsing the correction instruction into a model parameter adjustment command; S6. Based on the parametric adjustment command, the associated areas in the corrected model are incrementally solved and re-rendered to generate the final three-dimensional model, and the verification results of the physical simulation engine are updated.
2. The three-dimensional automatic modeling method based on AI large model technology according to claim 1 is characterized in that: Converting multimodal data into feature vectors in a unified encoding format includes the following steps: Input channels are separated, and the input stream is assigned to the corresponding parsing module according to the type of multimodal data. Data dimensions are aligned, and features from different sources are converted into latent vectors of the same dimension using a cross-modal embedding layer. The feature vectors after dimension alignment are normalized and fused to eliminate the dimensional differences between different modal data.
3. The three-dimensional automatic modeling method based on AI large model technology according to claim 1 is characterized in that: The dynamic arbitration mechanism is used to resolve conflicts between feature vectors, including the following steps: The predefined priority rule library calculates the credibility weights of different modalities. The priority rule library includes keyword trigger rules, device type preference rules, and scene default physical rules. According to the priority rule base, the credibility weights between different modal data are calculated. The credibility weights are positively correlated with the clarity of data and the clarity of user instructions. When the credibility weight difference exceeds the set threshold, the conflicting content in the low-weight modal data is automatically shielded, otherwise a visual conflict prompt interface is generated for users to manually select.
4. The three-dimensional automatic modeling method based on AI large model technology according to claim 1 is characterized in that: The initial 3D model is generated by jointly training a large language model and a neural radiation field, including the following steps: The text semantic vector and the image feature vector are aligned with cross-modal attention to establish a weight mapping matrix between semantic and spatial features; The decoding process is jointly optimized using multiple constraint loss functions, which include geometric structure loss, semantic matching loss, and physical pre-verification loss.
5. The three-dimensional automatic modeling method based on AI large model technology according to claim 4 is characterized in that: The initial 3D model is input into the physical simulation engine for real-time physical verification, including the following steps: Discretizing the initial three-dimensional model into grid units based on finite element analysis; Virtual mechanical loads are applied to each grid unit. The load type is automatically selected according to the scene. Indoor scenes are loaded with gravity and wind loads, and industrial parts scenes are loaded with torque and impact loads. When the model deformation exceeds the safety threshold, a red highlight mark is generated at the corresponding grid position.
6. The three-dimensional automatic modeling method based on AI large model technology according to claim 5 is characterized in that: Parsing the correction instructions into model parameterization adjustment commands includes the following steps: The intent recognition model is used to map natural language instructions into parameter modification types, including size adjustment, material replacement, and structure addition and deletion. Instructions with precise values are directly standardized into model parameters, and interpolation style candidates are generated for instructions with fuzzy descriptions.
7. The three-dimensional automatic modeling method based on AI large model technology according to claim 6 is characterized in that: Correct the associated areas in the model for incremental solution and re-render to generate the final 3D model, including the following steps: The local solution area is delineated according to the influence range of the parameter adjustment command, and the correlation strength of adjacent voxels is calculated to determine the solution boundary; the constraint equations of displacement gradient matching and strain energy conservation are set at the boundary to ensure the continuity of the new and old geometric structures.
8. The three-dimensional automatic modeling method based on AI large model technology according to claim 7 is characterized in that: Re-rendering to generate the final 3D model includes the following steps: A multi-resolution rendering strategy is adopted to first present the solution result with a low-precision mesh, and at the same time asynchronously calculate the physical properties of the high-precision mesh. After the user confirms, the surface parameterization reconstruction is triggered to complete the final rendering.
9. The three-dimensional automatic modeling method based on AI large model technology according to claim 8 is characterized in that: Updating the verification results of the physical simulation engine includes the following steps: Only incremental finite element analysis is performed on the solution area, and simulation data of the unmodified area is reused; The weight parameters in the multi-constraint loss function are dynamically adjusted according to the verification results to optimize the subsequent model generation process.
10. The three-dimensional automatic modeling system based on AI large model technology is characterized by: Includes the following modules: A multimodal data acquisition and conversion module, used to acquire at least two data types among images, texts, and videos input by users, and convert them into feature vectors in a unified coding format; Dynamic arbitration module, used to detect content conflicts between different modal feature vectors and generate conflict arbitration results based on user-preset priority rules or default physical rationality rules; The joint model generation module, based on the conflict arbitration results, jointly trains a large language model to parse semantic constraints, combines the neural radiation field to reconstruct the three-dimensional volume of the image spatial information, and outputs the initial three-dimensional model; The physical simulation verification module is used to input the initial 3D model into the physical simulation engine for real-time physical verification, including balance detection, material load-bearing simulation, dynamic collision test, and output a revised model containing structural defect markers; A dynamic correction instruction parsing module is used to receive dynamic correction instructions input by the user through natural language or gestures, and parse the correction instructions into model parameter adjustment commands; The incremental solution and rendering module, based on the parametric adjustment commands, corrects the associated areas in the model for incremental solution and re-renders to generate the final 3D model, and updates the verification results of the physical simulation engine.
Citation Information
Patent Citations
Multi-modal spatial data fusion method based on large model technology
CN119475253A
Robotic tactile sensing
US20220318459A1
Systems and methods for building material based determinations
WO2024152019A1
Systems, methods, devices, and platforms for industrial internet of things
WO2024155584A1
Multimodal user interfaces for interacting with digital model files
WO2025029976A2
Cited By
Light and shadow scheme optimization method and system based on digital twinborn and artificial intelligence
CN120337778A
A light and shadow scheme optimization method and system based on digital twin and artificial intelligence
CN120337778B
Building design scene automatic generation method and system based on artificial intelligence
CN120372782A
Method for generating scene atmosphere graph and three-dimensional model and constructing scene atmosphere graph and three-dimensional model in unreal engine
CN120451351A
Large model reasoning efficiency dynamic optimization and hardware sensing compression method
CN120494006A