Three-dimensional model intelligent generation method and system based on natural language

By performing semantic analysis of the natural language description input by the user and the use of a three-dimensional model generation network, the problem of inability to understand the semantics and context of natural language in the prior art is solved, and the generation and dynamic adjustment of high-quality three-dimensional models are achieved, which lowers the generation threshold and improves modeling efficiency and ease of use.

CN120147551AActive Publication Date: 2025-06-13BEIJING DONGFANG AIDIPU DIGITAL TECH CO LTD

Patent Information

Application Number
CN202510324628.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-06-13
Estimated Expiration
2045-03-19

AI Technical Summary

Technical Problem

The existing technology cannot effectively understand the semantics and context of natural language, resulting in the generation of three-dimensional models relying on professional and technical personnel, the model quality and efficiency are not high, the details are missing, the language instruction comprehension ability is insufficient, and it is difficult to dynamically adjust the model.

Method used

By semantic analysis of the natural language descriptions input by the received user, the network generates the initial three-dimensional model based on the customized three-dimensional model data structure and three-dimensional model generation network, and receives the user's model optimization instructions in real time, updates the initial model, and generates a dynamic interactive model.

Benefits of technology

It realizes the generation of high-quality three-dimensional models through simple language descriptions, and supports real-time dynamic adjustment and diversified expression according to user needs, which reduces the threshold for the generation and use of three-dimensional models and improves the efficiency, accuracy, flexibility and ease of use of modeling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147551A_ABST
    Figure CN120147551A_ABST
Patent Text Reader

Abstract

The invention discloses a three-dimensional model intelligent generation method and system based on a natural language, and the method comprises the steps: receiving and describing the natural language inputted by a user in real time, based on the three-dimensional model data structure, the de-noising symbol set and the stop word list, key information data are extracted according to natural language description and mapped into a three-dimensional model parameter set; according to the scene type described by the natural language and the three-dimensional model parameter set, calling a three-dimensional model generation network corresponding to the scene type described by the natural language to generate an initial three-dimensional model; and receiving and analyzing a model optimization instruction input by a user in real time, updating the geometric structure, detail features and physical attributes of the initial three-dimensional model according to an analysis result, generating a dynamic interaction model, and outputting a three-dimensional model file of the dynamic interaction model. By applying the method and the system provided by the invention, the efficiency, the accuracy, the flexibility, the usability and the user experience of three-dimensional model modeling are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of computer graphics image processing and natural language data processing. More specifically, it relates to a method and system for intelligent generation of three-dimensional models based on natural language. Background Art

[0002] Three-dimensional models have extensive applications in fields such as virtual reality (VR), augmented reality (AR), game development, industrial design, architectural design, and medical simulation. However, traditional three-dimensional modeling methods usually rely on professional software, requiring users to have professional modeling skills and consuming a large amount of time. In addition, the means of generating three-dimensional models based on existing three-dimensional scanning or parametric modeling methods also have problems such as complex operations and poor flexibility.

[0003] In recent years, with the rapid development of deep learning and natural language processing (NLP) technologies, it has gradually become possible to generate three-dimensional models through natural language descriptions. However, due to the inability to understand the semantics and context relationships of natural language, the generation of three-dimensional models depends on professional technical personnel, resulting in limitations in the quality, efficiency, and flexibility of natural language understanding in the existing technology. There are problems such as missing details in the generated three-dimensional models, insufficient language instruction understanding ability, and difficulty in dynamically adjusting the characteristics of the models.

[0004] Based on this, it is necessary to introduce a new method and system that can not only generate high-quality three-dimensional models through simple language descriptions, but also support the dynamic adjustment and diversified expression of three-dimensional models according to the personalized needs of users, so as to solve the technical problems in the existing technology, such as the inability to understand the semantics and context relationships of natural language, the dependence of three-dimensional model generation on professional technical personnel, and the missing details of three-dimensional models, insufficient language instruction understanding ability, and difficulty in dynamically adjusting the models, thereby reducing the threshold of three-dimensional model generation and use, and improving the efficiency, accuracy, flexibility, usability, and user experience of three-dimensional model modeling. Summary of the Invention

[0005] In view of the above-mentioned technical problems, the present invention provides a method and system for intelligently generating a three-dimensional model based on natural language. By performing semantic parsing on the natural language description input by the user received, an initial three-dimensional model is generated based on a custom three-dimensional model data structure and a three-dimensional model generation network trained and generated according to different scene types. The initial three-dimensional model is rendered and visually displayed based on an interactive window. At the same time, based on the interactive window, model optimization instructions input by the user are received in real time and used to update the geometric structure, detailed features, and physical properties of the initial three-dimensional model, generating a dynamic interactive model, and outputting a three-dimensional model file of the dynamic interactive model. This not only realizes the generation of high-quality three-dimensional models through simple language descriptions but also realizes the dynamic adjustment and diversified expression of three-dimensional models in real time according to the personalized needs of users, solving the technical problems existing in the prior art, such as the inability to understand the semantics and context relationship of natural language, the dependence of three-dimensional model generation on professional technicians, the lack of details in three-dimensional models, insufficient language instruction understanding ability, and difficulty in dynamically adjusting models, thereby reducing the threshold for the generation and use of three-dimensional models, while improving the efficiency, accuracy, flexibility, usability, and user experience of three-dimensional model modeling.

[0006] The present invention provides a method for intelligently generating a three-dimensional model based on natural language. The method includes: S1, defining a three-dimensional model data structure and a denoising symbol set, and defining and constructing a stop word list and a common sense knowledge base according to the scene type; S2, receiving and parsing in real time the natural language description input by the user, and parsing the natural language description based on the three-dimensional model data structure, the denoising symbol set, and the stop word list, extracting key information data for generating a three-dimensional model according to the parsing result, and mapping the key information data to a three-dimensional model parameter set; S3, training and generating a three-dimensional model generation network according to the scene type and the data structure of the three-dimensional model network, calling the three-dimensional model generation network corresponding to the scene type of the natural language description to generate an initial three-dimensional model according to the scene type of the natural language description and the three-dimensional model parameter set, and performing real-time rendering and visual display on the initial three-dimensional model based on a three-dimensional model rendering engine; S4, based on the initial three-dimensional model, receiving and parsing in real time model optimization instructions input by the user, and updating the geometric structure, detailed features, and physical properties of the initial three-dimensional model according to the parsing result, generating a dynamic interactive model, and generating and outputting a three-dimensional model file of the dynamic interactive model according to the user request;

[0007] Among them, the noise symbol set is used to store punctuation marks in natural language, the stop word list is used to store word segments that make up natural language in the form of strings, the word segments include pronouns, conjunctions, prepositions, numeral-classifiers, modal particles, high-frequency conjunctions, and general words, and the common sense knowledge base is used to store default rules for the spatial position relationships between entities in various scenarios; the three-dimensional model generation network is a network structure of a three-dimensional model that is predefined and trained in advance according to the scenario type, as well as a data storage space and a data transmission channel, which are used to store and transmit key information data of the three-dimensional model; the data structure of the three-dimensional model network includes scenario type, entity object ID, name of the entity object, network level, geometric data, texture data, and material data, and the network level corresponds to the entity relationship of each entity object.

[0008] Preferably, in step S1, the three-dimensional model data structure includes scenario type, entity object, entity relationship, attribute parameters, detailed features, and physical attributes; among them, the entity relationship includes entity level and topological structure; the attribute parameters include spatial position, shape parameters, size parameters, material parameters, color, height, length, width, height, and thickness; the detailed features include texture, material, smoothness, and concavo-convex features; the physical attributes include light reflection characteristics, hardness characteristics, and transparency.

[0009] Preferably, in step S2, the step of parsing the natural language description based on the three-dimensional model data structure, the noise symbol set, and the stop word list, and extracting key information data for generating a three-dimensional model according to the parsing result further includes: S21, obtaining text: parsing the natural language description of the user, determining the type of the natural language description, and based on the three-dimensional model data structure, performing text conversion processing on the natural language description according to the parsing result and the type of the natural language description to obtain the text of the natural language description; S22, text preprocessing: based on the noise symbol set and the stop word list, performing word segmentation and noise removal processing on the text of the natural language description, removing punctuation marks and word segments in the text to obtain preprocessed text; S23, syntax parsing: constructing a syntax tree according to the preprocessed text, and based on the constructed syntax tree and the preprocessed text, determining the entity object, scene backbone, and components of the natural language description, and annotating the part-of-speech of the natural language description; S24, semantic parsing: based on the entity object, scene backbone, and components of the natural language description, as well as the syntax tree and the preprocessed text, performing semantic extraction, identifying and determining the scene type of the natural language description, the attribute parameters of the entity object, and the entity relationship between entity objects, generating the key information data, and mapping it to a three-dimensional model parameter set;

[0010] Among them, the types of the natural language descriptions include text and speech; the text includes words, punctuation marks, conjunctions, and repeated words; the key information data includes: the scene type of the natural language description, entity objects, attribute parameters of the entity objects, and entity relationships between the entity objects.

[0011] Preferably, the step of S21 further includes a step of speech-text conversion and parsing, specifically: if the type of the natural language description is text, then according to the parsing result of the natural language description, obtain the text of the natural language description; if the type of the natural language description is speech, then convert the speech corresponding to the natural language description according to the parsing result of the natural language description to obtain the text of the natural language description.

[0012] Preferably, in step S22, the step of text preprocessing further includes: S22-1, denoising processing: decompose the text of the natural language description into words, phrases, and punctuation marks, and convert them into strings respectively. Traverse each of the converted strings based on the denoising symbol set. If there is a string in the converted strings that matches the punctuation marks in the denoising symbol set, then delete the matching string from the string corresponding to the natural language description to generate a denoised string; S22-2, stop word filtering: traverse the denoised string based on the stop word list. If there is a string in the denoised string that matches the word segmentation in the stop word list, then delete the string that matches the word segmentation in the stop word list from the denoised string to obtain the preprocessed text.

[0013] Preferably, in step S23, the step of syntax parsing further includes: S23-1, constructing a syntax tree: construct a syntax tree according to the preprocessed text, and based on the syntax tree, label the parts of speech of the words and phrases in the preprocessed text to determine the scene backbone of the natural language description and the components corresponding to each scene backbone, where the parts of speech include: core verbs and object nouns; S23-2, entity recognition: recognize the scene type corresponding to the natural language description according to the object nouns in the preprocessed text, and determine entity objects according to the scene backbone of the natural language description and the components corresponding to each scene backbone, and label the parts of speech that make up the natural language description.

[0014] Preferably, in step S24, the semantic parsing step further includes: S24-1, attribute extraction: according to the name of the entity object, extract the attribute data of each entity object corresponding to the natural language description, where the attribute data includes attribute parameters, detailed features, and physical attributes; S24-2, relationship recognition: according to the attribute parameters of each entity object corresponding to the natural language description, determine the entity relationship between each entity object. If there are null values in the attribute parameters of the entity object, then according to the scene type corresponding to the natural language description and the scene type of the common sense knowledge base, call the common sense knowledge base corresponding to the natural language description to determine the entity relationship between each entity object corresponding to the natural language description; S24-3, parameter generation: according to the scene type, entity object, attribute data of the entity object, and the entity relationship between each entity object of the natural language description, generate the key information data, and map the key information data to the three-dimensional model parameter set according to the scene type, the name and attributes of the entity object; where the three-dimensional model parameter set is an intermediate parameter set that automatically converts the key information data of the natural language description into a three-dimensional model. Based on the scene type, the name and attributes of the entity object, the entity relationship and attribute data between each entity object are hierarchically mapped in the order of spatial position, shape, material, and color.

[0015] Preferably, in step S3, the step of training and generating a three-dimensional model generation network according to the scene type and the data structure of the three-dimensional model network further includes: S311, data processing: obtain a data set of natural language descriptions corresponding to each scene type, and based on the obtained data set, perform network standard conversion and feature extraction according to each scene type to obtain three-dimensional model generation network data and network feature data with a unified data format, as well as three-dimensional model network features; S312, network model training: generate an initial generation network model based on the three-dimensional model generation network data and the three-dimensional model network features; S313, network model evaluation: optimize the initial generation network model based on the loss function and the network feature data to generate the three-dimensional model generation network;

[0016] Wherein, the formula of the loss function is:

[0017]

[0018] In the formula, L is the difference between the prediction result and the true result of the generation network model, n is the number of three-dimensional model network features, n is an integer greater than 1, and x i represents the i-th true three-dimensional model network feature, represents the i-th three-dimensional model network feature predicted by the generation network model, represents the L of the vector 2The square of the norm, which is used to calculate the square of the Euclidean distance between the true value and the predicted value of the generation network model. λ is a hyperparameter that controls the regularization strength, and R(θ) is a regularization term with respect to the model parameter θ.

[0019] Preferably, in step S3, the step of calling the three-dimensional model generation network corresponding to the scene type described in the natural language according to the scene type described in the natural language and the three-dimensional model parameter set to generate an initial three-dimensional model further includes: S321, global structure generation: input the three-dimensional model parameter set into the three-dimensional model generation network corresponding to the scene type described in the natural language, and determine the coordinates and geometric shapes of the spatial positions of each entity object according to the entity level of the entity object corresponding to the natural language description, and the entity relationship between entity objects, in the order from high to low entity level, the name and attribute parameters of the entity object, to generate the global geometric structure of the initial three-dimensional model; S322, local detail generation: based on the global structure of the initial three-dimensional model, use the detail features in the three-dimensional model parameter set to add the texture, material, smoothness and concavity and convexity features of each entity object corresponding to the natural language description, to generate the local enhanced geometric structure of the initial three-dimensional model; S322, physical property enhancement: based on the local detail enhanced structure of the initial three-dimensional model, use the physical properties in the three-dimensional model parameter set to add the light reflection characteristics, hardness characteristics and transparency of each entity object corresponding to the natural language description, to generate the initial three-dimensional model.

[0020] Correspondingly, the present invention also provides a three-dimensional model intelligent generation system based on natural language. The system includes an initialization module, a data processing module, an initial model construction module and a dynamic model interaction module;

[0021] Among them, the initialization module is used to define the three-dimensional model data structure and the denoising symbol set, and define and construct the stop word list and the common sense knowledge base according to the scene type; the data processing module is used to receive and parse the natural language description input by the user in real time, and based on the three-dimensional model data structure, the denoising symbol set and the stop word list, parse the natural language description, extract and generate the key information data of the three-dimensional model according to the parsing result, and map the key information data to the three-dimensional model parameter set; the initial model construction module is used to train and generate the three-dimensional model generation network according to the scene type and the data structure of the three-dimensional model network, call the three-dimensional model generation network corresponding to the scene type of the natural language description according to the scene type of the natural language description and the three-dimensional model parameter set to generate the initial three-dimensional model, and perform real-time rendering and visual display on the initial three-dimensional model based on the three-dimensional model rendering engine; the dynamic model interaction module is used to receive and parse the model optimization instruction input by the user in real time based on the initial three-dimensional model, and update the geometric structure, detail features and physical properties of the initial three-dimensional model according to the parsing result to generate a dynamic interaction model, and generate and output the three-dimensional model file of the dynamic interaction model according to the user request.

[0022] Among them, the noise symbol set is used to store punctuation marks in natural language, the stop word list is used to store word segments that make up natural language in the form of strings, the word segments include pronouns, conjunctions, prepositions, numeral classifiers, modal particles, high-frequency conjunctions and general words, and the common sense knowledge base is used to store the default rules of the spatial position relationship between entities in each scene; the three-dimensional model generation network is the network structure, data storage space and data transmission channel of the three-dimensional model that are predefined and trained in advance according to the scene type, and is used to store and transmit the key information data of the three-dimensional model; the data structure of the three-dimensional model network includes scene type, entity object ID, name of the entity object, network level, geometric data, texture data and material data, and the network level corresponds to the entity relationship of each entity object.

[0023] By applying the above technical solutions, the present invention semantically analyzes the natural language description of the received user input, generates an initial 3D model based on a custom 3D model data structure and a 3D model generation network trained according to different scenario types, and renders and visually displays the initial 3D model based on an interactive window. At the same time, based on the interactive window, it receives in real time and updates the geometric structure, detailed features, and physical properties of the initial 3D model according to the model optimization instructions input by the user to generate a dynamic interactive model, and outputs the 3D model file of the dynamic interactive model. It not only realizes the generation of high-quality 3D models through simple language descriptions, but also realizes the dynamic adjustment and diversified expression of 3D models according to the personalized needs of users in real time. It solves the technical problems existing in the prior art, such as the inability to understand the semantics and context relationship of natural language, the dependence of 3D model generation on professional technicians, the lack of 3D model details, the insufficient ability to understand language instructions, and the difficulty in dynamically adjusting the model, thereby reducing the threshold for 3D model generation and use, and improving the efficiency, accuracy, flexibility, usability, and user experience of 3D model modeling at the same time. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0025] Figure 1 FIG. shows a schematic flow chart of a method for intelligent generation of a 3D model based on natural language proposed in an embodiment of the present invention;

[0026] Figure 2 FIG. shows a schematic structural diagram of a system for intelligent generation of a 3D model based on natural language proposed in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0027] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0028] The present invention provides a method for intelligent generation of a 3D model based on natural language, as Figure 1 shown, the method includes the following steps:

[0029] S1. Define the three-dimensional model data structure and the denoising symbol set, and define and construct the stop word list and the common sense knowledge base according to the scene type.

[0030] Among them, the denoising symbol set is used to store punctuation marks in natural language, the stop word list is used to store word segments that make up natural language in the form of strings, the word segments include pronouns, conjunctions, prepositions, numeral-classifiers, modal particles, high-frequency conjunctions and general words, and the common sense knowledge base is used to store the default rules of the spatial position relationship between entities in each scene.

[0031] In this embodiment, in step S1, the three-dimensional model data structure includes scene type, entity object, entity relationship, attribute parameters, detail features, and physical attributes;

[0032] Among them,

[0033] The entity relationship includes entity level and topological structure;

[0034] The attribute parameters include spatial position, shape parameters, size parameters, material parameters, color, height, length, width, height and thickness;

[0035] The detail features include texture, material, smoothness and concavo-convex features;

[0036] The physical attributes include light reflection characteristics, hardness characteristics and transparency.

[0037] S2. Receive and parse the natural language description input by the user in real time, and based on the three-dimensional model data structure, the denoising symbol set and the stop word list, parse the natural language description, extract and generate the key information data of the three-dimensional model according to the parsing result, and map the key information data to the three-dimensional model parameter set.

[0038] In this embodiment, in step S2, the step of parsing the natural language description based on the three-dimensional model data structure, the denoising symbol set and the stop word list, and extracting and generating the key information data of the three-dimensional model according to the parsing result further includes:

[0039] S21. Obtain the text: Parse the natural language description of the user, determine the type of the natural language description, and based on the three-dimensional model data structure, perform text conversion processing on the natural language description according to the parsing result and the type of the natural language description to obtain the text of the natural language description;

[0040] S22. Text preprocessing: Based on the denoising symbol set and the stop word list, perform word segmentation and denoising processing on the text of the natural language description, remove punctuation marks and word segments in the text to obtain the preprocessed text;

[0041] S23, Syntax analysis: Construct a syntax tree based on the preprocessed text, and determine the entity objects, scene backbones, and components described by the natural language based on the constructed syntax tree and the preprocessed text, and label the parts of speech that make up the natural language description;

[0042] S24, Semantic analysis: Based on the entity objects, scene backbones, and components described by the natural language, as well as the syntax tree and the preprocessed text, perform semantic extraction, identify and determine the scene type of the natural language description, the attribute parameters of the entity objects, and the entity relationships between the entity objects, generate the key information data, and map it to a three-dimensional model parameter set;

[0043] Among them,

[0044] The types of the natural language description include text and speech;

[0045] The text includes characters, punctuation marks, conjunctions, and repeated words;

[0046] The key information data includes: the scene type of the natural language description, entity objects, the attribute parameters of the entity objects, and the entity relationships between the entity objects.

[0047] In this embodiment, the steps of S21 further include the steps of speech-to-text conversion and parsing, specifically:

[0048] If the type of the natural language description is text, then according to the parsing result of the natural language description, obtain the text of the natural language description;

[0049] If the type of the natural language description is speech, then convert the speech corresponding to the natural language description according to the parsing result of the natural language description, and obtain the text of the natural language description.

[0050] In this embodiment, in step S22, the steps of text preprocessing further include:

[0051] S22-1, Denoising processing: Decompose the text of the natural language description into words, phrases, and punctuation marks, and convert them into strings respectively. Traverse each converted string based on the denoising symbol set. If there is a string in the converted string that matches the punctuation mark in the denoising symbol set, then delete the matching string from the string corresponding to the natural language description, and generate a denoised string;

[0052] Through this step, the punctuation marks in the text of the natural language description are removed;

[0053] S22-2, Stop word filtering: Based on the stop word list, traverse the denoised string. If there is a string in the denoised string that matches the word segmentation in the stop word list, then delete the string that matches the word segmentation in the stop word list from the denoised string to obtain the preprocessed text.

[0054] Through this step, the word segmentation in the text of the natural language description is removed, including pronouns, conjunctions, prepositions, numeral-classifiers, modal particles, high-frequency conjunctions, and general words.

[0055] In this embodiment, in step S23, the step of syntax parsing further includes:

[0056] S23-1, Construct a syntax tree: Construct a syntax tree according to the preprocessed text, and based on the syntax tree, annotate the part-of-speech of the words and phrases in the preprocessed text to determine the scene backbone of the natural language description and the components corresponding to each scene backbone.

[0057] Among them, the part-of-speech includes: core verbs and object nouns;

[0058] S23-2, Entity recognition: Identify the scene type corresponding to the natural language description according to the object nouns in the preprocessed text, and determine the entity objects according to the scene backbone of the natural language description and the components corresponding to each scene backbone, and annotate the part-of-speech that constitutes the natural language description.

[0059] In this embodiment, in step S24, the step of semantic parsing further includes:

[0060] S24-1, Attribute extraction: Extract the attribute data of each entity object corresponding to the natural language description according to the name of the entity object. The attribute data includes attribute parameters, detailed features, and physical attributes.

[0061] S24-2, Relationship recognition: Determine the entity relationship between each entity object according to the attribute parameters of each entity object corresponding to the natural language description. If there is a null value in the attribute parameters of the entity object, then according to the scene type corresponding to the natural language description and the scene type of the common sense knowledge base, call the common sense knowledge base corresponding to the natural language description to determine the entity relationship between each entity object corresponding to the natural language description.

[0062] S24-3, Parameter generation: Generate the key information data according to the scene type, entity objects, attribute data of the entity objects, and the entity relationship between each entity object, and map the key information data to the three-dimensional model parameter set according to the scene type, name, and attributes of the entity objects.

[0063] Among them,

[0064] The three-dimensional model parameter set is a set of intermediate parameters for automatically converting key information data described in natural language into a three-dimensional model. Based on the scene type, the names and attributes of entity objects, the entity relationships and attribute data between entity objects are hierarchically mapped in the order of spatial position, shape, material, and color.

[0065] It should be noted that this step provided by the present invention can not only parse the natural language description input by the user and extract the key information required for three-dimensional model generation, such as shape, size, material, texture, detailed features, etc.

[0066] It can also support multi-language input and complex description parsing, such as "generate a round dining table with wood grain, about 1 meter in height and 80 centimeters in diameter".

[0067] S3. Train and generate a three-dimensional model generation network according to the scene type and the data structure of the three-dimensional model network. Call the three-dimensional model generation network corresponding to the scene type described in the natural language according to the scene type described in the natural language and the three-dimensional model parameter set to generate an initial three-dimensional model, and perform real-time rendering and visual display on the initial three-dimensional model based on the three-dimensional model rendering engine.

[0068] Among them,

[0069] The three-dimensional model generation network is a network structure of a three-dimensional model that is predefined and trained according to the scene type, as well as a data storage space and a data transmission channel for storing and transmitting key information data of the three-dimensional model;

[0070] The data structure of the three-dimensional model network includes scene type, entity object ID, name of the entity object, network level, geometric data, texture data, and material data, and the network level corresponds to the entity relationship of each entity object.

[0071] In this embodiment, in step S3, the step of training and generating a three-dimensional model generation network according to the scene type and the data structure of the three-dimensional model network further includes:

[0072] S311. Data processing: Obtain a data set of natural language descriptions corresponding to each scene type, and perform network standard conversion and feature extraction according to each scene type based on the obtained data set to obtain three-dimensional model generation network data and network feature data with a unified data format, as well as three-dimensional model network features;

[0073] S312. Network model training: Generate an initial generation network model based on the three-dimensional model generation network data and the three-dimensional model network features;

[0074] S313, Network model evaluation: Optimize the initial generation network model based on the loss function and the network feature data to generate the 3D model generation network;

[0075] Among them, the formula of the loss function is:

[0076]

[0077] In the formula, L is the difference between the prediction result and the real result of the generation network model, n is the number of 3D model network features, n is an integer greater than 1, x i represents the i-th real 3D model network feature, represents the i-th 3D model network feature predicted by the generation network model, represents the square of the L 2 norm of the vector, which is used to calculate the square of the Euclidean distance between the real value and the predicted value of the generation network model. λ is a hyperparameter that controls the regularization strength, and R(θ) is a regularization term regarding the model parameter θ;

[0078] represents the square of the L 2 norm of the vector, which is used to calculate the square of the Euclidean distance between the real value and the predicted value of the generation network model, so as to measure the difference degree between the two.

[0079] In this embodiment, in step S3, the step of calling the 3D model generation network corresponding to the scene type described in the natural language and the 3D model parameter set to generate the initial 3D model further includes:

[0080] S321, Global structure generation: Input the 3D model parameter set into the 3D model generation network corresponding to the scene type described in the natural language. According to the entity level of the entity object corresponding to the natural language description and the entity relationship between entity objects, determine the coordinates of the spatial position and the geometric shape of each entity object in the order from high to low entity level, the name and attribute parameters of the entity object, and generate the global geometric structure of the initial 3D model;

[0081] S322, Local detail generation: Based on the global structure of the initial 3D model, use the detail features in the 3D model parameter set to add textures, materials, smoothness, and bump features to each entity object corresponding to the natural language description, and generate the local enhanced geometric structure of the initial 3D model;

[0082] S322, Physical property enhancement: Based on the local detail enhancement structure of the initial three-dimensional model, add the light reflection characteristics, hardness characteristics, and transparency of the entity objects corresponding to each of the natural language descriptions using the physical properties in the set of three-dimensional model parameters to generate the initial three-dimensional model.

[0083] S4, Based on the initial three-dimensional model, receive and parse the model optimization instructions input by the user in real time, and update the geometric structure, detail features, and physical properties of the initial three-dimensional model according to the parsing results to generate a dynamic interaction model, and generate and output the three-dimensional model file of the dynamic interaction model according to the user request.

[0084] To enable those skilled in the art to more accurately understand the above technical solutions provided by the present invention, the above technical solutions will be further supplemented and explained by way of examples.

[0085] Step 1. The natural language description received from the user is "Generate a living room scene, including a sofa, a coffee table, and a carpet". At this time, start the modeling process and record the requirement description input by the user.

[0086] Step 2. Parse the natural language description, extract the key information related to the generation of the three-dimensional model, and generate key information data;

[0087] The steps of the natural language parsing include:

[0088] 2.1 Perform syntactic parsing on the natural language description and distinguish the input types

[0089] Text parsing: Directly process the text description input by the user, which is "Generate a living room scene, including a sofa, a coffee table, and a carpet".

[0090] Voice parsing: Convert the user's voice into text through speech recognition (ASR) technology and then perform subsequent processing.

[0091] 2.2 Text preprocessing

[0092] Denoising processing: Remove unnecessary symbols or repeated words in the user input

[0093] Word segmentation: Decompose the input text into words or phrases. That is, ["Generate", "a", "living room", "scene", ",", "including", "sofa", "、", "coffee table", "and", "carpet", "."]

[0094] Stop word filtering: Remove irrelevant words. That is, ["Generate", "living room", "scene", "sofa", "coffee table", "carpet"]

[0095] 2.3 Syntactic parsing, extract syntactic structures

[0096] Use a dependency parsing tool (such as SpaCy or Stanford NLP) to construct the syntax tree of the input text, determine the main stem of the sentence, and determine the main stem structure of the sentence.

[0097] The syntax parsing result is:

[0098] Scene main stem: living room;

[0099] Components: sofa, coffee table, carpet.

[0100] 2.4 Perform semantic extraction to identify entities, attributes, and relationships related to 3D modeling;

[0101] Entity recognition: Identify the scene type and objects ("living room", "sofa", "coffee table", "carpet").

[0102] Attribute extraction: If the entity object has specific attribute parameters, extract or default set the attributes of the object.

[0103] Relationship recognition: If the entity relationships between the entity objects have been determined, obtain and determine the relationships between the objects or default to the conventional layout.

[0104] Construct a coordinate system based on the names, attributes, and entity relationships of the entity objects, and determine the specific position coordinates of each entity object in this coordinate system.

[0105] 2.5 Map the semantic information to a set of parameters required for 3D model generation, including shape parameters, size parameters, material parameters, etc.

[0106] Semantic information includes the name of the object, the attribute description of the object, the relationship description between the parts of the object, and the description related to the scene, etc. This information can comprehensively describe the characteristics and details of an object or scene, providing a basis for subsequent 3D model generation.

[0107] The mapping of semantic information is based on the semantic parsing algorithm of natural language processing, combined with the pre-trained model generation network. Through the semantic hierarchical mapping of parameters such as shape and material, the automatic conversion from natural language to 3D modeling parameters is realized.

[0108] The set of parameters required for 3D model generation is an intermediate value. To meet the technical requirements of 3D model generation, the key information related to 3D model generation is the progressive relationship from the original information to the specific data available for model generation, providing a direct input for the generation network.

[0109] That is:

[0110]

[0111]

[0112] Step 3. Invoke the 3D model generation network based on the key information to generate an initial 3D model;

[0113] The 3D model generation network is a generative model based on deep learning, developed by combining the diffusion model and the generative adversarial network (GAN). Its core function is to generate a 3D geometric model consistent with the description from the input semantic parameters. Invoke the pre-trained generation network through the API interface and input the parsed parameter set (such as shape, size, material, etc.). The generation network decodes layer by layer, maps the high-dimensional semantic embedding to the 3D geometric space, and finally outputs the initial model.

[0114] The generation network is based on the diffusion model or the generative adversarial network (GAN), and gradually generates the 3D model through a multi-level generation strategy, including:

[0115] The multi-level generation strategy is an important part of the 3D model generation network. By gradually generating and optimizing, it ensures that the generated 3D model has high quality and high detail level. It is mainly divided into the following steps:

[0116] a. Global geometric structure generation stage

[0117] Establish the basic contour and spatial distribution framework of the 3D model. Generate the global geometric shape based on the input semantic information. Determine the basic parameters (entity name, attributes) and the spatial position of the model.

[0118] The implementation method is to use the diffusion model or the generative adversarial network (GAN) to generate a low-resolution global geometric structure from the high-dimensional semantic embedding.

[0119] Finally, output a low-resolution 3D model that contains the basic shape but lacks details.

[0120] b. Local detail generation stage

[0121] Enrich the surface details of the model and enhance the realism. Add detail features, including texture, bump features, surface smoothness, etc. Refine the material (such as wood grain, metallic luster) and color.

[0122] The implementation method is to use texture mapping and a deep learning-based detail generation module:

[0123] Texture generation: Generate high-resolution surface textures according to the input parameters.

[0124] Local optimization: Add detail features to complex surfaces (such as carving, decoration).

[0125] Finally, output a medium-resolution 3D model that contains rich surface features.

[0126] c. Physical property enhancement stage

[0127] Endow the model with physical properties to enhance its practical usability and realism. Add lighting reflection characteristics, such as metallic luster and transparency.

[0128] The implementation method is to use physically based rendering (PBR) technology and deep learning methods to endow the model with material and lighting characteristics.

[0129] Finally, a high-resolution 3D model with real materials and physical properties is output.

[0130] d. Integrity verification stage

[0131] Ensure that the generated 3D model meets the quality standards in terms of geometric structure and properties. Check geometric integrity. Verify the consistency of detailed characteristics and physical parameters.

[0132] The implementation method is to use automated detection algorithms, geometric repair tools (such as Meshlab) or model verification networks (deep learning).

[0133] Finally, a complete 3D model that passes the quality inspection is output, ready for output or further adjustment.

[0134] Generate the scene structure:

[0135] 3.1 Generate the global geometric structure of the 3D model based on semantic information;

[0136] Global layout generation: Layout the spatial positions of the sofa, coffee table, and carpet.

[0137] 3.2 Add detailed features, including local textures and complex surface properties;

[0138] Sofa details

[0139] Material: Fabric or leather.

[0140] Texture: The fabric can be smooth or have a simple weave pattern; the leather has a smooth surface, possibly with fine pores or faux leather patterns.

[0141] Coffee table details

[0142] Material: Wood or glass.

[0143] Texture: Wood has natural wood grain; glass is completely smooth and may be transparent or translucent.

[0144] Carpet details

[0145] Material: Wool or chemical fiber.

[0146] Texture: Wool carpets are usually fluffy and have a soft touch; chemical fiber carpets are flatter, easier to clean, and may mimic the appearance of other materials.

[0147] 3.3 Verify the integrity of the generated results.

[0148] Ensure that the generated model is complete and consistent in terms of geometric structure and visual performance.

[0149] Generate the initial model kitchen_model = generate_3d_scene(parameters)

[0150] Generate a three-dimensional scene based on the parsed semantic parameters:

[0151] The parameters are:

[0152]

[0153]

[0154]

[0155] Step 4. Post-process and optimize the initial three-dimensional model, and generate a dynamic interactive model. The content of the post-processing optimization includes geometric repair, texture enhancement, and physical property assignment;

[0156] Regarding the technical details in the post-processing optimization, the goal of geometric repair is to ensure that there are no errors in the model during rendering or use.

[0157] Non-manifold geometry repair;

[0158] Repair of duplicate faces and broken faces.

[0159] Different from the global geometric structure in Step 3, the model geometric structure is the core content in the post-processing optimization stage. It focuses on the integrity and topological structure of the model and is the object of quality optimization after generation. It is necessary to ensure the functionality, rendering effect, and geometric accuracy of the generated model in subsequent use.

[0160] The global geometric structure is the basic contour of the three-dimensional model, focusing on the overall shape, size, and spatial distribution. It is the rough stage of generating the model. It mainly provides a framework and reference for adding detailed features and physical properties subsequently.

[0161] The post-processing optimization steps include:

[0162] 1) Perform topological optimization on the model geometric structure, including repairing non-manifold geometry, duplicate faces, and broken faces;

[0163] The main content of topological optimization is the following three

[0164] Repair non-manifold geometry:

[0165] Non-manifold geometry refers to geometric shapes with abnormal topological structures (such as cases where more than two faces share an edge).

[0166] Non-manifold structures may lead to rendering errors or the inability to apply physical simulations.

[0167] Repair duplicate faces:

[0168] Detect and remove redundant patches to avoid redundant calculations and data storage problems in the model.

[0169] Repair broken faces:

[0170] Patch the open areas (such as holes or unclosed edges) in the model to ensure that the geometric structure is completely closed.

[0171] The implementation method mainly uses geometric analysis algorithms to detect non-manifold structures and broken faces, and re-links the patches based on topological automatic repair methods. Use automated software to provide non-manifold geometry detection and repair functions or use dedicated modeling tools to manually repair the broken face structure.

[0172] 2) Optimize the smoothness of the model surface to reduce noise and defects;

[0173] Improve the visual effect and application value of the model by optimizing the surface quality.

[0174] Mainly remove the extra small protrusions or edges generated during the generation process to make the surface smoother. Repair the surface irregular areas caused by generation errors. Ensure that the surface shows continuity visually and tactilely.

[0175] Implementation method:

[0176] Based on the curvature optimization smoothing algorithm, use a deep learning model to avoid defect detection and repair.

[0177] Use 3D modeling tools to smooth the surface automatically or manually.

[0178] 3) Add or adjust the physical properties of the model, including light reflection characteristics, hardness, and transparency.

[0179] Endow the model with real materials and optical properties to make it suitable for rendering, physical simulation, and other practical applications.

[0180] Light reflection characteristics: Simulate the optical effects in the real world (such as specular reflection, diffuse reflection)

[0181] Hardness characteristics: Define the rigidity or flexibility of the material, such as the brittleness of glass or the strength of metal.

[0182] Transparency: Simulate the visual effects of transparent or translucent materials, such as glass and liquid.

[0183] Implementation method:

[0184] Physically Based Rendering (PBR) technology: Define optical properties such as reflectivity, roughness, and metallicity.

[0185] Material mapping: Use texture maps (color maps, normal maps) to define the surface characteristics of materials.

[0186] Transparency adjustment: Set the Alpha value (transparency parameter) of the material to achieve transparent or translucent effects.

[0187] Support users to dynamically adjust 3D models through multiple rounds of language input. The dynamic adjustment includes:

[0188] 1) Receive user input and parse language instructions, and parse them into executable adjustment commands.

[0189] (1) User append description: "Change the carpet to gray"

[0190] (2) User append description: "Add a pot of green plants on the coffee table"

[0191] 2) Adjust the shape, size, color, and texture of the model, and modify the geometric and material properties of the model in real time according to the parsing results.

[0192] "Change the carpet to gray" parses dynamic input and updates carpet properties:

[0193]

[0194]

[0195] 3) Add or delete local structures of the model, support users to dynamically add or delete certain components of the model, and optimize the structure of the model.

[0196] "Add a pot of green plants on the coffee table" The system parses and executes the scene update:

[0197]

[0198]

[0199] 4. Render the model update results in real time. After each adjustment instruction is input by the user, the changes of the model are presented in time to enhance the interaction experience.

[0200] 5. Support multi-round interaction and iterative adjustment, allowing users to make multiple iterative modifications when adjusting the model until a model that meets the requirements is generated.

[0201] Step 5. Output the final 3D model file, supporting the user to select the required file format and resolution.

[0202] The 3D model supports output in multiple file formats, including but not limited to OBJ, STL, and FBX file formats, and supports the following functions:

[0203] 1. User-defined export options, such as whether to include materials and textures:

[0204] Format: OBJ, FBX;

[0205] Whether to include materials: Yes;

[0206] Resolution: Export a high-resolution version for rendering and a low-resolution version for real-time display.

[0207] 2. Generate different resolution versions of the model as needed, including high-resolution and lightweight low-polygon models:

[0208] High-resolution file:

[0209] living_room_scene_with_gray_carpet_and_plant_high_res.obj

[0210] Low-resolution file:

[0211] living_room_scene_with_gray_carpet_and_plant_low_poly.fbx

[0212] 3. Export the scene model:

[0213] export_model(kitchen_model, format = "OBJ", resolution = "high", filename = "living_room_scene_with_gray_carpet_and_plant_high_res.obj")

[0214] export_model(kitchen_model, format = "FBX", resolution = "low", filename = "living_room_scene_with_gray_carpet_and_plant_low_poly.fbx")

[0215] Export the 3D scene model to the specified format and resolution.

[0216]

[0217] Example 1: Generate a static model.

[0218] The user inputs the description "Generate a red cylindrical vase, 30 cm in height and 10 cm in diameter". Parse the natural language description, extract the model features, call the generation network to generate the 3D model of the vase, and after completion, export it in OBJ format for the user to download.

[0219] 1. User input: "Generate a red cylindrical vase, 30 cm in height and 10 cm in diameter"

[0220] 2. Parameters extracted after parsing:

[0221] parameters = {

[0222] "shape": "cylinder",

[0223] "dimensions": {"height": 30, "diameter": 10},

[0224] "material": {"color": "red"}

[0225] }

[0226] 3. Call the generation network:

[0227] model = generate_3d_model(parameters).

[0228] 4. Perform optimization processing:

[0229] optimized_model = optimize_model(model).

[0230] 5. Export the model:

[0231] export_model(optimized_model, format = "OBJ", filename = "red_vase.obj")

[0232] Example 2: Generate a dynamic interactive model.

[0233] The user inputs the description "Generate a blue sports car". After generating the initial model, the user adds the descriptions "Change the wheels to black" and "Add a sunroof to the roof". Update the model in real time and export it in FBX format.

[0234] Step 1: Receive user input

[0235] Initial input: "Generate a blue sports car"

[0236] Record the initial description and start the modeling process.

[0237] Step 2: Parse the natural language description.

[0238] Syntax parsing:

[0239] Subject: sports car;

[0240] Attribute: blue;

[0241] Semantic mapping:

[0242] Shape parameter: {"type": "car"};

[0243] Material parameter: {"color": "blue"};

[0244] Output parameter set:

[0245] parameters = {

[0246] "type": "car",

[0247] "color": "blue"

[0248] }

[0249] Step 3: Call the 3D model generation network.

[0250] Initial generation:

[0251] Call the pre-trained model based on the Generative Adversarial Network (GAN).

[0252] Global characteristics: Construct the geometric structure of the car body (overall appearance, proportion).

[0253] Detail generation: Add local characteristics (wheels, windows).

[0254] Integrity verification: Ensure the geometric consistency of the car body.

[0255] Output the initial model:

[0256] car_model = generate_3d_model(parameters)

[0257] Step 4: Dynamic interaction adjustment.

[0258] User append description: "Change the wheels to black".

[0259] Parse the dynamic input and update the model:

[0260] updated_parameters = {

[0261] "wheels": {"color": "black"}

[0262] }

[0263] car_model = update_3d_model(car_model, updated_parameters)

[0264] User additional description: "Add a sunroof to the car roof"

[0265] Parse and execute model update:

[0266] updated_parameters = {

[0267] "roof": {"feature": "sunroof"}

[0268] }

[0269] car_model = update_3d_model(car_model, updated_parameters)

[0270] Render the dynamic model update result in real time and display it to the user.

[0271] Step 5: Export the final model.

[0272] The user selects the export option:

[0273] Format: FBX;

[0274] Whether to include materials: Yes;

[0275] File name: "blue_car_with_black_wheels_and_sunroof.fbx".

[0276] Export the model:

[0277] export_model(car_model, format = "FBX", filename = "blue_car_with_black_wheels_and_sunroof.fbx").

[0278] Example 3: Modeling of complex scenes.

[0279] The user inputs the description "Generate a living room scene including a sofa, a coffee table and a carpet". The system generates a complete scene according to the description and supports further refinement of the description, such as "Change the carpet to gray" and "Add a potted plant on the coffee table".

[0280] Corresponding to the method for intelligent generation of 3D models based on natural language in an embodiment of the present invention, the present invention also discloses a system for intelligent generation of 3D models based on natural language. As Figure 2 shown, the system includes an initialization module, a data processing module, an initial model construction module, and a dynamic model interaction module;

[0281] Among them,

[0282] The initialization module is used to define the 3D model data structure and the denoising symbol set, and define and construct a stop word list and a common sense knowledge base according to the scene type;

[0283] The data processing module is used to receive and parse the natural language description input by the user in real time, and parse the natural language description based on the 3D model data structure, the denoising symbol set, and the stop word list, extract the key information data for generating the 3D model according to the parsing result, and map the key information data to a set of 3D model parameters;

[0284] It should be noted that the data processing module provided by the present invention can not only parse the natural language description input by the user and extract the key information required for 3D model generation, such as shape, size, material, texture, detail features, etc. It can also support multi-language input and complex description parsing, such as "generate a round dining table with wood grain, about 1 meter in height and 80 centimeters in diameter".

[0285] The initial model construction module is used to train and generate a 3D model generation network according to the scene type and the data structure of the 3D model network, call the 3D model generation network corresponding to the scene type of the natural language description according to the scene type of the natural language description and the set of 3D model parameters to generate an initial 3D model, and perform real-time rendering and visual display on the initial 3D model based on a 3D model rendering engine.

[0286] It should be noted that the initial model construction module provided by the present invention is further used for:

[0287] Based on the parsing result, call a pre-trained 3D model generation network to generate an initial 3D model, where the 3D model generation network is a 3D generator based on a diffusion model or a generative adversarial network.

[0288] Perform multi-level mapping on the input semantics, and generate a 3D model layer by layer from global features (such as shape) to local details (such as texture).

[0289] Global feature extraction: Generate the basic geometric structure and global feature framework of the model according to the input semantic information.

[0290] Extract semantic information related to global shape, size, and proportion.

[0291] Call geometric generation tools to generate a basic framework, such as cylinders and cubes.

[0292] Local detail generation: On the basis of the global framework, add surface textures, detailed structures, and characteristic parameters.

[0293] Extract local features (such as materials and textures) and add them to the model surface.

[0294] Use texture mapping technology to optimize detail features and adjust lighting and material properties.

[0295] Multi-layer mapping integration: Combine global features and local details to generate a complete 3D model.

[0296] Combine global features and local details to generate a complete 3D model.

[0297] Through iterative optimization and consistency verification, ensure that the model conforms to the semantic description.

[0298] The dynamic model interaction module is used to, based on the initial 3D model, receive and parse model optimization instructions input by the user in real time, and update the geometric structure, detail features, and physical properties of the initial 3D model according to the parsing results to generate a dynamic interaction model, and generate and output a 3D model file of the dynamic interaction model according to the user request.

[0299] Among them,

[0300] The noise symbol set is used to store punctuation marks in natural language, the stop word list is used to store word segments that make up natural language in the form of strings, the word segments include pronouns, conjunctions, prepositions, numeral-classifiers, modal particles, high-frequency conjunctions, and general words, and the common sense knowledge base is used to store default rules for the spatial position relationships between entities in various scenarios;

[0301] The 3D model generation network is a network structure of a 3D model that is predefined and trained in advance according to the scene type, as well as a data storage space and a data transmission channel, and is used to store and transmit key information data of the 3D model;

[0302] The data structure of the 3D model network includes scene type, entity object ID, name of the entity object, network level, geometric data, texture data, and material data, and the network level corresponds to the entity relationship of each entity object.

[0303] It should be noted that the dynamic model interaction module is also used to perform real-time adjustment on the initial 3D model or dynamic interaction model. Users can modify the generated 3D model by inputting new natural language descriptions. For example, "Change the color of the table to dark brown" and "Add chair matching". It supports multi-round conversations with users to gradually optimize the model.

[0304] Performing real-time adjustment on the initial 3D model or dynamic interaction model also includes topology structure correction, surface detail enhancement, and model compression. After real-time adjustment, models of different quality levels are output, such as high-resolution and low-polygon models.

[0305] Among them, the goal of model compression is to reduce the file size, improve the rendering efficiency, and maintain the visual quality.

[0306] The main steps are as follows:

[0307] 1) Geometric simplification:

[0308] Polygon reduction: Reduce the model complexity through algorithms such as Edge Collapse and Vertex Clustering.

[0309] LOD generation: Create model versions with different resolutions to adapt to different viewing distances and device performance requirements.

[0310] 2) Texture optimization and compression:

[0311] Compress texture files using formats such as ASTC to significantly reduce the storage volume.

[0312] Generate normal maps to retain high-resolution details while reducing the number of polygons.

[0313] 3) Material and shader optimization:

[0314] Merge material instances to reduce rendering state switching.

[0315] Simplify shader logic to improve the rendering efficiency on low-performance devices.

[0316] 4) Remove redundant information:

[0317] Delete unused materials, textures, and hidden faces.

[0318] Simplify animation curves to reduce the keyframe data volume.

[0319] 5) File format conversion and optimization:

[0320] Convert to GLTF / GLB or FBX format to adapt to different platform requirements.

[0321] Reasonably select the management method of embedded resources and external links.

[0322] By applying the above technical solutions, through semantic parsing of the natural language description of the received user input, an initial 3D model is generated based on a custom 3D model data structure and a 3D model generation network trained according to different scenario types, and the initial 3D model is rendered and visually displayed based on an interactive window. At the same time, based on the interactive window, model optimization instructions input by the user are received in real time and used to update the geometric structure, detailed features, and physical properties of the initial 3D model, generating a dynamic interactive model and outputting a 3D model file of the dynamic interactive model. This not only realizes the generation of high-quality 3D models through simple language descriptions but also realizes the dynamic adjustment and diverse expression of 3D models according to the personalized needs of users in real time, solving the technical problems existing in the prior art, such as the inability to understand the semantics and context relationship of natural language, the dependence of 3D model generation on professional technicians, and the lack of details in 3D models, insufficient language instruction understanding ability, and difficulty in dynamically adjusting models, thereby reducing the threshold for 3D model generation and use while improving the efficiency, accuracy, flexibility, usability, and user experience of 3D model modeling.

[0323] In addition, by parsing the basic description information in the user input text, basic features can be identified and the underlying semantic meaning and context relationship can be further analyzed. The method provided by the present invention can more accurately convert the user's intention into specific 3D geometric features.

[0324] Users are allowed to provide feedback in real time during the generation process. Based on this feedback, the generation strategy is automatically adjusted to optimize the output result. As the number of uses increases, it will accumulate more data for improving its own performance. This mechanism enhances the self-adaptability and personalized service ability of the system.

[0325] By using a wide range of data resources for training, the model can improve the quality and diversity of the generation results while ensuring high efficiency. The training data covers various types of object descriptions, which helps the system to stably output high-quality 3D models when dealing with complex or rare object descriptions.

[0326] Traditional 3D modeling usually requires multiple steps, including manually drawing sketches and setting parameters, etc. In contrast, the method and system provided by the present invention allow users to quickly obtain the corresponding 3D model by simply describing the object needed in natural language, simplifying the modeling process and improving the efficiency.

[0327] Considering the differences in the needs of different user groups, especially that non-professional users may lack the necessary modeling knowledge and technical background, the method and system provided by the present invention pay special attention to the design of user experience. Users only need to express their ideas in natural language form, without any professional knowledge, and can easily create 3D models, improving the ease of use.

[0328] The method and system provided by the present invention support a multi-round dialogue mode. After the initial generation, users can further refine the requirements according to the actual situation, such as adjusting the proportion of specific parts, changing the material texture, etc. This flexibility ensures that even application scenarios with strict requirements for details can be met, providing flexible and dynamic adjustment capabilities.

[0329] When the method and system provided by the present invention were designed, cross-domain compatibility and portability were taken into account. Whether it is in the fields of game development, film and television production, education and scientific research, etc., as long as 3D content creation is involved, this technology can be seamlessly docked. This provides a convenient 3D content generation solution for various industries, with wide applicability and good scalability.

[0330] Each embodiment in this specification is described in a related manner. For the same or similar parts between the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments.

[0331] The above are only the preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention are included in the protection scope of the present invention.

Claims

1. A method for intelligently generating a three-dimensional model based on natural language, characterized in that: The method comprises: S1, define the 3D model data structure and the denoising symbol set, define and build the stop word list and common sense knowledge base according to the scene type; S2, receiving in real time and analyzing the natural language description input by the user, and based on the 3D model data structure, the denoising symbol set and the stop word list, extracting and generating key information data of the 3D model according to the analysis result, and mapping the key information data into a 3D model parameter set; S3, training and generating a 3D model generation network according to the scene type and the data structure of the 3D model network, calling the 3D model generation network corresponding to the scene type described in the natural language according to the scene type described in the natural language and the 3D model parameter set to generate an initial 3D model, and performing real-time rendering and visual display of the initial 3D model based on a 3D model rendering engine; S4, based on the initial three-dimensional model, receiving and parsing the model optimization instructions input by the user in real time, and updating the geometric structure, detail features and physical properties of the initial three-dimensional model according to the parsing results, generating a dynamic interactive model, and generating and outputting a three-dimensional model file of the dynamic interactive model according to the user's request; in, The noise symbol set is used to store punctuation marks in natural language, the stop word list is used to store the segmented words constituting natural language in the form of strings, and the segmented words include pronouns, conjunctions, prepositions, quantifiers, modal particles, high-frequency conjunctions and common words, and the common sense knowledge base is used to store the default rules of the spatial position relationship between entities in various scenes; The three-dimensional model generation network is a network structure of a three-dimensional model that is pre-defined and trained according to the scene type, as well as a data storage space and a data transmission channel, which are used to store and transmit key information data of the three-dimensional model; The data structure of the three-dimensional model network includes scene type, entity object ID, name of entity object, network level, geometric data, texture data and material data, and the network level corresponds to the entity relationship of each entity object.

2. The method according to claim 1, characterized in that In step S1, the three-dimensional model data structure includes scene type, entity object, entity relationship, attribute parameter, detail feature, and physical attribute; in, The entity relationship includes entity level and topology; The attribute parameters include spatial position, shape parameters, size parameters, material parameters, color, height, length, width, height and thickness; The detailed features include texture, material, smoothness and concave-convex features; The physical properties include light reflection properties, hardness properties and transparency.

3. The method according to claim 1, characterized in that In step S2, the step of parsing the natural language description based on the three-dimensional model data structure, the denoising symbol set and the stop word list, and extracting and generating key information data of the three-dimensional model according to the parsing result further includes: S21, obtaining text: parsing the user's natural language description to determine the type of the natural language description, and based on the three-dimensional model data structure, performing text conversion processing on the natural language description according to the parsing result and the type of the natural language description to obtain the text of the natural language description; S22, text preprocessing: based on the denoising symbol set and the stop word list, performing word segmentation and denoising processing on the text described in the natural language, removing punctuation marks and word segmentation in the text, and obtaining a preprocessed text; S23, syntax analysis: constructing a syntax tree according to the preprocessed text, and determining the entity objects, scene trunks and components of the natural language description based on the constructed syntax tree and the preprocessed text, and marking the parts of speech constituting the natural language description; S24, semantic analysis: performing semantic extraction based on the entity objects, scene trunks and components described in the natural language, the syntax tree and the preprocessed text, identifying and determining the scene type described in the natural language, the attribute parameters of the entity objects, and the entity relationship between the entity objects, generating the key information data, and mapping it into a three-dimensional model parameter set; in, The type of the natural language description, including text and voice; The text includes words, punctuation marks, conjunctions and repeated words; The key information data includes: the scene type, entity objects, attribute parameters of entity objects, and entity relationships between entity objects described in the natural language.

4. The method according to claim 3, characterized in that The step S21 also includes the step of voice-to-text conversion and analysis, which is specifically: If the type of the natural language description is text, obtaining the text of the natural language description according to the parsing result of the natural language description; If the type of the natural language description is speech, the speech corresponding to the natural language description is converted according to the parsing result of the natural language description to obtain the text of the natural language description.

5. The method according to claim 3, characterized in that In step S22, the text preprocessing step further includes: S22-1, denoising processing: decomposing the text of the natural language description into words, phrases, and punctuation marks, and converting them into character strings respectively, traversing each converted character string based on the denoising symbol set, and if there is a character string matching the punctuation mark in the denoising symbol set in the converted character string, deleting the matching character string from the character string corresponding to the natural language description, and generating a denoised character string; S22-2, stop word filtering: based on the stop word list, traverse the denoised character string, if there is a character string in the denoised character string that matches the word in the stop word list, delete the character string that matches the word in the stop word list from the denoised character string to obtain the preprocessed text.

6. The method according to claim 3, characterized in that In step S23, the syntax parsing step further includes: S23-1, constructing a syntax tree: constructing a syntax tree according to the preprocessed text, and marking the parts of speech of words and phrases in the preprocessed text based on the syntax tree, determining the scene trunk described in the natural language, and the components corresponding to each scene trunk, Wherein, the parts of speech include: core verbs and object nouns; S23-2, entity recognition: identifying the scene type corresponding to the natural language description based on the object nouns in the preprocessed text, and determining the entity object based on the scene trunk of the natural language description and the components corresponding to each scene trunk, and marking the parts of speech that constitute the natural language description.

7. The method according to claim 3, characterized in that In step S24, the semantic analysis step further includes: S24-1, attribute extraction: extracting attribute data of each entity object corresponding to the natural language description according to the name of the entity object, the attribute data including attribute parameters, detailed features and physical properties; S24-2, relationship identification: determining the entity relationship between the entity objects according to the attribute parameters of the entity objects corresponding to the natural language description, and if there is a null value in the attribute parameters of the entity objects, calling the common sense knowledge base corresponding to the natural language description according to the scene type corresponding to the natural language description and the scene type of the common sense knowledge base, and determining the entity relationship between the entity objects corresponding to the natural language description; S24-3, parameter generation: generating the key information data according to the scene type, entity object, attribute data of the entity object, and entity relationship between the entity objects described in the natural language, and mapping the key information data into the three-dimensional model parameter set according to the scene type, name and attribute of the entity object; in, The three-dimensional model parameter set is an intermediate parameter set for automatically converting key information data described in natural language into a three-dimensional model. Based on the scene type, the name and attributes of the entity object, the entity relationship and attribute data between each entity object are hierarchically mapped in the order of spatial position, shape, material and color.

8. The method according to claim 1, characterized in that In step S3, the step of training and generating a 3D model generation network according to the scene type and the data structure of the 3D model network further includes: S311, data processing: obtaining a data set of natural language descriptions corresponding to each scene type, and performing network standard conversion and feature extraction according to each scene type based on the obtained data set to obtain three-dimensional model generation network data and network feature data in a unified data format, as well as three-dimensional model network features; S312, network model training: generating an initial generation network model based on the three-dimensional model generation network data and the three-dimensional model network features; S313, network model evaluation: optimizing the initial generation network model based on the loss function and the network feature data to generate the three-dimensional model generation network; Among them, the formula of the loss function is: Where L is the difference between the predicted result of the generated network model and the actual result, n is the number of network features of the 3D model, n is an integer greater than 1, and x i represents the i-th real 3D model network feature, represents the network features of the i-th 3D model predicted by the generated network model, It represents the square of the L2 norm of the vector, which is used to calculate the square of the Euclidean distance between the true value and the predicted value of the generative network model. λ is a hyperparameter that controls the strength of regularization. R(θ) is the regularization term about the model parameter θ.

9. The method according to claim 1, characterized in that In step S3, the step of calling a 3D model generation network corresponding to the scene type described in the natural language and the 3D model parameter set to generate an initial 3D model further includes: S321, global structure generation: input the 3D model parameter set into the 3D model generation network corresponding to the scene type described in the natural language, determine the coordinates and geometric shapes of the spatial positions of the various entity objects according to the entity levels of the entity objects corresponding to the natural language description and the entity relationships between the entity objects, and generate the global geometric structure of the initial 3D model in accordance with the entity levels from high to low, the names of the entity objects and the attribute parameters; S322, local detail generation: based on the global structure of the initial three-dimensional model, using the detail features in the three-dimensional model parameter set to add the texture, material, smoothness and concave-convex features of the entity objects corresponding to the natural language descriptions, to generate a local enhanced geometric structure of the initial three-dimensional model; S322, physical property enhancement: Based on the local detail enhancement structure of the initial three-dimensional model, the physical properties in the three-dimensional model parameter set are used to add the lighting reflection characteristics, hardness characteristics and transparency of each entity object corresponding to the natural language description to generate an initial three-dimensional model.

10. A system for implementing the natural language-based intelligent generation method of three-dimensional models according to claim 1, characterized in that: The system includes an initialization module, a data processing module, an initial model building module and a dynamic model interaction module; in, The initialization module is used to define the 3D model data structure and the denoising symbol set, and to define and construct a stop word list and a common sense knowledge base according to the scene type; The data processing module is used to receive and process the natural language description input by the user in real time, and parse the natural language description based on the three-dimensional model data structure, the denoising symbol set and the stop word list, extract and generate key information data of the three-dimensional model according to the parsing result, and map the key information data into a three-dimensional model parameter set; The initial model construction module is used to train and generate a 3D model generation network according to the scene type and the data structure of the 3D model network, call the 3D model generation network corresponding to the scene type described in the natural language according to the scene type described in the natural language and the 3D model parameter set to generate an initial 3D model, and perform real-time rendering and visual display of the initial 3D model based on the 3D model rendering engine; The dynamic model interaction module is used to receive and analyze the model optimization instructions input by the user in real time based on the initial three-dimensional model, and update the geometric structure, detail features and physical properties of the initial three-dimensional model according to the analysis results to generate a dynamic interaction model, and generate and output the three-dimensional model file of the dynamic interaction model according to the user request; in, The noise symbol set is used to store punctuation marks in natural language, the stop word list is used to store the segmented words constituting natural language in the form of strings, and the segmented words include pronouns, conjunctions, prepositions, quantifiers, modal particles, high-frequency conjunctions and common words, and the common sense knowledge base is used to store the default rules of the spatial position relationship between entities in various scenes; The three-dimensional model generation network is a network structure of a three-dimensional model that is pre-defined and trained according to the scene type, as well as a data storage space and a data transmission channel, which are used to store and transmit key information data of the three-dimensional model; The data structure of the three-dimensional model network includes scene type, entity object ID, name of entity object, network level, geometric data, texture data and material data, and the network level corresponds to the entity relationship of each entity object.

Citation Information

Patent Citations

  • Large language model training method and device, training data construction method and device, equipment and medium

    CN118014011A

  • Translating method for translating a natural-language description into a computer-language description

    US20150242396A1

  • Generating three-dimensional digital content from natural language requests

    US20200151277A1

Cited By

  • 3D modeling feature extraction method and system combined with natural language processing

    CN120451369A

  • 3D modeling feature extraction method and system combined with natural language processing

    CN120451369B

  • Text-driven 3D model generation method and system for game development

    CN120526060A

  • Text-driven 3D model generation method and system for game development

    CN120526060B

  • Functional scene modeling and large model cross-dimension content generation method and system

    CN120676222A