A 3D design assisting method based on multi-modal input and intelligent optimization
Through multimodal input and intelligent optimization methods, combined with text and image data, using GPT-4O, CLIP-ViT, BLIP-2 and Point-E models, a knowledge graph is constructed to solve the accuracy and efficiency problems of 3D design under single modal input, and achieve efficient and accurate 3D model generation and optimization.
Patent Information
- Application Number
- CN202411913269.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2044-12-24
AI Technical Summary
Existing 3D design methods rely on single modal input, which makes it difficult to accurately express complex design requirements. Model evaluation standards lack consistency, resulting in insufficient quality and efficiency of generated models.
Adopting multimodal input and intelligent optimization methods, by combining text and image data, using GPT-4O, CLIP-ViT, BLIP-2 and Point-E models for feature extraction and fusion, building a knowledge graph for model evaluation, and generating and optimizing 3D models.
It significantly improves the flexibility and accuracy of 3D model generation, enhances adaptability to complex design requirements, and improves generation efficiency and the comprehensive performance of the model.
Smart Images

Figure CN119832157B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of 3D design assistance technology, and in particular relates to a 3D design assistance method based on multimodal input and intelligent optimization. Background Art
[0002] In current 3D design, how to efficiently and accurately generate models that meet user needs is still an urgent problem to be solved. Traditional methods mostly rely on single modal input (such as text or images), which has limitations when dealing with complex design requirements: text input is difficult to accurately express design intent, and image input is limited by conditions such as quality, angle and lighting, resulting in unstable recognition results. In addition, the model evaluation standards lack consistency, making it difficult to ensure the quality of the generated model and design efficiency. The present invention combines multimodal input of text and images, adopts multimodal learning technology and knowledge graphs, and can significantly improve the flexibility and accuracy of 3D model generation, and optimize the intelligence and efficiency of the model screening process. Summary of the Invention
[0003] The purpose of the present invention is to provide a 3D design assistance method based on multimodal input and intelligent optimization to solve the problems existing in the above-mentioned prior art.
[0004] To achieve the above objectives, the present invention provides a 3D design assistance method based on multimodal input and intelligent optimization, comprising:
[0005] Obtaining text data and image data to be processed;
[0006] Performing feature extraction on the text data to be processed and the image data to be processed respectively to obtain text features and image features; the text features include size description, functional description, material description and structure description; the image features include image structure features;
[0007] Performing feature fusion on the text features and the image features to obtain fused features;
[0008] constructing a plurality of 3D models based on the fused features;
[0009] A knowledge graph is constructed based on preset design rules and the text features, each of the 3D models is evaluated based on the knowledge graph, and an optimal 3D model is determined based on the evaluation results.
[0010] Optionally, the process of acquiring the text features specifically includes:
[0011] The text data to be processed is parsed based on the GPT-4O model to obtain the text features.
[0012] Optionally, the image structure features include edge features, texture features and shape features of objects in the image data to be processed.
[0013] Optionally, the process of obtaining the image structure features specifically includes:
[0014] inputting the image data to be processed into the CLIP-ViT model, extracting the contours and boundaries of objects in the image data to be processed, and obtaining the edge features;
[0015] encoding local pixel differences of the image data to be processed according to a local binary pattern algorithm, and obtaining the texture features;
[0016] performing frequency domain representation on the contours and boundaries of objects based on Fourier transform, and obtaining the shape features.
[0017] Optionally, the process of fusing the text features and the image features specifically includes:
[0018] projecting the image features and the text features into a unified dimensional space through an adaptive matrix, and fusing the image features and the text features based on a multi-head attention mechanism to obtain the fused features.
[0019] Optionally, the process of constructing a plurality of 3D models based on the fused features specifically includes:
[0020] inputting the fused features into a Point-E model, generating a plurality of initial models through point cloud technology, and rendering and optimizing each initial model to obtain a plurality of 3D models.
[0021] Optionally, the process of constructing a knowledge graph based on the preset design rules and the text features specifically includes:
[0022] converting preset industry basic design rules, size descriptions and the functional descriptions into screening rules, and constructing a knowledge graph based on the screening rules.
[0023] Optionally, the process of evaluating each 3D model based on the knowledge graph specifically includes:
[0024] scoring each 3D model based on the knowledge graph, and selecting a 3D model with the highest score as an optimal 3D model.
[0025] The technical effects of the present application are:
[0026] The method provided by this invention, through the combination of multimodal learning technology, pre-trained models, and knowledge graphs, can effectively fuse text and image inputs. Through feature fusion, it improves adaptability to complex design requirements and enhances the model's comprehensive performance capabilities. Constructing a knowledge graph that supports dynamic updates enables the system to flexibly integrate the design rules and functional requirements required by users, guiding the model screening process. Combining these advantages, this method significantly improves the efficiency and accuracy of 3D model generation, providing powerful intelligent support for design innovation. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0028] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of this application. The exemplary embodiments and descriptions of this application are intended to explain this application and do not constitute an improper limitation on this application. In the accompanying drawings:
[0029] Figure 1 This is a flow chart of a 3D design assistance method based on multimodal input and intelligent optimization in an embodiment of the present invention;
[0030] Figure 2 Schematic diagram of a 3D model generation process driven by multimodal input in an embodiment of the present invention;
[0031] Figure 3 Schematic diagram of the fusion of image and text feature vectors during the feature fusion process in an embodiment of the present invention;
[0032] Figure 4 This is a schematic diagram of the embodiment of the present invention in which GPT-4O and knowledge graph jointly assist in model screening. DETAILED DESCRIPTION
[0033] Various exemplary embodiments of the present invention will now be described in detail. This detailed description should not be considered as limiting the present invention, but rather as a more detailed description of certain aspects, features, and embodiments of the present invention.
[0034] It is to be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the present application. Additionally, for a range of values of, for example, concentrations, solvent amounts, and the like, there can be endpoints other than the absolute terminals. These endpoints are provided as a convenience for the reader, but are not intended to limit the scope of the present application. The disclosure of any endpoints start and end points are meant to be interchangeable to encompass ranges including the recited endpoints.
[0035] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. Although methods similar or equivalent to those described herein can be used in the practice or testing of the present application, the preferred methods are described. All publications mentioned in this specification are herein incorporated by reference to disclose and describe the methods and / or materials in connection with which the publications are cited. The citation of any reference is not construed as an admission that it is prior art with respect to the present application.
[0036] Many modifications and variations of this application can be made in the light of the above teachings without departing from the spirit and scope thereof. Additional implementations of the application will be apparent to those skilled in the art from consideration of the specification and practice of the application disclosed herein. The specification and examples given are exemplary only and are not intended to be limiting.
[0037] As used herein, the terms "comprises", "comprising", "includes", "including", "has", "having" and the like are open-ended terms that are intended to permit but not to limit the description and / or claims hereafter.
[0038] It should be noted that the embodiments and features of the embodiments in the present application can be combined with each other on the premise of no conflict. The present application will be described in detail below with reference to the accompanying drawings and in conjunction with the embodiments.
[0039] Embodiment One
[0040] As shown in the Figure 1 - Figure 4 The embodiment provides a 3D design assistance method based on multi-modal input and intelligent optimization, which comprises the following steps: obtaining to-be-processed text data and to-be-processed image data; performing feature extraction on the to-be-processed text data and the to-be-processed image data respectively to obtain text features and image features; the text features comprise size description, functional description, material description and structure description; the image features comprise image structure features; performing feature fusion on the text features and the image features to obtain fused features; constructing a plurality of 3D models based on the fused features; constructing a knowledge graph based on preset design rules and the text features, evaluating each 3D model based on the knowledge graph, and determining an optimal 3D model based on an evaluation result.
[0041] This embodiment discloses a 3D design assistance method based on multimodal input and intelligent optimization, including the following steps: S100: multimodal input processing; parse text descriptions through pre-trained GPT-4O to extract design intent and functional requirements, and use CLIP-ViT to extract image structural features. S200: multimodal fusion; use BLIP-2 to fuse text and image features to generate a unified feature vector as 3D generation input. S300: preliminary 3D model generation; use Point-E and other models to generate 10 preliminary 3D models, and use Blender for preliminary rendering to ensure that they meet the design intent. S400: requirement analysis and model screening; use GPT-4O to parse user descriptions, convert semantics into screening rules, construct a knowledge graph to assist screening, use semantic similarity analysis to evaluate the match between the model and requirements, score according to design rules and screen the model with the highest score. This embodiment improves the efficiency and accuracy of 3D design and supports design innovation through multimodal input fusion and intelligent optimization.
[0042] This embodiment provides a 3D design assistance method and system based on multimodal input and intelligent optimization. The method integrates text and image input and utilizes advanced multimodal learning technology to achieve efficient 3D model generation and screening. The method specifically includes the following steps:
[0043] like Figure 1 As shown, this embodiment provides a 3D design assistance method based on multimodal input and intelligent optimization. This method combines image and text input, uses multimodal fusion technology and intelligent generation algorithms to achieve efficient 3D model generation and optimization. The specific steps are as follows:
[0044] Step S100: Multimodal input processing: By processing multimodal input image and text data, the system can simultaneously obtain visual features and language descriptions. Image input is processed by the CLIP-ViT model to extract structural features from the image, while text input is parsed by GPT-4O to obtain design intent and functional requirements. Specifically, the following steps are included: S101-S102:
[0045] Step S101: Image input; use the CLIP-ViT model to extract features from the input image. The main goal is to obtain geometric features and structural information from the image for subsequent multimodal fusion and 3D model generation. Image feature extraction specifically includes the following aspects:
[0046] S101-1: Edge features: The CLIP-ViT model first analyzes the edge information in the image and extracts the outline and boundary of the object through convolution operation. These edge features are represented by vectors E(e1,e2,...,e n ), where e nRepresents the edge strength value at the nth pixel. Edge features are the core information describing the geometric shape of an object.
[0047] S101-2: Texture features: In addition to edge information, CLIP-ViT also supports the extraction of image surface texture features. The local binary pattern (LBP) algorithm is used to encode the local pixel differences of the image and generate a texture vector T (t1, t2, ..., t m ), where t m Indicates the texture mode within the local window.
[0048] S101-3: Shape features: By further analyzing the edges and contours of the image, the overall geometric shape features of the object are extracted. The object contour is represented in the frequency domain using Fourier transform to obtain the shape descriptor S(s1,s2,...,s i ), where s i Frequency components that represent the outline of an object.
[0049] Furthermore, the image features extracted by CLIP-ViT can be expressed as:
[0050] G={E,T,S}
[0051] Step S102: Text input: GPT-4O parses the user-provided text input, primarily extracting functional requirements and design constraints related to 3D model generation from the design description. The specific text feature extraction process is as follows:
[0052] S102-1: Size description: GPT-4O parses size-related information in the text, such as length, width, height, or diameter. By performing word segmentation and entity recognition on the sentence containing the size description, the size value D (d1, d2, ..., d j ) extracted, d j Represents the j-th size parameter.
[0053] S102-2: Functional description: Identify the functional requirements of the model from the text, such as the purpose of the design object or its operating characteristics in the actual scenario. The functional requirements are generated through GPT-4O to generate a function vector F(f1,f2,...,f k ), f k Indicates a specific functional requirement, such as strength, bearing capacity, etc.
[0054] S102-3: Material and structure description: Further analyze the material selection and structure requirements in the text. Use word vectorization technology to convert the material properties in the text (such as steel, aluminum, etc.) into vectors M (M1, M2, ..., M l ), and generate the constraint vector C(C1,C2,...,Cq ).
[0055] Furthermore, the text features are expressed as:
[0056] T={D,F,M,C}
[0057] The above features are used together with image features for multimodal fusion to ensure that the generated 3D model meets the user's design requirements.
[0058] Step S200: Multimodal fusion; Figure 2 As shown in Figure 1, the BLIP-2 model is used to fuse image features and text features to generate a unified feature vector. The specific fusion steps include:
[0059] S200-1: Feature Adaptation: First, the image feature G and the text feature T are dimensionally matched to ensure that the two feature vectors are aligned in the same dimensional space. Assume that the image feature dimension is n and the text feature dimension is m. The feature dimension conversion is completed by the adaptation matrix A, which is expressed as:
[0060] R=A·G+A·T
[0061] Where A is the adaptation matrix, which is used to project image and text features into a unified dimensional space.
[0062] S200-2: Attention mechanism fusion: image features and text features are fused through the multi-head attention mechanism. Let the attention weight be α i , the fused feature vector R is expressed as:
[0063]
[0064] where α i is the attention weight calculated based on the feature importance, G i and T i are the i-th components of image and text features respectively.
[0065] S200-3: Fusion feature output: Further, the generated fusion vector R (r1, r2, ..., r s ) represents a unified multimodal feature. This feature will be used as input to generate a 3D model in subsequent steps.
[0066] Step S300: Preliminary 3D model generation: After multimodal fusion is complete, the system uses the Point-E model to generate multiple preliminary 3D models. These models are generated based on the fused unified feature vector. The Point-E model is particularly good at generating high-quality point cloud models and is suitable for representing complex geometric structures. Specifically, the following steps are included: S301-S302:
[0067] Step S301: Generate a 3D model using Point-E; Figure 3 As shown, the fused feature vector $R$ is input to the Point-E model, which generates multiple preliminary 3D models through point cloud technology. Each 3D model is represented as a set of points:
[0068] R=A·G+A·TP={(x 1, y1,z1),(x 12 y2,z2),...,(x n, y n ,z n )}
[0069] Where (x n, y n ,z n ) is the coordinate of each point in the point cloud, representing the geometric structure of the model.
[0070] During the model generation process, attention is paid to details and structural integrity to ensure that the generated 3D model meets the user's design requirements.
[0071] Step S302: Use Blender to render the 3D model. After generating the preliminary point cloud model, use Blender to further render and optimize the model. Rendering includes setting lighting effects, applying material mapping, and adjusting the camera perspective to ensure that the generated model visually meets the design requirements. The rendering result is:
[0072] M=R(P)+Material Map+Lighting
[0073] Among them, R(P) is the rendered 3D model, which combines lighting and material information.
[0074] Step S400: User requirements analysis and model screening: The system uses GPT-4O to analyze the user's design requirements, converts them into screening rules, and intelligently screens the generated 3D models based on the design rules in the knowledge graph. This includes the following steps S401-S403:
[0075] Step S401: GPT-4O analyzes user requirements. The system analyzes the functional requirements and geometric constraints entered by the user and converts them into actionable screening rules. These rules include wing load capacity, fuselage streamlined design, material requirements, etc.
[0076] Step S402: Construct a knowledge graph; Figure 4 As shown in the figure, a knowledge graph containing basic design rules, functional requirements, and geometric constraints is constructed. This graph is used to assist in matching the generated model with user requirements during the model screening process. The specific process is as follows:
[0077] S402-1: Basic Design Rules: Extract basic design rules from industry standards and existing design experience, covering the physical parameters, material specifications, and basic functional requirements required by the model. These rules serve as nodes in the graph to verify the basic compliance of the generated model.
[0078] S402-2: Functional Requirements Matching: Match the functional requirements entered by the user with the historical design experience and rules in the knowledge graph. For example, if the user wants the generated model to have the function of "wind resistance", the knowledge graph will return models with corresponding functional designs for comparison.
[0079] Expressed as: Function matching matrix {M f1 ,M f2 ,...,M fn}, the matrix compares the functional characteristics of each model with user requirements to ensure that the design goals are met.
[0080] S402-3: Geometric constraint verification: Through the geometric constraint nodes, the knowledge graph can verify whether the generated model meets the geometric constraints such as size and shape. It is represented as: Geometric constraint node network {G n1 ,G n2 ,...,G nm}, used to determine whether the generated model meets the user's geometric requirements.
[0081] Step S403: Model scoring and selection; Figure 4 As shown, according to the above rules, the generated 3D models are scored based on the structural integrity, functional compatibility and innovation. The system automatically evaluates each model and selects the model with the highest score as the further recommended design solution. The specific scoring steps are as follows:
[0082] S403-1: Structural Integrity: First, the generated 3D model is scored for its geometric structural integrity, primarily to assess whether the model complies with the user's geometric constraints and ensure that all geometric features of the model are complete.
[0083] Scoring rules: The geometric deviation calculation formula ΔG = |actual geometric value - ideal geometric value|, and the scoring standard is 0-10 points.
[0084] S403-2: Functional matching: By matching the user's functional requirements (such as strength, durability, etc.), the system scores the functional characteristics of each model. The score is based on the functional matching matrix in the knowledge graph.
[0085]
[0086] Scoring rules: Function matching scoring formula F s=∑User demand characteristics*Model characteristics, ensuring the adaptability of each model to functional requirements.
[0087] S403-3: Innovation: Evaluate the innovative design of the model. The system will score the model's innovation in function, form or material.
[0088] Scoring rules: The innovation score is based on the difference between the historical design and the existing model in the knowledge graph. The greater the difference, the higher the score. The standard is 0-10 points.
[0089] Furthermore, the system will give an overall score to all generated models based on the above scoring criteria, and recommend the model with the highest score to the user as a further design solution.
[0090] The following is an embodiment of this embodiment (combined with Figure 3 and Figure 4 ).
[0091] Step S101: Image input; In this embodiment, the user first provides an image of the aircraft's appearance. The system uses the CLIP-ViT model to extract key geometric features from the image. These features include:
[0092] S101-1: Geometry: Contour information of the wing, fuselage, and tail: The system identifies the chord length and curvature of the wing, the streamlined design of the fuselage, and the angle and layout of the tail as the basic geometric features for subsequent 3D generation.
[0093] S101-2: Detailed Features: The system further extracts information about important design components such as the engine air intake and landing gear location.
[0094] S101-3: Texture and Material Characteristics: Extracts the texture details and reflective properties of the aircraft surface. This information will be used to reflect the visual realism in the subsequent rendering process.
[0095] The extracted features are converted into image feature vectors G(g1,g2,...,g n ), where g1 represents the geometric features of the wing, g2 represents the length features of the fuselage, and so on. This vector is used for the next step of multimodal fusion processing.
[0096] Step S102: Text input: The user also provides a detailed text description of the aircraft design. The system parses the text using GPT-4O to extract specific information about the design goals and functional requirements. The extracted text features include:
[0097] S102-1: Functional requirement: Enhance the load-bearing capacity of the wing while maintaining aerodynamic performance.
[0098] S102-2: Geometric constraints: Maintain the aerodynamic design of the aircraft, avoiding significant geometric changes to the wings and tail.
[0099] S102-3: Structural optimization: Reduce the overall weight of the aircraft while ensuring structural strength, especially in the wing area.
[0100] S102-4: Material selection: Use carbon fiber materials for the wings and lightweight aluminum alloys for the fuselage.
[0101] S102-5: Special requirements: Use high-temperature-resistant materials in the intake area to ensure safety under high-temperature airflow.
[0102] GPT-4O converts the above information into a text feature vector, denoted as T(t1, t2,..., t m ), where t1 represents functional requirements, t2 represents geometric constraints, and so on.
[0103] Step S200: Multimodal fusion; as shown in Figure 3 , the system fuses image features and text features through the BLIP-2 model to generate a unified feature vector R(r1, r2,..., r s ). The specific fusion steps include:
[0104] S200-1: Feature adaptation processing: To ensure that image features and text features are processed in the same vector space, BLIP-2 first standardizes image features G(g1, g2,..., g n ) and text features T(t1, t2,..., t m ) through an adaptation layer to ensure that both feature vectors have the same dimension.
[0105] S200-2: Feature fusion generation: After feature splicing and weight adjustment, a unified feature vector R(r1, r2,..., r s ) is generated. This vector includes the geometric shape of the aircraft and also reflects the user's design goals and functional requirements.
[0106] Step S300: 3D model generation; after the generation of the unified feature vector after fusion, the system uses the Point-E model to generate a 3D model, with the following specific steps:
[0107] Step S301: Point-E generates 3D model: The Point-E model receives the unified feature vector R(r1, r2,..., r s) as input, generating multiple 3D point cloud models. The model generation focuses on accurately representing the complex geometric structures of the wing, fuselage, and tail, while also focusing on the wing's load-bearing capacity and structural optimization. The model embodies a streamlined design and lightweight fuselage structure.
[0108] Step S302: Blender rendering: The generated point cloud model is imported into Blender for further rendering. The Blender rendering process includes:
[0109] S302-1: Lighting settings: simulate the reflection effect of real light to ensure the realistic appearance of the model under different lighting conditions.
[0110] S302-2: Material Application: Apply the specified carbon fiber or aluminum alloy material to different parts of the model.
[0111] S302-3: View angle adjustment: generating multiple rendering images of the 3D model according to different view angle settings.
[0112] Step S400: User demand analysis and model screening, such as Figure 4 As shown in the figure, after generating a preliminary 3D model, the system uses GPT-4O to analyze user needs and filters the generated model based on the knowledge graph. The specific steps are as follows:
[0113] Step S401: GPT-4O parses user requirements: The system uses GPT-4O to parse the user's design description and convert it into screening rules. These rules cover the structural integrity, functional compatibility, and geometric constraints of the aircraft, and are converted into computer-executable screening criteria in text form.
[0114] Step S402: Knowledge Graph Construction: The system constructs a knowledge graph that includes basic design rules, functional requirements, and geometric constraints. During the model screening process, the knowledge graph matching process can further demonstrate the specific matching method and judgment criteria.
[0115] Through various data sources such as design documents, industry databases, standard documents and expert interviews, information related to basic design rules, functional requirements and geometric constraints is systematically collected to ensure the comprehensiveness and accuracy of the knowledge graph. Furthermore, named entity recognition and relationship extraction techniques are used to perform in-depth knowledge extraction on the collected text data, identify specific design parameters and their relationships, and model the information in the form of nodes and edges to form a structured knowledge graph. Then, after the graph is established, data is regularly collected from the latest industry standards, research results and user feedback to dynamically update the knowledge graph, ensure that its content is consistent with actual design requirements, and introduce new design standards and domain knowledge. Furthermore, in the process of model scoring and selection, the constructed knowledge graph is used to judge the compliance and adaptability of the generated model, and assist in screening out the 3D model that best meets user needs.
[0116] Step S403: Model scoring and selection: The system scores the generated 3D models according to the design rules. The scoring criteria include:
[0117] S403-1: Structural integrity: Based on specific quantitative parameters such as the geometric continuity of the model, stress distribution between nodes, and stiffness index.
[0118] S403-2: Functional matching: Introduce quantitative evaluation criteria for specific functions, such as load bearing capacity and numerical analysis of aerodynamic performance, to judge whether the model meets the user's functional requirements.
[0119] S403-3: Innovation: This part can analyze the uniqueness and improvements of the model in design by comparing it with existing knowledge graph models.
[0120] It is feasible that in step S403, the specific operation steps are as follows: First, according to the design rules, the scoring criteria are set, including evaluation indicators such as structural integrity, functional matching and innovation, among which structural integrity evaluates the completeness and accuracy of the model in terms of geometric structure, functional matching evaluates the ability of the model to meet the user's design needs, and innovation evaluates the uniqueness of the model in design ideas and forms. Further, using the set scoring criteria, the 10 preliminary 3D models generated in step S300 are comprehensively scored. The specific method is to conduct a quantitative evaluation for each indicator, for example, using a scoring scale of 0 to 10 to ensure the objectivity and consistency of the evaluation process. Then, the scoring results of each indicator are combined to calculate the total score of each model, ensuring that each indicator is weighted according to its importance in the design to reflect the comprehensive performance of the model. Further, based on the scoring results, the model with the highest score is selected as the basis for subsequent optimization and adjustment.
[0121] This embodiment proposes an innovative 3D design assistance method and system by combining multimodal learning technology, pre-trained GPT-4O model, CLIP-ViT model and knowledge graph. The system effectively integrates text and image input, uses GPT-4O to extract design intent and functional requirements, and analyzes image structural features through CLIP-ViT, laying the foundation for subsequent 3D model generation. The feature fusion achieved by the BLIP-2 model improves the adaptability to complex design requirements and enhances the comprehensive performance of the model. In the generation stage, the Point-E model is used to generate a high-quality preliminary 3D model, and it is rendered through Blender to ensure that the model meets the design intent. In addition, the construction of a knowledge graph that supports dynamic updates and the use of the GPT-4O model enable the system to flexibly integrate the design rules and functional requirements required by users to guide the model screening process. Combining the above advantages, this method significantly improves the generation efficiency and accuracy of 3D models, providing strong intelligent support for design innovation.
[0122] This embodiment provides a 3D design assistance system based on multimodal input and intelligent optimization, the system comprising:
[0123] Multimodal Input Module: This module uses pre-trained GPT-4O to parse user text input, extract design intent and functional requirements, and uses CLIP-ViT to extract structural features from images for subsequent processing.
[0124] Feature fusion module: uses BLIP-2 to fuse text and image features to generate a unified feature vector, which provides input for 3D model generation;
[0125] 3D model generation module: Generate a preliminary 3D model using the Point-E model and render it through Blender to ensure that the generated model meets the design intent;
[0126] User Requirements Parsing and Rule Matching Module: This module uses GPT-4O to analyze the semantics of design inputs, builds and maintains a knowledge graph containing basic design rules, functional requirements, and geometric constraints, dynamically updates the knowledge graph to ensure the validity of its content, and assists in determining the compliance and adaptability of the generated model during the model screening process.
[0127] Rule-based intelligent model scoring and optimization module: Based on the set scoring criteria, the model's structural integrity, functional matching and innovation are evaluated. After comprehensive scoring, the model with the highest score is selected for further optimization and adjustment.
[0128] Any process or method described in the flowchart of this embodiment or in other ways herein can be understood as representing a module, segment, or portion of code including one or more executable instructions for implementing specific logical functions or process steps, and can be implemented in any computer-readable medium for use by an instruction execution system, device, or apparatus. The computer-readable medium can be any medium that stores, communicates, propagates, or transmits a program for use by an execution system, device, or apparatus, including read-only memory, magnetic disks, or optical disks.
[0129] The above description is merely a preferred embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A 3D design assistance method based on multimodal input and intelligent optimization, characterized in that: include: Obtaining text data and image data to be processed; Performing feature extraction on the text data to be processed and the image data to be processed respectively to obtain text features and image features; the text features include size description, functional description, material description and structure description; the image features include image structure features; Performing feature fusion on the text features and the image features to obtain fused features; The specific steps of fusing text features and image features include: Feature adaptation: Dimension matching is performed on the image feature G and the text feature T to ensure that the two feature vectors can be aligned in the same dimensional space. Assume that the image feature dimension is n and the text feature dimension is m. The conversion of feature dimensions is completed through the adaptation matrix A, which is expressed as: R=A·G+A·T Where A is the adaptation matrix, which is used to project image and text features into a unified dimensional space; Attention mechanism fusion: Through the multi-head attention mechanism, the image features and text features are fused: let the attention weight be α i , expressed as: In the formula, R is the fused feature vector, α i is the attention weight calculated based on the feature importance, G i and T i are the i-th components of image and text features respectively; Output fusion features: The generated fusion vector R represents a unified multimodal feature, which is used as input for generating a 3D model; Input the fused features into the Point-E model, generate several initial models using point cloud technology, and perform rendering optimization on each initial model to obtain several 3D models; Analyze user requirements: Analyze the design requirements input by users and convert them into actionable screening rules. The screening rules include wing load capacity, fuselage streamline design and material requirements; Build a knowledge graph: Build a knowledge graph that includes basic design rules, functional requirements, and geometric constraints. Use this graph to help match the generated model with user needs during the model screening process. Model scoring and selection: The generated 3D models are scored based on the design rules in the knowledge graph, and the 3D model with the highest score is selected as the optimal 3D model; the scoring criteria include structural integrity, functional matching, and innovation.
2. The 3D design assistance method based on multimodal input and intelligent optimization according to claim 1, characterized in that: The process of obtaining the text features specifically includes: The text data to be processed is parsed based on the GPT-4O model to obtain the text features.
3. The 3D design assistance method based on multimodal input and intelligent optimization according to claim 1, characterized in that: The image structural features include edge features, texture features and shape features of objects in the image data to be processed.
4. The 3D design assistance method based on multimodal input and intelligent optimization according to claim 3, characterized in that: The process of acquiring the image structural features specifically includes: Inputting the image data to be processed into the CLIP-ViT model, extracting the outline and boundary of the object in the image data to be processed, and obtaining the edge feature; encoding local pixel differences of the image data to be processed according to a local binary pattern algorithm to obtain the texture features; The shape feature is obtained by performing frequency domain representation on the outline and boundary of the object based on Fourier transform.
Citation Information
Patent Citations
Database alarm intelligent diagnosis method and system based on multi-modal knowledge graph fusion and small sample learning
CN118467229A
Knowledge graph and rule constraint combined data intelligent analysis method and system
CN118606440A