3D scene generation system and method, medium and equipment
Through the 3D scene generation system based on 3D model generation technology and visual big model, the existing 3D model construction methods are solved, and efficient, fast and diversified 3D scene generation is achieved, improving the model quality and detailed performance.
Patent Information
- Application Number
- CN202510383145.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-06-27
AI Technical Summary
The existing 3D model construction methods are time-consuming, costly, and highly subjective. The existing 3D data sets are limited in scale, insufficient diversity, and poor quality and consistency.
A 3D scene generation system based on 3D model generation technology and visual big model is adopted. The system extracts objects and their position information through image recognition, combines visual big model to generate similar images, and converts the image into 3D objects through 3D model generation technology, and finally loads and optimizes the scene in a 3D simulation environment.
It realizes efficient and rapid generation of high-quality and diverse 3D scenes, reduces the cost of 3D model construction, avoids the subjectivity of manual modeling, and improves the details and texture performance of the model.
Smart Images

Figure CN120219635A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of 3D scene modeling, and in particular to a 3D scene generation system, method, medium and device based on 3D model generation technology and vision large model. Background Art
[0002] In the process of constructing a three-dimensional model, traditional methods include manual modeling by professional modelers and scanning real scenes using devices such as radar. However, these methods have limitations.
[0003] I. Manual Modeling:
[0004] 1) Time-consuming and laborious: Manual modeling requires professionals to invest a large amount of time and effort, especially when facing complex models, the workload is even more substantial. 2) High cost: The cost of manual modeling is relatively high. For projects with limited budgets, relying on manual modeling is not realistic. 3) Strong subjectivity: The quality and detail level of the model may vary depending on the experience and skill level of the modeler, lacking consistency. Differences in the style and method of manual modeling may lead to inconsistencies in the details and accuracy of the output model, affecting the final output quality.
[0005] II. Scanning Real Scenes Using Professional Data Acquisition Devices:
[0006] 1) Expensive equipment: High-quality professional data acquisition devices are costly and require professional personnel to operate. In addition, the operation and maintenance of these devices also require professional technicians, further increasing the cost. 2) Environmental limitations: The scanning process may be affected by environmental factors such as insufficient light, narrow space, reflective surfaces or moving objects, resulting in difficult data acquisition. For example, scanning complex indoor scenes in low-light environments or conducting data acquisition in narrow underground passages may face numerous difficulties. Reflective surfaces can cause interference to lidar signals, while moving objects may lead to data inconsistencies. These environmental limitations will affect the scanning effect and reduce the data quality. 3) Complex data processing: The large amount of data generated by scanning requires complex processing and optimization, increasing the time and technical requirements. Scanning devices usually generate point cloud data, which often has problems such as missing details and noise. Point cloud data may not be able to accurately capture the small parts of objects. In addition, the process of processing point cloud data, including point cloud alignment, noise reduction and three-dimensional reconstruction, requires high-performance computing resources and professional knowledge. The data processing process is complex and time-consuming, which may affect the project progress. The resolution and accuracy of point cloud data are also limited by the device performance and may not meet the requirements of high-precision three-dimensional models.
[0007] At the same time, existing 3D model datasets also have corresponding disadvantages:
[0008] I. Limited dataset size: The currently widely used 3D dataset Objaverse-xl only contains 10 million, which is much smaller compared to the large-scale datasets for language, image, and video tasks. The limited dataset size restricts the ability to train powerful and generalizable 3D generation models.
[0009] II. Insufficient data diversity: Smaller datasets usually lack the diversity required to capture all 3D shapes and textures, resulting in the model performing well on seen objects but having poor generalization on unseen objects.
[0010] III. Poor data quality and consistency: Many existing datasets contain complex scenes with multiple objects, which complicates the training process and makes it difficult to generate 3D model datasets individually. Inconsistent or incomplete annotations in existing datasets may lead to inaccuracies in model training and evaluation. Existing datasets usually lack high-quality textures and detailed 3D models, which are crucial for realistic 3D generation. Summary of the Invention
[0011] The purpose of the present invention is to overcome the deficiencies of the prior art and propose a 3D scene generation system, method, medium, and device. The system first extracts the objects and their position information in the user-input picture through image recognition and generates similar images in combination with a large vision model. Then, it uses 3D model generation technology to convert the images into 3D objects and screens high-quality models through scoring and similarity comparison. Finally, the screened 3D models are loaded into the 3D simulation environment according to the original position information to generate high-quality and diverse 3D scenes.
[0012] The purpose of the present invention is achieved through the following technical solutions: A 3D scene generation system based on 3D model generation technology and a large vision model, including:
[0013] An input processing module, including an input text processing module and an input picture processing module, processes the multi-modal data input by the user according to the large vision model and the image-based object recognition model, extracts and converts the input data, and obtains an object picture and an object pose record file; wherein, the image-based object recognition model is based on the Transformer architecture and combines the DINO object detection algorithm and the GLIP model to identify the objects in the input picture and generate an object pose record file; the object pose record file includes the category name of the object, the bounding box information, and the image set of the recognized object; the large vision model analyzes the text information, generates a list of target objects, and combines the output of the object recognition model to screen out the final set of target objects;
[0014] The 3D model generation module expands the input single object image, generates a set of 3D models based on the expanded images, evaluates the quality and similarity of the model set, and selects the set of 3D models that meet the preset conditions, and performs the following steps:
[0015] a. Image expansion: Based on the image generation ability of the vision large model, generate multiple expanded images from a single input image;
[0016] b. 3D model generation: Based on 3D model generation technology, use the expanded image set as input to generate a corresponding set of 3D models;
[0017] c. 3D model quality evaluation: Use the 3D simulation environment to obtain the images of each 3D model set, and evaluate the quality of the images by the vision large model; Set evaluation criteria and divide the scoring levels. The evaluation criteria include the appearance logic, appearance representativeness and monomericity of the 3D model. According to the scoring results, retain the 3D models with scores higher than the set threshold, and eliminate the 3D models with scores lower than the set threshold;
[0018] d. 3D model similarity evaluation: Based on the DINO algorithm, establish a visual similarity model. Through self-supervised learning, use the Transformer architecture to extract image features, and perform similarity calculation in the feature space to compare the similarity between the images of the 3D models and the input images, and eliminate the 3D models with similarity lower than the set threshold to ensure that the finally generated 3D models are highly consistent with the input images;
[0019] e. Output the final 3D model set: Select a preset number of 3D models with the highest scores from the filtered 3D model set as the final output, and select the set of 3D models that meet the preset conditions;
[0020] The 3D scene generation module builds a scene in the 3D simulation environment according to the object pose record file and the 3D model set;
[0021] The 3D scene optimization module evaluates the rationality of the 3D scene according to the general reasoning ability of the vision large model, improves the unreasonable parts of the 3D scene, and finally obtains the optimized 3D scene.
[0022] Furthermore, the input processing module includes:
[0023] The input text processing module uses the vision large model to read the text information input by the user, including natural language and any other readable text, and the vision large model obtains a list of objects after analyzing the text information:
[0024] D = {X1, X2,..., X n};
[0025] Among them, D represents a list of objects, and X n represents an object in the list, and n is the object number;
[0026] The input picture processing module uses an image-based object recognition model to recognize the objects in the input picture and obtains the following output:
[0027]
[0028] d i =(c i , B i );
[0029] B i =[x min , y min , x max , y max ;
[0030] Among them, is the information set in the object pose record file, and d i represents the information of each recognized object, represents the set of recognizable types of this image-based object recognition model, and B i represents the bounding box, x min represents the minimum x coordinate of the bounding box, y min represents the minimum y coordinate of the bounding box, x max represents the maximum x coordinate of the bounding box, y max represents the maximum y coordinate of the bounding box;
[0031] Combined with the object list D={X1, X2,..., X n} obtained by the input text processing module, the target set is as follows:
[0032]
[0033] Finally, the input picture is processed to obtain the object picture set
[0034]
[0035] Among them:
[0036]
[0037] represents passing through Bi the process of intercepting the corresponding image area according to the bounding box information in
[0038] Furthermore, the image augmentation includes the following steps:
[0039] a1. Input image parsing: Extract features from the input image based on a large vision model, and analyze the visual information of the image features. The visual information includes type, shape, material, and style;
[0040] a2. Latent space encoding and semantic understanding: Map the image features and natural language descriptions to the latent space for multi-modal modeling;
[0041] a3. Image transformation and conditional generation: Generate multiple transformed images based on the diffusion model through the DALL·E 3 model to ensure a high degree of similarity between the image style and image features;
[0042] a4. Output image set: Introduce diversity to form an augmented image set that meets the preset conditions.
[0043] Furthermore, the 3D model quality assessment includes:
[0044] The evaluation criteria include the appearance logic, appearance representativeness, and monomericity of the 3D model; Appearance logic means that the appearance of the 3D model needs to conform to physical logic, that is, to judge whether there are missing key components or structural abnormalities in the appearance of the 3D model; Appearance representativeness means that the appearance of the 3D model needs to conform to the description context, that is, to judge the matching situation between the 3D model and the style and era characteristics of the input image; Monomericity means that the 3D model needs to belong to a single object, rather than a combined object, that is, to judge the situation where the generated 3D model includes related but independent objects;
[0045] Corresponding scores are given according to the evaluation criteria, and the scoring levels are divided. The scoring levels include high level, medium level and low level, and a corresponding scoring interval corresponds to a scoring level. During the process of evaluating the image quality, if there is no missing key component or abnormal structure in the appearance of the 3D model, the 3D model directly matches the style and era characteristics of the input image, or the 3D model does not include relevant but independent objects, the corresponding scoring interval is 8-10 points, and the scoring level is high level; if there are some missing key components or abnormal structures in the appearance of the 3D model, the 3D model partially matches the style and era characteristics of the input image, or the 3D model includes some relevant but independent objects, the corresponding scoring interval is 4-7 points, and the scoring level is medium level; if there are all missing key components or abnormal structures in the appearance of the 3D model, the 3D model does not match the style and era characteristics of the input image, or the 3D model includes multiple relevant but independent objects, the corresponding scoring interval is 1-3 points, and the scoring level is low level; according to the scoring results, 3D models with scores higher than 8 points are retained, and 3D models with scores lower than 8 points are removed, that is, 3D models with high-level scores are retained.
[0046] Furthermore, the 3D model generation module includes:
[0047] Based on the image-to-image generation ability of the vision large model, expand the object picture set Obtain the expanded picture set
[0048]
[0049] Among them, where u represents a preset expansion coefficient. The larger the expansion coefficient, the more pictures similar to the object pictures are obtained. Correspondingly, the more computing resources and generation time are consumed;
[0050] Based on the image-to-3D process in the 3D model generation technology, use each picture in the picture set as input to obtain the corresponding 3D model output, and obtain the output 3D model set Among them:
[0051]
[0052] f represents the process of using a picture as input in the 3D model generation technology to obtain the corresponding 3D model. Multiple perspective pictures are obtained from a single input picture, and then a 3D model is obtained from the multiple perspective pictures. The appearance characteristics of the original input picture are all reflected in the finally generated 3D model;
[0053] For the obtained 3D model set Give a score to each of the 3D models, and eliminate the 3D models that do not meet the requirements. The evaluation criteria include the appearance logic, appearance representativeness, and monomericity of the 3D models. The appearance logic means that the appearance of the 3D model needs to conform to physical logic, that is, to judge whether there are missing key components or structural abnormalities in the appearance of the 3D model; the appearance representativeness means that the appearance of the 3D model needs to conform to the description context, that is, to judge the matching situation between the 3D model and the style and era characteristics of the input image; the monomericity means that the 3D model needs to belong to a single object, rather than a combined object, that is, to judge the situation where the generated 3D model includes related but independent objects; perform corresponding scoring according to the evaluation criteria, retain the 3D models with scores higher than the set threshold, and eliminate the 3D models with scores lower than the set threshold, and select high-quality 3D models from the 3D model set to obtain a set
[0054] Use a visual similarity model to evaluate the similarity between the 3D model and the set of object pictures and screen out the 3D model that is closest to the object picture from it, so that the finally generated 3D scene is as close as possible to the 3D scene in the input picture; for each 3D model in the set , use a simulated camera in the 3D simulation environment to obtain an image of the 3D model, input the image and the image in the corresponding set of object pictures into the visual similarity model to obtain a score z. All 3D models with a score z lower than the set threshold are eliminated, and finally the 3D model set
[0055] Furthermore, the 3D scene generation module includes:
[0056] Load the 3D model into the 3D simulation environment according to the object pose record file and the 3D model set , use the collision detection mechanism of the simulation environment to detect collisions between different 3D models. If there is a collision, scale the colliding objects until there is no collision, then adjust the pose according to the corresponding records in the object pose record file, and finally add default lighting to build the final 3D scene.
[0057] Furthermore, the 3D scene optimization module includes:
[0058] Scene understanding and object recognition: The large vision model analyzes the input 3D scene image, recognizes the objects and their categories in the scene, and combines its preset common sense knowledge base to understand the semantic information of the scene;
[0059] Orientation rationality evaluation: Based on the common sense knowledge base, the large vision model evaluates whether the orientation of each 3D model in the scene conforms to common sense rules. If the orientations of all 3D models are evaluated as reasonable, the task ends; otherwise, proceed to the improvement plan generation step;
[0060] Improvement plan generation: For 3D models with unreasonable orientations, the large vision model generates specific improvement plans, including adjusting the rotation angle or position;
[0061] Plan execution and scene update: The 3D simulation environment receives the improvement plan generated by the large vision model, adjusts the orientation of the 3D model, updates the 3D scene, and simultaneously jumps to the evaluation of orientation rationality.
[0062] A 3D scene generation method based on 3D model generation technology and large vision model, which is implemented by a processor calling the input processing module, 3D model generation module, 3D scene generation module, and 3D scene optimization module in the 3D scene generation system based on 3D model generation technology and large vision model described above.
[0063] A non - transitory computer - readable medium storing instructions, which, when executed by a processor, execute the 3D scene generation method based on 3D model generation technology and large vision model described above.
[0064] A computing device, including a processor and a memory for storing processor - executable programs. When the processor executes the programs stored in the memory, it implements the 3D scene generation method based on 3D model generation technology and large vision model described above.
[0065] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0066] 1. The present invention can efficiently and quickly generate 3D scenes. Compared with the prior art, the instant generation process of the present invention is applicable to real - time or interactive applications;
[0067] 2. The present invention can automatically generate 3D scenes. Without manual modeling, it only needs to input text descriptions or images to automatically generate 3D models, which is easy to use and expands the application scope;
[0068] 3. The present invention does not require professional data acquisition equipment for scanning, reducing the cost of 3D model construction.
[0069] 4. The present invention can significantly improve the quality of the generated 3D models, making the models perform better in terms of details and textures;
[0070] 5. The present invention constructs 3D scenes based on large vision models, enhancing the ability to model the deep semantic associations between images and texts, enabling it to perform more complex reasoning tasks;
[0071] 6. The present invention can optimize the image generation model, improving the authenticity, coherence, and detail richness of images;
[0072] 7. The present invention can reduce the model calculation cost, improve the inference speed, and make the vision large model more suitable for real-time application scenarios;
[0073] 8. The vision similarity model of the present invention can enhance the model's ability to recognize object features under different angles, scales, and lighting conditions, and improve the robustness;
[0074] 9. The vision similarity model of the present invention can reduce the computational complexity and make large-scale image matching tasks easier to deploy in resource-constrained environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] Figure 1 FIG. is an architecture diagram of a 3D scene generation system based on 3D model generation technology and a vision large model.
[0076] Figure 2 FIG. is a flowchart of the operation of a 3D scene generation system based on 3D model generation technology and a vision large model.
[0077] Figure 3 FIG. is an architecture diagram of the input processing module of a 3D scene generation system based on 3D model generation technology and a vision large model. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0078] The present invention will be further described below with reference to specific embodiments.
[0079] Embodiment 1
[0080] Referring to Figures 1 to 2 shown, a 3D scene generation system based on 3D model generation technology and a vision large model provided in this embodiment, developed using the Python language, includes:
[0081] 1) An input processing module, including an input text processing module and an input image processing module, which processes the multimodal data input by the user according to the vision large model and the image-based object recognition model, extracts and transforms the input data, and obtains an object picture and an object pose record file; among them, referring to Figure 3 shown, the image-based object recognition model is based on the Transformer architecture and combines the DINO object detection algorithm and the GLIP model to identify the objects in the input picture and generate an object pose record file; the object pose record file includes the category name of the object, the bounding box information, and the image set of the recognized object; the vision large model VLM analyzes the text information, generates a list of target objects, and combines the output of the object recognition model to screen out the final set of target objects;
[0082] In this embodiment, the input data includes a scene picture in RGB format and a piece of text provided by the user. The picture depicts the target scene that the user expects to generate and serves as a visual reference. The text provided by the user is used to describe the requirements, which involve the screening of objects in the scene, the attention to specific features, or the supplementary description of the overall environment. The text is in natural language or other readable expressions.
[0083] After receiving the input data, the vision large model mainly performs the following tasks: image understanding and object recognition, scene analysis, and text parsing and target object screening; the vision large model first analyzes the input picture to identify the objects contained therein and extracts the following key information: the category of the objects, the number of the objects, and the appearance features of the objects; analyzes the overall scene to obtain the following information: the scene category, the overall atmosphere of the scene, including lighting conditions and style features. Further parse the text provided by the user and conduct a comprehensive analysis in combination with the image content. The main role of the user text is to guide the vision large model to focus on specific objects or scene elements. The user uploads a picture of a living room scene and clearly requests to focus on the desktop items in the text. The vision large model needs to comprehensively analyze the picture and the text and extract the object information that meets the requirements. Finally, the model will screen out the target objects and generate a list of objects, which contains all the object categories that meet the user's needs. When there are multiple objects of the same category in the input picture, the list may contain duplicate category names to accurately reflect the distribution of objects in the scene.
[0084] The object recognition model based on images analyzes the input pictures to detect all the objects therein and generates corresponding record files. The record files contain the following key information: 1. Object categories and confidence evaluation: The object recognition model extracts the object categories in the pictures and assigns a confidence score to each category. The names of all object categories form a category set, and each element in this set is attached with a confidence value between 0 and 1. For an identification result "apple:0.77", where "apple" represents the object category and 0.77 represents the recognition confidence of this category. The object recognition model sets a confidence threshold, and the recognition results below this threshold will be discarded to ensure the reliability of the recognition results. This threshold is usually a built-in parameter of the model and can also be customized and adjusted by the user. Therefore, the final record file only contains the object categories with confidence higher than the set threshold and their corresponding confidence values. 2. Bounding boxes of the objects: The object recognition model also generates bounding boxes for each detected object, annotating the position and size of the object in the image in the form of a rectangular box. The coordinates of the bounding box are represented by pixel points, ensuring that the position of the object can be accurately located and allowing subsequent cropping and further processing of the image. 3. Image set of the recognized objects: Based on the bounding box information, the corresponding object regions are cropped from the input image to construct an object image set. Each image in this set corresponds to a certain object category recognized in the record file.
[0085] By combining the analysis results of the large vision model and the object recognition model, the target objects that meet the user's needs are further screened. Specifically, first, two category sets are constructed. One is the set of target object categories recognized by the large vision model, and the other is the set of object categories extracted by the object recognition model. The intersection of the two represents the set of categories that are recognized by both models and are considered target objects. The recognition information of all objects in this intersection will be retained for subsequent processing and 3D model generation.
[0086] The specific execution operations of the input processing module are as follows:
[0087] The input text processing module uses the large vision model to read the text information input by the user, including natural language and any other readable text. After analyzing the text information, the large vision model obtains a list of objects:
[0088] D = {X1, X2,..., X n};
[0089] where D represents the list of objects, X n represents the objects in the list, and n is the object number;
[0090] The input picture processing module uses the object recognition model based on images to recognize the input pictures For the objects in, the following output is obtained:
[0091]
[0092] d i = (c i , B i );
[0093] B i = [x min , y min , x max , y max ;
[0094] Wherein, is the information set in the object pose record file, and d i represents the information of each recognized object, represents the set of recognizable types of this image-based object recognition model, and B i represents the bounding box, x min represents the minimum x coordinate of the bounding box, y min represents the minimum y coordinate of the bounding box, x max represents the maximum x coordinate of the bounding box, and y max represents the maximum y coordinate of the bounding box;
[0095] Combined with the object list D = {X1, X2,..., X n} obtained by the input text processing module, the target set is as follows:
[0096]
[0097] Finally, the input picture is processed to obtain the object picture set
[0098]
[0099] Wherein:
[0100]
[0101] represents the process of intercepting the corresponding picture area of i through the bounding box information in B .
[0102] 2) 3D Model Generation Module: Expand the input single-object image, generate a 3D model set based on the expanded images, evaluate the quality and similarity of the model set, and select a 3D model set that meets the preset conditions. The 3D model generation module performs the following steps:
[0103] a. Image Expansion: Based on the image generation ability of the large vision model, generate 20 expanded images from a single input image, including the following steps:
[0104] a1. Input Image Parsing: Extract features from the input image based on the large vision model, and analyze the visual information of the image features. The visual information includes type, shape, material, and style.
[0105] a2. Latent Space Encoding and Semantic Understanding: Map the image features and natural language descriptions to the latent space for multi-modal modeling.
[0106] a3. Image Transformation and Conditional Generation: Through the DALL·E 3 model, generate multiple transformed images based on the diffusion model to ensure a high degree of similarity in image style and image features.
[0107] a4. Output Image Set: Introduce diversity to form an expanded image set that meets the preset conditions. All generated images are in RGB format to ensure compatibility with subsequent processing flows.
[0108] Expand the object image set based on the image generation ability of the large vision model Obtain the expanded image set
[0109]
[0110] Among them, where u represents the preset expansion coefficient. The larger the expansion coefficient, the more images similar to the object image are obtained. Correspondingly, more computing resources and generation time are consumed. Similar pictures, and the corresponding consumption of computing resources and generation time is also more;
[0111] b. 3D Model Generation: Based on 3D model generation technology, use the expanded image set as input to generate the corresponding 3D model set; based on the image-to-3D process in 3D model generation technology, use each image in the image set as input to obtain the corresponding 3D model output, and obtain the output 3D model set Among them:
[0112]
[0113] Let \(f\) denote the process in 3D model generation technology where an image is used as input to obtain the corresponding 3D model. Multiple perspective images are obtained from a single input image, and then a 3D model is obtained from the multiple perspective images. All the appearance features of the original input image are reflected in the finally generated 3D model.
[0114] c. 3D model quality assessment: Use a 3D simulation environment to obtain images of each set of 3D models, and evaluate the quality of the images by a large vision model. The evaluation criteria include the appearance logic, appearance representativeness, and monomericity of the 3D models. Appearance logic means that the appearance of the 3D model should conform to physical logic, that is, judge whether there are missing key components or abnormal structures in the appearance of the 3D model. Appearance representativeness means that the appearance of the 3D model should conform to the description context, that is, judge the matching situation between the 3D model and the style and era characteristics of the input image. Monomericity means that the 3D model should belong to a single object, rather than a combined object, that is, judge the situation where the generated 3D model includes relevant but independent objects. Corresponding scores are given according to the evaluation criteria, and the score levels are divided. The score levels include high level, medium level, and low level, and a corresponding score interval corresponds to a score level. During the process of evaluating the image quality, if there are no missing key components or abnormal structures in the appearance of the 3D model, the 3D model directly matches the style and era characteristics of the input image, or the 3D model does not include relevant but independent objects, then the corresponding score interval is 8 - 10 points, and the score level is high level. If there are some missing key components or abnormal structures in the appearance of the 3D model, the 3D model partially matches the style and era characteristics of the input image, or the 3D model includes some relevant but independent objects, then the corresponding score interval is 4 - 7 points, and the score level is medium level. If there are all missing key components or abnormal structures in the appearance of the 3D model, the 3D model does not match the style and era characteristics of the input image, or the 3D model includes multiple relevant but independent objects, then the corresponding score interval is 1 - 3 points, and the score level is low level. According to the scoring results, retain the 3D models with scores higher than 8 points and eliminate the 3D models with scores lower than 8 points, that is, retain the 3D models with high-level scores from the set of 3D models to obtain a set
[0115] d. 3D model similarity assessment: Establish a visual similarity model based on the DINO algorithm. Through self-supervised learning, use the Transformer architecture to extract image features, and calculate the similarity in the feature space to compare the similarity between the image of the 3D model and the input picture, and screen out the 3D model closest to the object picture from them, so that the finally generated 3D scene is as close as possible to the 3D scene in the input picture. For each 3D model in the set use a simulated camera in the 3D simulation environment to obtain the rendered image of the 3D model, and compare this image with the corresponding set of object pictures The visual similarity model for the input image in it obtains a score z. All 3D models with a score z lower than the set threshold of 0.7 are eliminated, and finally a set of 3D models is obtained.
[0116] e. Output the final set of 3D models: Select a preset number of 3D models with the highest scores from the filtered set of 3D models as the final output, and select a set of 3D models that meet the preset conditions.
[0117] 3) 3D scene generation module. According to the object pose record file and the set of 3D models, build a scene in the 3D simulation environment; according to the object pose record file and the set of 3D models Load the 3D models into the 3D simulation environment. First, place all the 3D models in an initial layout according to preset rules. Subsequently, use the collision detection mechanism of the simulation environment to determine whether the meshes of different models overlap. If a collision is detected, scale the colliding objects by a certain proportion and re-detect until there are no more collisions in the scene, thus ensuring a reasonable object layout. Then adjust the pose according to the corresponding records in the object pose record file, and finally add default lighting to build the final 3D scene.
[0118] 4) 3D scene optimization module. Evaluate the rationality of the 3D scene according to the general reasoning ability of the large visual model, improve the unreasonable parts of the 3D scene, and finally obtain the optimized 3D scene. The 3D scene optimization module includes:
[0119] Scene understanding and object recognition: The large visual model analyzes the input 3D scene image, recognizes the objects and their categories in the scene, and combines its preset common sense knowledge base to understand the semantic information of the scene;
[0120] Orientaion rationality evaluation: Based on the common sense knowledge base, the large visual model evaluates whether the orientation of each 3D model in the scene conforms to the common sense rules. If the orientations of all 3D models are evaluated as reasonable, the task ends; otherwise, proceed to the improvement plan generation step;
[0121] Improvement plan generation: For 3D models with unreasonable orientations, the large visual model generates specific improvement plans, including adjusting the rotation angle or position;
[0122] Plan execution and scene update: The 3D simulation environment receives the improvement plan generated by the large visual model, adjusts the orientation of the 3D models, updates the 3D scene, and at the same time jumps to the orientation rationality evaluation.
[0123] Embodiment 2
[0124] This embodiment discloses a 3D scene generation method based on 3D model generation technology and a vision large model. This method is implemented by a processor invoking the input processing module, 3D model generation module, 3D scene generation module, and 3D scene optimization module in the 3D scene generation system based on 3D model generation technology and a vision large model described in Embodiment 1.
[0125] Embodiment 3
[0126] This embodiment discloses a non-transitory computer-readable medium storing instructions. When the instructions are executed by a processor, the steps of the 3D scene generation method based on 3D model generation technology and a vision large model described in Embodiment 2 are executed.
[0127] The non-transitory computer-readable medium in this embodiment can be a medium such as a magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), USB flash drive, or mobile hard disk.
[0128] Embodiment 4
[0129] This embodiment discloses a computing device, including a processor and a memory for storing processor-executable programs. When the processor executes the programs stored in the memory, the 3D scene generation method based on 3D model generation technology and a vision large model described in Embodiment 2 is implemented.
[0130] The computing device described in this embodiment can be a desktop computer, laptop computer, smart phone, PDA handheld terminal, tablet computer, programmable logic controller (PLC), or other terminal devices with processor functions.
[0131] The above-described embodiments are only the preferred embodiments of the present invention, and do not limit the scope of implementation of the present invention. Therefore, any changes made according to the shape and principle of the present invention should be covered by the protection scope of the present invention.
Claims
1. A 3D scene generation system based on 3D model generation technology and visual large models, characterized by: include: The input processing module includes an input text processing module and an input image processing module, which processes the multimodal data input by the user according to the visual big model and the image-based object recognition model, extracts and transforms the input data, and obtains the object image and the object posture record file; wherein the image-based object recognition model is based on the Transformer architecture and combines the DINO target detection algorithm and the GLIP model, and is used to identify the object in the input image and generate the object posture record file; the object posture record file includes the type name of the object, the bounding box information and the image set of the identified object; the visual big model performs text information analysis, generates a target object list, and combines the output of the object recognition model to screen out the final target object set; The 3D model generation module expands the image according to the input single object image, generates a 3D model set based on the expanded image, and performs quality and similarity evaluation on the model set, selects the 3D model set that meets the preset conditions, and performs the following steps: a. Image expansion: Based on the image generation capability of the large visual model, multiple expanded images are generated from a single input image; b. 3D model generation: Based on the 3D model generation technology, the expanded image set is used as input to generate the corresponding 3D model set; c. 3D model quality assessment: Use the 3D simulation environment to obtain images of each 3D model set, and use the visual large model to evaluate the quality of the images; set evaluation standards and divide the scoring levels. The evaluation standards include the appearance logic, appearance representativeness and monomer of the 3D model. According to the scoring results, retain the 3D models with scores higher than the set threshold, and eliminate the 3D models with scores lower than the set threshold; d. 3D model similarity evaluation: A visual similarity model is established based on the DINO algorithm. Through self-supervised learning, the Transformer architecture is used to extract image features, and similarity is calculated in the feature space. The image of the 3D model is compared with the similarity of the input image, and 3D models with similarity lower than the set threshold are eliminated to ensure that the final generated 3D model is highly consistent with the input image. e. Output the final 3D model set: select a preset number of 3D models with the highest scores from the screened 3D model set as the final output, and select a 3D model set that meets the preset conditions; 3D scene generation module, which builds scenes in a 3D simulation environment based on object posture record files and 3D model sets; The 3D scene optimization module evaluates the rationality of the 3D scene based on the general reasoning ability of the visual big model, improves the unreasonable parts of the 3D scene, and finally obtains the optimized 3D scene.
2. The 3D scene generation system based on 3D model generation technology and visual large model according to claim 1 is characterized in that: The input processing module comprises: The input text processing module uses the visual big model to read the text information input by the user, including natural language and other readable texts, wherein the visual big model analyzes the text information to obtain the object list: D={X1,X2,...,X n }; Where D represents the object list, X n Represents an object in the list, n is the object number; The input image processing module uses an image-based object recognition model to identify the input image The following output is obtained: d i =(c i ,B i ); B i =[x min ,y min ,x max ,y max ]; in, is the information set in the object pose record file, d i Represents the information of each object being identified, represents the set of recognizable types of the image-based object recognition model, B i represents the bounding box, x min Indicates the minimum x-coordinate and y-coordinate of the bounding box. min Indicates the minimum y coordinate of the bounding box, x max Indicates the maximum x coordinate, y coordinate of the boundingbox max Indicates the maximum y coordinate of the bounding box; Combined with the input text processing module, the object list D = {X1, X2, ..., X n }, and finally get the target set as follows: Finally, the input image Process and get a collection of object pictures in: Indicates that B i The bounding box information in the The process of corresponding image area.
3. The 3D scene generation system based on 3D model generation technology and visual large model according to claim 1 is characterized in that: Image augmentation includes the following steps: a1. Input image analysis: Extract features from the input image based on the visual big model and analyze the visual information of the image features, including type, shape, material and style; a2. Latent space encoding and semantic understanding: Map image features and natural language descriptions into latent space for multimodal modeling; a3. Image transformation and conditional generation: Through the DALL·E 3 model, multiple transformed images are generated based on the diffusion model to ensure the high similarity of image style and image features; a4. Output image set: Introduce diversity to form an expanded image set that meets preset conditions.
4. The 3D scene generation system based on 3D model generation technology and visual large model according to claim 1 is characterized in that: 3D model quality assessment includes: The evaluation criteria include the appearance logic, appearance representativeness and monomericity of the 3D model; appearance logic means that the appearance of the 3D model must conform to physical logic, that is, to judge whether the appearance of the 3D model has key components missing or structural abnormalities; appearance representativeness means that the appearance of the 3D model must conform to the description context, that is, to judge whether the 3D model matches the style and era characteristics of the input image; monomericity means that the 3D model must belong to a single object rather than a combination of objects, that is, to judge whether the generated 3D model includes related but independent objects; According to the evaluation criteria, corresponding scores are scored and divided into scoring levels, which include high level, medium level and low level, and the corresponding scoring interval corresponds to one scoring level; in the process of evaluating the image quality, if the appearance of the 3D model does not have any missing key components or abnormal structure, the 3D model directly matches the style and era characteristics of the input image, or the 3D model does not include related but independent objects, the corresponding scoring interval is 8-10 points, and the scoring level is high; if the appearance of the 3D model has some missing key components or abnormal structure, the 3D model partially matches the style and era characteristics of the input image, or the 3D model includes some related but independent objects, the corresponding scoring interval is 4-7 points, and the scoring level is medium; if the appearance of the 3D model has all missing key components or abnormal structure, the 3D model does not match the style and era characteristics of the input image, or the 3D model includes multiple related but independent objects, the corresponding scoring interval is 1-3 points, and the scoring level is low; according to the scoring results, 3D models with scores higher than 8 points are retained, and 3D models with scores lower than 8 points are eliminated, that is, 3D models with high scores are retained.
5. The 3D scene generation system based on 3D model generation technology and visual large model according to claim 1 is characterized in that: The 3D model generation module includes: Expand the collection of object images based on the image generation capability of the visual big model Get the expanded image collection in, u represents the preset expansion coefficient. The larger the expansion coefficient, the more images with the object will be obtained. Similar images consume more computing resources and take more time to generate. Based on the image generation process in 3D model generation technology, Take each picture in as input, get the corresponding 3D model output, and get the output 3D model set in: f represents the process of taking an image as input and obtaining a corresponding 3D model in the 3D model generation technology, obtaining multi-view images from a single input image, and then obtaining a 3D model from the multi-view images. The appearance features of the original input image are reflected in the final generated 3D model; For the obtained 3D model collection A score is given to each 3D model, and 3D models that do not meet the requirements are eliminated. The evaluation criteria include the appearance logic, appearance representativeness and monomericity of the 3D model. Appearance logic means that the appearance of the 3D model must conform to physical logic, that is, it is judged whether the appearance of the 3D model has key components missing or structural abnormalities; appearance representativeness means that the appearance of the 3D model must conform to the description context, that is, it is judged whether the 3D model matches the style and era characteristics of the input image; monomericity means that the 3D model must belong to a single object rather than a combination of objects, that is, it is judged that the generated 3D model includes related but independent objects; corresponding scores are given according to the evaluation criteria, and 3D models with scores higher than the set threshold are retained, and 3D models with scores lower than the set threshold are eliminated. Select high-quality 3D models to get a collection Using visual similarity models to evaluate 3D models and object image collections The similarity of the image is obtained by filtering out the 3D model closest to the object image, so that the final generated 3D scene is as close as possible to the 3D scene in the input image; For each 3D model in the 3D simulation environment, use the simulated camera in the 3D simulation environment to obtain the image of the 3D model, and compare the image with the corresponding object picture set The image in is input into the visual similarity model to obtain a score z. The 3D models with a score z lower than the set threshold are eliminated, and finally a 3D model set is obtained.
6. The 3D scene generation system based on 3D model generation technology and visual large model according to claim 5 is characterized in that: The 3D scene generation module includes: Record files and 3D model collections based on object poses The 3D model is loaded into the 3D simulation environment, and the collision detection mechanism of the simulation environment is used to detect the collision between different 3D models. If a collision occurs, the collision object is scaled until there is no collision, and then the posture is adjusted according to the corresponding record in the object posture record file. Finally, the default lighting is added to build the final 3D scene.
7. The 3D scene generation system based on 3D model generation technology and visual large model according to claim 1 is characterized in that: The 3D scene optimization module includes: Scene understanding and object recognition: The visual big model analyzes the input 3D scene image, identifies the objects and their categories in the scene, and understands the semantic information of the scene in combination with its preset common sense knowledge base; Orientation rationality assessment: Based on the common sense knowledge base, the visual big model evaluates whether the orientation of each 3D model in the scene conforms to the common sense rules. If the orientation of all 3D models is evaluated as reasonable, the task ends. Otherwise, the improvement plan generation step is carried out. Improvement plan generation: For 3D models with unreasonable orientation, the visual large model generates specific improvement plans, including adjusting the rotation angle or position; Solution execution and scene update: The 3D simulation environment receives the improvement solution generated by the visual large model, adjusts the orientation of the 3D model, updates the 3D scene, and jumps to the orientation rationality assessment.
8. A 3D scene generation method based on 3D model generation technology and visual large model, characterized in that: The method is implemented by a processor calling an input processing module, a 3D model generation module, a 3D scene generation module and a 3D scene optimization module in a 3D scene generation system based on 3D model generation technology and a visual large model as described in any one of claims 1-7.
9. A non-transitory computer-readable medium storing instructions, characterized in that: When the instructions are executed by a processor, the 3D scene generation method based on 3D model generation technology and visual large model according to claim 8 is executed.
10. A computing device comprising a processor and a memory for storing a program executable by the processor, characterized in that: When the processor executes the program stored in the memory, the 3D scene generation method based on 3D model generation technology and visual large model described in claim 8 is implemented.