Multi-view image fused 3D scene model generation method and device, and storage medium

By extracting entities and spatial relationships from scene description text, a structured semantic representation text is generated. The perspective is determined by combining semantic segmentation and spatial geometry theory. A text-generated graph model and a self-supervised layout optimizer are used to generate a logically reasonable 3D scene model, which solves the problems of object layout deviation and semantic deviation in existing technologies and achieves high-precision 3D scene reconstruction.

CN121767550APending Publication Date: 2026-03-31SHENZHEN YIDAO DIGITAL TECHNOLOGY R&D CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-21
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing text-driven 3D scene generation methods suffer from object layout and semantic biases, making it difficult to accurately reproduce the detailed descriptions in the text. Furthermore, their reliance on scarce 3D data resources results in weak model generalization ability, making it impossible to generate logically consistent 3D scenes.

Method used

By extracting entities, attributes, and spatial relationships from the scene description text, a structured semantic representation text is generated. Combining semantic segmentation and spatial geometry theory, multiple target perspectives are determined. A single-view image is generated using a text-generated graph model, and the geometric model of the sub-objects is recovered through a 3D generation model. Finally, a self-supervised layout optimizer adjusts the pose parameters to generate a logically reasonable 3D scene model.

Benefits of technology

It achieves semantically accurate, geometrically reasonable, and physically consistent 3D scene model generation, solves the problems of object layout deviation and semantic deviation in existing technologies, and improves the generalization ability and generation accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121767550A_ABST
    Figure CN121767550A_ABST
Patent Text Reader

Abstract

The invention discloses a 3D scene model generation method and device fusing a multi-view image, and a storage medium, and relates to the technical field of 3D image generation, and the method comprises the steps: determining a plurality of target views based on a structured semantic representation text; based on the structured semantic representation text and the single-view-angle information corresponding to the target view angle, generating a single-view-angle image through a text generation graph model; determining category information and 3D position information of the sub-object, and generating a sub-object 3D geometric model through a 3D generation model based on the category information, the position information and each single-view image; and performing attitude estimation on the sub-object 3D geometric models, obtaining attitude parameters of the sub-object 3D geometric models in the 3D scene, assembling the sub-object 3D geometric models, controlling the corresponding sub-object 3D geometric models according to the attitude parameters, and generating a target 3D scene model. The technical problem that in the prior art, object layout deviation and semantic deviation exist in a 3D scene model generated based on texts is solved, and accurate 3D scene model generation is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of 3D image generation technology, and in particular to a method, device and storage medium for generating 3D scene models by fusing multi-view images. Background Technology

[0002] There are two main types of existing text-driven 3D scene generation methods. One relies on a pre-built fixed 3D asset library and a large language model. By leveraging the natural language processing capabilities of the large language model, it first parses the text semantics and then generates an object layout scheme. Although it can parse the text semantics, it lacks fine-grained constraints on visual geometry and spatial relationships. The generated object layouts often exhibit problems that violate physical principles, such as spatial overlap, functional conflicts, or floating. The other method is based on a 2D diffusion model supervising a 3D neural radiation field. However, it is limited by relying on manually annotated depth maps or pre-built 3D model libraries, resulting in weak generalization ability and difficulty in accurately reproducing the detailed descriptions in the text, leading to inaccurate 3D scenes. Summary of the Invention

[0003] The main purpose of this application is to provide a method, device and storage medium for generating 3D scene models by integrating multi-view images, which aims to solve the technical problems of object layout deviation and semantic deviation in existing text-based 3D scene models.

[0004] To achieve the above objectives, this application proposes a method for generating a 3D scene model by fusing multi-view images. The method includes: Extract entities, attributes, and spatial relationships from scene description text to generate structured semantic representation text; Based on the structured semantic representation of the text, multiple target perspectives of the target 3D scene model are determined through semantic segmentation algorithms and spatial geometry theory; Based on the structured semantic representation text and the single-view information corresponding to each of the target viewpoints, a text-generated image model is used to generate single-view images corresponding to each of the target viewpoints. The category information and 3D position information of the sub-objects in each of the single-view images are determined, and based on the category information and position information of the sub-objects and each of the single-view images, a 3D geometric model of the sub-object corresponding to the sub-object is generated by a 3D generation model. The pose of the sub-object 3D geometric model is estimated to obtain the pose parameters of the sub-object 3D geometric model in the 3D scene. The sub-object 3D geometric model is assembled, and the pose of the corresponding sub-object 3D geometric model in the 3D scene is controlled according to the pose parameters to generate an initial 3D scene model. The pose parameters of the sub-object 3D geometric models in the initial 3D scene model are adjusted by a self-supervised layout optimizer to generate the target 3D scene model corresponding to the scene description text.

[0005] Furthermore, to achieve the above objectives, this application also proposes a 3D scene model generation device that integrates multi-view images, the 3D scene model generation device integrating multi-view images comprising: The text parsing module is used to extract entities, attributes, and spatial relationships from scene description text and generate structured semantic representation text. The multi-view planning module is used to determine multiple target views of the target 3D scene model based on the structured semantic representation text, through semantic segmentation algorithms and spatial geometry theory. The image generation module is used to generate single-view images corresponding to each target viewpoint based on the structured semantic representation text and the single-view information corresponding to each target viewpoint through a text-generated image model. The 3D model generation module is used to determine the category information and 3D position information of the sub-objects in each of the single-view images, and generate the corresponding 3D geometric model of the sub-object based on the category information and position information of the sub-objects and each of the single-view images through the 3D generation model. The pose estimation module is used to estimate the pose of the sub-object 3D geometric model, obtain the pose parameters of the sub-object 3D geometric model in the 3D scene, assemble the sub-object 3D geometric model, and control the pose of the corresponding sub-object 3D geometric model in the 3D scene according to the pose parameters to generate an initial 3D scene model. The layout optimization module is used to adjust the pose parameters of the sub-object 3D geometric models in the initial 3D scene model through a self-supervised layout optimizer, and generate the target 3D scene model corresponding to the scene description text.

[0006] Furthermore, to achieve the above objectives, this application also proposes a 3D scene model generation device that integrates multi-view images. The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor. The computer program is configured to implement the steps of the 3D scene model generation method that integrates multi-view images as described above.

[0007] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the 3D scene model generation method for fusing multi-view images as described above.

[0008] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the 3D scene model generation method for fusing multi-view images as described above.

[0009] The one or more technical solutions proposed in this application have at least the following technical effects: Extracting entities, attributes, and spatial relationships from scene description text to generate structured semantic representation text, accurately extracting entities and spatial relationships; based on the structured semantic representation text, determining multiple target perspectives of the target 3D scene model through semantic segmentation algorithms and spatial geometry theory, ensuring that the subsequently generated 2D images cover key spatial information, providing sufficient geometric clues for 3D reconstruction, and overcoming the limitation of insufficient single-view information; based on the structured semantic representation text and the single-view information corresponding to each target perspective, generating single-view images corresponding to each target perspective through a text-generated image model, providing rich visual basis for 3D modeling; determining the category information and 3D position information of sub-objects in each single-view image, and based on the category information and position information of the sub-objects and each single-view image, generating the corresponding 3D geometric model of the sub-object through a 3D generation model, transforming 2D visual information into a concrete 3D geometric model, performing pose estimation on the sub-object 3D geometric model, obtaining the pose parameters of the sub-object 3D geometric model in the 3D scene, and assembling the sub-object 3D geometric model. This paper establishes the basic framework of the scene model and controls the pose of the corresponding sub-object 3D geometric models in the 3D scene according to the pose parameters to generate an initial 3D scene model. The pose parameters of the sub-object 3D geometric models in the initial 3D scene model are adjusted by a self-supervised layout optimizer to generate a target 3D scene model that conforms to the scene description text description. By integrating NLP, computer vision and 3D generation technologies, semantic constraints and multi-view geometric verification are deeply integrated to solve the technical problems of object layout deviation and semantic deviation in the existing text-based 3D scene models. This achieves the generation of 3D scene models that are semantically accurate, geometrically reasonable and physically consistent. The 3D scene model generation method fused with multi-view images provided in this application extracts entities, attributes, and spatial relationships by performing fine-grained analysis on scene description text, generating structured semantic representation text, which provides clear semantic priors for subsequent geometric modeling; based on the structured semantic representation text, it intelligently plans multiple target perspectives (such as forward view and top view) by combining semantic segmentation and spatial geometry theory, providing rich and complementary geometric clues for 3D reconstruction; and uses a text-generated graph model to generate consistent multi-view 2D images (single-view images corresponding to each target perspective) under the dual constraints of semantics and perspective. Each single-view image not only retains the visual details of the scene description text, but also implicitly contains cross-view geometric consistency.Based on this, through depth estimation and 3D generation models, the 3D geometric model of each sub-object is accurately recovered from multi-view images (single-view images corresponding to each target viewpoint); the pose of each sub-object 3D model is estimated and assembled into an initial scene, thus initially realizing the 3D mapping of text spatial relationships; and a self-supervised layout optimizer is introduced to automatically adjust the pose parameters of the sub-object 3D geometric model, eliminating unreasonable layouts such as floating and overlapping, so that the final target 3D scene model is highly unified in terms of geometric structure, semantic alignment and physical rationality. Attached Figure Description

[0010] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0011] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is a flowchart illustrating a first embodiment of the method for generating a 3D scene model by fusing multi-view images in this application. Figure 2 This is a schematic diagram of the module structure of the 3D scene model generation device that integrates multi-view images according to an embodiment of this application; Figure 3 This is a schematic diagram of the device structure of the hardware operating environment involved in the 3D scene model generation method that integrates multi-view images in the embodiments of this application.

[0013] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0014] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0015] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0016] The main solution of this application embodiment is as follows: extract entities, attributes, and spatial relationships from the scene description text to generate structured semantic representation text; based on the structured semantic representation text, determine multiple target perspectives of the target 3D scene model through semantic segmentation algorithms and spatial geometry theory; based on the structured semantic representation text and the single-view information corresponding to each target perspective, generate single-view images corresponding to each target perspective through a text-generated graph model; determine the category information and 3D position information of sub-objects in each single-view image, and generate the corresponding sub-object 3D geometric model through a 3D generation model based on the category information and position information of the sub-objects and each single-view image; perform pose estimation on the sub-object 3D geometric model to obtain the pose parameters of the sub-object 3D geometric model in the 3D scene, assemble the sub-object 3D geometric model, and control the pose of the corresponding sub-object 3D geometric model in the 3D scene according to the pose parameters to generate an initial 3D scene model; adjust the pose parameters of the sub-object 3D geometric model in the initial 3D scene model through a self-supervised layout optimizer to generate the target 3D scene model corresponding to the scene description text.

[0017] In this embodiment, for ease of description, the following description will focus on a 3D scene model generation system that recognizes and fuses multi-view images.

[0018] Existing text-driven image generation technologies use a diffusion model as their core architecture. Through deep feature learning on massive image datasets, they map text descriptions to 2D images. Their core mechanism is based on Markov chain Monte Carlo theory, gradually reconstructing high-fidelity images that conform to the text's semantics from a Gaussian noise distribution during continuous denoising iterations. This allows for accurate capture of visual elements such as artistic style, color scheme, and object details. However, the output of this technology is limited to two-dimensional planar visual information, lacking crucial three-dimensional scene data such as geometric coordinates describing spatial depth and relative positions between objects. Directly applying this technology to 3D scene construction requires additional reliance on manually annotated depth maps or pre-built 3D model libraries. This indirect generation method not only easily leads to semantic discrepancies between the generated results and the text description in terms of spatial layout but also suffers from insufficient generalization ability when adapting to diverse scene styles, making it difficult to meet the needs of complex 3D creation. Existing text-to-3D scene generation technologies can be mainly divided into two categories: one is represented by pre-trained 2D text-to-image diffusion models (DreamFusion), which employs a fractional distillation sampling (SDS) loss mechanism to supervise and optimize the 3D neural radiance fields (3D NeRF) through the 2D diffusion model. This type of technology relies on a large amount of labeled 3D training data. In practical applications, due to the high cost of 3D data acquisition and annotation, its generalization ability is weak, and it can only generate scenes with simple structures and little detail. Furthermore, in multi-view rendering, inconsistencies in object shapes and lighting effects often occur. Another type, layout generation models based on pre-trained Transformers (LayoutGPT), leverage the powerful natural language processing capabilities of Large Language Models (LLMs) to first parse text semantics and generate object layout schemes. However, this type of technology relies on a pre-built fixed 3D asset library, preventing users from flexibly customizing the appearance of objects. Furthermore, due to the lack of visual geometric constraints, the generated layouts are prone to illogical phenomena such as overlapping objects and contradictory functional logic; for example, a generated furniture layout might prevent a door from opening properly. Both of these technologies are hampered by the scarcity of 3D data resources and the information gap in the conversion process between 2D representation and 3D space, respectively. They struggle to achieve an ideal balance between scene diversity, style consistency, and layout rationality, thus restricting the widespread practical application of text-to-3D scene generation technology. This application addresses three core problems existing in current technologies: First, existing text-to-3D methods heavily rely on scarce 3D data resources, resulting in weak model generalization ability. For example, taking a niche scene like a "Barbie-themed clinic" as an example, there are almost no matching samples in current mainstream 3D datasets, making it difficult for algorithms to learn and generate 3D scenes with unique aesthetic features and themes. Second, the precision of text command response is insufficient. For commands containing detailed quantity and shape descriptions, such as "3 round coffee tables," accurate conversion is impossible. This is attributed to the lack of effective 2D visual intermediaries in existing technologies, which cannot strictly constrain details such as the number and geometry of objects, leading to deviations in the number of objects or distortions in shape during 3D generation. Third, scene layouts often exhibit physical inconsistencies, such as objects floating in the air or overlapping entities. Existing technologies struggle to accurately estimate depth and pose based on 2D images and cannot accurately capture the spatial relationships between objects, ultimately resulting in spatially inconsistent 3D scenes.

[0019] This application provides a solution that extracts entities, attributes, and spatial relationships from scene description text through fine-grained parsing, generating structured semantic representation text to provide clear semantic priors for subsequent geometric modeling. Based on the structured semantic representation text, multiple target perspectives (such as forward view and top view) are intelligently planned using semantic segmentation and spatial geometry theory, providing rich and complementary geometric cues for 3D reconstruction. A text-generated graph model is used to generate consistent multi-view 2D images (single-view images corresponding to each target perspective) under both semantic and perspective constraints. Each single-view image not only retains the visual details of the scene description text but also implicitly contains cross-view geometric consistency. On this basis, through depth estimation and a 3D generation model, the 3D geometric model of each sub-object is accurately recovered from the multi-view images (single-view images corresponding to each target perspective). The pose of each sub-object 3D model is estimated and assembled into an initial scene, initially realizing the 3D mapping of text spatial relationships. A self-supervised layout optimizer is introduced to automatically adjust the pose parameters of the sub-object 3D geometric models, eliminating unreasonable layouts such as floating and overlapping, so that the final target 3D scene model is highly unified in terms of geometric structure, semantic alignment, and physical rationality.

[0020] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or a 3D scene model generation device that can achieve the above functions by fusing multi-view images. The following description uses a 3D scene model generation system that fuses multi-view images as an example to illustrate this embodiment and the subsequent embodiments.

[0021] Based on this, embodiments of this application provide a method for generating a 3D scene model by fusing multi-view images, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the 3D scene model generation method that integrates multi-view images according to this application.

[0022] In this embodiment, the method for generating a 3D scene model by fusing multi-view images includes steps 101-106: Step 101: Extract entities, attributes, and spatial relationships from the scene description text to generate structured semantic representation text.

[0023] Specifically, scene description text is natural language text input by the user, used to define the core elements of the 3D scene. For example, in a modern-style living room, a white wooden coffee table is placed in front of a gray sofa, near a floor-to-ceiling window. Entities are the core objects with specific forms and functions in the scene, and are the basic elements constituting the 3D scene, such as the living room, coffee table, sofa, and floor-to-ceiling window in the scene description text example above. Attributes are the inherent or additional characteristics of entities, including color, material, shape, quantity, etc., such as white, wood (attribute of the coffee table), gray (attribute of the sofa), modern style (attribute of the living room). Spatial relationships are the relative positions or orientations between entities, and are key information for constructing the spatial layout of the scene, such as relationships described by phrases like "placed in front of" or "near". Structured semantic representation text is semantic data presented in a standardized and ordered format (such as key-value pairs, tree structures, structured dictionaries, etc.) after the scene description text has been parsed and processed. It can clearly mark the correspondence between entities, attributes, and spatial relationships, making it convenient for subsequent algorithms to directly read and use.

[0024] In some embodiments, the user-input scene description text (such as "There is a red L-shaped sofa in the living room, facing the TV") can be segmented, part-of-speech tagged, and dependency parsed to identify object nouns in the scene description text as entities (such as sofa, TV). By combining a pre-trained language model and knowledge graph, adjectives or phrases modifying entities can be extracted as associated attributes, and prepositional structures or locative words can be parsed to determine spatial relationships (such as facing). The obtained entities, attributes, and spatial relationships are organized into a machine-readable structured data format, thereby providing accurate semantic constraints and prior guidance for subsequent multi-view planning, image generation, and 3D layout.

[0025] Step 102: Based on the structured semantic representation of the text, determine multiple target perspectives of the target 3D scene model through semantic segmentation algorithms and spatial geometry theory.

[0026] Specifically, spatial geometry theory utilizes geometric relationships in three-dimensional space, such as perspective (camera position and orientation), object projection range, and visibility (e.g., perspective projection principles and visibility criteria), to calculate the coverage and information integrity of entities from different perspectives. The target 3D scene model is the final 3D scene model to be generated, conforming to the text description. The core entities and spatial relationships of the target 3D scene model need to be accurately presented through multi-view images. Multiple target perspectives are a set of perspectives (e.g., front view, top view, side view, oblique view, etc.) that can relatively completely cover the key entities of the scene and clearly present the spatial relationships between entities, satisfying the principles of maximizing the visibility of key entities and maximizing the recognizability of spatial relationships.

[0027] In some embodiments, based on structured semantic representation text, a semantic segmentation algorithm can be used to virtually divide the spatial distribution area of ​​core entities in the scene according to the entities, attributes, and spatial relationships in the structured semantic representation text. Coverage priority weights are assigned based on the importance of entities (e.g., core furniture takes precedence over decorative objects). At the same time, based on spatial geometry theory, multiple sets of candidate viewpoints with different azimuth and pitch angles are preset, and the visibility score and projected area of ​​each entity under each set of candidate viewpoints are evaluated. By combining the object importance weights, visibility scores, and projected areas under each set of candidate viewpoints, an optimization algorithm (e.g., greedy selection or combination scoring) is used to select the candidate viewpoint combination that maximizes the expression of scene information, forming multiple target viewpoints of the target 3D scene model, providing accurate viewpoint parameters for subsequent image generation.

[0028] In addition, following engineering drawing standards, front views, top views, and side views can be directly selected as multiple target perspectives to display the core information of the scene from different dimensions. For complex scenes, such as large buildings or spaces with intricate internal structures, auxiliary perspectives such as oblique views and axonometric drawings can be added based on factors such as scene complexity and object distribution characteristics. Taking a living room scene as an example, when constructing multiple views, the front view is selected facing the main furniture placement area to ensure that the frontal layout of core objects such as sofas, coffee tables, and televisions is fully presented; the top view is taken perpendicular to the ground to clearly present the floor plan of the entire living room space, as well as the positional relationships and spatial proportions of each object.

[0029] Step 103: Based on the structured semantic representation of the text and the single-view information corresponding to each target viewpoint, generate single-view images corresponding to each target viewpoint through the text-generated image model.

[0030] Specifically, text-to-image generation models are diffusion-based models, such as the Stable Diffusion model and the third-generation text-to-image generation model (DALL·E3), which can generate realistic images based on text prompts. The single-view information corresponding to the target viewpoint includes specific parameters for each target viewpoint, which may include the camera position. Camera orientation Camera focal length A single-view image is a 2D image of a scene observed from the perspective of a single target. It needs to accurately represent the appearance, relative position, spatial relationship, shape, and size information of the entity from that target's perspective.

[0031] In some embodiments, for each target viewpoint, a text-based image generation model based on a diffusion model is used to generate 2D images. The core principle of the diffusion model is to gradually add noise to the data distribution, construct a Markov chain, and then gradually recover the original data from the noisy data by learning the inverse process. In the single-view image generation process, the structured semantic representation text and the single-view information corresponding to the target viewpoint are used as inputs to the text-based image model. The specific parameters of the target viewpoint may include the camera position. Camera orientation Camera focal length These parameters are obtained through the transformation matrix. Perform a projection transformation on the scene to obtain the view condition vector corresponding to the target viewpoint. Simultaneously, the structured semantic representation text is encoded into semantic vectors. semantic vector Viewpoint condition vector related to target viewpoint The data is then fused and input into the generator of the text-based image model. During the diffusion process, the initial noisy image is assumed to be... It follows a standard normal distribution. ,go through The final image is generated through a second iteration of denoising. This refers to the single-view image corresponding to the target's viewpoint. Each iteration can be represented as:

[0032] in, , This is a noise scheduling parameter used to control the amount of noise added in each iteration, typically increasing with the number of iterations. The increase, Gradually increase the size to ensure that the image is gradually restored from the noisy image to the true image; To follow a standard normal distribution The generator learns the noise through training. The mapping relationship between text description (structured semantic representation of text) and perspective information (single-view information corresponding to the target perspective) is achieved using neural networks. Predict the noise that needs to be removed in each iteration ,Right now .

[0033] During the generation process, the image quality and content accuracy can be optimized by adjusting the generator's control parameters (such as noise intensity, number of iterations, etc.), so that the generated 2D scene image can better meet the input text and perspective requirements.

[0034] By using structured semantic representation of text as content constraints and combining the geometric context of each target viewpoint (i.e., single-view information), the text-generated image model is driven to generate 2D images that are consistent with the target viewpoint, semantically accurate, and rich in detail. These are single-view images corresponding to the target viewpoint. Each single-view image not only reflects the content described in the text but also conforms to the perspective relationship and spatial layout under a specific target viewpoint, which can provide high-quality, multi-view aligned visual input for subsequent 3D reconstruction.

[0035] Step 104: Determine the category information and 3D position information of the sub-objects in each single-view image, and generate the corresponding 3D geometric model of the sub-object based on the category information and position information of the sub-objects and each single-view image through the 3D generation model.

[0036] Specifically, sub-objects are independent entity units that constitute a 3D scene, i.e., specific objects that can be distinguished in a single-view image, such as sofas, coffee tables, and lamps. The category information of a sub-object refers to the object category to which it belongs (e.g., sofa, lamp, coffee table), serving as the basic morphological and attribute characteristics used to determine the 3D geometric model of the sub-object. The 3D position information of a sub-object is its three-dimensional coordinate data (X, Y, Z axis coordinates) in the world coordinate system, used to accurately define the sub-object's specific position in 3D space. This 3D position information is crucial for constructing a reasonable spatial layout. The 3D generative model is a generative model capable of inferring 3D geometric shapes from 2D information, such as DreamFusion++. 3D generative models are typically based on Generative Adversarial Networks (GANs) or neural radiation field techniques, generating 3D models based on 2D image features and spatial information. The 3D geometric model of a sub-object is a three-dimensional digital representation of the sub-object with a realistic geometric topology, including information such as the object's outline, size, and surface details, serving as a core asset that can be directly used for scene assembly.

[0037] In some embodiments, instance segmentation can be performed on each single-view image to identify sub-objects in the single-view image and determine the category information of the sub-objects (such as sofa, coffee table). Simultaneously, based on known camera parameters (camera intrinsic and extrinsic matrices) and the single-view images, the position coordinates of the sub-objects in the 3D space of the world coordinate system are calculated to obtain the 3D position information of the sub-objects. Then, the category information of the sub-objects, the 3D position information of the sub-objects, and each single-view image are used as conditional inputs and fed into a 3D generative model. The 3D generative model can utilize multi-view appearance consistency constraints to reconstruct a high-fidelity 3D geometric model of the sub-objects, i.e., the sub-object 3D geometric model, and output a set of independent 3D assets with semantic labels and spatial locations, providing a structured and editable geometric foundation for subsequent pose estimation and scene assembly.

[0038] Step 105: Perform pose estimation on the 3D geometric model of the sub-object, obtain the pose parameters of the 3D geometric model of the sub-object in the 3D scene, assemble the 3D geometric model of the sub-object, and control the pose of the corresponding 3D geometric model of the sub-object in the 3D scene according to the pose parameters to generate the initial 3D scene model.

[0039] Specifically, the pose parameters are quantified data describing the spatial pose of the sub-objects, mainly including rotation angles (such as Euler angles and rotation matrices, defining the object's orientation) and translation vectors (such as 3D coordinate offsets, defining fine-tuning of the object's position). The initial 3D scene model is the preliminary 3D scene obtained after assembling the sub-objects, containing the geometric shape, position, and pose information of all sub-objects.

[0040] In some embodiments, a 6D pose regression network can be used to process the 3D geometric models of each sub-object. The 6D pose regression network can be a Point-View Network++ (PVNet++) enhanced version. It calculates the pose parameters of the sub-object by using feature matching and spatial transformation algorithms, which include rotation angles (such as Euler angles and rotation matrices around the X, Y, and Z axes) and translation vectors (3D coordinate offsets). Subsequently, the 3D engine can spatially adjust the 3D geometric models of each sub-object based on the pose parameters of the sub-object. The orientation of the sub-object's 3D geometric model can be adjusted by rotating the angle (e.g., rotating the sofa 180° to face the TV), and the 3D geometric model can be fine-tuned by translation vectors (e.g., translating the vase 20cm along the Y-axis to center it on the coffee table), so that the pose of each sub-object conforms to the functional logic (e.g., the legs of the table and chair are perpendicular to the ground). Finally, all the adjusted sub-object 3D geometric models are integrated into a unified world coordinate system according to their 3D position information, and assembled into an initial 3D scene model containing complete entities and preliminary spatial layout, providing a basic framework for subsequent layout optimization.

[0041] Step 106: Adjust the pose parameters of the 3D geometric models of sub-objects in the initial 3D scene model using a self-supervised layout optimizer to generate the target 3D scene model corresponding to the scene description text.

[0042] Specifically, the self-supervised layout optimizer is an algorithm module that can learn and optimize targets autonomously based on scene features (such as object collisions and spatial relationships) without the need for manually labeled supervised data. In this application, the core is to construct a self-supervised loss function and adjust the pose parameters of the 3D geometric model of the sub-object through iterative optimization.

[0043] In some embodiments, based on an initial 3D scene model, a self-supervised layout optimizer constructs a scene layout loss function, which may include sub-object overlap loss, space utilization loss, and distance loss between sub-objects. Subsequently, starting with the initial pose parameters (the pose parameters of the 3D geometric models of the sub-objects in the initial 3D scene model), a gradient descent algorithm is used for iterative optimization: the rotation angle of the sub-object 3D geometric models is fine-tuned according to the gradient of the scene layout loss function (e.g., adjusting the orientation of cabinet doors to avoid collisions) and the translation vector (e.g., moving suspended objects to reasonable positions). Simultaneously, multi-view image projection constraints can be combined to prevent deviation from the visual baseline. When the scene layout loss function is lower than a preset threshold (e.g., scene layout loss < 0.01), the optimization converges, outputting a target 3D scene model whose spatial relationships between sub-objects conform to physical logic and are consistent with the scene description text. Without relying on manual annotation or real 3D scene labels, by introducing physical rules, geometric constraints and semantic consistency as supervision signals, the pose parameters (position and orientation) of the 3D geometric models of each sub-object in the initial 3D scene model are automatically optimized, and unreasonable layouts (such as overlap, floating, and functional conflicts) are eliminated to generate a target 3D scene model that conforms to both text description and physical logic.

[0044] By introducing a physics engine for collision detection and gravity constraints, combined with a scene function reasoning module, problems such as object suspension and overlap in traditional methods can be solved. By adjusting the position of objects through a self-supervised layout optimizer, small objects (such as cups) are ensured to be stably placed on furniture surfaces, and the object overlap rate is reduced to below 2% (approximately 30% for LayoutGPT).

[0045] Optionally, the step of adjusting the pose parameters of the sub-object 3D geometry models in the initial 3D scene model using a self-supervised layout optimizer includes: Based on the pose parameters of the sub-object 3D geometric models in the initial 3D scene model, the sub-object overlap loss, space utilization loss and sub-object distance loss corresponding to the initial 3D scene model are calculated. The sub-object overlap loss is obtained by calculating the proportion of the intersection volume between sub-objects. The space utilization loss is used to evaluate the uniformity of space filling of the target 3D scene model. The sub-object distance loss is used to evaluate the reasonable spacing between sub-objects. A scene layout loss function is constructed based on sub-object overlap loss, space utilization loss, and distance loss between sub-objects. The pose parameters of the sub-object 3D geometric model are adjusted by optimizing the scene layout loss function.

[0046] Specifically, the self-supervised layout optimizer, based on a reinforcement learning framework, uses scene layout rationality as the optimization objective. Through continuous interaction and learning between the agent and the scene environment, it automatically adjusts the position and pose of objects. This is achieved by defining a scene layout loss function. Optimize, Taking into account factors such as sub-object overlap and space utilization, it can be represented as:

[0047] in, The sub-object overlap loss is measured by calculating the proportion of the intersection volume between sub-objects. The sub-object overlap loss can be achieved by detecting the geometric intersection volume between any two sub-object 3D geometric models and normalizing it to the proportion of the smaller object's volume, which is used to penalize unreasonable penetration. For space utilization loss, the space utilization loss divides the 3D scene model space into a voxel mesh. By calculating the variance or entropy value of the distribution of occupied voxels, the uniformity of the overall space filling is evaluated to avoid local overcrowding or large areas of emptiness. The distance loss between sub-objects can be calculated by comparing the actual center distance of related objects with a reasonable range based on the spatial relationships defined in the structured semantic representation, and using Huber or L2 loss to measure the deviation. Equal to the weighting coefficients, the degree of influence of each factor on layout optimization can be adjusted according to actual needs.

[0048] In optimizing the pose parameters of the 3D geometric models of sub-objects, the pose parameters (position and rotation) of each sub-object's 3D geometric model are used as optimization variables. Gradient descent is used for iterative updates to gradually reduce object overlap, balance spatial distribution, and correct object spacing. Finally, the optimized pose parameters are applied to each sub-object's 3D model to generate a target 3D scene model that is geometrically conflict-free, space-efficient, with spacing conforming to human cognitive habits, and strictly aligned with text semantics, thus generating a natural and reasonable scene layout.

[0049] Based on the 3D scene model generation method fused with multi-view images provided in this application, entities, attributes, and spatial relationships are extracted from scene description text to generate structured semantic representation text, accurately extracting entities and spatial relationships. Based on the structured semantic representation text, multiple target viewpoints of the target 3D scene model are determined through semantic segmentation algorithms and spatial geometry theory, ensuring that the subsequently generated 2D images cover key spatial information, providing sufficient geometric clues for 3D reconstruction and overcoming the limitation of insufficient single-view information. Based on the structured semantic representation text and the single-view information corresponding to each target viewpoint, single-view images corresponding to each target viewpoint are generated through a text-generated image model, providing rich visual evidence for 3D modeling. The category information and 3D position information of sub-objects in each single-view image are determined, and based on the category information and position information of the sub-objects and each single-view image, a 3D generation model is used to generate the corresponding 3D geometric model of the sub-object. The 2D visual information is transformed into a concrete 3D geometric model, and the pose of the sub-object 3D geometric model is estimated to obtain the pose parameters of the sub-object 3D geometric model in the 3D scene. The sub-object 3D geometric model is then assembled. This paper establishes the basic framework of the scene model and controls the pose of the corresponding sub-object 3D geometric models in the 3D scene according to the pose parameters to generate an initial 3D scene model. The pose parameters of the sub-object 3D geometric models in the initial 3D scene model are adjusted by a self-supervised layout optimizer to generate a target 3D scene model that conforms to the scene description text description. By integrating NLP, computer vision and 3D generation technologies, semantic constraints and multi-view geometric verification are deeply integrated to solve the technical problems of object layout deviation and semantic deviation in the existing text-based 3D scene models. This achieves the generation of 3D scene models that are semantically accurate, geometrically reasonable and physically consistent. The 3D scene model generation method fused with multi-view images provided in this application extracts entities, attributes, and spatial relationships by performing fine-grained analysis on scene description text, generating structured semantic representation text, which provides clear semantic priors for subsequent geometric modeling; based on the structured semantic representation text, it intelligently plans multiple target perspectives (such as forward view and top view) by combining semantic segmentation and spatial geometry theory, providing rich and complementary geometric clues for 3D reconstruction; and uses a text-generated graph model to generate consistent multi-view 2D images (single-view images corresponding to each target perspective) under the dual constraints of semantics and perspective. Each single-view image not only retains the visual details of the scene description text, but also implicitly contains cross-view geometric consistency.Based on this, through depth estimation and 3D generation models, the 3D geometric model of each sub-object is accurately recovered from multi-view images (single-view images corresponding to each target viewpoint); the pose of each sub-object 3D model is estimated and assembled into an initial scene, thus initially realizing the 3D mapping of text spatial relationships; and a self-supervised layout optimizer is introduced to automatically adjust the pose parameters of the sub-object 3D geometric model, eliminating unreasonable layouts such as floating and overlapping, so that the final target 3D scene model is highly unified in terms of geometric structure, semantic alignment and physical rationality.

[0050] In some embodiments, the step of extracting entities, attributes, and spatial relationships from scene description text to generate structured semantic representation text includes: The scene description text is segmented using a word segmentation function to obtain the corresponding word segmentation results; The word segmentation results are tagged with part-of-speech tags using a part-of-speech tagging function to obtain the corresponding tagging results; Dependency parsing is performed on the annotation results using a syntactic analysis function to construct a graph of grammatical dependency relationships between the corresponding words. Based on the grammatical dependency graph, the set of sub-object entities, the set of sub-object attributes, and the set of spatial relationships between sub-objects are extracted from the grammatical dependency graph through a large language model; Query the set of associated attributes of each sub-object in the knowledge graph within the sub-object entity set, add the set of associated attributes to the attribute set of the sub-object, and obtain the expanded attribute set of the sub-object. The set of sub-object entities, the set of attributes of the expanded sub-objects, and the set of spatial relationships are defined as the structured semantic representation text.

[0051] Specifically, word segmentation functions are tools that cut continuous scene description text into independent semantic units (words), such as the word segmentation interfaces of Jieba and NLTK. Part-of-speech tagging functions are tools that assign part-of-speech tags (such as nouns and adjectives) to the word segmentation results, such as the part-of-speech tagging module of SpaCy. Syntactic analysis functions are tools that parse the grammatical relationships (such as subject-verb and verb-object relationships) between words, such as Stanford Parser. Grammatical dependency graphs are graphical representations that intuitively present the dependency logic between words using nodes (words) and edges (grammatical relationships). Large language models: pre-trained models with context understanding capabilities, such as the GPT series and Bidirectional Encoder Representations from Transformers (BERT), are used to extract semantic information from grammatical relationships. Knowledge graphs are structured databases that store entities and associated attributes, such as WordNet and ConceptNet, used to supplement the attributes of sub-objects.

[0052] As an example, in the word segmentation stage, based on linguistic morphology rules, the continuous scene description text can be divided into the smallest semantic units using a word segmentation function. For instance, the scene description text "There is a red sofa in the living room" can be precisely broken down into words such as living room, have, one, red, and sofa, which is the word segmentation result mentioned above. In the part-of-speech tagging stage, each word is assigned a corresponding part-of-speech label using a part-of-speech tagging function, such as noun (NN), verb (VB), adjective (JJ), etc., to obtain the corresponding tagging results. In the syntactic analysis process, dependency parsing algorithms are used to construct a complete syntax tree structure, clarifying the grammatical dependency relationships between words. For example, "living room" as the subject forms a subject-verb relationship with the predicate verb "have" (root), and "red sofa" as the object forms a verb-object structure with "have". Through systematic text deconstruction, a solid grammatical foundation can be provided for subsequent semantic understanding. In mathematical expression, let the scene description text be... ,in Indicates the first Each vocabulary unit; through word segmentation function Obtain word segmentation results Part-of-speech tagging function The word segmentation results are labeled to obtain ,in for The corresponding part-of-speech tags; then the syntactic analysis function is used. Constructing a syntax dependency graph ,in Represents the structure of the grammatical dependency graph.

[0053] Next, leveraging the contextual understanding capabilities of the large language model, the set of sub-object entities, sub-object attributes, and spatial relationships between sub-objects are extracted from the grammatical dependency graph. This identifies key entities describing the scene in the text, such as sub-object entities like "sofa" and "coffee table," sub-object attributes like "red" and "blue," and spatial relationships like "in..." and "near." Each extracted sub-object entity is then matched with object entities in the knowledge graph to obtain the associated attribute set of each sub-object entity. For example, the knowledge graph reveals that sofas typically have seats and backrests, common sofa shapes include L-shapes and U-shapes, and sizes fall within a certain range. This enriches and expands the semantic information of each sub-object entity, not only transforming the text from grammatical structure to semantic content but also providing detailed semantic basis for subsequent scene construction. In terms of mathematical description, let the set of sub-object entities extracted by the large language model be... The sub-object property set is The set of spatial relationships is The semantic parsing algorithm for large language models is... The input text is The semantic representation obtained after semantic parsing is: Meanwhile, let the knowledge graph be... Sub-object entity The set of attributes associated in a knowledge graph is Through knowledge graph matching function It can associate entities with attributes in a knowledge graph, that is... This expands the attribute information of the entity, adds the set of associated attributes to the attribute set of the sub-object, and obtains the expanded attribute set of the sub-object; and determines the sub-object entity set, the expanded attribute set of the sub-object, and the spatial relationship set as the structured semantic representation text.

[0054] In some embodiments, the step of determining multiple target perspectives of a target 3D scene model based on structured semantic representation text using semantic segmentation algorithms and spatial geometry theory includes: The semantic segmentation algorithm identifies the set of sub-object entities, the set of spatial relationships between sub-objects, and the set of attributes of sub-objects in the structured semantic representation text. For each sub-object in the sub-object entity collection, obtain the typical shape parameters corresponding to the sub-object, and modify the typical shape parameters through the attributes of the sub-object to obtain the initial shape parameters corresponding to the sub-object; Based on the set of spatial relationships between sub-objects and the rules for setting the origin of the scene coordinate system, the relative positions of each sub-object in three-dimensional space are logically deduced to generate the initial position coordinates of each sub-object in the 3D scene. For each viewpoint, based on the initial shape parameters and initial position coordinates of each sub-object, calculate the projected area and visibility ratio of each sub-object under the viewpoint, and determine the importance weight of each sub-object. Based on the projected area, visibility ratio, and importance weight, an objective function for viewpoint selection is constructed. The objective function is optimized by selecting the combination of perspectives that maximizes the objective function value from multiple perspectives as multiple target perspectives for the target 3D scene model.

[0055] Specifically, typical shape parameters are quantified data of common shapes of sub-objects (e.g., the default length × width × height of a regular sofa is 2.2m × 0.9m × 0.8m). Typical shape parameters are preset based on the statistical characteristics of similar objects. Initial shape parameters are shape parameters after being modified by the attributes of the sub-objects; for example, the typical length of a small sofa is reduced to 1.8m.

[0056] As an example, a semantic segmentation algorithm is used to parse the structured semantic representation text, accurately identifying the set of sub-object entities, the set of spatial relationships between sub-objects, and the set of attributes of sub-objects within the structured semantic representation text, such as sofa: {color=gray, size=small}, coffee table: {shape=round}. Pre-defined typical shape parameters are matched to each sub-object, and then corrected based on the sub-object's attributes to obtain the initial shape parameters for each sub-object. Subsequently, according to the rules for setting the origin of the scene coordinate system, such as taking the bottom left corner of the room as the origin, with the X-axis along the length direction, the Y-axis along the width direction, and the Z-axis along the height direction, the initial position coordinates are deduced based on spatial relationships, generating the initial position coordinates of each sub-object in the 3D scene. If there are... Each sub-object, The set of position coordinates of each sub-object in three-dimensional space is: Each sub-object With a set of shape parameters When determining the target viewpoint, an objective function for viewpoint selection can be constructed by calculating indicators such as the projected area and visibility ratio of sub-objects under different viewpoint combinations. :

[0057] in, For child objects The importance weight can be set based on factors such as the function and area of ​​the sub-object in the scene; Represents sub-objects Visibility from a specific perspective; if fully visible... Partially visible values ​​are determined by the proportion of visibility; invisible values ​​are determined by... ; For child objects The projected area from this viewpoint. The objective function selected based on the viewpoint. The method selects the viewpoint combination that maximizes the objective function value as multiple target viewpoints for the target 3D scene model. Through viewpoint selection that combines semantics and geometry, the selected multiple target viewpoints can fully cover the core sub-objects, clearly present the sub-object shapes and spatial relationships, provide accurate parameters for subsequent multi-view image generation, avoid blind spots in single-view information, and lay a visual foundation for the accurate generation of 3D scenes.

[0058] In some embodiments, the step of generating single-view images corresponding to each target viewpoint using a text-generated image model based on structured semantic representation text and single-view information corresponding to each target viewpoint includes: Text encoding is performed on the structured semantic representation text to obtain the corresponding semantic vector; Obtain camera parameters and determine the view parameters corresponding to each target view. Then, concatenate the view parameters corresponding to each target view with the camera parameters to obtain the single view information corresponding to each target view. By performing projection transformation on each single-view information, the view condition vector corresponding to each target view is obtained. The viewpoint condition vectors corresponding to each target viewpoint are input into the U-Net network of the text-based image model along with the semantic vectors. The U-Net network of the text-based image model predicts noise and gradually removes noise to generate single-view images corresponding to each target viewpoint.

[0059] Specifically, for each target viewpoint, a text-based image model based on a diffusion model is used to generate 2D images. During single-view image generation, structured semantic representation text, camera parameters, and viewpoint parameters corresponding to the target viewpoint are used as inputs to the text-based image model. Camera parameters may include the camera position C. Camera focal length The viewpoint parameters corresponding to the target viewpoint can be... By concatenating the viewpoint parameters corresponding to each target viewpoint with the camera parameters, we can obtain the single-viewpoint information for each target viewpoint. This single-viewpoint information is then processed using a transformation matrix. Perform a projection transformation on the scene to obtain the view condition vector corresponding to the target viewpoint. Simultaneously, the structured semantic representation text is encoded into semantic vectors. semantic vector Viewpoint condition vector related to target viewpoint The data is then fused and input into the generator of the text-based image model. During the diffusion process, the initial noisy image is assumed to be... It follows a standard normal distribution. ,go through The final image is generated through a second iteration of denoising. That is, the single-view image corresponding to the target's viewpoint.

[0060] During training, the mean square error loss between predicted noise and actual noise is minimized. Optimize the generator parameters of the text graph model :

[0061] Optionally, during training, to ensure the coherence and accuracy between multi-view images (single-view images corresponding to each target viewpoint), a viewfinder is introduced. Figure 1 Consistency loss function Constraint optimization is performed on 2D images generated from different viewpoints. For feature matching, a feature extraction network from deep learning, such as a Convolutional Neural Network (CNN), is used to extract features of sub-objects from 2D images from different viewpoints, including color, shape, and position. Let the set of 2D images from different viewpoints be denoted as . Through feature extraction function Extract feature vectors from each 2D image , .See Figure 1 Sexual damage The formula used to measure the degree of difference in features of the same sub-object in 2D images generated from different viewpoints is as follows:

[0062] in, The feature distance metric function can be Euclidean distance. ,in and for 3D feature vector; It can also be the cosine distance. This is used to measure the similarity between features of the same sub-objects in 2D images generated from different viewpoints. During the optimization process, the viewpoint... Figure 1 Sexual damage With mean squared error loss (as mentioned above) These factors, when combined, constitute the overall loss function. :

[0063] in, These are hyperparameters used to balance the quality of single-view generation with the view itself. Figure 1 Weights between consistency values. The total loss function is then calculated using the backpropagation algorithm. The generator of the palindrome image model is passed and its parameters are adjusted to ensure that images from different perspectives maintain consistency in the features of the same sub-objects, thereby improving the coherence and accuracy of multi-view images and providing a reliable 2D image foundation for subsequent 3D scene construction.

[0064] In some embodiments, the step of determining the category information and location information of sub-objects in each single-view image includes: Extract a set of sub-object entities from the structured semantic representation text; the set of sub-object entities contains multiple sub-object entities. Each single-view image is input into an attention-based segmentation model. The segmentation model extracts features from the single-view images to obtain feature maps corresponding to each single-view image. Each feature map is input into the segmentation head of the segmentation model. Based on the segmentation head, the feature maps are processed to obtain the probability distribution matrix corresponding to each feature map containing the pixel points belonging to different object categories. Based on a preset probability threshold, each probability distribution matrix is ​​converted into a corresponding binary mask matrix; Based on a binary mask matrix and a single-view image, the local feature vector of the sub-object corresponding to the sub-object in the single-view image is extracted by an object recognition algorithm. Based on the local feature vector of the sub-object, the various types of objects in the object category feature library are matched to obtain the initial category of the sub-object. The object category feature library contains multiple types of objects and reference feature vectors of each type of object. The initial category of the sub-object is matched with the sub-object entity by calculating semantic similarity to determine the category information of the sub-object; Each single-view image is input into a multi-scale depth fusion model. The multi-scale depth fusion model is used to estimate the depth of each single-view image to obtain the depth map corresponding to each single-view image. For each pixel in a single-view image, the 3D coordinates of the pixel in the camera coordinate system are calculated based on the camera intrinsic parameter matrix, the depth value of the pixel in the depth map, and the image coordinates of the pixel. The 3D coordinates of the pixel in the camera coordinate system are then converted into global coordinates in the world coordinate system using the camera extrinsic parameter matrix. The 3D position information of the sub-object is determined based on the global coordinates of each pixel in the world coordinate system.

[0065] Specifically, the object category feature library is a database storing reference feature vectors for objects of various categories (such as sofas and coffee tables) for category matching. Multi-scale depth fusion models: These are models that fuse multi-scale image features to estimate depth, such as Metric3D, which can generate high-precision depth maps. A depth map is an image showing the scene depth (distance from the object to the camera) for each pixel.

[0066] As an example, a set of sub-object entities is extracted from the structured semantic representation of the text, serving as the benchmark for subsequent category matching. For each single-view image (e.g., front view, top view), the single-view image input is introduced into an attention-based segmentation model, such as Attention U-Net. The segmentation model focuses on regions in the single-view image that may contain sub-objects (e.g., areas with clearly defined furniture outlines) through a self-attention module. This effectively captures long-distance dependencies in the image, thereby enhancing the ability to recognize object boundaries in complex scenes, suppressing background interference, and performing multi-scale feature extraction on the single-view image, outputting feature maps containing information such as texture, edges, and semantics. ,in , , These represent the height, width, and number of channels of the feature map, respectively.

[0067] Then, the feature map The input segmentation model's segmentation head (which can consist of convolutional layers and a softmax activation function) is used to analyze the feature map. Perform pixel-level classification processing and output the probability distribution matrix of each pixel belonging to different object categories. , The probability distribution matrix represents the total number of object categories. Each element represents the probability value (range 0-1) of a corresponding pixel belonging to a different object category (e.g., sofa, coffee table, background). This probability is based on a preset probability threshold. Convert each probability distribution matrix into a corresponding binary mask matrix. ,Right now .

[0068] A binary mask matrix can accurately define the pixel range of a sub-object in an image (such as the image area occupied by a sofa). In the sub-object category recognition stage, an object recognition algorithm based on a convolutional neural network (CNN) is used to segment the binary mask matrix. With single-view images Combine and extract local feature vectors of sub-objects (Includes information such as shape, texture, and color distribution).

[0069] The extracted local feature vectors of sub-objects are matched with a pre-trained object category feature library. This library pre-stores reference feature vectors for multiple object categories (such as furniture and appliances) (obtained through training with a large number of samples). Feature similarity is calculated. ,in As a reference feature for a certain type of object in the feature library, when When the threshold is exceeded, the object category (such as sofa, coffee table, etc.) can be identified, and the initial category of the sub-object can be obtained. The identification result (the initial category of the sub-object) is then matched and verified with the sub-object entities extracted from text parsing; semantic similarity can be calculated. (This can be calculated using models such as Word2Vec or BERT), for example, when semantic similarity... When the value is ≥0.8, the initial category is confirmed as the final category information of the sub-object to ensure the accuracy of the recognition results.

[0070] Simultaneously, to determine the location information of sub-objects, the single-view image is input into a multi-scale depth fusion model. This model improves the accuracy of depth estimation by fusing image features from different scales. The input to the multi-scale depth fusion model is the single-view image. After a series of convolutional layers, pooling layers, and upsampling layers in the multi-scale depth fusion model, the depth map corresponding to the single-view image is output. The objective function for depth estimation in a multi-scale deep fusion model employs smoothing. Loss function:

[0071] in, The total number of pixels in the image. The depth values ​​predicted by the multi-scale depth fusion model. This is the actual depth value. The function is defined as:

[0072] After acquiring the depth information of each pixel in the single-view image, the image coordinates in the 2D image (i.e., the single-view image) are transformed into 3D coordinates in the camera coordinate system by combining camera parameters and multi-view geometry principles. Let the camera intrinsic parameter matrix be... ,in For the camera Focal length of direction, The coordinates of the image center are; the extrinsic parameter matrix is... , for The rotation matrix, for The translation vector. The 2D image coordinates are... Depth value The corresponding 3D coordinates in the camera coordinate system It can be calculated using the following formula:

[0073] Then, the 3D coordinates of the pixels in the camera coordinate system are converted to global coordinates in the world coordinate system using the camera extrinsic parameter matrix.

[0074] The 3D position information of a sub-object in the world coordinate system can be obtained by averaging the global coordinates of all pixels of the same sub-object in the world coordinate system. Through attention segmentation and semantic matching, it is ensured that the sub-object category is strictly consistent with the entity described in the text, avoiding misclassification. At the same time, through depth estimation and coordinate transformation, the pixel information in the 2D image (single-view image) is accurately mapped to the 3D spatial position, providing dual constraints of category and space for the subsequent generation of the 3D geometric model of the sub-object and scene assembly. It is the core link connecting 2D visual information and 3D scene construction, and can be used to ensure the entity accuracy and spatial layout rationality of the final 3D scene.

[0075] In some embodiments, the step of generating a 3D geometric model of the sub-object corresponding to the sub-object using a 3D generative model based on the category information and location information of the sub-object and each single-view image includes: The category information of the sub-object, the location information of the sub-object, and the single-view images are input into the 3D generation model based on the generative adversarial network. The 3D generation model generates a 3D point cloud model of the sub-object that matches the input conditions. A quadratic error metric algorithm is used to simplify the mesh of the 3D point cloud model of the sub-object. Vertices in the mesh of the initial 3D geometric model of the sub-object are iteratively merged to obtain the corresponding simplified 3D geometric model of the sub-object. The simplified 3D geometric model of the sub-object is smoothed using the Laplacian smoothing algorithm to obtain the corresponding smoothed 3D geometric model of the sub-object. Extract the attribute set of sub-objects from the structured semantic representation text; the attribute set of each sub-object contains the attributes of each sub-object. Determine the 2D image of the sub-object corresponding to the sub-object based on the single-view image, and establish the texture mapping relationship between the 2D image of the sub-object and the 3D geometric model of the sub-object; Based on the attributes, position information, and texture mapping relationship of the sub-object, the color and texture information in the 2D image of the sub-object is mapped to each facet of the corresponding 3D geometric model of the sub-object through a material mapping algorithm, thus obtaining the 3D geometric model of the sub-object.

[0076] Specifically, the 3D generative model is a model built based on Generative Adversarial Networks (GANs). It can generate corresponding 3D point cloud models based on input data (such as the category information, location information, and single-view images of sub-objects). The 3D generative model includes a generator and a discriminator. Based on the idea of ​​GANs, the 3D generative model uses adversarial training to make the generated 3D models as realistic as possible. The sub-object 3D point cloud model is a digital representation of the sub-object composed of a large number of three-dimensional coordinate points, which can initially reflect the geometric shape of the sub-object.

[0077] As an example, during the generation process, the generator takes the category information of the sub-object, the location information of the sub-object, and various single-view images as input, and uses a 3D generative model to generate a 3D point cloud model of the sub-object that matches the input conditions. , The number of point clouds is [not specified]. To improve the quality and rendering efficiency of the sub-object 3D model, the generated sub-object 3D point cloud model needs to be optimized. For mesh simplification, the Quadratic Error Metrics (QEM) algorithm can be used to calculate the error matrix of each vertex (reflecting the impact of vertex deletion on the model surface). Iterative merging of vertex pairs with the smallest errors (such as vertices that are close together and have little impact on the overall shape) reduces the number of faces, preserving key structures while reducing model complexity. By iteratively merging vertices in the mesh, the number of meshes is reduced while maintaining the basic shape of the model, resulting in the simplified sub-object 3D geometric model. For smoothing, the Laplacian smoothing algorithm is used, adjusting vertex positions to make the model surface smoother. The specific Laplacian smoothing formula is:

[0078] in, As vertices The original location, The position after smoothing. For smoothing coefficients, As vertices The set of adjacent vertices, The number of adjacent vertices is used. For each vertex of the simplified sub-object 3D geometric model mesh, the average coordinates of its neighboring vertices are calculated. This vertex is then fine-tuned towards the average coordinates to eliminate the sharp edges caused by triangulation, resulting in a smooth simplified sub-object 3D geometric model. Simultaneously, the set of sub-object attributes is extracted from the structured semantic representation text. A 2D image of the sub-object is cropped from a single-view image, and a texture mapping relationship between 2D image pixels and 3D model surface patches is established using a UV mapping algorithm. This is used to define the specific region of the 2D image corresponding to each facet. Based on the sub-object's attributes, position information, and texture mapping relationships, corresponding material textures are assigned to the smoothed 3D geometric model of the sub-object. Through a material mapping algorithm, material features in the 2D image are mapped to the surface of the 3D model. According to the texture mapping relationships, the color and texture information in the 2D image are mapped to each triangular facet of the 3D model, thus assigning material textures and obtaining the 3D geometric model of all sub-objects. This is achieved through multi-view... Figure 2 2D image information is converted into specific 3D assets, providing basic elements for scene assembly. At the same time, accurate sub-object segmentation, depth estimation, and 3D model generation ensure the consistency between 3D assets, text descriptions, and 2D images. The generated 3D assets (sub-object 3D geometric models) will be used for the final scene construction during the scene assembly stage.

[0079] In some embodiments, the step of performing pose estimation on the 3D geometric model of the sub-object and obtaining the pose parameters of the 3D geometric model of the sub-object in a 3D scene includes: The 2D image of the sub-object, the 3D point cloud model of the sub-object, and the 3D position information of the sub-object are input into the 6D pose regression network. The pose of the 3D geometric model of the sub-object is estimated based on the 6D pose regression network, and the pose parameters of the 3D geometric model of the sub-object in the 3D scene are obtained. The pose parameters include rotation angle and translation vector.

[0080] Specifically, the 6D pose regression network is a deep learning network that can directly predict the 6-DOF pose of an object, such as the pose diffusion model PoseDiffusion. "6D" refers to 3 rotational degrees of freedom (describing orientation) and 3 translational degrees of freedom.

[0081] As an example, a 2D image of a sub-object (providing appearance texture and 2D contour), a 3D point cloud model of the sub-object, and the 3D position information of the sub-object are input into a 6D pose regression network. The 6D pose regression network uses a convolutional neural network (CNN) to extract visual features (such as high-dimensional vectors of edges and textures) from the 2D image of the sub-object, and a point cloud feature extractor to extract geometric features (such as distances and angular relationships between vertices) from the 3D point cloud model of the sub-object. It also encodes the 3D position information of the sub-object into spatial feature vectors. Subsequently, the feature fusion module of the 6D pose regression network (including an attention mechanism) fuses the visual, geometric, and spatial features, focusing on key correlation information. The fused features are then input into the regression head to predict the 6D pose parameters of the sub-object's 3D geometric model in the 3D scene. Specifically, for a set of sub-object 3D geometric models... 3D geometric model of each sub-object attitude parameters middle, This represents the rotation angle about the x, y, and z axes. Let be the translation vector in three-dimensional space.

[0082] Optionally, during the training of the 6D pose regression network, the network takes a large amount of 3D object data with labeled poses as input and minimizes the mean squared error (MSE) loss function between the predicted pose and the true pose. The optimization formula is as follows:

[0083] in, These are the pose parameters predicted by the network. These are the actual attitude parameters.

[0084] In some embodiments, after the step of generating the target 3D scene model corresponding to the scene description text, the method further includes: Style features are extracted from single-view images corresponding to each target viewpoint using a convolutional neural network to obtain the corresponding style feature vectors. The 3D geometric model of the sub-object is digitally analyzed to obtain the corresponding material feature vector; The style feature vector, material feature vector, and sub-object 3D geometric model are input into a style transfer network based on generative adversarial network. The output is a style-unified sub-object 3D geometric model, which is then assembled to obtain the second-stage target 3D scene model. Based on the second-stage target 3D scene model, the lighting re-rendering algorithm is used to re-render the second-stage target 3D scene model to obtain the third-stage target 3D scene model. Shadow mapping technology is used to generate shadows on the third-stage target 3D scene model, resulting in a fourth-stage target 3D scene model that includes shadow effects that conform to physical laws.

[0085] Specifically, the material feature vector is the surface attribute encoding parsed from the 3D geometric model and texture of the sub-object, including roughness, metallicity, albedo, etc., which can be obtained through a material classification network or physical property estimation module.

[0086] Style transfer networks are models built on generative adversarial networks that can inject 2D style features into the textures and shading of 3D scenes to achieve cross-view style unification.

[0087] The second-stage target 3D scene model is a 3D scene model that has undergone style transfer and has a unified visual style, but whose lighting has not yet been optimized. The third-stage target 3D scene model is a 3D scene model with unified lighting after re-rendering, in which the lighting direction, intensity, and color temperature are consistent. The fourth-stage target 3D scene model is the final output 3D scene model, which has geometric correctness, stylistic consistency, reasonable lighting, and physically based shadows.

[0088] As an example, multi-view stereo vision technology is used to acquire single-view images corresponding to each target viewpoint. A convolutional neural network is then used to extract color and texture style features from these single-view images, resulting in corresponding style feature vectors. ,in Representing the Each sub-object has several feature dimensions. Simultaneously, for each sub-object's 3D geometric model, its texture map or surface normal / roughness map is analyzed. The corresponding material feature vector can be extracted using a lightweight material classification network or rule engine. Subsequently, the style feature vector, material feature vector, and target 3D scene model are input into the style transfer network. The style transfer network is based on a generative adversarial network (GAN) architecture, consisting of a generator... and discriminator Composition. Generator Material feature vectors of 3D assets and style feature vector As input, the output is the material feature with a unified style. ,Right now Discriminator Then it is responsible for distinguishing the generated material features By using material features that realistically match the target style, and through adversarial training, the generator can produce 3D assets (sub-object 3D geometric models) that visually match the target style. After assembly, a second-stage target 3D scene model with a unified style is obtained. Following style unification, post-processing operations are performed on the assembled second-stage target 3D scene model, including lighting settings and shadow generation. Lighting settings employ Physically Based Rendering (PBR) technology, simulating the real-world light propagation process and calculating the interaction between light and object surfaces, including diffuse reflection and specular reflection. Its lighting calculation model can be represented as:

[0089] in, For the final light intensity, and These are the diffuse reflection and specular reflection coefficients, respectively. and For diffuse and specular reflection light intensity, The surface normal vector, Let be the direction vector of the light ray. The view direction vector. The reflection direction vector, This refers to the glossiness index. Based on the second-stage target 3D scene model, a relighting algorithm is used to re-render the second-stage target 3D scene model, that is, to set the lighting for the second-stage target 3D scene model, thus obtaining the third-stage target 3D scene model.

[0090] Shadow generation employs shadow mapping technology. This involves rendering a depth map of the scene from the light source's perspective, then comparing it to a scene rendered from the camera's perspective to determine if objects are in shadow, thus generating realistic shadow effects. Shadow mapping is then used to generate shadows on the third-stage target 3D scene model, resulting in a fourth-stage target 3D scene model with physically accurate shadow effects. After these processes, a complete 3D scene model that matches the input scene description is finally generated. This model supports exporting wave front object files in multiple formats, including OBJ, FBX, and GLTF, facilitating its use in various 3D application scenarios.

[0091] In the scene assembly stage, the independently generated 3D assets are integrated into a complete 3D scene. Through operations such as pose estimation, layout optimization, and style unification, the final 3D scene model is made to meet the requirements of text description and visual aesthetics. The final output 3D scene model, as the result of the 3D scene model generation method of fusion of multi-view images provided in this application, can be directly applied to downstream fields to complete the entire process from text input to 3D scene output.

[0092] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the method for generating 3D scene models by fusing multi-view images in this application. Any simple transformations based on this technical concept are within the protection scope of this application.

[0093] This application also provides a 3D scene model generation device that integrates multi-view images; please refer to [reference needed]. Figure 2 The 3D scene model generation device that integrates multi-view images includes: The text parsing module 201 is used to extract entities, attributes and spatial relationships from the scene description text and generate structured semantic representation text; The multi-view planning module 202 is used to determine multiple target views of the target 3D scene model based on structured semantic representation text, through semantic segmentation algorithms and spatial geometry theory. The image generation module 203 is used to generate single-view images corresponding to each target viewpoint based on the structured semantic representation text and the single-view information corresponding to each target viewpoint through the text-generated image model. The 3D model generation module 204 is used to determine the category information and 3D position information of the sub-objects in each single-view image, and generate the corresponding 3D geometric model of the sub-object based on the category information and position information of the sub-objects and each single-view image through the 3D generation model. The pose estimation module 205 is used to estimate the pose of the 3D geometric model of the sub-object, obtain the pose parameters of the 3D geometric model of the sub-object in the 3D scene, assemble the 3D geometric model of the sub-object, and control the pose of the corresponding 3D geometric model of the sub-object in the 3D scene according to the pose parameters to generate an initial 3D scene model. The layout optimization module 206 is used to adjust the pose parameters of the 3D geometric models of sub-objects in the initial 3D scene model through a self-supervised layout optimizer, and generate the target 3D scene model corresponding to the scene description text.

[0094] The 3D scene model generation apparatus for fusing multi-view images provided in this application, employing the 3D scene model generation method for fusing multi-view images in the above embodiments, can solve the technical problems of object layout deviation and semantic deviation in text-based 3D scene models in the prior art. Compared with the prior art, the beneficial effects of the 3D scene model generation apparatus for fusing multi-view images provided in this application are the same as those of the 3D scene model generation method for fusing multi-view images provided in the above embodiments, and other technical features in the 3D scene model generation apparatus for fusing multi-view images are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0095] This application provides a 3D scene model generation device that integrates multi-view images. The 3D scene model generation device that integrates multi-view images includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the 3D scene model generation method that integrates multi-view images in the above embodiment 1.

[0096] The following is for reference. Figure 3 This document illustrates a structural schematic diagram of a 3D scene model generation device suitable for implementing embodiments of this application, which integrates multi-view images. The 3D scene model generation device integrating multi-view images in embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (PADs), portable media players (PMPs), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 3 The 3D scene model generation device that integrates multi-view images shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0097] like Figure 3As shown, the 3D scene model generation device that integrates multi-view images may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 1002 or a program loaded from storage device 1003 into random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the 3D scene model generation device that integrates multi-view images. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. Communication device 1009 allows the 3D scene model generation device fusing multi-view images to wirelessly or wiredly communicate with other devices to exchange data. Although the figure shows a 3D scene model generation device fusing multi-view images with various systems, it should be understood that it is not required to implement or possess all the systems shown. More or fewer systems can be implemented alternatively.

[0098] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0099] The 3D scene model generation device integrating multi-view images provided in this application, employing the 3D scene model generation method integrating multi-view images in the above embodiments, can solve the technical problems of object layout deviation and semantic deviation in text-based 3D scene models in the prior art. Compared with the prior art, the beneficial effects of the 3D scene model generation device integrating multi-view images provided in this application are the same as those of the 3D scene model generation method integrating multi-view images provided in the above embodiments, and other technical features in this 3D scene model generation device integrating multi-view images are the same as those disclosed in the previous embodiment method, and will not be repeated here.

[0100] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0101] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0102] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the 3D scene model generation method for fusing multi-view images in the above embodiments.

[0103] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0104] The aforementioned computer-readable storage medium may be included in a 3D scene model generation device that integrates multi-view images; or it may exist independently and not be assembled into a 3D scene model generation device that integrates multi-view images.

[0105] The aforementioned computer-readable storage medium carries one or more programs that, when executed by a 3D scene model generation device that fuses multi-view images, cause the 3D scene model generation device to: extract entities, attributes, and spatial relationships from scene description text to generate structured semantic representation text; determine multiple target views of the target 3D scene model based on the structured semantic representation text using semantic segmentation algorithms and spatial geometry theory; generate single-view images corresponding to each target view based on the structured semantic representation text and the single-view information corresponding to each target view using a text-generated image model; and determine the... The system obtains the category and 3D position information of the sub-objects, and based on the category and position information of the sub-objects and each single-view image, generates the corresponding 3D geometric model of the sub-object through a 3D generative model. It then performs pose estimation on the 3D geometric model of the sub-object to obtain its pose parameters in the 3D scene, assembles the 3D geometric model of the sub-object, and controls the pose of the corresponding 3D geometric model of the sub-object in the 3D scene according to the pose parameters to generate an initial 3D scene model. Finally, it adjusts the pose parameters of the 3D geometric model of the sub-object in the initial 3D scene model through a self-supervised layout optimizer to generate the target 3D scene model corresponding to the scene description text.

[0106] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0107] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0108] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0109] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described method for generating 3D scene models by fusing multi-view images. This solves the technical problems of object layout and semantic deviations in text-based 3D scene models in the prior art. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the 3D scene model generation method for fusing multi-view images provided in the above embodiments, and will not be repeated here.

[0110] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described method for generating a 3D scene model by fusing multi-view images.

[0111] The computer program product provided in this application can solve the technical problems of object layout deviation and semantic deviation in text-based 3D scene models in the prior art. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the 3D scene model generation method that fuses multi-view images provided in the above embodiments, and will not be repeated here.

[0112] The above are only some embodiments of this application and do not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A method for generating 3D scene models by fusing multi-view images, characterized in that, The method for generating 3D scene models by fusing multi-view images includes: Extract entities, attributes, and spatial relationships from scene description text to generate structured semantic representation text; Based on the structured semantic representation of the text, multiple target perspectives of the target 3D scene model are determined through semantic segmentation algorithms and spatial geometry theory; Based on the structured semantic representation text and the single-view information corresponding to each of the target viewpoints, a text-generated image model is used to generate single-view images corresponding to each of the target viewpoints. The category information and 3D position information of the sub-objects in each of the single-view images are determined, and based on the category information and position information of the sub-objects and each of the single-view images, a 3D geometric model of the sub-object corresponding to the sub-object is generated by a 3D generation model. The pose of the sub-object 3D geometric model is estimated to obtain the pose parameters of the sub-object 3D geometric model in the 3D scene. The sub-object 3D geometric model is assembled, and the pose of the corresponding sub-object 3D geometric model in the 3D scene is controlled according to the pose parameters to generate an initial 3D scene model. The pose parameters of the sub-object 3D geometric models in the initial 3D scene model are adjusted by a self-supervised layout optimizer to generate the target 3D scene model corresponding to the scene description text.

2. The method for generating a 3D scene model by fusing multi-view images as described in claim 1, characterized in that, The steps of extracting entities, attributes, and spatial relationships from the scene description text to generate structured semantic representation text include: The scene description text is segmented using a word segmentation function to obtain the corresponding word segmentation results; The word segmentation results are tagged with part-of-speech tags using a part-of-speech tagging function to obtain the corresponding tagging results; The annotation results are subjected to dependency parsing using a syntactic analysis function to construct a graph of grammatical dependencies between words. Based on the grammatical dependency graph, the set of sub-object entities, the set of sub-object attributes, and the set of spatial relationships between sub-objects in the grammatical dependency graph are extracted using a large language model; Query the set of associated attributes of each sub-object in the knowledge graph in the sub-object entity set, and add the set of associated attributes to the attribute set of the sub-object to obtain the expanded attribute set of the sub-object; The set of sub-object entities, the expanded set of attributes of the sub-objects, and the set of spatial relationships are determined as the structured semantic representation text.

3. The method for generating a 3D scene model by fusing multi-view images as described in claim 1, characterized in that, The steps of determining multiple target perspectives of the target 3D scene model based on the structured semantic representation text, using semantic segmentation algorithms and spatial geometry theory, include: The semantic segmentation algorithm identifies the set of sub-object entities, the set of spatial relationships between sub-objects, and the set of attributes of sub-objects in the structured semantic representation text. For each sub-object in the sub-object entity set, obtain the typical shape parameters corresponding to the sub-object, and modify the typical shape parameters through the attributes of the sub-object to obtain the initial shape parameters corresponding to the sub-object; Based on the set of spatial relationships between the sub-objects and the rules for setting the origin of the scene coordinate system, the relative positions of each sub-object in the three-dimensional space are logically deduced to generate the initial position coordinates of each sub-object in the 3D scene. For each viewpoint, based on the initial shape parameters and initial position coordinates of each sub-object, the projected area and visibility ratio of each sub-object under the viewpoint are calculated, and the importance weight of each sub-object is determined. Based on the projected area, the visibility ratio, and the importance weight, an objective function for viewpoint selection is constructed; The objective function is optimized, and the combination of perspectives that maximizes the value of the objective function is selected from multiple perspectives as multiple target perspectives of the target 3D scene model.

4. The method for generating a 3D scene model by fusing multi-view images as described in claim 1, characterized in that, The step of generating single-view images corresponding to each target viewpoint using a text-generated image model based on the structured semantic representation text and the single-view information corresponding to each target viewpoint includes: The structured semantic representation text is encoded to obtain the corresponding semantic vector; The camera parameters are obtained, and the view parameters corresponding to each of the target viewpoints are determined. The view parameters corresponding to each of the target viewpoints are then concatenated with the camera parameters to obtain the single view information corresponding to each of the target viewpoints. The single-view information is projected and transformed to obtain the view condition vector corresponding to each target view. The viewpoint condition vectors corresponding to each of the target viewpoints are input into the U-Net network of the text-based image model along with the semantic vectors. The U-Net network of the text-based image model predicts noise and gradually removes noise to generate the single-view image corresponding to each of the target viewpoints.

5. The method for generating a 3D scene model by fusing multi-view images as described in claim 1, characterized in that, The step of determining the category information and location information of sub-objects in each of the single-view images includes: Extract the set of sub-object entities from the structured semantic representation text, wherein the set of sub-object entities contains multiple sub-object entities; Each of the single-view images is input into a segmentation model with an attention mechanism. The segmentation model is used to extract features from the single-view images to obtain feature maps corresponding to each single-view image. Each of the feature maps is input into the segmentation head of the segmentation model, and the feature maps are processed based on the segmentation head to obtain the probability distribution matrix corresponding to each of the feature maps, which contains the pixel points belonging to different object categories; Based on a preset probability threshold, each of the probability distribution matrices is converted into a corresponding binary mask matrix; Based on the binary mask matrix and the single-view image, the local feature vector of the sub-object corresponding to the sub-object in the single-view image is extracted by an object recognition algorithm; Based on the local feature vector of the sub-object, various types of objects in the object category feature library are matched to obtain the initial category of the sub-object. The object category feature library contains multiple types of objects and reference feature vectors of each type of object. The initial category of the sub-object is matched with the sub-object entity by calculating semantic similarity to determine the category information of the sub-object; Each of the single-view images is input into a multi-scale depth fusion model, and the depth of each of the single-view images is estimated by the multi-scale depth fusion model to obtain the depth map corresponding to each of the single-view images. For each pixel in the single-view image, the 3D coordinates of the pixel in the camera coordinate system are calculated based on the camera intrinsic parameter matrix, the depth value of the pixel in the depth map, and the image coordinates of the pixel. The 3D coordinates of the pixel in the camera coordinate system are then converted into global coordinates in the world coordinate system using the camera extrinsic parameter matrix. The 3D position information of the sub-object is determined based on the global coordinates of each pixel in the world coordinate system.

6. The method for generating a 3D scene model by fusing multi-view images as described in claim 1, characterized in that, The step of generating a 3D geometric model of the sub-object corresponding to the sub-object based on the category information and location information of the sub-object and each of the single-view images includes: The category information of the sub-object, the location information of the sub-object, and each single-view image are input into the 3D generation model based on the generative adversarial network. The 3D generation model generates a 3D point cloud model of the sub-object that matches the input conditions. A quadratic error metric algorithm is used to simplify the mesh of the sub-object 3D point cloud model. Vertices in the mesh of the initial sub-object 3D geometric model are iteratively merged to obtain the corresponding simplified sub-object 3D geometric model. The simplified 3D geometric model of the sub-object is smoothed using the Laplacian smoothing algorithm to obtain the corresponding smoothed 3D geometric model of the sub-object. Extract the attribute set of the sub-objects in the structured semantic representation text, wherein the attribute set of the sub-objects contains the attributes of each sub-object; Based on the single-view image, determine the 2D image of the sub-object corresponding to the sub-object, and establish the texture mapping relationship between the 2D image of the sub-object and the 3D geometric model of the sub-object; Based on the attributes of the sub-object, the position information of the sub-object, and the texture mapping relationship, the color and texture information in the 2D image of the sub-object is mapped to each facet of the corresponding 3D geometric model of the sub-object through a material mapping algorithm, thereby obtaining the 3D geometric model of the sub-object corresponding to the sub-object.

7. The method for generating a 3D scene model by fusing multi-view images as described in claim 1, characterized in that, The step of adjusting the pose parameters of the sub-object 3D geometric model in the initial 3D scene model using a self-supervised layout optimizer includes: Based on the pose parameters of the sub-object 3D geometric models in the initial 3D scene model, the sub-object overlap loss, space utilization loss and sub-object distance loss corresponding to the initial 3D scene model are calculated. The sub-object overlap loss is obtained by calculating the proportion of the intersection volume between the sub-objects. The space utilization loss is used to evaluate the uniformity of space filling of the target 3D scene model. The sub-object distance loss is used to evaluate the reasonable spacing between the sub-objects. Based on the sub-object overlap loss, the space utilization loss, and the distance loss between sub-objects, a scene layout loss function is constructed; The pose parameters of the sub-object 3D geometric model are adjusted by optimizing the scene layout loss function.

8. The method for generating a 3D scene model by fusing multi-view images as described in claim 1, characterized in that, After the step of generating the target 3D scene model corresponding to the scene description text, the method further includes: Style features are extracted from the single-view images corresponding to each target viewpoint using a convolutional neural network to obtain the corresponding style feature vectors. The 3D geometric model of the sub-object is digitally analyzed to obtain the corresponding material feature vector; The style feature vector, the material feature vector, and the sub-object 3D geometric model are input into a style transfer network based on a generative adversarial network. The output is a style-unified sub-object 3D geometric model, which is then assembled to obtain the second-stage target 3D scene model. Based on the second-stage target 3D scene model, the lighting re-rendering algorithm is used to re-render the second-stage target 3D scene model to obtain the third-stage target 3D scene model. Shadow mapping technology is used to generate shadows on the third-stage target 3D scene model, resulting in a fourth-stage target 3D scene model that includes shadow effects that conform to physical laws.

9. A 3D scene model generation device that integrates multi-view images, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the method for generating a 3D scene model by fusing multi-view images as described in any one of claims 1 to 8.

10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the method for generating a 3D scene model by fusing multi-view images as described in any one of claims 1 to 8.