A high-quality 3D model generation method for garden education
By inserting a low-rank matrix into a general 3D generative model and constructing a dedicated loss function, the problems of low modeling efficiency and unstable quality in landscape education are solved, and efficient and high-quality landscape 3D model generation is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HENAN WEIDA ELECTRONIC TECH CO LTD
- Filing Date
- 2026-02-09
- Publication Date
- 2026-06-02
Smart Images

Figure CN122134923A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of model generation, and in particular to a method for generating high-quality 3D models for garden education. Background Technology
[0002] Currently, in fields such as garden education, landscape teaching, virtual simulation, and digital cultural tourism, 3D model generation technology has gradually become an important tool for content creation and teaching demonstrations. Existing 3D model generation applications mainly rely on the following two technical paths: The first is to carry out 3D model construction based on software tools (such as 3ds Max, Blender, SketchUp, etc.) for manual modeling. This method requires professionals to manually build the model structure, which places high demands on the modeler's spatial understanding, artistic foundation, and software operation skills. It has a long learning cycle, a complex production process, and low modeling efficiency. Moreover, in complex garden scenes, it requires a lot of repetitive and detailed work, which significantly increases time and labor costs. The second is an automated 3D generation method based on general AI large models, such as automatically inferring the 3D geometry and texture of objects using a single image. However, in real-world scenarios, such general-purpose models still have significant limitations in vertical fields such as landscape education: First, the accuracy of shape reconstruction is insufficient: 3D models generated from a single image often only perform well from the frontal view, while unobserved areas such as the sides and back often exhibit incomplete structures, topological errors, or texture distortion; Second, the ability to represent domain-specific objects is weak: General-purpose models lack the ability to understand the structure of professional objects such as garden plants, rocks, and ancient building components, making it difficult to accurately generate details unique to the corresponding domain (such as natural textures, plant hierarchical structures, and the proportions and relationships of traditional garden components).
[0003] Using 2D generation as an intermediary to regenerate 3D still fails to address issues such as the lack of strict domain features, weak 3D geometric consistency, inability to fine-tune for specific vertical domains (e.g., gardening), inability to guarantee fine-grained local structure in the generated results, and inability to inject domain geometric priors during the 2D to 3D transition. Even the latest training-free 3D scene generation techniques (such as ArtiScene) have not solved the problem of insufficient geometric detail and texture consistency for gardening objects; their generation relies on 2D intermediary images and cannot perform fine-tuning based on domain features.
[0004] In summary, existing technologies in landscape education generally suffer from drawbacks such as low modeling efficiency, unstable quality of automatically generated models, and lack of domain features. There is an urgent need for a method to generate 3D models of professional objects that can reconstruct them with high quality for the field of landscape education. Summary of the Invention
[0005] To address the problems existing in the background technology, this invention proposes a method for generating high-quality 3D models for garden education.
[0006] A method for generating high-quality 3D models for landscape education, comprising the following steps:
[0007] S100: Acquire a 2D target image and perform optimization processing on the 2D target image;
[0008] S200. For the optimized single 2D target image, construct multiple pose images with consistent cross-viewpoints;
[0009] S300. Based on the general 3D guided large model, a fine-tuning framework based on a low-rank matrix is constructed, and constraints are imposed on the fine-tuning framework to serve as a high-quality 3D model.
[0010] S400: Input multiple pose images into the high-quality 3D model and generate the required 3D garden model.
[0011] Based on the above, in step S300, the general 3D guided large model includes a 2D feature encoding adaptation layer, a 3D geometry generation adaptation layer, and a texture mapping adaptation layer. A low-rank matrix pair that can be updated is configured for each of the 2D feature encoding adaptation layer, the 3D geometry generation adaptation layer, and the texture mapping adaptation layer.
[0012] Based on the above, in the 2D feature encoding adaptation layer, low-rank matrix pairs are inserted into the general ViT model. ,in and These represent the input and output feature dimensions of the original feature encoding layer, respectively. To inject the rank of a low-rank matrix; to inject low-rank matrix pairs into the general model of the 3D geometry generation adapter layer. ,in and These represent the input and output feature dimensions of the original 3D geometry generation adaptation layer, respectively. To inject the rank of a low-rank matrix; insert low-rank matrix pairs into the general model of the texture mapping adaptation layer. ,in and These represent the input and output feature dimensions of the original general 3D texture mapping model, respectively. The rank of the injected low-rank matrix.
[0013] Based on the above, in step S300, the constraint condition is a dedicated loss function constructed for the fine-tuning framework of the large model.
[0014] Based on the above, the overall loss function is:
[0015]
[0016] in, These represent the weighting coefficients for each type of loss. These represent geometric consistency loss, texture fidelity loss, multi-view reconstruction constraint loss, and domain prior loss, respectively.
[0017] Based on the above, geometric consistency loss for:
[0018]
[0019]
[0020] in, Represents the chamfer distance function. Indicates the generated 3D model The point set obtained by sampling in the middle, Indicates from the real target The point set obtained by sampling, x and y represent respectively and One of the points, It is an L2 norm.
[0021] Based on the above, texture fidelity loss for:
[0022]
[0023] in, Indicates from generative model An image rendered from a certain viewpoint. Indicates from the real target Real images from the same perspective, Indicates pixel-level loss. This represents the perceived level loss.
[0024] Based on the above, multi-view reconstruction constraint loss for:
[0025]
[0026] Where V is the set of all viewpoints used for constraints. This indicates the rendering loss at a certain perspective.
[0027] Based on the above, domain prior loss for:
[0028]
[0029]
[0030]
[0031] in, Represents the sparse loss of a low-rank matrix. Represents the geometric smoothness loss. and They represent in The low-rank matrix injected in the adaptation layer, Denotes the F-norm, The value of the implicit representation function of the 3D model at point p. express The second derivative at point p.
[0032] This invention has outstanding substantive features and significant progress compared to the prior art, specifically:
[0033] (1) The present invention encloses the key parameter matrix of the general 3D generation model in an updatable low-rank matrix pair, and can express the fine-grained geometric and texture information of garden objects with a very small number of parameter updates. This solves the problems of high cost and difficulty in deployment of full fine-tuning in the prior art, as well as the defects of general models in capturing garden details and the problem of overfitting in small sample training.
[0034] (2) The present invention injects LoRA adaptation modules into the 2D feature encoding layer, the 3D geometry generation layer and the texture mapping layer respectively, so that the model can learn the unique geometric morphology (leaf order, tree crown structure, mountain and rock blocks), texture and material features (bark texture, moss, stone texture), scene semantics and teaching style (type differentiation, landscape structure) of garden objects at different depth levels, forming a bottom-up, layered adaptation capability from structure to texture to semantics;
[0035] (3) This invention constructs a composite loss function that integrates geometric consistency loss, texture fidelity loss, multi-view reconstruction constraint, low-rank sparsity loss and geometric smoothness loss, which is used to accurately guide the optimization of third-order LoRA. This makes the model significantly better than the existing technology in terms of structural integrity, cross-view consistency, local texture accuracy and domain semantic rationality, and achieves high-precision adaptation to small sample data in the field of garden education. Attached Figure Description
[0036] Figure 1 This is a flowchart illustrating the process of this invention. Detailed Implementation
[0037] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0038] like Figure 1 As shown, a method for generating a high-quality 3D model for garden education includes the following steps: S100, acquiring a 2D target image and optimizing the 2D target image; S200, constructing multiple pose images with consistent cross-viewpoints for the optimized single 2D target image; S300, constructing a fine-tuning framework based on a low-rank matrix on the basis of a general 3D guided large model, and constructing constraints on the fine-tuning framework to obtain a high-quality 3D model; S400, inputting the multiple pose images into the high-quality 3D model and generating the required 3D garden model.
[0039] In reality, users can upload a 2D target image using real-life photos or illustrations from textbooks, or generate a 2D target image using a model. Then, they can edit and optimize the 2D target image. The editing and optimization process begins with background removal, accurately separating foreground garden objects (such as plants and rocks) from the complex background of the 2D target image. This is followed by denoising, cropping, normalization, and style unification, all using existing common methods, which will not be elaborated further. For the optimized single image, multiple images with precise poses and consistent cross-viewpoints are generated through synthesis / prediction, serving as supervisory signals for fine-tuning the 3D model.
[0040] Because garden 3D data has many fine-grained features (such as irregular leaf shapes and complex structural features) and a high proportion of small samples, this embodiment inserts a low-rank adaptation module into the key layer of the general 3D generation model based on the principle of low-rank matrix factorization. This achieves "small parameter updates → accurate adaptation to garden 3D generation", avoiding parameter redundancy and overfitting caused by full fine-tuning.
[0041] In this embodiment, the framework adopts a "three-level adaptation + domain constraint" architecture, with each level optimized for the characteristics of garden 3D data, as detailed below:
[0042] General large models typically include a 2D feature encoding adaptation layer, a 3D geometry generation adaptation layer, and a texture mapping adaptation layer when generating 3D models.
[0043] Since general-purpose large-scale models typically have a large number of parameters, fine-tuning training requires a significant amount of resources. Therefore, this embodiment uses the original parameter matrix of the large model. (m is the model input dimension, n is the model output dimension) Enclose a low-rank matrix pair and This is used to learn features specific to the field of landscape architecture. The final parameters are expressed as follows: ,in This represents the final parameters of the large model. Represents the parameters of the original large model. This represents the low-rank matrix that needs to be updated.
[0044] In constructing the fine-tuning framework design, this embodiment injects a low-rank matrix pair B and A, which can be updated, into each of the three adaptation layers. Specifically, in the 2D feature encoding layer, a low-rank matrix pair is inserted into the general ViT model. ,in and These represent the input and output feature dimensions of the original feature encoding layer, respectively. In this embodiment, to inject the rank of the low-rank matrix, The value is set to 64; similarly, low-rank matrix pairs are injected into the general model of the 3D geometry generation adaptation layer. ,in and These represent the input and output feature dimensions of the original 3D geometry generation adaptation layer, respectively. In this embodiment, to inject the rank of the low-rank matrix, The value is 64; a low-rank matrix pair is inserted into the general model of the texture mapping adaptation layer. ,in and These represent the input and output feature dimensions of the original general 3D texture mapping model, respectively. In this embodiment, to inject the rank of the low-rank matrix, The value is 64.
[0045] In this embodiment, the constraint is a dedicated loss function. The dedicated loss function for the large model fine-tuning framework is constructed as follows:
[0046] To address the characteristics of 3D generation tasks in the field of landscape architecture, such as complex structures, fine-grained local textures, and high requirements for cross-view consistency, this invention constructs a composite loss function that integrates geometric consistency, texture fidelity, multi-view reconstruction constraints, and domain priors. This loss function is specifically designed to guide the training of low-rank matrices B and A, enabling the fine-tuned large model to accurately capture 3D features of landscape architecture with minimal parameter overhead.
[0047] The overall loss function of this invention is:
[0048]
[0049] in, These are weighting coefficients for each loss item, which can be adjusted according to the proportion of the data size. These represent geometric consistency loss, texture fidelity loss, multi-view reconstruction constraint loss, and domain prior loss, respectively.
[0050] (1) Geometric consistency loss
[0051] Garden 3D models typically feature numerous irregular leaf surfaces, layered canopy structures, and curved surface components. Therefore, a geometric consistency loss is incorporated into the 3D geometry generation adaptation layer to constrain the correspondence between predicted and real geometry. This loss aims to constrain the generated 3D model. In shape and structure, it is similar to the real target To maintain consistency, especially for complex geometric structures in gardens (such as irregularly shaped branches and leaves), the Chamfer Distance (CD) is typically used to measure the differences between geometric shapes.
[0052]
[0053]
[0054] in Represents the chamfer distance function. Indicates the generated 3D model The point set obtained by sampling in the middle, Indicates from 3D model The point set obtained by sampling, x and y represent respectively and One of the points, It is the L2 norm, which measures the geometric distance between two points. This metric primarily ensures that the shape and detailed structure of the generated 3D model closely approximate the geometry of the real target object.
[0055] (2) Texture fidelity loss
[0056] This loss function aims to constrain the generated 3D model to match the real target in terms of local texture details in the rendered image, specifically targeting the fine-grained characteristics of local textures (such as leaf textures). The method used in this embodiment combines pixel-level loss and perceptual loss:
[0057]
[0058] in Indicates from generative model An image rendered from a certain viewpoint. Indicates from generative model Real images from the same perspective, Indicates pixel-level loss. This represents the perceived level loss.
[0059] (3) Multi-view reconstruction constraint loss
[0060] This loss focuses on characteristics with high requirements for cross-view consistency. It requires the model to be able to consistently interpret multi-view inputs or render realistic images from any viewpoint when generating 3D shapes and textures.
[0061]
[0062] Where V is the set of all viewpoints used for constraints. This indicates the rendering loss at a certain perspective.
[0063] (4) Domain prior loss
[0064] This loss is used to inject knowledge specific to the landscape architecture domain (including the smoothness of natural objects and the sparsity of low-rank structures) to guide the low-rank matrix. and Training them and preventing them from overfitting on small sample data.
[0065]
[0066] in Represents the sparse loss of a low-rank matrix. This represents the geometric smoothness loss, used to penalize drastic changes (high-frequency noise) on the model surface. The specific expression is as follows:
[0067]
[0068] in, and They represent in The low-rank matrix injected into the adaptation layers (including 2D feature encoding layers, 3D geometry generation, and texture mapping layers). The F-norm is the square root of the sum of the squares of all elements in the matrix. This loss is used to constrain... and The weight magnitude is adjusted to prevent it from becoming too large and indirectly encourage its sparsity. This allows the low-rank matrix to learn only the most critical, non-redundant features in the field of landscape architecture, thereby avoiding overfitting in scenarios with a small number of parameter updates.
[0069]
[0070] in, The value of the implicit representation function of the 3D model at point p. express The second derivative at point p. This loss penalizes drastic variations (high-frequency noise) on the 3D model surface, ensuring that the generated garden structure is natural and smooth.
[0071] Compared with existing technologies, the method for generating high-quality 3D garden models of the present invention has the following core advantages in terms of efficiency, accuracy, and resource consumption:
[0072] 1. Achieve high-precision domain adaptation
[0073] (1) The ability to capture fine-grained structures is significantly enhanced.
[0074] Based on a low-rank fine-tuning framework with third-order adaptation, this invention can accurately generate highly complex structures unique to garden scenes (such as irregular leaves, layered canopies, rock textures, and broken edges). Compared to general models, the geometric topology generated by this invention is more complete and the local details are more stable, effectively avoiding structural collapse, local missing parts, and topological errors commonly found in existing technologies.
[0075] (2) Cross-perspective consistency is significantly improved
[0076] By introducing multi-view reconstruction constraints, this invention significantly improves the texture distortion and geometric discontinuity problems of general single-image 3D reconstruction methods in non-observation areas such as the sides and back. The generated results maintain high consistency across multiple viewpoints, meeting the requirement in landscape architecture teaching that "structural details can be observed from any viewpoint."
[0077] 2. The high efficiency of LoRA fine-tuning
[0078] (1) Extremely low resource consumption, suitable for small sample training scenarios in gardening.
[0079] The low-rank matrix factorization method used in this invention only updates a very small number of parameters and does not require full fine-tuning of a large model. Training can be completed on a single GPU or a lightweight computing device, which significantly reduces the cost of model deployment for gardening education institutions.
[0080] (2) Significantly reduces the risk of overfitting and improves generalization ability.
[0081] By introducing low-rank sparse constraints and domain prior smoothing terms, this invention effectively suppresses overfitting under small sample conditions, enabling the fine-tuned model to maintain stable performance across different garden objects. Especially when plant objects exhibit significant morphological differences, this invention can still generate 3D models with natural structures and continuous textures.
[0082] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims.
Claims
1. A method for generating high-quality 3D models for garden education, characterized in that, Including the following steps: S100: Acquire a 2D target image and perform optimization processing on the 2D target image; S200. For the optimized single 2D target image, construct multiple pose images with consistent cross-viewpoints; S300. Based on the general 3D guided large model, a fine-tuning framework based on a low-rank matrix is constructed, and constraints are imposed on the fine-tuning framework to serve as a high-quality 3D model. S400: Input multiple pose images into the high-quality 3D model and generate the required 3D garden model.
2. The method for generating high-quality 3D models for garden education according to claim 1, characterized in that: In step S300, the general 3D guided large model includes a 2D feature encoding adaptation layer, a 3D geometry generation adaptation layer, and a texture mapping adaptation layer. A low-rank matrix pair that can be updated is configured for each of the 2D feature encoding adaptation layer, the 3D geometry generation adaptation layer, and the texture mapping adaptation layer.
3. The method for generating high-quality 3D models for garden education according to claim 2, characterized in that: In the 2D feature encoding adaptation layer, low-rank matrix pairs are inserted into the general ViT model. ,in and These represent the input and output feature dimensions of the original feature encoding layer, respectively. To inject the rank of a low-rank matrix; to inject low-rank matrix pairs into the general model of the 3D geometry generation adapter layer. ,in and These represent the input and output feature dimensions of the original 3D geometry generation adaptation layer, respectively. To inject the rank of a low-rank matrix; insert low-rank matrix pairs into the general model of the texture mapping adaptation layer. ,in and These represent the input and output feature dimensions of the original general 3D texture mapping model, respectively. The rank of the injected low-rank matrix.
4. The method for generating high-quality 3D models for garden education according to claim 1, characterized in that: In step S300, the constraint is a specific loss function constructed for the fine-tuning framework of the large model.
5. The method for generating high-quality 3D models for garden education according to claim 4, characterized in that: The overall loss function is: in, These represent the weighting coefficients for each type of loss. These represent geometric consistency loss, texture fidelity loss, multi-view reconstruction constraint loss, and domain prior loss, respectively.
6. The method for generating high-quality 3D models for garden education according to claim 5, characterized in that: Geometric consistency loss for: in, Represents the chamfer distance function. Indicates the generated 3D model The point set obtained by sampling in the middle, Indicates from the real target The point set obtained by sampling, x and y represent respectively and One of the points, It is an L2 norm.
7. The method for generating high-quality 3D models for garden education according to claim 5, characterized in that: Texture fidelity loss for: in, Indicates from generative model An image rendered from a certain viewpoint. Indicates from the real target Real images from the same perspective, Indicates pixel-level loss. This represents the perceived level loss.
8. The method for generating high-quality 3D models for garden education according to claim 5, characterized in that: Multi-view reconstruction constraint loss for: Where V is the set of all viewpoints used for constraints. This indicates the rendering loss at a certain perspective.
9. The method for generating high-quality 3D models for garden education according to claim 5, characterized in that: Domain Prior Loss for: in, Represents the sparse loss of a low-rank matrix. Represents the geometric smoothness loss. and They represent in The low-rank matrix injected in the adaptation layer, Denotes the F-norm, The value of the implicit representation function of the 3D model at point p. express The second derivative at point p.