Three-dimensional scene interactive generation method
An interactive modeling system was constructed using SAM image segmentation, LLM perspective analysis, and the three-distance perspective transformation algorithm. This system solved the problems of operational complexity and user experience in the generation of 3D scenes of Chinese landscape paintings, and achieved an efficient and interpretable 3D modeling process.
Patent Information
- Application Number
- CN202511156098.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-18
- Publication Date
- 2025-11-07
AI Technical Summary
Existing technologies for generating 3D scenes that recreate Chinese landscape paintings involve complex processes, require a high level of user modeling experience, lack intelligent assistance and interactive guidance, and struggle to guarantee stylistic consistency and spatial aesthetics.
An interactive modeling system based on primitive segmentation, perspective analysis, and 3D generation is constructed using SAM-based image segmentation, LLM-based perspective analysis, and a three-distance perspective transformation algorithm. This system provides interpretable perspective recommendations and real-time parameter adjustment.
It lowers the barrier to 3D modeling, improves the efficiency of digitizing traditional art, enhances user experience, and ensures stylistic consistency and spatial aesthetics.
Smart Images

Figure CN120912804A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of three-dimensional scene generation and digital humanities, in particular to an interactive three-dimensional scene generation method. BACKGROUND
[0002] As an important part of traditional oriental art, Chinese landscape painting has unique spatial composition and aesthetic logic, which poses a challenge to three-dimensional scene generation. Traditional landscape painting constructs multi-viewpoint space with point perspective, among which the three-distance method proposed by Guo Xi, namely high distance, deep distance and flat distance, is the most representative. According to the difference of viewing angle, Guo Xi divided the form of mountain landscape into three types: "from the foot of the mountain to the top of the mountain, called high distance; from the front of the mountain to the back of the mountain, called deep distance; from the near mountain to the far mountain, called flat distance". In the process of three-dimensional scene construction of Chinese landscape painting, in order to reproduce the artistic conception pursued by the painting, the modeler not only needs to master the basic technology of three-dimensional modeling software, but also needs to understand the unique perspective expression method of landscape painting.
[0003] There have been a series of works on the 3D modeling of Chinese landscape painting, which attempt to reconstruct its brushwork style and visual features. Earlier researches mostly focus on the digital expression of brushwork. Shen, L. 2014. Serenity: 3D Animation to Simulate Chinese Ink Painting. Master's thesis, Victoria University of Wellington; designed a system that can automatically simulate the texture of the stroke method; Li, X. and Qiu, G. 2016. Modeling, rendering and drawing method of digital 3D freehand ink landscape painting. China CN105678835A, filed November 23, 2015, issued June 15, 2016, realized the expression of ink style of rock scene through geometric analysis method; Yao, Z., Sun, Q., Liu, B., Lu, Y., Liu, G., Yang, X.-D., and Mi, H. 2024. InkBrush: A sketching tool for 3D ink painting. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, ACM, 1-15; proposed a freehand brush-driven 3D painting tool that can reproduce ink stroke features such as fly white and rough edge; these methods have made some progress in improving the digital expression of landscape painting, but mainly focus on the visual style, and pay less attention to the restoration of spatial perspective structure.To further improve the efficiency of modeling, three-dimensional generative models in computer graphics can be introduced into the modeling task of landscape paintings. The multi-view reconstruction method represented by Mildenhall, B., Srinivasan, P. P., Tancik, M., Barron, J. T., Ramamoorthi, R., and Ng, R. 2020. NeRF: Representing scenes as neural radiance fields for view synthesis. arXiv preprint arXiv:2003.08934 can reproduce natural scenes with high fidelity, but it relies heavily on camera parameters and performs poorly when dealing with single landscape paintings that lack perspective information. The latest feedforward image-driven three-dimensional generative model Tochilkin, D., Pankratz, D., Liu, Z., Huang, Z., Letts, A., Li, Y., and Cao, Y. P. 2024. Triposr: Fast 3D object reconstruction from a single image. arXiv preprint arXiv:2403.02151 proposes a more efficient three-dimensional generation path. This method combines image Transformer encoders and TriplaneNeRF decoders and performs well in general object reconstruction tasks. However, TripoSR uses an end-to-end structure and lacks visualization feedback and controllable mechanisms, making it difficult for users to intervene in perspective results and only suitable for three-dimensional generation of single objects in three-dimensional scenes. To avoid the distortion problem of linear perspective, some studies propose using auxiliary lines to generate height maps (Zhou, P., Li, K., Wei, W., Wang, Z., and Zhou, M. 2020. Fast generation method of 3D scene in Chinese landscape painting. Multimedia Tools and Applications 79, 23-24, 16441-16457.) or using oblique projection to simulate multi-view composition (Li, X., Chen, Z., Huang, Y., Zheng, J., and Deng, S. 2023. Digital ink and water landscape painting three-dimensional modeling and presentation in the form of metaverse. Journal of Wenzhou University (Natural Science Edition) 44, 4, 50-61.).
[0004] Although the above methods can restore the three-dimensional effect of the three-dimensional effect to some extent, the operation process is complex and requires high modeling experience for users, making it difficult to popularize to non-professional creators. In addition, in terms of user experience, these methods mostly lack intelligent assistance and interactive guidance, and the modeling results are difficult to guarantee style consistency and spatial aesthetics.
[0005] Therefore, an interactive generation method of a three-dimensional scene is proposed to solve the problems in the background. SUMMARY
[0006] The present application aims to provide an interactive generation method of a three-dimensional scene to solve the problems in the background.
[0007] To achieve the above object, the present application provides the following technical solution: an interactive generation method of a three-dimensional scene, comprising the following steps:
[0008] (1) Interactive image segmentation based on SAM: extracting key elements in the picture mainly based on mountains, and generating a segmentation silhouette of the corresponding elements;
[0009] (2) Perspective analysis of landscape painting images based on LLM: outputting the probability distribution and explainable analysis results of three perspective methods of high perspective, deep perspective and flat perspective, and recommending the optimal perspective method;
[0010] (3) Three-dimensional scene generation based on three-perspective method: proposing a three-dimensional layout algorithm based on three-perspective method, combining the perspective type and adjustment parameters selected by the user, arranging the three-dimensional model generated by the segmented elements, and constructing an interactive fine-tuning landscape painting three-dimensional scene;
[0011] (4) Analysis and evaluation based on the system: taking the scene modeling task of a typical landscape painting as an example, analyzing and objectively evaluating the three-dimensional scene generated by the system, collecting data such as generation efficiency and layout restoration degree, and verifying the effectiveness and usability of the method.
[0012] Preferably, the step (1) comprises the following sub-steps:
[0013] (1.1) Upload the landscape painting image and enter the interactive segmentation view; the user marks the foreground and background hint points by clicking, the system calls SAM for image segmentation in real time, and generates the corresponding mask area;
[0014] (1.2) After the user confirms the segmentation area, the system generates an element thumbnail, and saves the element number and position layout information for subsequent analysis and modeling input.
[0015] Preferably, the step (2) comprises the following sub-steps:
[0016] (2.1) The system calls the LLM model for three-distance method perspective analysis reasoning based on visual features such as overall structure of the image, arrangement of graph elements, hierarchical relationship, etc., to lay the analysis foundation for subsequent perspective judgment;
[0017] (2.2) The system outputs the score distribution of high-distance, deep-distance and flat-distance perspective methods, and selects the optimal recommended perspective method through the majority voting mechanism, and determines the preliminary perspective direction;
[0018] (2.3) The most representative model analysis explanation is selected in combination with the average value of voting, and is presented in JSON format, enhancing the consistency and interpretability of the results;
[0019] (2.4) The user can view the analysis results of each perspective method and choose whether to accept the recommended perspective method, and the system will pass the selected perspective type to the next stage.
[0020] Preferably, the step (3) comprises the following sub-steps:
[0021] (3.1) For each graph element, call the TripoSR graph generation three-dimensional model interface, generate a normalized three-dimensional grid model according to the graph element silhouette, and record its standard height value;
[0022] (3.2) According to the perspective type selected by the user and the system preset parameters, combined with the two-dimensional position of the graph element and the depth factor, use the perspective conversion formula to calculate its three-dimensional coordinates and scaling ratio;
[0023] (3.3) Arrange the three-dimensional models of all graph elements in a unified layout, apply adjustable layout parameters including horizontal sparsity, vertical rhythm, overall scaling coefficient, etc., to simulate the three-distance method space semantics;
[0024] (3.4) The user can adjust various parameters in real time through the slider in the generated view, and the system dynamically updates the three-dimensional scene to realize the interactive composition experience of what you see is what you get.
[0025] Preferably, the step (4) comprises the following sub-steps:
[0026] (4.1) Case analysis: In order to verify the operability and modeling integrity of the system in actual use scenarios, a typical user operation task flow is set, and a representative Chinese landscape painting is selected as experimental material. The user completes the interactive operation process of the three stages of graph element segmentation, perspective analysis and parameter adjustment, and three-dimensional scene generation according to the method;
[0027] In the perspective analysis stage, the system analyzes the perspective method according to the picture composition characteristics, and recommends the most matched three-far method; in the three-dimensional generation stage, the system completes the space layout according to the foregoing input, and allows the user to perform slider parameter fine tuning. Finally, the generated three-dimensional landscape painting scene is successfully exported to the Blender software environment for subsequent immersive scene display and space interaction building;
[0028] (4.2) Generation result analysis: To further evaluate the applicability and user experience of the system in assisting the construction of three-dimensional scenes of Chinese landscape paintings, a number of subjects are organized to complete the preset modeling tasks, covering a number of typical landscape composition cases. The system records the key operation behaviors of the participants and saves their final generated three-dimensional modeling results as an important basis for evaluating the modeling effect. By analyzing the three-dimensional scenes output by each group of users, the adaptability of the system to different styles of landscape paintings is demonstrated;
[0029] (4.3) Objective evaluation of generation results: To verify the consistency of the system-generated scene and the original landscape painting in terms of layout from the spatial structure level, the following two-dimensional quantitative evaluation indicators are proposed:
[0030] A. Two-dimensional mask reprojection comparison:
[0031] Select the key elements extracted by the user, and perform perspective reprojection on the three-dimensional scene generated by the system to obtain a two-dimensional mask image. Compare this image with the mask region manually annotated by experts in the original image at the pixel level, calculate the intersection over union (IoU) value, and measure the fitting degree of the elements in terms of relative position and boundary contour;
[0032] B. Element occlusion relationship matrix:
[0033] According to the composition order of the original image, a reference occlusion relationship matrix annotated by experts is constructed to count the front and back occlusion relationships between elements. Compare this matrix with the occlusion relationship matrix automatically derived from the three-dimensional scene generated by the system, calculate the occlusion matching rate, and evaluate the restoration accuracy of the system-generated scene in terms of depth and three-far perspective relationship.
[0034] Preferably, in step (1.1), the SAM model has zero-shot generalization capability, and the first click submits the original image path and the feature extraction of the hint point, and in subsequent iterations, the intermediate features in the previous round are used as prior inputs to reduce repeated calculations; the system outputs multiple candidate masks and provides an IoU quality score when first interacting, and automatically selects the highest score result.
[0035] Preferably, in step (2.1), the LLM is set as "Chinese landscape painting perspective analysis expert" through prompt word engineering, and is injected with three-far method definition and discrimination points, and is required to perform no less than three-step implicit thinking; in step (2.3), the JSON structure contains Analysis field and Perspective field, and the system adopts a concurrent calling strategy (temperature=0.7) to obtain multiple results and perform average aggregation.
[0036] Preferably, in step (3.2), the two-dimensional to three-dimensional reference mapping introduces a global scaling coefficient (0.1 cm / pixel), and the overall size of the three-dimensional scene is
[0037] ; the mountain reference three-dimensional coordinates are defined as: , wherein are the coordinates of the upper left corner and the lower right corner of the mountain two-dimensional bounding box.
[0038] Preferably, in step (3.3), the three-far method perspective correction is specifically:
[0039] High-far method: vertical scaling coefficient , the corrected coordinates , is a high-far method adjustable parameter, and the default is 1.5, wherein represents the depth of the mountain in the three-dimensional scene, that is, the bottom of the primitive in the three-dimensional landscape is higher, and the depth is greater ;
[0040] Deep-far method: scaling coefficient , the corrected coordinates , φ is the camera viewing angle, is a deep-far method adjustable parameter;
[0041] Flat-far method: a double-section scaling strategy is adopted, the foreground threshold θ=0.3, the near scene enhancement parameter is 2.5 by default, the far scene parameter , the corrected coordinates .
[0042] Preferably, in step (4.3), the IoU calculation formula is:
[0043]
[0044] , wherein is the mask area after projection of the three-dimensional model, is the real mask area of the corresponding primitive in the original image; in the occlusion relationship matrix, if the primitive Gi is in front of Gj and constitutes an occlusion, it is recorded as Otherwise, 0, the accuracy rate is calculated by comparing with the expert annotation matrix.
[0045] Compared with the prior art, the beneficial effects of the application are that the three-dimensional scene interactive generation method can integrate the SAM-based primitive segmentation, the LLM-based perspective analysis and the three-dimensional perspective conversion algorithm into one, and construct a three-stage interactive modeling system of "primitive segmentation-perspective analysis-three-dimensional generation". Through the interpretable three-distant method perspective recommendation and real-time slider parameter adjustment, the system helps users quickly master the key points of high-distant, deep-distant and flat-distant perspective layout, and solves the pain points of difficult space restoration and long adjustment time under the scatter perspective of landscape painting. The usability of the system is verified by using the system to carry out user experiments. The interactive generation method proposed in the application reduces the threshold of three-dimensional modeling and improves the digital efficiency of traditional art; BRIEF DESCRIPTION OF DRAWINGS
[0046] Figure 1 A three-dimensional scene interactive generation method based on the three-distant method of Chinese landscape painting is designed and a system schematic diagram is shown.
[0047] Figure 2 A SAM model segmentation schematic diagram is shown.
[0048] Figure 3 A perspective analysis prompt design framework diagram based on the prompt word engineering is shown.
[0049] Figure 4 A two-dimensional mountain to three-dimensional scene coordinate mapping schematic diagram is shown.
[0050] Figure 5 A system typical operation flowchart is shown.
[0051] Figure 6 A three-dimensional scene generation example diagram of the three-distant method perspective analysis of a typical landscape painting under the assistance of the system is shown. DETAILED DESCRIPTION
[0052] The technical solutions in the embodiments of the application will be clearly and completely described below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only part of the embodiments of the application, not all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the application.
[0053] Please refer to Figures 1-6 The application provides the following technical solutions: a three-dimensional scene interactive generation method.
[0054] The application provides a three-dimensional scene interactive generation method, including the following steps:
[0055] (1) With the help of the SAM model, the graphic element segmentation of human-computer collaboration is completed;
[0056] (2) The large language model is used to intelligently distinguish the perspective type of the picture and give an interpretable recommendation;
[0057] (3) Based on the unified perspective conversion algorithm and parameterized layout, real-time switching is realized between different perspective methods, and the three-dimensional scene is fine-tuned.
[0058] (4) Based on the system, analysis and evaluation are carried out, a three-dimensional scene interactive generation system of Chinese landscape painting based on the three-dimensional method is developed, typical landscape painting modeling tasks are set, case analysis and objective evaluation of the three-dimensional scene generated by the system are carried out, data such as generation efficiency and layout restoration degree are collected, and the effectiveness and usability of the invention are verified.
[0059] The module design diagram and system interface schematic diagram of the method are as shown in Figure 1 .
[0060] The details are as follows:
[0061] (1) Pre-computation and interactive generation of graphic element segmentation: This step is used to realize the initial extraction of the three-dimensional modeling target area in Chinese landscape painting. By combining SAM and interactive visualization interface, the user-guided graphic element segmentation operation is realized, as shown in Figure 2 . Specifically, the following sub-steps are included:
[0062] (1.1) Introduce SAM model as the basic segmentation engine: Considering the characteristics of complex image structure and lack of labeled samples in Chinese landscape painting, the invention selects SAM as the core tool of the image segmentation module. SAM has good zero-shot generalization ability and can efficiently perform target area segmentation in complex image structures based on a small amount of hints, which is suitable for the extraction task of elements such as mountains in Chinese landscape painting.
[0063] (1.2) Interactive prompt point labeling mechanism: Users can add foreground or background prompt points by clicking on the image with the mouse. Left-click to mark the foreground (green dot), which is the mountain area that the user wants to keep; right-click to mark the background (orange dot), which is the area that the user wants to exclude from the segmentation.
[0064] The system listens to the mouse click event, obtains the normalized coordinates in real time, and adds the click type (1 for foreground, 0 for background) to the prompt sequence. When the user clicks for the first time, the system submits the original image path and the prompt points to complete a one-time feature extraction. Subsequently, every time the user adds a prompt point, the system re-invokes the mask prediction module to generate an updated segmentation result. The pseudo code of interactive prompt point collection and mask prediction is shown in Algorithm 1.
[0065]
[0066] (1.3) Real-time segmentation feedback and mask display: To improve the efficiency of interactive response, the system will output multiple candidate masks at the same time during the first interaction, and provide an IoU-based quality score for each candidate result. The system automatically selects the mask with the highest score as the current segmentation result and caches its internal feature identifier for subsequent use. In the subsequent iteration process, the system uses the intermediate features generated by the previous segmentation as prior input for the model, thereby reducing repeated calculations, focusing on boundary refinement and local correction, and improving segmentation inference speed and computational efficiency under multiple iterations.
[0067] (2) Perspective analysis and explainability recommendation: This step is used to assist users in automatically identifying the three-distant perspective type embodied in the input landscape painting, and provides recommendation results and explanation information in a structured and visualized manner. The overall framework of the prompt word design strategy is shown in Figure 3 , which includes the following sub-steps:
[0068] (2.1) Automatic perspective analysis module design: Since the perspective judgment of traditional Chinese landscape painting highly depends on expert experience, the analysis process is time-consuming and highly subjective. Therefore, the invention designs and introduces an automatic perspective analysis module based on LLM to support users in quickly judging the dominant perspective expression form of the input image in the high-distant, deep-distant, or flat-distant method.
[0069] (2.2) Prompt word construction and role setting: The invention sets the role of the large language model as "Chinese landscape painting perspective analysis expert" through prompt word engineering, and injects the definition and discrimination points of the three-distant method to improve the model's ability to identify perspective types. At the same time, the prompt word requires the model to perform an implicit thinking process of not less than three steps internally, but prohibits explicit output of thinking steps to guide the model to complete the reasoning and judgment.
[0070] (2.3) Task decomposition and structured output: The prompt word refines the perspective recognition task into four sub-tasks: view direction judgment, spatial level analysis, scene arrangement mode inspection, and perspective method matching degree scoring, and requires the model to output results in JSON format to ensure uniform structure and complete information. The prompt word also provides standard templates and a small number of examples to improve output consistency and controllability.
[0071] The top-level fields of the JSON structure are Analysis and Perspective, respectively. Among them, the Analysis field lists the percentage of the possibility of high, deep, and level perspective methods and the corresponding Chinese interpretation, providing a fine-grained and intuitive basis for judgment; the Perspective field directly gives the final judgment result based on the overall analysis of the model. In addition, considering that there may be special cases in the actual picture that do not conform to the typical three-perspective layout, the Perspective field also allows the model to output a special value other to better reflect the uncertainty and particularity in the real analysis situation.
[0072] (2.4) Concurrent reasoning and self-consistency integration: To improve the stability and robustness of the analysis results, the system uses a concurrent calling strategy for the GPT-4o model, setting the relative higher randomness parameter temperature=0.7 to obtain multiple analysis results with diversity. The system then determines the dominant perspective type by majority voting on all results and averages the probability values of each perspective method.
[0073] (2.5) Representative explanation screening and result transmission: Finally, the system selects the explanation closest to the average judgment result in all analyses as the representative output and presents it in a structured and visualized form on the front end, including the recommended perspective type, corresponding likelihood score, and model explanation content, while transmitting the recommended result to the three-dimensional modeling module to trigger the corresponding perspective conversion process.
[0074] (3) Perspective conversion and three-dimensional modeling: This step is used to realize the effective conversion of the three-perspective method perspective logic in Chinese landscape painting to the three-dimensional modeling space, aiming to restore the subjective spatial image expressed by the landscape painting, which includes the following sub-steps:
[0075] (3.1) Two-dimensional to three-dimensional reference mapping: First, map the input Chinese landscape painting image to the basic three-dimensional modeling space to establish a unified three-dimensional coordinate system. As shown in Figure 4 , the system determines the position and size of each mountain in the image space based on its two-dimensional bounding box and represents its bottom position by the center point coordinates. Let the size of the two-dimensional image be WxH (unit: pixels), and the left-top and right-bottom coordinates of the two-dimensional bounding box of each mountain be . To convert pixel units to physical units in the three-dimensional modeling space, introduce a global scaling factor (unit: cm / pixel), which is set to 0.1 according to the experience of three-dimensional designers. Accordingly, the overall size of the three-dimensional scene is:
[0076]
[0077] To ensure the alignment of the lower edge of the picture, a unified three-dimensional initial layout is constructed, and the bottom of the mountain is aligned to the same horizontal reference surface. The reference three-dimensional coordinates are defined as:
[0078]
[0079] To control the actual three-dimensional size of each mountain, a local normalization scaling factor is introduced to map the normalized generated three-dimensional model to the actual scale. The calculation method is as follows:
[0080]
[0081] where is the maximum value of the normalized model generated by the th primitive in the height direction in three-dimensional space, which is automatically obtained from the vertex coordinates of the model. This strategy ensures consistent scale coordination between different mountains and provides flexible regulation space for subsequent perspective correction.
[0082] (3.2) Three-far perspective correction: Based on the above unified layout, an interactive perspective correction mechanism is proposed to better simulate the three-far space characteristics in Chinese landscape painting. This mechanism supports users to adjust the relative position and scale of the mountains in the three-dimensional scene according to the identified "three-far" type (high-far, deep-far, and flat-far), thereby restoring the aesthetic logic of scattered point perspective in traditional painting.
[0083] According to expert opinions, the key to three-far perspective is the relative position and relative scale of the mountains. The system first calculates the depth factor of the primitive in the two-dimensional image based on the vertical bottom position to represent its depth in three-dimensional space. The higher the primitive, the farther it is, and the larger the value.
[0084] Further, the system introduces three global parameters to control the three-dimensional layout characteristics under different perspective methods:
[0085] A, : control the horizontal distribution density in the X-axis direction;
[0086] B, : control the vertical depth level rhythm in the Z-axis direction;
[0087] C, : control the overall three-dimensional scaling ratio, affecting the spatial tension and visual softness.
[0088] (3.3) Local scale correction of three-dimensional scene: On the basis of the above unified layout and global parameter control, the system further corrects the scale of each mountain according to the visual performance law of the three-distance method, as follows:
[0089] A, high-distance method: emphasizing the vertical stretching of the mountain, showing the effect of looking up at the towering mountain. Define the vertical scaling coefficient wherein is the adjustable parameter provided by the system (default 1.5), indicating the maximum stretching strength, and the corrected coordinates are:
[0090]
[0091] B, deep-distance method: focusing on the staggered front and back levels, using a sine function to control the scale perturbation to avoid scale mutation. Define the scaling coefficient wherein is the camera's tilt angle. γD is the adjustable parameter provided by the system (default 1), simulating the near and far layers, guiding the model to produce a slight stagger in vision rather than an extreme height difference. The correction formula is:
[0092]
[0093] C, flat-distance method: using a two-section scaling strategy to enhance the contrast between the near and far scenes, simulating the horizontal space tension. The scaling coefficient is defined as:
[0094]
[0095] wherein the foreground judgment threshold is set to 0.3, and the near scene enhancement parameter default value is 2.5.
[0096] wherein is the foreground threshold, and the value 0.3 represents the foreground in the flat-distance method. is the adjustable parameter provided by the system (default 2.5), controlling the closer the foreground, the larger the scale, taking the basic scaling ratio, controlling the far scene to be appropriately enlarged. The three-dimensional scale correction is:
[0097]
[0098] (4) Analysis and evaluation based on the system: by developing a three-dimensional scene interactive generation system for Chinese landscape paintings based on the three-distance method, and setting up typical landscape modeling tasks, the three-dimensional scenes generated by the system are analyzed and objectively evaluated, and data such as generation efficiency and layout restoration degree are collected to verify the effectiveness and usability of the invention. Specifically, the following sub-steps are included:
[0099] (4.1) Case analysis: To verify the operability and modeling integrity of the system in actual use scenarios, a typical user operation task flow is set, and a representative Chinese landscape painting is selected as experimental material. The user completes the interactive operation flow of the three stages of graph segmentation, perspective analysis and parameter adjustment, three-dimensional scene generation in turn according to the method, and the typical operation flow is as shown in Figure 5
[0100] A, in the graph segmentation stage, the user uses the left mouse button to click on the typical area of the mountain to be segmented, adds a green foreground prompt point, and the system immediately calls the SAM model and refreshes the segmentation mask in real time. The area not recognized as foreground is marked with a semi-transparent mask with a transparency of 190 / 255 to help the user distinguish foreground and background. If you want to delete the misselected area, you can add an orange background prompt point by right-clicking the mouse, and the system will automatically update the segmentation result. When the fine segmentation of a single graph is completed, the user clicks the "Confirm Segmentation" button in the toolbar, and the system will automatically generate a thumbnail for the segmentation result and display it in the thumbnail panel on the interface.
[0101] B, in the perspective analysis stage, the system displays the perspective recommendation probability pie chart of the three far method, and the probability of the three perspective types (high far, deep far, and flat far) is 0% initially. After the user clicks the "Start Analysis" button, the interface will prompt "Waiting for analysis results…", and the system will immediately call the backend LLM to analyze the perspective type of the input image. After the analysis is completed, the interface is updated to the recommended result view, which includes the probability distribution graph of the three perspective methods, the detailed judgment basis and analysis explanation of each perspective type. The user can click the explanation button corresponding to different perspective methods to switch to read the specific explanation content, and improve the understanding of the recommended logic.
[0102] C, in the three-dimensional generation stage, the user needs to complete the three-dimensional modeling of each mountain graph in turn before generating the complete overall scene. The system interface clearly divides the "single model generation" and "overall scene construction" operation areas, making it easy for users to flexibly switch and independently operate between local modeling and global layout.
[0103] C1, when generating a single model, the user needs to click to select the target graph from the thumbnail list of the segmentation view, and the selected graph will be displayed in the thumbnail list with a darkened border. Then, click the "Generate Model" button in the "Single Model" area, and the system will call the three-dimensional generation module to generate the corresponding three-dimensional model based on the silhouette and prompt information of the graph.
[0104] C2, when building the overall scene, the user clicks the "generate model" button in the "overall scene" area, and the system will layout all the graphic models in the three-dimensional space according to the perspective type recommended in the perspective analysis stage to generate a preliminary scene. The system defaults to the perspective method recommended in the perspective analysis stage to build the overall space structure. Users can adjust the overall layout according to specific creative needs. The system provides parameter sliders to control perspective parameters and supports switching different three-far method types to observe the differences in composition.
[0105] (4.2) Generation result analysis: To further evaluate the applicability and user experience of the system in assisting the construction of three-dimensional scenes of Chinese landscape paintings, a number of participants were organized to complete the preset modeling tasks, covering a number of typical landscape painting composition cases. The system records the key operation behaviors of the participants and saves their final generated three-dimensional modeling results as an important basis for evaluating the modeling effect.
[0106] Figure 6 For the three-far perspective analysis results completed by some participants and the corresponding three-dimensional scene generation examples, the conversion process of classical landscape paintings from two-dimensional pictures to three-dimensional space layout under the assistance of the system is shown.
[0107] In this experiment, the average task time of 32 participants was 15.23 minutes. Compared with the traditional manual modeling, which often takes several hours to build a basic scene, the interactive scene generation method proposed in this study has higher time efficiency. Especially for beginners, the traditional modeling method not only has complex operation and steep learning curve, but also it is difficult to produce ideal results before mastering the perspective structure specific to landscape paintings. The system reduces the entry threshold in terms of guided learning.
[0108] In terms of task stability, the average number of user errors was 2.06, which was at a relatively low level. This result shows that the interactive graphic segmentation tool based on the SAM model in the system has good design rationality, and users can usually obtain ideal segmentation results through a small amount of interaction without repeated attempts, thereby effectively improving the smoothness of task execution.
[0109] Regarding perspective switching, participants switched an average of 4.19 times, indicating a certain exploratory tendency. Five users, based on the initial automatically recommended perspective, determined their final viewpoint with no more than three switches, while the remaining users tended to compare and judge multiple perspective schemes more thoroughly. Brief interviews revealed that some users wanted to actively explore the impact of different perspective methods on the overall spatial layout of the image through switching, to aid their composition decisions; while users with experience in landscape painting modeling preferred to directly adopt the system's recommendations and then fine-tune detailed parameters to efficiently achieve a final composition that suited their personal aesthetic. Although some users switched perspectives during the process, all users ultimately chose the system-recommended perspective scheme as the basis for modeling, reflecting the excellent reliability of the system's recommendations.
[0110] (4.3) Objective evaluation of the generated results: To verify the consistency between the system-generated scene and the original landscape painting in terms of layout from the perspective of spatial structure, the following two dimensions of quantitative evaluation indicators are proposed:
[0111] A. Two-dimensional mask reprojection comparison:
[0112] For 2D mask comparison, the Intersection over Union (IoU) metric is used to perform pixel-level matching of key primitives in the original image and the reprojected image of the system-generated scene, thereby evaluating the similarity of primitives in terms of relative position and contour shape. Specifically, after the user generates a satisfactory 3D scene, the 3D model is projected according to the current viewpoint to generate a corresponding 2D mask image, which is then compared with the primitive masking region obtained by SAM segmentation in the original image. The formula for calculating IoU is as follows:
[0113]
[0114] in, This refers to the masked area after the 3D model is projected. This represents the actual masked area of the corresponding primitive in the original image. The average IoU of the main primitives in the user-generated artwork reaches 0.761, indicating that the system has high consistency in restoring the spatial position and two-dimensional contour of the primitives.
[0115] B. Primitive Occlusion Relationship Matrix:
[0116] In terms of primitive occlusion relationship analysis, this paper employs manual annotation to construct a reference occlusion relationship matrix. Specifically, evaluators with a background in Chinese landscape painting, based on the compositional logic of the original image, the relative positions of primitives, and visual hierarchy, analyze each pair of primitives. and Determine the occlusion relationship from the current viewpoint. If we consider the primitive view... The upper part of the senses is located If the front is blocked and the body is blocked, it is recorded as , otherwise . The annotation result is the true occlusion relationship, which is compared with the artificially annotated occlusion matrix in the final three-dimensional scene view to calculate the accuracy of the three-dimensional scene occlusion relationship and quantify its performance in spatial structure restoration.
[0117] The occlusion relationship matrix generated by the system is consistent with the artificially annotated result, with an average accuracy of 99.4%.
[0118] Although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent replacements for part of the technical features, and any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A method for interactive generation of a three-dimensional scene, characterized in that, Comprising the following steps: Step (1) SAM-based interactive image segmentation: Extract key primitives in the picture dominated by mountains, and generate segmentation silhouettes corresponding to the primitives; Step (2) LLM-based three-distance method perspective analysis of landscape painting images: Output the probability distribution and interpretable analysis results of the three perspective methods of high distance, deep distance, and flat distance, and recommend the optimal perspective method; Step (3) Three-dimensional scene generation based on three-distance method perspective: Propose a three-dimensional layout algorithm based on the three-distance method, combine the perspective type selected by the user and the adjustment parameters, and arrange the three-dimensional models generated by the segmented primitives to construct an interactive fine-tunable three-dimensional scene of landscape paintings; Step (4) Analysis and evaluation based on the system: Taking the scene modeling task of a typical landscape painting as an example, analyze and objectively evaluate the three-dimensional scene generated by the system, collect data such as generation efficiency and layout restoration degree, and verify the effectiveness and usability of the method.
2. The method of claim 1, wherein, The step (1) comprises the following sub-steps: Step (1.1) Upload the landscape painting image and enter the interactive segmentation view; the user marks the foreground and background hint points by clicking, and the system calls SAM for image segmentation in real time to generate the corresponding mask region; Step (1.2) After the user confirms the segmentation region, the system generates primitive thumbnail images and saves the primitive number and position layout information for subsequent analysis and modeling input.
3. The method of claim 1, wherein: The step (2) comprises the following sub-steps: Step (2.1) The system calls the LLM model to perform three-distance method perspective analysis reasoning based on the overall structure of the image, primitive arrangement, hierarchical relationship, and other visual features, laying the foundation for subsequent perspective judgment; Step (2.2) The system outputs the score distribution of the three perspective methods of high distance, deep distance, and flat distance, and selects the optimal recommended perspective method through a majority voting mechanism, clearly defining the preliminary perspective direction; Step (2.3) Select the most representative model analysis explanation based on the average value of the votes, and present it in JSON format to enhance the consistency and interpretability of the results; Step (2.4) The user can view the analysis results of each perspective method and choose whether to accept the recommended perspective method, and the system passes the selected perspective type to the next stage.
4. The method of claim 1, wherein: The step (3) comprises the following sub-steps: Step (3.1) For each primitive, call the TripoSR graph generation three-dimensional model interface, generate a normalized three-dimensional mesh model according to the primitive silhouette, and record its standard height value; Step (3.2) According to the perspective type selected by the user and the system preset parameters, combine the two-dimensional position of the primitive and the depth factor to calculate its three-dimensional coordinates and scaling ratio using the perspective conversion formula; Step (3.3) Arrange all primitive three-dimensional models in a unified layout, apply adjustable layout parameters including horizontal sparsity, vertical rhythm, and overall scaling coefficient, etc., to simulate the three-distance method space semantics; Step (3.4) The user can adjust various parameters in real time through the slider in the generated view, and the system dynamically updates the three-dimensional scene to realize the interactive composition experience of what you see is what you get.
5. The method of claim 1, wherein: The step (4) comprises the following sub-steps: Step (4.1) case analysis: To verify the operability and modeling integrity of the system in actual use scenarios, a typical user operation task flow is set, and a representative Chinese landscape painting is selected as experimental material; the user completes the interactive operation process of the three stages of graphic segmentation, perspective analysis and parameter adjustment in sequence according to the method; In the graphic segmentation stage, the user extracts the foreground mountain area based on the prompt point control through the interactive interface; in the perspective analysis stage, the system analyzes the perspective method according to the composition characteristics of the picture, and recommends the most matched three-far method; in the three-dimensional generation stage, the system completes the spatial layout according to the input, and allows the user to make sliding parameter fine tuning; finally, the generated three-dimensional landscape painting scene is successfully exported to the Blender software environment for subsequent immersive scene display and space interaction building; Step (4.2) generation result analysis: To further evaluate the applicability and user experience of the system in assisting the construction of three-dimensional scenes of Chinese landscape paintings, a number of subjects are organized to complete the preset modeling task, covering a number of typical landscape painting composition cases; the system records the key operation behavior of the participants and saves their final three-dimensional modeling results as an important basis for evaluating the modeling effect; by analyzing the three-dimensional scenes output by each group of users, the adaptability of the system to different styles of landscape paintings is demonstrated; Step (4.3) objective evaluation of the generated results: To verify the consistency of the system-generated scene and the original landscape painting in terms of layout from the spatial structure level, the following two-dimensional quantitative evaluation indicators are proposed: A. Two-dimensional mask projection comparison: Select the key graphics extracted by the user, and perform view projection on the three-dimensional scene generated by the system to obtain a two-dimensional mask; compare the mask with the mask area annotated by experts in the original image at the pixel level, and calculate the intersection over union (IoU) value to measure the fitting degree of the graphics in terms of relative position and boundary contour; B. Graphic occlusion relationship matrix: According to the composition order of the original picture, a reference occlusion relationship matrix annotated by experts is constructed to count the front and back occlusion relationships between graphics; compare this matrix with the occlusion relationship matrix automatically derived from the three-dimensional scene generated by the system, and calculate the occlusion matching rate to evaluate the restoration accuracy of the system-generated scene in terms of depth and three-far perspective relationship.
6. The method of claim 2, wherein: In step (1.1), the SAM model has zero-shot generalization capability, and the first click submits the original picture path and the prompt point to complete feature extraction. In subsequent iterations, the intermediate features from the previous round are used as prior inputs to reduce repeated calculations. The system outputs multiple candidate masks and provides an IoU quality score during the first interaction, and automatically selects the highest scoring result.
7. The method of claim 3, wherein: In step (2.1), the LLM is set to "Chinese landscape painting perspective analysis expert" through prompt word engineering, which injects the definition and discrimination points of the three-far method and requires at least three-step implicit thinking; in step (2.3), the JSON structure includes the Analysis field and the Perspective field, and the system uses a concurrent calling strategy (temperature=0.7) to obtain multiple results and perform average aggregation.
8. The method of claim 4, wherein: In step (3.2), the two-dimensional to three-dimensional reference mapping introduces a global scaling factor (0.1 cm / pixel), and the overall size of the three-dimensional scene is ; the reference three-dimensional coordinates of the mountain are defined as: wherein are the coordinates of the upper left and lower right corners of the two-dimensional bounding box of the mountain.
9. The method of claim 4, wherein: In the step (3.3), the three-far method perspective correction is specifically: HighFar: vertical scale factor , corrected coordinates , is a HighFar tunable parameter, default 1.5, where represents the depth of the mountain in the 3D scene, i.e. the lower the bottom of the primitive in the 3D landscape, the greater the depth ; Deepth method: scaling factor , corrected coordinates , φ is the camera tilt angle, is the Deepth method adjustable parameter; Far distance method: adopt double section zoom strategy, , foreground threshold θ = 0.3, near scene enhancement parameter Default 2.5, far scene parameter , corrected coordinates .
10. The method of claim 5, wherein: In the step (4.3), the IoU calculation formula is: ; wherein is a projected mask region of the three-dimensional model, is a true mask region of the corresponding primitive in the original drawing; In the occlusion relationship matrix, if the primitive Gi is in front of Gj and constitutes occlusion, it is recorded as , otherwise it is 0. The accuracy rate is calculated by comparing it with the expert annotation matrix.