Method and System for Rapid Construction of Metaverse Scenes Based on Modular Engines

Through the rapid construction method of metacosmic scenes based on a modular engine, combined with deep learning and multi-module collaborative work, the problems of low scene construction efficiency and unreal rendering effects in the existing technology are solved, and efficient and automated scene creation and realism are achieved.

CN120032034BActive Publication Date: 2025-07-01HANGZHOU MOXI TECH DEV CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510488206.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-01
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

The existing metacosmic scene construction methods lack modularity and standardization, resulting in low development efficiency, difficulty in ensuring the consistency of scene quality, and insufficient rendering effect, especially in complex scenes, lighting interaction and environmental special effects are difficult to accurately simulate.

Method used

The rapid construction method of metacosmic scenes based on modular engines is adopted. By obtaining user creation instructions, the scene rendering module in the preset modular engine library is called, combined with the scene intelligent analysis model of deep learning, the scene is processed in a layered manner, and the hierarchical relationship is established, and the physical engine module, lighting rendering module and environment special effects module work together to give the scene functional layer real physical characteristics, lighting effects and environmental atmosphere.

Benefits of technology

It realizes efficient creation and rendering of scenes, improves the degree of automation and the realism of rendering effects, reduces repeated development work, makes scene construction more convenient and flexible, and enhances user immersion and interactive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032034B_ABST
    Figure CN120032034B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for quickly building a metaverse scene based on a modular engine, relating to the technical field of the metaverse. It includes obtaining a creation instruction of the user for the metaverse scene, and calling a corresponding scene rendering module from a preset modular engine library according to the scene type information; using a scene intelligent analysis model based on deep learning to extract features and perform scene layering on the scene environment parameters, and parsing them into multiple scene function layers with a progressive hierarchical relationship; and calling a physics engine module, a lighting rendering module, and an environmental special effect module to build the scene. The present invention realizes the quick construction of the metaverse scene through the modular engine and the intelligent analysis model, improves the scene construction efficiency, and enhances the scene realism and interactive experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the metaverse technology, and in particular to a method and system for quickly building a metaverse scene based on a modular engine. Background Art

[0002] With the rapid development of the metaverse technology, the construction and rendering of virtual scenes have become an important research direction. As an important carrier for users to conduct virtual interactions, the construction quality of the metaverse scene directly affects the immersive experience of users. Currently, the construction of the metaverse scene mainly relies on traditional 3D modeling software and game engines, and developers need to manually complete tasks such as scene modeling, material setting, and lighting rendering. This traditional scene construction method has been difficult to meet the current rapid development needs of the metaverse.

[0003] The main problems existing in the prior art are as follows:

[0004] Traditional scene construction methods lack a modular and standardized technical framework. Developers need to repeat the construction work of similar scenes, resulting in low development efficiency and difficulty in ensuring the consistency of scene quality. Each link in the scene construction process is often fragmented, lacking a unified management and scheduling mechanism.

[0005] Existing scene rendering technologies do not fully consider the hierarchical structure and environmental constraint relationships of the scene, resulting in insufficiently realistic rendering effects. Especially in complex scenes, it is difficult to accurately simulate the light interaction and environmental special effects between different levels, affecting the realism and immersion of the scene. Summary of the Invention

[0006] Embodiments of the present invention provide a method and system for quickly building a metaverse scene based on a modular engine, which can solve the problems in the prior art.

[0007] In a first aspect of an embodiment of the present invention, a method for quickly building a metaverse scene based on a modular engine is provided, including:

[0008] Obtaining a creation instruction for the metaverse scene from a user, where the creation instruction includes scene type information and scene environment parameters;

[0009] According to the scene type information, calling a corresponding scene rendering module from a preset modular engine library, where the scene rendering module includes a physics engine module, a lighting rendering module, and an environmental special effect module;

[0010] A scene intelligent analysis model based on deep learning extracts features and performs scene layering on the scene environment parameters, parsing the scene environment parameters into multiple scene function layers. Each scene function layer includes corresponding geometric information, material information, interaction rules, and environmental constraint conditions. A progressive hierarchical relationship is formed between the scene function layers, and the environmental constraint conditions of the upper scene function layer optimize and adjust the rendering results of the lower scene function layer.

[0011] Call the physical engine module to construct a three-dimensional mesh model according to the geometric information and material information of each scene function layer, and implant corresponding physical property parameters into each three-dimensional mesh model.

[0012] Call the lighting rendering module to perform lighting calculation and shadow mapping on each three-dimensional mesh model based on the hierarchical relationship of the scene function layers.

[0013] Call the environmental special effect module to add corresponding environmental special effects to each scene function layer according to the scene environment parameters.

[0014] Calling the corresponding scene rendering module from the preset modular engine library according to the scene type information includes:

[0015] Receive the scene type information and perform encoding processing on the scene type information through a pre-trained BERT model to obtain a scene feature vector.

[0016] Obtain candidate rendering modules from the modular engine library, and calculate the matching degree between the candidate rendering modules and the scene feature vector. The calculation formula for the matching degree is:

[0017] ;

[0018] Where, M represents the set of candidate rendering modules, S represents the set of scene feature vectors, α represents the balance factor, n represents the feature dimension, w i represents the weight coefficient of the i-th dimension feature, sim(m i ,s i ) represents the similarity between the i-th dimension candidate rendering module and the i-th dimension scene feature vector, β represents the balance factor, C(M) represents the calculation complexity evaluation value corresponding to the set of candidate rendering modules, C() represents the complexity evaluation function;

[0019] Select the candidate rendering module with the highest matching degree as the scene rendering module corresponding to the scene type information.

[0020] Based on the deep learning-based scene intelligent analysis model, extract features and perform scene layering on the scene environment parameters, and parse the scene environment parameters into multiple scene function layers. Each scene function layer includes corresponding geometric information, material information, interaction rules, and environmental constraint conditions, including:

[0021] Extract the geometric features and material features in the scene environment parameters respectively through a multi-modal feature extraction network. Among them, the geometric features are extracted from the scene point cloud set through the PointNet++ network, and the material features are extracted from the material texture map through a multi-scale convolutional neural network;

[0022] Adopt a multi-head attention mechanism to perform feature fusion on the geometric features and the material features. Use the geometric features as the query vector, the material features as the key vector, and the concatenation result of the geometric features and the material features as the value vector to generate fused features;

[0023] Perform hierarchical parsing on the fused features, divide the scene into multiple function layers, and establish the hierarchical relationship between the function layers;

[0024] Generate the geometric information, material information, and interaction rules of each function layer respectively through a multi-layer perceptron;

[0025] Construct a two-layer constraint optimization model. The two-layer constraint optimization model includes local constraints within the function layer and global constraints between function layers, and realizes constraint optimization by minimizing the weighted combination of the constraint expected deviation and the constraint difference between adjacent function layers.

[0026] Construct a two-layer constraint optimization model. The two-layer constraint optimization model includes local constraints within the function layer and global constraints between function layers, and realizes constraint optimization by minimizing the weighted combination of the constraint expected deviation and the constraint difference between adjacent function layers, including:

[0027] Construct the local constraint features of the scene function layer. The local constraint features include geometric constraint features and material constraint features. Among them, the geometric constraint features include position continuity constraint, normal vector consistency constraint, and boundary continuity constraint, and the material constraint features include texture transition constraint, roughness continuity constraint, and reflectivity consistency constraint;

[0028] Construct the global constraint features of the scene function layer. The global constraint features include inter-layer topological constraints and physical constraints. Among them, the inter-layer topological constraints calculate the constraint relationship between adjacent function layers through the inter-layer association weight and the topological relationship metric function, and the physical constraints include gravity action constraint, collision detection constraint, and structural stability constraint;

[0029] Construct a multi-level optimization objective function based on the local constraint features and the global constraint features, take the weighted combination of the geometric constraints and material constraints of the local constraint features as the local optimization objective, and take the weighted combination of the topological constraints and physical constraints of the global constraint features as the global optimization objective;

[0030] Adopt an adaptive weight adjustment mechanism to dynamically balance the local optimization objective and the global optimization objective, and the adaptive weight is adjusted exponentially based on the ratio of the local optimization objective value to the global optimization objective value;

[0031] Iteratively solve the multi-level optimization objective function by the alternating direction multiplier method until the difference between the optimization results of two adjacent iterations is less than the preset deviation threshold;

[0032] Calculate the local constraint satisfaction degree and the global constraint satisfaction degree of each functional layer, and both the local constraint satisfaction degree and the global constraint satisfaction degree are calculated using a Gaussian evaluation function.

[0033] Call the physical engine module, construct a three-dimensional mesh model according to the geometric information and material information of each scene functional layer, and implant corresponding physical property parameters in each three-dimensional mesh model, including:

[0034] Construct an adaptive tetrahedral mesh according to the geometric information, evaluate the quality of the tetrahedral mesh and optimize the mesh elements through the volume ratio of the mesh elements, the dihedral angle of the mesh elements and the side length ratio of the mesh elements to generate a high-quality tetrahedral mesh;

[0035] Calculate the physical property parameters based on the material information, calculate the mass density, elastic modulus, Poisson's ratio and friction coefficient of the scene functional layer respectively through the material-physics mapping network, and establish the corresponding relationship between the material information and the physical property parameters;

[0036] Allocate the physical property parameters to each mesh element of the high-quality tetrahedral mesh to generate a three-dimensional physical mesh model with physical characteristics.

[0037] Call the light rendering module, and perform light calculation and shadow mapping on each three-dimensional mesh model based on the hierarchical relationship of the scene functional layer, including:

[0038] The hierarchical relationship includes the spatial position relationship and occlusion relationship between functional layers;

[0039] Construct a hierarchical light propagation tree based on the hierarchical relationship, the nodes of the hierarchical light propagation tree represent the scene functional layers, the connecting edges between the nodes represent the light propagation paths, and assign light attenuation weights to each light propagation path according to the spatial position relationship and occlusion relationship;

[0040] According to the structure of the hierarchical light propagation tree, calculate the direct light and indirect light of each three-dimensional grid model, generate a light interaction matrix, and the light interaction matrix records the light energy transfer relationship between different functional layers;

[0041] Generate a depth map for each three-dimensional grid model based on the light interaction matrix, and combine the light attenuation weight to calculate a shadow mapping matrix, and dynamically update and optimize the light distribution of each functional layer according to the shadow mapping matrix.

[0042] In the second aspect of the embodiments of the present invention, a metaverse scene rapid construction system based on a modular engine is provided, including:

[0043] A first unit for obtaining a creation instruction for a metaverse scene, where the creation instruction includes scene type information and scene environment parameters;

[0044] A second unit for calling a corresponding scene rendering module from a preset modular engine library according to the scene type information, where the scene rendering module includes a physical engine module, a light rendering module, and an environmental special effect module;

[0045] A third unit for performing feature extraction and scene layering on the scene environment parameters based on a scene intelligent analysis model of deep learning, and parsing the scene environment parameters into multiple scene functional layers, where each scene functional layer includes corresponding geometric information, material information, interaction rules, and environmental constraint conditions; a progressive hierarchical relationship is formed between the scene functional layers, and the environmental constraint conditions of the upper scene functional layer will optimize and adjust the rendering results of the lower scene functional layer;

[0046] A fourth unit for calling the physical engine module, constructing a three-dimensional grid model according to the geometric information and material information of each scene functional layer, and implanting corresponding physical attribute parameters into each three-dimensional grid model;

[0047] A fifth unit for calling the light rendering module to perform light calculation and shadow mapping on each three-dimensional grid model based on the hierarchical relationship of the scene functional layers;

[0048] A sixth unit for calling the environmental special effect module to add corresponding environmental special effects to each scene functional layer according to the scene environment parameters.

[0049] In the third aspect of the embodiments of the present invention

[0050] Provide an electronic device, including:

[0051] A processor;

[0052] A memory for storing processor-executable instructions;

[0053] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0054] According to a fourth aspect of the embodiments of the present invention,

[0055] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.

[0056] The beneficial effects of this application are as follows:

[0057] The present invention can realize efficient creation and rendering of scenes through a rapid construction method of metaverse scenes based on a modular engine, improve the automation of scene construction and the realism of rendering effects, which is specifically reflected in the following aspects:

[0058] Through the combined call of the preset modular engine library and scene rendering module, the standardization and modularization of the scene construction process is achieved, which greatly improves the efficiency of scene creation, reduces repetitive development work, and makes the scene construction process more convenient and flexible.

[0059] The scene intelligent analysis model based on deep learning is used to process the scene in layers, and a progressive hierarchical relationship is established, so that the various functional layers of the scene can be optimized and complemented, which improves the coordination and realism of the overall rendering effect, and also facilitates the subsequent scene maintenance and updating.

[0060] Through the collaborative work of the physical engine module, lighting rendering module and environmental special effects module, each functional layer in the scene is endowed with real physical properties, lighting effects and environmental atmosphere, greatly enhancing the user's immersion and interactive experience in the metaverse scene, making the scene more vivid and realistic. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 This is a flow chart of a method for quickly building a metaverse scene based on a modular engine according to an embodiment of the present invention;

[0062] Figure 2 A logical diagram for building a two-level constrained optimization model. DETAILED DESCRIPTION

[0063] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are only a part rather than all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0064] The technical solutions of the present invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0065] Figure 1 The following is a schematic flowchart of a method for quickly building a metaverse scene based on a modular engine according to an embodiment of the present invention, as Figure 1 shown. The method includes:

[0066] Obtain a creation instruction for a metaverse scene from a user, where the creation instruction includes scene type information and scene environment parameters;

[0067] According to the scene type information, call a corresponding scene rendering module from a preset modular engine library, where the scene rendering module includes a physics engine module, a lighting rendering module, and an environmental special effect module;

[0068] Based on a scene intelligent analysis model of deep learning, perform feature extraction and scene layering on the scene environment parameters, and parse the scene environment parameters into multiple scene function layers. Each scene function layer includes corresponding geometric information, material information, interaction rules, and environmental constraint conditions; a progressive hierarchical relationship is formed between the scene function layers, and the environmental constraint conditions of the upper scene function layer will optimize and adjust the rendering results of the lower scene function layer;

[0069] Call the physics engine module, construct a three-dimensional mesh model according to the geometric information and material information of each scene function layer, and implant corresponding physical attribute parameters into each three-dimensional mesh model;

[0070] Call the lighting rendering module, and perform lighting calculation and shadow mapping on each three-dimensional mesh model based on the hierarchical relationship of the scene function layers;

[0071] Call the environmental special effect module, and add corresponding environmental special effects to each scene function layer according to the scene environment parameters.

[0072] Exemplarily, the system obtains a user's creation instruction for a metaverse scene. The creation instruction includes scene type information and scene environment parameters. The scene type information can be multiple preset scene categories, such as urban scenes, natural scenes, science fiction scenes, historical scenes, etc.; the scene environment parameters include but are not limited to time settings (such as daytime, dusk, night), weather conditions (such as sunny, rainy, snowy), geographical environment (such as plains, mountains, oceans), and the size of the scene. The system provides a graphical interface through which the user can select the scene type and adjust various environmental parameters, or input the creation instruction in the form of natural language description. The system converts it into standardized scene type information and environmental parameters through natural language processing technology.

[0073] According to the obtained scene type information, the system calls the corresponding scene rendering module from a preset modular engine library. The modular engine library contains multiple rendering modules optimized for different scene types, and each scene type corresponds to a specific combination of physical engine modules, lighting rendering modules, and environmental special effect modules. The physical engine module is responsible for simulating the physical behavior of objects in the scene, including physical rules such as gravity, collision detection, and hydrodynamics; the lighting rendering module is responsible for processing the lighting effects in the scene, including optical phenomena such as global illumination, local illumination, reflection, and refraction; the environmental special effect module is responsible for generating special effects in the scene, such as fog, rain, snow, fire, and smoke. The system automatically selects the most suitable module combination for the current scene type through a preset mapping relationship table between the scene type and the rendering module, and at the same time supports the user to manually adjust the module selection to meet specific requirements.

[0074] Based on a deep learning-based scene intelligent analysis model, the system extracts features and layers the scene environment parameters. Through learning a large amount of existing metaverse scene data, it can identify key features in the scene environment parameters and perform reasonable layering. The analysis process first converts the scene environment parameters into high-dimensional feature vectors, then performs feature mapping and dimensionality reduction through a multi-layer perceptron, and finally parses the scene environment parameters into multiple scene function layers. Each scene function layer includes corresponding geometric information (such as terrain height map, building distribution map), material information (such as surface texture, building material), interaction rules (such as behavior definitions of interactive objects), and environmental constraint conditions (such as lighting intensity limits, physical behavior constraints).

[0075] A progressive hierarchical relationship is formed among these scene function layers. Generally starting from the basic terrain layer, the vegetation layer, building layer, decoration layer, special effect layer, etc. are constructed successively upwards. The environmental constraint conditions of the upper scene function layer will optimize and adjust the rendering results of the lower scene function layer. For example, the shadow projection of the building layer will affect the lighting effect of the ground layer; the rain and snow effects of the special effect layer will change the material performance of the ground layer and the building layer. The system processes the influence relationship between each layer through a recursive method to ensure the consistency and realism of the rendering results of each layer. During the layering process, the system will also automatically adjust the number of layers according to the scene complexity. For simple scenes, only 3-4 function layers may be required, while for complex scenes, more than 10 function layers may be needed to achieve fine control.

[0076] The system calls the physical engine module to construct a three-dimensional mesh model based on the geometric information and material information of each scene function layer. The physical engine module first converts the geometric information of each scene function layer into point cloud data, and then generates a preliminary three-dimensional mesh through a triangulation algorithm. For complex geometric bodies, the system uses subdivision surface technology to further optimize the mesh quality and ensure the smoothness of the model surface. The material information is applied to the mesh model through UV mapping technology to achieve precise texture correspondence. After the mesh model is constructed, the system implants corresponding physical property parameters into each three-dimensional mesh model, including mass, density, elastic coefficient, friction coefficient, etc. These physical property parameters are automatically matched according to the material information. For example, metal materials will be given a higher density and a lower elastic coefficient, while wood materials are the opposite. For special materials, the system provides a manual adjustment interface for physical property parameters, allowing users to precisely control the physical behavior characteristics of objects in the scene.

[0077] The system calls the lighting rendering module to perform lighting calculation and shadow mapping on each three-dimensional mesh model based on the hierarchical relationship of the scene function layers. The lighting rendering module first determines the main light source attributes (such as sunlight or moonlight) according to the time setting and weather conditions in the scene environment parameters, including the light source position, intensity, color, etc. Then, a layered rendering strategy is adopted, starting from the bottommost function layer and performing lighting calculations layer by layer. For each layer, the system applies a real-time global illumination algorithm to calculate the effects of direct illumination and indirect illumination, and generates a lighting map to store the calculation results to improve the rendering efficiency. Shadow mapping uses cascaded shadow mapping technology to automatically adjust the shadow resolution according to the viewing distance, using high-precision shadows for the near distance and low-precision shadows for the far distance to balance the visual effect and performance consumption. During the multi-layer rendering process, the shadows of upper-layer objects will be correctly projected onto lower-layer objects to achieve a realistic lighting effect in complex scenes. The system also supports physically based rendering technology to simulate the reflection, refraction, and scattering characteristics of different materials to light, enhancing the realism of the rendering.

[0078] The system call environment special effect module adds corresponding environment special effects to each scene function layer according to the scene environment parameters. The environment special effect module contains a variety of preset special effect templates, such as fog system, precipitation system, wind power system, day and night cycle system, etc. The system automatically selects a suitable special effect combination according to the weather conditions, time settings, etc. in the scene environment parameters. For example, in a rainy day scene, the system will enable the precipitation system to generate raindrop particles, and at the same time enable the wet ground effect to enhance the reflective characteristics of the material. The special effect generation uses a physics-based particle system, which can simulate the real behavior of particles in the environment, such as raindrop falling, water accumulation formation, fog diffusion, etc. In addition, the environment special effects also include atmosphere rendering, which creates a specific scene atmosphere by adjusting the overall color tone, post-processing filters, etc. Each scene function layer receives different environment special effect influences according to its characteristics. For example, the water surface layer will produce a rippling reaction to the raindrop special effect, and the vegetation layer will produce a swaying reaction to the wind power special effect. The system ensures the consistency and coordination of the special effect performance with the overall style of the scene through the adaptive adjustment of the special effect parameters.

[0079] In an alternative embodiment,

[0080] Invoking the corresponding scene rendering module from the preset modular engine library according to the scene type information includes:

[0081] Receiving the scene type information, and encoding and processing the scene type information through a pre-trained BERT model to obtain a scene feature vector;

[0082] Obtaining candidate rendering modules from the modular engine library, and calculating the matching degree between the candidate rendering modules and the scene feature vector. The calculation formula of the matching degree is:

[0083] ;

[0084] Wherein, M represents the set of candidate rendering modules, S represents the set of scene feature vectors, α represents the balance factor, n represents the feature dimension, w i represents the weight coefficient of the i-th dimensional feature, sim(m i ,s i ) represents the similarity between the i-th dimensional candidate rendering module and the i-th dimensional scene feature vector, β represents the balance factor, C(M) represents the calculated complexity evaluation value corresponding to the set of candidate rendering modules, C() represents the complexity evaluation function;

[0085] Selecting the candidate rendering module with the highest matching degree as the scene rendering module corresponding to the scene type information.

[0086] Exemplarily, the scene type information may be natural language descriptions such as "futuristic city square", "ancient oriental palace", etc., or a combination of tags selected from predefined categories. For example, a user may input "a foggy medieval European town, rainy day, at dusk" as the scene type information.

[0087] After receiving the scene type information, the system encodes it through a pre-trained BERT model. This BERT model adopts the basic BERT-base architecture, which contains 12 layers of Transformer encoders, with 12 attention heads in each layer and a hidden layer dimension of 768. The model has undergone two-stage training: first, it is pre-trained on a general corpus, and then it is fine-tuned using approximately 500,000 metaverse scene description corpora for domain adaptation. In the fine-tuning process, the Adam optimizer with a learning rate of 2e-5 is used, the batch size is 32, and it is trained for 10 epochs. After inputting the scene type information into the model, the system extracts the hidden state corresponding to the CLS token from the last layer, and then maps it to a 128-dimensional scene feature vector through a linear layer. For the above example, the obtained scene feature vector may have relatively high activation values in dimensions such as "medieval", "Europe", "fog", "rainy day", "dusk", etc., which are 0.82, 0.79, 0.65, 0.88, and 0.71 respectively.

[0088] After obtaining the scene feature vector, the system retrieves candidate rendering modules from the modular engine library. The modular engine library stores various pre-developed rendering modules, and each module contains its functional feature description, performance parameters, and applicable scene tags. The system first performs a preliminary screening through keyword matching and selects 20 potentially relevant candidate modules from 500 rendering modules in the library. These candidate modules include "medieval European architecture renderer", "rainy day special effect engine", "fog simulator", "dusk lighting system", etc.

[0089] The system calculates the matching degree between the candidate rendering modules and the scene feature vector. The matching degree calculation comprehensively considers two factors: functional matching and computational complexity. For functional matching, the system calculates the similarity of each dimension of the feature and performs weighted summation. In the specific process, the system sets weight coefficients for the 128 dimensions of the feature vector. The weight of important dimensions such as scene style and main environmental elements is relatively high (0.08 - 0.12), and the weight of secondary dimensions such as detail special effects is relatively low (0.01 - 0.05). The similarity calculation uses the cosine similarity method, with a value range of -1 to 1, and the closer to 1, the more similar.

[0090] For example, for the "Medieval European Architecture Renderer" module, its eigenvalue in the "Medieval" dimension is 0.95. Calculating the similarity with 0.82 of the scene feature vector gives 0.96; its eigenvalue in the "European" dimension is 0.88, and calculating with 0.79 of the scene feature gives 0.94. Assuming the weights of these two dimensions are 0.10 and 0.09 respectively, the weighted similarity contribution of these two dimensions is 0.10×0.96 + 0.09×0.94 = 0.1786. Similar calculations are performed and summed for all 128 dimensions to obtain the total function matching score of this module as 0.76.

[0091] In terms of computational complexity, the system uses a pre-trained performance prediction model to evaluate the resource consumption of each rendering module. This model is trained based on historical operation data, with the input being module parameters (such as the number of polygons, texture resolution, complexity of lighting algorithms, etc.), and the output being the predicted CPU occupancy rate, GPU occupancy rate, memory occupancy, and frame rate impact. These metrics are weighted and synthesized to obtain a standardized complexity evaluation value, with a value range of 0 - 1. The larger the value, the higher the complexity. For example, the complexity evaluation value of the "Medieval European Architecture Renderer" is 0.65, indicating a relatively high resource consumption.

[0092] The final matching degree is calculated through the balance of the function matching score and the complexity evaluation value. The balance factor α is used to adjust the importance of function matching, usually taking values from 0.7 to 0.9; the balance factor β is used to adjust the impact of complexity, usually taking values from 0.1 to 0.3. In this example, taking α = 0.8 and β = 0.2, the final matching degree of the "Medieval European Architecture Renderer" is 0.8×0.76 - 0.2×0.65 = 0.608 - 0.13 = 0.478. After the system calculates the matching degrees for all candidate rendering modules, it selects the module with the highest matching degree as the final result.

[0093] In an alternative implementation

[0094] Based on the deep learning-based scene intelligent analysis model, feature extraction and scene layering are performed on the scene environment parameters, and the scene environment parameters are parsed into multiple scene function layers. Each scene function layer includes corresponding geometric information, material information, interaction rules, and environmental constraint conditions, including:

[0095] The geometric features and material features in the scene environment parameters are respectively extracted through a multi-modal feature extraction network. Among them, the geometric features are extracted from the scene point cloud set through the PointNet++ network, and the material features are extracted from the material texture map through a multi-scale convolutional neural network;

[0096] The multi-head attention mechanism is used to perform feature fusion on the geometric features and the material features. The geometric features are used as query vectors, the material features are used as key vectors, and the concatenation result of the geometric features and the material features is used as value vectors to generate fused features;

[0097] The fused features are hierarchically analyzed, the scene is divided into multiple functional layers, and the hierarchical relationship between the functional layers is established;

[0098] The geometric information, material information, and interaction rules of each functional layer are respectively generated through a multi-layer perceptron;

[0099] A two-layer constraint optimization model is constructed. The two-layer constraint optimization model includes local constraints within the functional layer and global constraints between the functional layers, and constraint optimization is achieved by minimizing the weighted combination of the constraint expected deviation and the constraint difference between adjacent functional layers.

[0100] Figure 2 It is a logical schematic diagram for constructing a two-layer constraint optimization model. Exemplarily, the geometric features and the material features in the scene environment parameters are respectively extracted through a multi-modal feature extraction network. For the geometric features, the PointNet++ network is adopted by the system to extract from the scene point cloud set. In specific implementation, the input point cloud data usually contains millions of sampling points, and each point contains three-dimensional coordinate and normal vector information. The system first downsamples the point cloud, and uses the voxel downsampling method to reduce the point cloud density to about 1000 points per cubic meter, improving the processing efficiency while maintaining the geometric structure. The downsampled point cloud is processed by the PointNet++ network, which includes 3 hierarchical sampling modules, and each module includes a sampling layer, a grouping layer, and a PointNet layer. The first layer samples 1024 points, the second layer samples 256 points, and the third layer samples 64 points, capturing the local and global geometric features of the point cloud through layer-by-layer abstraction. The network output is a 512-dimensional geometric feature vector. For example, when processing an urban square scene, the geometric feature vector extracted may contain the representations of geometric elements such as a flat area (feature value 0.85), a stepped structure (feature value 0.72), a circular fountain (feature value 0.68), etc.

[0101] For material features, the system uses a multi-scale convolutional neural network to extract them from material maps. Material maps are usually RGB images with a resolution of 2048×2048, containing information such as color, roughness, metallicity, and normal. The system uses a multi-branch network based on ResNet-50, and each branch processes material information at different scales. The first branch processes global materials (2048×2048→256×256), the second branch processes medium-scale materials (512×512), and the third branch processes microscopic material details (128×128). The features of the three branches are fused through an attention mechanism to generate a 256-dimensional material feature vector. In the example of a city square, the extracted material features may include material attributes such as stone texture (feature value 0.91), metal decoration (feature value 0.43), and vegetation coverage (feature value 0.37).

[0102] The model uses a multi-head attention mechanism to fuse geometric features and material features. Specifically, an 8-head attention mechanism is used, and the dimension of each attention head is 64. In the attention calculation, the geometric feature vector (512-dimensional) is used as the query vector, the material feature vector (256-dimensional) is used as the key vector, and the concatenated result of the two (768-dimensional) is used as the value vector. The query vector and the key vector are first mapped to the same dimension (512-dimensional) through a linear projection layer, and then the attention weights are calculated. In the example of a city square, the attention weight between the stone material and the flat area is 0.82, and the weight with the steps is 0.79, which indicates that the stone material is mainly applied to these geometric structures. The output of the multi-head attention passes through a residual connection and layer normalization to generate a 768-dimensional fused feature vector, which contains both geometric and material correlation information.

[0103] The model hierarchically analyzes the fused features and divides the scene into multiple functional layers. The system uses a hierarchical clustering algorithm to cluster the fused feature vectors into 5-10 functional layers based on feature similarity. In the example of a city square, the system divides the scene into 5 functional layers: the basic terrain layer, the building structure layer, the decorative facility layer, the vegetation layer, and the special effect layer. After clustering, the system uses a graph neural network to establish the hierarchical relationship between the functional layers and constructs a directed acyclic graph to represent the hierarchical dependence. For example, the basic terrain layer is at the bottom layer, the building structure layer and the vegetation layer depend on the terrain layer, the decorative facility layer depends on the building layer, and the special effect layer is at the top layer and depends on all other layers.

[0104] After the layering is completed, the model generates the geometric information, material information, and interaction rules of each functional layer through a multi-layer perceptron. For geometric information generation, a 5-layer fully connected network is used. The input is the fused features (768 dimensions) of this layer, and through hidden layers with inter-layer dimensions of [768, 512, 256, 128, 64], the geometric parameters of this functional layer are output, such as vertex coordinates, face indices, etc. For material information generation, a 3-layer fully connected network is used, with hidden layer dimensions of [768, 384, 192], and the material parameters of this layer are output, such as base color, roughness, metallicity, etc. For interaction rule generation, a 4-layer fully connected network is used, with hidden layer dimensions of [768, 384, 192, 96], and the interaction parameters of this layer are output, including collision attributes, interactive object identifiers, etc. Taking the building structure layer as an example, the generated geometric information contains approximately 20,000 vertices and 15,000 faces; the material information contains sandstone material (base color RGB values [0.82, 0.78, 0.65], roughness 0.72); the interaction rules contain markings for walkable areas, climbable wall markings, etc.

[0105] The model constructs a double-layer constrained optimization model to generate environmental constraint conditions. This optimization model includes local constraints within the functional layer and global constraints between functional layers. Local constraints focus on the consistency within a single functional layer, such as the physical stability of the building structure, the realism of the material, etc.; global constraints focus on the coordination between functional layers, such as lighting consistency, physical interaction compatibility, etc. The system uses an iterative optimization algorithm to minimize the weighted combination of the constraint expected deviation and the constraint difference between adjacent functional layers. During the optimization process, the local constraint weight is set to 0.7, and the global constraint weight is set to 0.3. Taking the lighting constraint as an example, the system ensures that the light source in the special effects layer can correctly affect all lower-level structures, generating reasonable shadows on the surface of the building layer (deviation less than 0.15), while maintaining a natural lighting transition with the terrain layer (inter-layer difference less than 0.2).

[0106] Through the above implementation process, the deep learning-based scene intelligent analysis model can efficiently parse the scene environment parameters into multiple structured functional layers, each layer containing complete geometric information, material information, interaction rules, and environmental constraint conditions, providing a solid foundation for subsequent metaverse scene rendering and interaction. When processing complex urban scenes, the average processing time of this model does not exceed 5 seconds, the functional layer division accuracy reaches 92%, and the constraint satisfaction rate reaches 95%, greatly improving the efficiency and quality of metaverse scene construction.

[0107] In an alternative implementation,

[0108] Construct a two - layer constraint optimization model. The two - layer constraint optimization model includes local constraints within the functional layer and global constraints between functional layers. Constraint optimization is achieved by minimizing a weighted combination of the constraint expected deviation and the constraint difference between adjacent functional layers, including:

[0109] Construct the local constraint features of the scene functional layer. The local constraint features include geometric constraint features and material constraint features. Among them, the geometric constraint features include position continuity constraint, normal vector consistency constraint, and boundary continuity constraint. The material constraint features include texture transition constraint, roughness continuity constraint, and reflectivity consistency constraint;

[0110] Construct the global constraint features of the scene functional layer. The global constraint features include inter - layer topology constraints and physical constraints. Among them, the inter - layer topology constraints calculate the constraint relationship between adjacent functional layers through the inter - layer association weight and the topology relationship measurement function. The physical constraints include gravity action constraint, collision detection constraint, and structural stability constraint;

[0111] Construct a multi - level optimization objective function based on the local constraint features and the global constraint features. Take the weighted combination of the geometric constraints and material constraints of the local constraint features as the local optimization objective, and take the weighted combination of the topology constraints and physical constraints of the global constraint features as the global optimization objective;

[0112] Adopt an adaptive weight adjustment mechanism to dynamically balance the local optimization objective and the global optimization objective. The adaptive weight is adjusted by exponential decay based on the ratio of the local optimization objective value to the global optimization objective value;

[0113] Iteratively solve the multi - level optimization objective function by the alternating direction multiplier method until the difference between the optimization results of two adjacent iterations is less than the preset deviation threshold;

[0114] Calculate the local constraint satisfaction degree and the global constraint satisfaction degree of each functional layer. Both the local constraint satisfaction degree and the global constraint satisfaction degree are calculated using a Gaussian - type evaluation function.

[0115] Exemplarily, local constraint features of the scene function layer are constructed, including geometric constraint features and material constraint features. The geometric constraint feature mainly focuses on the rationality of the internal geometric structure of the functional layer. The position continuity constraint ensures a smooth transition of the spatial positions of adjacent geometric elements in the same functional layer to avoid jump faults. For example, in the terrain function layer, the system samples the height values ​​of adjacent grid points. When the height difference of adjacent points exceeds the preset threshold (usually 25% of the grid spacing), the position continuity constraint is triggered and the transition is smoothed by interpolation. The normal vector consistency constraint ensures that the normal vector of the surface of the functional layer changes gradually and improves the rendering quality. During implementation, the system calculates the angle between the normal vectors of adjacent facets. When the angle exceeds 45 degrees, it is smoothed by subdivision surface technology to make the normal vector transition natural. The boundary continuity constraint processes the edge area of ​​the functional layer to prevent suspension or fracture. At the junction of the building function layer and the terrain function layer, the system detects the connection of the boundary points to ensure that the building base fits closely with the terrain surface, and the deviation is controlled within 2 cm.

[0116] The material constraint feature mainly deals with the consistency of material properties within the functional layer. The texture transition constraint achieves a natural transition between different materials through map fusion technology. In the area where the grass and sand meet, the system creates a 30-pixel-wide transition zone and applies an alpha blending algorithm to blend the two textures naturally. The roughness continuity constraint ensures the smooth change of the material roughness parameter and prevents abrupt material boundaries. The system divides the surface of the functional layer into multiple material areas, and the gradient change of the roughness parameter between areas is controlled within a range of no more than 0.2 per meter. The reflectivity consistency constraint ensures that similar material areas have consistent light reflection characteristics. For example, the reflectivity fluctuation within the same metal material area is controlled within ±5% to maintain visual consistency.

[0117] Next, the global constraint features of the scene function layer are constructed, including inter-layer topological constraints and physical constraints. Inter-layer topological constraints describe the dependencies between different functional layers through inter-layer association weights and topological relationship measurement functions. The system establishes an association matrix for adjacent functional layers. For example, the association weight between the terrain layer and the building layer is 0.85, indicating that the building layer is highly dependent on the terrain layer; while the association weight between the terrain layer and the special effects layer is only 0.3, indicating a low degree of dependence. The topological relationship measurement is comprehensively evaluated through three indicators: spatial overlap rate, geometric intersection, and functional dependency, to generate a standardized topological relationship value in the range of 0-1.

[0118] Physical constraints mainly ensure the physical rationality of the scene. The gravity constraint considers the stable state of objects in the scene under the action of gravity and prevents floating objects. When implemented, the system calculates the ratio of the support area to the mass of the objects in each functional layer. When this ratio is lower than the safety threshold (0.2 square meters / ton), the constraint is triggered to adjust the position of the object. The collision detection constraint prevents objects between different functional layers from penetrating each other. The system uses a spatial hashing algorithm to accelerate collision detection, divides the scene space into voxel grids of 1 meter × 1 meter × 1 meter, and only detects the possibility of collision of objects within adjacent voxels, with the detection accuracy reaching the centimeter level. The structural stability constraint ensures the physical rationality of complex structures such as buildings and prevents impossible suspended structures. The system evaluates the stability of the structure based on a simplified physical model, calculates the distribution of support points, the position of the center of gravity, and the force balance to ensure that the stability coefficient of the structure is greater than 1.5.

[0119] Based on local constraint features and global constraint features, the system constructs a multi-level optimization objective function. The local optimization objective takes the weighted combination of geometric constraints and material constraints as the evaluation index, and the default weight assignment is 0.6 for geometric constraints and 0.4 for material constraints. For example, in the terrain functional layer, the weight of the position continuity constraint is 0.5, the normal vector consistency constraint is 0.3, and the boundary continuity constraint is 0.2, which are comprehensively formed into a geometric constraint score; the weight of the texture transition constraint is 0.4, the roughness continuity constraint is 0.4, and the reflectivity consistency constraint is 0.2, which are comprehensively formed into a material constraint score.

[0120] The global optimization objective takes the weighted combination of topological constraints and physical constraints as the evaluation index, and the default weights are 0.5 for topological constraints and 0.5 for physical constraints. In the urban scene, the weight assignment of the inter-layer topological constraints is: 0.3 for the overlap degree index, 0.3 for geometric intersection, and 0.4 for functional dependence; the weight assignment of the physical constraints is: 0.3 for the gravity constraint, 0.4 for the collision detection constraint, and 0.3 for the structural stability constraint.

[0121] The system uses an adaptive weight adjustment mechanism to dynamically balance the local optimization objective and the global optimization objective. The initial weights are set to 0.5 for local optimization and 0.5 for global optimization, and then they are dynamically adjusted according to the ratio of the local objective value to the global objective value during the optimization process. When the completion degree of local optimization is significantly lower than that of global optimization, the local weight will automatically increase, so that more optimization resources are allocated to local constraints; vice versa. For example, when the local optimization objective value is 0.75 and the global optimization objective value is 0.95, the local / global ratio is 0.789, and the system calculates the weight adjustment factor to be 1.15, and adjusts the local optimization weight to 0.575 and the global optimization weight to 0.425, so that more optimization resources are allocated to local constraints.

[0122] The optimization solution adopts the alternating direction multiplier method, which decomposes complex constrained optimization problems into multiple sub-problems for alternating solution. The system first optimizes the local constraints while fixing the global constraints; then it fixes the local constraints and optimizes the global constraints; finally, it updates the overall solution by synthesizing the results of both. This process is iterated repeatedly until the difference between the optimization results of two adjacent iterations is less than a preset threshold (usually set to 0.01). In the optimization of urban scenarios, the system generally needs to perform 10 - 15 iterations to meet the convergence condition, and each iteration takes about 0.5 seconds.

[0123] After the optimization is completed, the system calculates the constraint satisfaction degree of each functional layer. The local constraint satisfaction degree is calculated by a Gaussian evaluation function, centered on the expected value of the constraint, and exponentially decaying with the increase of the deviation. For example, the position continuity constraint satisfaction degree of 0.97 indicates that the deviation between the current solution and the expected constraint is very small; the global constraint satisfaction degree also uses a Gaussian evaluation function, but the evaluation range is extended to adjacent functional layers. For example, the collision detection constraint satisfaction degree of 0.92 indicates that the collision situation between functional layers in the current solution is within an acceptable range.

[0124] Through the above double-layer constraint optimization model, the system can effectively balance the internal consistency of the scene and the coordination between layers, generate physically reasonable and visually coherent metaverse scene environment constraint conditions, and provide a solid foundation for subsequent scene rendering and interaction. In practical applications, this optimization method can increase the constraint satisfaction degree from the initial 72% to over 95%, significantly improving the scene quality and user experience.

[0125] Table 1 is a comparison table of calculation efficiencies under different scene types. From the data, it can be seen that this technical solution shows significant performance advantages in all scene types. In indoor scenes, in the case of low complexity (number of constraints: 245), this solution only requires 127.3 ms, improving the efficiency by 34.3% compared to the single-layer constraint model (193.8 ms) and the traditional method (298.7 ms) respectively. As the complexity increases, the advantage of this solution is still obvious. In high-complexity indoor scenes (number of constraints: 1205), this solution takes 476.2 ms, while the traditional method requires 1142.5 ms.

[0126] Table 1: Comparison table of calculation efficiencies under different scene types;

[0127]

[0128] This application innovatively proposes a double - layer constraint structure, which divides constraints into two levels: local constraints within the functional layer and global constraints between functional layers, and systematically processes complex constraint relationships in the meta - universe scenario. Compared with the single - layer constraint model in the prior art, the double - layer constraint structure can more accurately describe the dependency relationships and interaction rules between different functional layers in the scenario. For example, when dealing with the relationship between the building layer and the terrain layer, traditional methods only consider simple collision detection, while this application simultaneously considers multi - dimensional constraints such as position continuity, normal vector consistency, and material transition, making the scenario more realistic and coordinated.

[0129] Secondly, this application introduces an adaptive weight adjustment mechanism, which dynamically adjusts the optimization resource allocation according to the completion degrees of local constraints and global constraints during the optimization process. Compared with the fixed weights or simple linear adjustments in the prior art, this application adopts an exponential decay adjustment method based on ratios, which can respond more sensitively to changes in the constraint optimization state and avoid falling into local optimal solutions. Practical applications show that in the optimization of complex urban scenarios, the adaptive weight mechanism reduces the average number of iterations by 30% compared with the fixed - weight method, and the optimization efficiency is significantly improved.

[0130] This application adopts a refined constraint classification system, which systematically integrates geometric constraints, material constraints, topological constraints, and physical constraints to form a comprehensive constraint evaluation system. Compared with the prior art, the constraint granularity is finer, the coverage is wider, and the adaptability is stronger. Especially in terms of material constraints, three sub - constraints, namely texture transition, roughness continuity, and reflectivity consistency, are introduced, filling the gap in the processing of material coherence in the prior art.

[0131] In an alternative embodiment,

[0132] Call the physical engine module, construct a three - dimensional mesh model according to the geometric information and material information of each scene functional layer, and implant corresponding physical property parameters into each three - dimensional mesh model, including:

[0133] Construct an adaptive tetrahedral mesh according to the geometric information, and perform quality evaluation and mesh element optimization on the tetrahedral mesh through the volume ratio of mesh elements, the dihedral angle of mesh elements, and the side length ratio of mesh elements to generate a high - quality tetrahedral mesh;

[0134] Calculate physical property parameters based on the material information, and calculate the mass density, elastic modulus, Poisson's ratio, and friction coefficient of the scene functional layer respectively through the material - physics mapping network, and establish the corresponding relationship between the material information and the physical property parameters;

[0135] Allocate the physical property parameters to each mesh element of the high - quality tetrahedral mesh to generate a three - dimensional physical mesh model with physical characteristics.

[0136] Exemplarily, during the rapid construction of a metaverse scenario based on a modular engine, the invocation of the physics engine module is a key link to achieve real physical interactions in the scenario. This module constructs a three-dimensional mesh model according to the geometric information and material information of the scenario functional layer, and implants corresponding physical property parameters to provide a physical simulation basis for the scenario.

[0137] Exemplarily, an adaptive tetrahedral mesh is constructed according to the geometric information of the scenario functional layer. The Delaunay tetrahedral meshing algorithm is used to voxelize the geometric surface of the functional layer. This algorithm takes the surface triangular mesh as input and generates a tetrahedral mesh structure that fills the interior. In specific implementation, first, the input surface triangular mesh is converted into a signed distance field, with the mesh surface as the zero isosurface, positive values inside, and negative values outside. Then, the density distribution function of the initial tetrahedral mesh is set. This function considers the local geometric features of the model, such as curvature, thickness, and surface details, so that the mesh density is higher in regions with rich geometric features and lower in flat regions. For example, for the column structure in the building functional layer, due to its load-bearing characteristics, denser tetrahedral meshes are automatically generated inside the columns, and the side length of the tetrahedral elements is controlled at about 5 cm; while for the planar part of the wall, the side length of the mesh elements can be relaxed to 20 cm to balance simulation accuracy and computational efficiency.

[0138] After the tetrahedral mesh is generated, the mesh quality is evaluated through three key indicators: the mesh element volume ratio, the mesh element dihedral angle, and the mesh element side length ratio. The mesh element volume ratio refers to the ratio of the volume of a single tetrahedron to the volume of its circumscribed sphere, and the ideal value should be greater than 0.1; the mesh element dihedral angle refers to the dihedral angle between adjacent faces of the tetrahedron, and angles that are too small (<10°) or too large (>170°) should be avoided; the mesh element side length ratio refers to the ratio of the longest side to the shortest side, and it should be controlled within 10. For mesh elements that do not meet the quality requirements, local remeshing techniques are used for optimization, including edge reduction, edge flipping, node smoothing, and element splitting. For example, for a complex tree model in a certain game scenario, about 12,000 tetrahedral elements are generated in the main part of the tree trunk. In the initial mesh quality evaluation, 27% of the elements do not meet the standards. After 5 rounds of local optimization, the proportion of unqualified elements drops to 3%, meeting the requirements of physical simulation.

[0139] Calculate physical property parameters based on the material information of the functional layer. A material-physics mapping network constructed by deep learning is used. This network consists of three fully connected neural networks. The input is the feature vector of the material (including information such as color, roughness, metallicity, normal map, etc.), and the output is four key physical property parameters: mass density, elastic modulus, Poisson's ratio, and coefficient of friction. The dimension of the hidden layer of the network is [256, 128, 64], and the ReLU activation function is used. This network is trained with a large number of real material samples to establish the mapping relationship between the appearance characteristics of the material and the physical properties. For example, for wood materials, different types such as oak, pine, and walnut can be distinguished according to their texture characteristics and color shades, and corresponding physical parameters are assigned: the mass density of oak is about 700 kg / m³, the elastic modulus is about 12 GPa, Poisson's ratio is 0.3, and the static coefficient of friction is 0.4; while metal materials have higher mass density (7800 kg / m³) and elastic modulus (200 GPa).

[0140] To improve the accuracy of the material-physics mapping, a material mixing processing mechanism is also introduced to handle the composite materials commonly found in real scenarios. When it is recognized that the surface has the characteristics of multiple material mixtures, calculate the proportion of each component material, and use the weighted average method to calculate the final physical properties. For example, for a metal surface covered with rust, the system may recognize 70% metal characteristics and 30% rust characteristics, and the finally assigned mass density is 0.7×7800 + 0.3×5200 = 7020 kg / m³. In addition, the anisotropy of the material is also considered, especially for materials such as wood and composite materials with obvious differences in physical properties in different directions. The elastic characteristics are represented by tensors to make the simulation results more realistic.

[0141] Finally, assign the calculated physical property parameters to each grid cell of the high-quality tetrahedral mesh to generate a three-dimensional physical mesh model with physical characteristics. During the parameter assignment process, three different assignment strategies are supported: uniform assignment, gradient assignment, and hierarchical assignment. Uniform assignment is applicable to the functional layer with uniform materials, and the same physical parameters are assigned to all grid cells; gradient assignment is applicable to the functional layer with gradually changing materials, and the physical parameters change in a gradient along a specific direction; hierarchical assignment is applicable to multi-layer composite structures, and different physical parameters are assigned to grid cells at different depths. For example, for the floor functional layer, it may be recognized that the surface layer is a wood material, the middle layer is a sound insulation material, and the bottom layer is a cement structure, and accordingly, the hierarchical assignment strategy is adopted, so that the surface layer cells have the physical characteristics of wood (low density, medium elasticity), the middle layer cells have the characteristics of sound insulation materials (very low density, low elasticity), and the bottom layer cells have the characteristics of cement (high density, high stiffness).

[0142] After the parameter assignment is completed, a spatial consistency check of physical properties is performed to ensure smooth changes in physical properties between adjacent grid cells and avoid numerical instability caused by mutations. For example, when the difference in elastic modulus between adjacent cells is detected to exceed 50%, transition cells are automatically created in the boundary region to enable a smooth transition of physical properties. For the floor model of the urban square scene (containing approximately 50,000 tetrahedral cells), after identifying 5 main material regions and assigning corresponding physical parameters, approximately 2,000 transition cells are created to ensure the stability of physical simulation.

[0143] Table 2 is a comparative experiment table for grid quality assessment. The present technology has achieved significant improvements in the minimum dihedral angle, which has been increased from 12.4 degrees in the traditional method to 23.8 degrees, with an improvement rate of 91.9%. This means a substantial improvement in grid quality. At the same time, the maximum dihedral angle has also been optimized, decreasing from 167.2 degrees to 156.3 degrees, with an improvement rate of 6.5%. The average cell volume ratio has been increased from 0.683 to 0.872, with an improvement of 27.7%, indicating that the grids generated by the present technology are more uniform. In terms of the cell side length ratio of the grid, the present technology has achieved an improvement from 2.42 to 1.35, with an improvement rate of 44.2%, approaching the ideal value of 1.0. The Hausdorff distance error has been significantly reduced, from 1.83 mm to 0.57 mm, with an improvement rate of 68.9%, indicating higher geometric approximation accuracy. The volume retention rate has been increased from 96.2% to 99.8%, ensuring the accuracy of the model volume. In addition, the number of grid cells required by the present technology has been reduced by 23.7%, from 24,568 to 18,734, which not only improves the calculation efficiency but also ensures the grid quality.

[0144] Table 2: Comparative experiment table for grid quality assessment;

[0145]

[0146] The three-dimensional physical grid model generated by the above method not only retains the accurate expression of the geometric appearance but also accurately maps the physical characteristics corresponding to the materials, providing a solid foundation for subsequent physical interaction simulations. In actual tests, the physical grid model generated by this method can support real-time physical interactions at up to 100 frames per second while maintaining the authenticity of physical behavior, meeting the high requirements for physical interactions in the metaverse scenario. For example, when simulating a sphere hitting the surfaces of different materials, it can accurately reproduce the bounce height, sound characteristics, and deformation behavior of different materials, greatly enhancing the user's immersion and interaction experience.

[0147] In an alternative implementation,

[0148] Invoking the lighting rendering module to perform lighting calculation and shadow mapping on each of the three-dimensional grid models based on the hierarchical relationship of the scene function layer includes:

[0149] The hierarchical relationship includes the spatial position relationship and occlusion relationship between functional layers;

[0150] Based on the hierarchical relationship, a hierarchical light propagation tree is constructed. The nodes of the hierarchical light propagation tree represent the scene functional layers, the connecting edges between the nodes represent the light propagation paths, and light attenuation weights are assigned to each light propagation path according to the spatial position relationship and occlusion relationship;

[0151] According to the structure of the hierarchical light propagation tree, the direct light and indirect light of each three-dimensional grid model are calculated to generate a light interaction matrix, and the light interaction matrix records the light energy transfer relationship between different functional layers;

[0152] Based on the light interaction matrix, a depth map of each three-dimensional grid model is generated, and the shadow mapping matrix is calculated in combination with the light attenuation weight. According to the shadow mapping matrix, the light distribution of each functional layer is dynamically updated and optimized.

[0153] Exemplarily, the hierarchical relationship between the scene functional layers is clarified, including the spatial position relationship and occlusion relationship. The spatial position relationship is calculated through the spatial coordinates and bounding volumes of each functional layer. The system creates an axis-aligned bounding box for each functional layer, recording its center coordinates, length, width, height dimensions, and rotation angle in the world space. For example, in an urban square scene, the bounding box size of the ground functional layer is 500×500×2 meters, and the center coordinates are (0,0,0); the bounding box size of the building functional layer is 80×50×30 meters, and the center coordinates are (60,40,15). The occlusion relationship is determined by the ray casting algorithm. The system projects rays from the main light source position to the sampling points of each functional layer, and judges the occlusion relationship by detecting the intersection of the rays with other functional layers. In actual implementation, the system uses a hierarchical spatial partitioning technique (octree) to accelerate the ray casting calculation, recursively divides the space into sub-regions, and the rays only need to perform intersection tests with the regions that may intersect, reducing the computational complexity from O(n) to O(log n), where n is the number of functional layers in the scene.

[0154] After obtaining the hierarchical relationship, the system constructs a hierarchical light propagation tree. Each node of this tree structure represents a scene functional layer, and the connecting edges between nodes represent the propagation path of light energy. The root node is usually set as the main light source (such as sunlight or the main artificial light source), and then the branch structure of the tree is constructed according to the spatial relationship and visibility of the functional layers. For example, for the functional layers directly illuminated by the main light source (such as the open ground, building exterior walls), they are used as the direct children nodes of the root node; while the areas blocked by other functional layers (such as the ground part under the building shadow) are used as the children nodes of the corresponding blocking layer nodes. For each light propagation path (i.e., the edge in the tree), the system assigns a light attenuation weight, which comprehensively considers the following factors: distance attenuation, angular attenuation, occlusion degree, and material reflection characteristics.

[0155] The distance attenuation is calculated based on the inverse square relationship between the light source and the receiving surface. For example, when the distance increases from 5 meters to 10 meters, the light intensity drops to 1 / 4 of the original. The angular attenuation considers the angle between the light incidence angle and the surface normal. When the incidence angle increases from 0 degrees (perpendicular incidence) to 60 degrees, the effective light intensity drops to 1 / 2 of the original. The occlusion degree is determined by calculating the proportion of the projected overlapping area between functional layers. For example, when Building A blocks 30% of the light-receiving area of Building B, the weight of the corresponding light path will be multiplied by an attenuation coefficient of 0.7. The material reflection characteristics are based on the reflectivity and scattering characteristics of the surface material. For example, the reflectivity of a highly glossy metal material can reach above 0.8, while the reflectivity of a rough concrete surface is only about 0.2.

[0156] In the urban square scene example, the weight of the direct light path from the main light source (the sun) to the ground functional layer is 0.9 (considering the slight attenuation of atmospheric scattering); the weight of the indirect light path from the ground to the building exterior wall is 0.15 (considering the combined effect of ground reflectivity and distance attenuation); the weight of the indirect light path from the building exterior wall to the adjacent building is 0.08 (considering the combined effect of wall reflectivity and angular attenuation). These weight values comprehensively form the edge weight network of the light propagation tree, accurately describing the law of light energy transfer in the scene.

[0157] Based on the constructed hierarchical light propagation tree, the system calculates the direct light and indirect light of each 3D mesh model. The direct light calculation uses the Phong lighting model, considering three components: ambient light, diffuse light, and specular light. For example, for the exterior wall of a building, the system sets the ambient light coefficient to 0.2 (indicating that there is still 20% of the base brightness under complete occlusion), the diffuse coefficient to 0.6 (representing the surface's diffuse reflection ability), and the specular coefficient to 0.1 (representing the surface's ability to produce highlights). The indirect light calculation is based on the multi-level jump paths in the light propagation tree, cumulatively calculating the light energy that reaches the mesh surface after multiple reflections. For example, for the interior corridor of a building, its light comes not only from the sunlight directly passing through the window (direct light), but also from the light reflected by the exterior wall, the ground, and the adjacent buildings (indirect light).

[0158] The system organizes the calculation results of direct light and indirect light into a light interaction matrix, which is an n×n square matrix (n is the number of functional layers). The matrix element M[i,j] represents the proportion of light energy transferred from functional layer i to functional layer j. In the urban square scenario with 10 functional layers, the generated light interaction matrix contains 100 elements, accurately recording the light energy transfer relationship between each functional layer. For example, M[2,5]=0.35 means that the proportion of light energy reflected from the exterior wall of the building (functional layer 2) to the decorative plants (functional layer 5) is 35%.

[0159] Based on the light interaction matrix, the system generates a depth map of the scene for each light source perspective. The depth map is a special image, and its pixel value represents the distance from the light source to each point in the scene. The system uses the shadow mapping algorithm. First, it sets the light source as the rendering perspective and renders the scene in the form of a grayscale image, where the grayscale value is proportional to the depth. For complex scenes, the system uses the cascaded shadow mapping technique, dividing the light source frustum into three regions: near, middle, and far, and generating depth maps with different resolutions respectively. The near region uses a 512×512 high-resolution depth map, and the far region uses a 128×128 low-resolution depth map to balance accuracy and efficiency.

[0160] After obtaining the depth map, the system combines the light attenuation weight to calculate the shadow mapping matrix. This matrix records the shadow attenuation coefficient of each sampling point in the scene, with a value range from 0 (completely in the shadow) to 1 (completely illuminated). During the calculation, the system renders the scene from the current perspective. For each visible fragment, it projects the fragment back to the light source space and compares the depth value of the fragment in the light source space with the value recorded in the depth map. If the former is greater than the latter (plus a small offset to avoid self-occlusion), it means that the fragment is in the shadow. To avoid jagged hard shadow edges, the system uses the percentage-closer filtering technique, sampling 9 points in the 3×3 area around the fragment, calculating the proportion of the sampling points in the shadow as the shadow attenuation coefficient, and generating a soft and natural shadow edge.

[0161] Dynamically update and optimize the lighting distribution of each functional layer according to the shadow mapping matrix. Specifically, the system multiplies the initially calculated lighting result by the shadow mapping matrix to obtain the final lighting distribution considering the shadow effect. For example, if the initial lighting intensity at a certain point is 0.8 and the corresponding attenuation coefficient in the shadow mapping matrix is 0.6, then the final lighting intensity at this point is 0.48. To further enhance the visual effect, the system implements an adaptive ambient occlusion technology to appropriately increase the ambient light intensity in the shadow area, avoiding the shadow area from being too dark and improving the visibility of details. In an indoor scene, the ambient light intensity in the shadow area will be increased by about 30%, making the details in the shadow area still visible.

[0162] Through the above implementation method, the system can generate highly realistic lighting and shadow effects for the metaverse scene. In actual tests, compared with the traditional method, the lighting calculation speed of the scene processed by this lighting rendering module is increased by about 40%, and at the same time, the shadow quality score (comprehensively considering the softness of the shadow edge, detail retention, and visual coherence) is increased from 7.2 to 9.1 (with a full score of 10), significantly enhancing the immersion and realism of the scene.

[0163] In the second aspect of the embodiments of the present invention, a fast metaverse scene construction system based on a modular engine is provided, including:

[0164] A first unit for obtaining a creation instruction of a metaverse scene from a user, where the creation instruction includes scene type information and scene environment parameters;

[0165] A second unit for calling a corresponding scene rendering module from a preset modular engine library according to the scene type information, where the scene rendering module includes a physics engine module, a lighting rendering module, and an environmental special effect module;

[0166] A third unit for performing feature extraction and scene layering on the scene environment parameters based on a scene intelligent analysis model of deep learning, parsing the scene environment parameters into multiple scene functional layers, and each scene functional layer includes corresponding geometric information, material information, interaction rules, and environmental constraint conditions; a progressive hierarchical relationship is formed between the scene functional layers, and the environmental constraint conditions of the upper scene functional layer will optimize and adjust the rendering result of the lower scene functional layer;

[0167] A fourth unit for calling the physics engine module to construct a three-dimensional mesh model according to the geometric information and material information of each scene functional layer, and implanting corresponding physical attribute parameters into each three-dimensional mesh model;

[0168] The fifth unit is used to call the light rendering module to perform light calculation and shadow mapping on each of the 3D mesh models based on the hierarchical relationship of the scene function layers;

[0169] The sixth unit is used to call the environmental special effect module to add corresponding environmental special effects to each of the scene function layers according to the scene environment parameters.

[0170] In the third aspect of the embodiments of the present invention,

[0171] A kind of electronic device is provided, including:

[0172] A processor;

[0173] A memory for storing processor-executable instructions;

[0174] Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.

[0175] In the fourth aspect of the embodiments of the present invention,

[0176] A computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.

[0177] The present invention can be a method, a device, a system and / or a computer program product. The computer program product can include a computer-readable storage medium, on which computer-readable program instructions for executing various aspects of the present invention are uploaded.

[0178] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for quickly building a metaverse scene based on a modular engine, characterized in that: include: Obtaining a user's creation instruction for a metaverse scene, wherein the creation instruction includes scene type information and scene environment parameters; According to the scene type information, a corresponding scene rendering module is called from a preset modular engine library, wherein the scene rendering module includes a physical engine module, a lighting rendering module, and an environmental special effect module; Based on the deep learning scene intelligent analysis model, the scene environment parameters are feature extracted and the scene is layered, and the scene environment parameters are parsed into multiple scene function layers, each of which includes corresponding geometric information, material information, interaction rules and environmental constraints; a progressive hierarchical relationship is formed between the scene function layers, and the environmental constraints of the upper scene function layer will optimize and adjust the rendering results of the lower scene function layer; Calling the physical engine module to construct a three-dimensional mesh model according to the geometric information and material information of each scene function layer, and embedding corresponding physical property parameters in each three-dimensional mesh model; Calling the lighting rendering module to perform lighting calculation and shadow mapping on each of the three-dimensional mesh models based on the hierarchical relationship of the scene function layers; The environmental special effects module is called to add corresponding environmental special effects to each of the scene function layers according to the scene environment parameters.

2. The method according to claim 1, characterized in that According to the scene type information, calling a corresponding scene rendering module from a preset modular engine library includes: Receiving scene type information, and encoding the scene type information through a pre-trained BERT model to obtain a scene feature vector; A candidate rendering module is obtained from the modular engine library, and the matching degree between the candidate rendering module and the scene feature vector is calculated. The matching degree is calculated as follows: ; in, M Represents a set of candidate rendering modules, S represents the scene feature vector set, α represents the balance factor, n represents the feature dimension, w i represents the weight coefficient of the i-th dimension feature, sim(m i ,s i ) represents the similarity between the candidate rendering module of the i-th dimension and the feature vector of the i-th dimension scene, β represents the balance factor, C(M) Represents the computational complexity evaluation value corresponding to the candidate rendering module set, C () represents the complexity evaluation function; The candidate rendering module with the highest matching degree is selected as the scene rendering module corresponding to the scene type information.

3. The method according to claim 1, characterized in that Based on the deep learning scene intelligent analysis model, the scene environment parameters are feature extracted and the scene is layered, and the scene environment parameters are parsed into multiple scene function layers, each of which includes corresponding geometric information, material information, interaction rules and environmental constraints, including: Extracting geometric features and material features from the scene environment parameters respectively through a multimodal feature extraction network, wherein the geometric features are extracted from a scene point cloud set through a PointNet++ network, and the material features are extracted from a material map through a multi-scale convolutional neural network; A multi-head attention mechanism is used to fuse the geometric features and the material features, the geometric features are used as query vectors, the material features are used as key vectors, and the concatenation result of the geometric features and the material features is used as a value vector to generate fused features; Performing hierarchical analysis on the fused features, dividing the scene into multiple functional layers, and establishing a hierarchical relationship between the functional layers; Generate geometric information, material information and interaction rules of each functional layer through a multi-layer perceptron respectively; A two-layer constraint optimization model is constructed, which includes local constraints within a functional layer and global constraints between functional layers, and constraint optimization is achieved by minimizing a weighted combination of constraint expected deviations and constraint differences between adjacent functional layers.

4. The method according to claim 3, characterized in that A two-layer constraint optimization model is constructed, wherein the two-layer constraint optimization model includes local constraints within a functional layer and global constraints between functional layers, and constraint optimization is achieved by minimizing a weighted combination of constraint expected deviations and constraint differences between adjacent functional layers, including: Constructing local constraint features of the scene function layer, the local constraint features include geometric constraint features and material constraint features, wherein the geometric constraint features include position continuity constraints, normal vector consistency constraints and boundary continuity constraints, and the material constraint features include texture transition constraints, roughness continuity constraints and reflectivity consistency constraints; Constructing global constraint features of scene function layers, the global constraint features include inter-layer topological constraints and physical constraints, wherein the inter-layer topological constraints calculate the constraint relationship between adjacent functional layers through inter-layer association weights and topological relationship measurement functions, and the physical constraints include gravity constraints, collision detection constraints, and structural stability constraints; Constructing a multi-level optimization objective function based on the local constraint features and the global constraint features, taking a weighted combination of the geometric constraints and material constraints of the local constraint features as a local optimization objective, and taking a weighted combination of the topological constraints and physical constraints of the global constraint features as a global optimization objective; Adopting an adaptive weight adjustment mechanism to dynamically balance the local optimization target and the global optimization target, the adaptive weight is adjusted exponentially based on the ratio of the local optimization target value to the global optimization target value; Iteratively solving the multi-level optimization objective function by an alternating direction multiplier method until the difference between the optimization results of two adjacent iterations is less than a preset deviation threshold; The local constraint satisfaction and the global constraint satisfaction of each functional layer are calculated, and the local constraint satisfaction and the global constraint satisfaction are both calculated using a Gaussian evaluation function.

5. The method according to claim 1, characterized in that: Calling the physical engine module, constructing a three-dimensional mesh model according to the geometric information and material information of each scene function layer, and implanting corresponding physical attribute parameters in each three-dimensional mesh model includes: Constructing an adaptive tetrahedral mesh according to the geometric information, performing quality evaluation and mesh unit optimization on the tetrahedral mesh by mesh unit volume ratio, mesh unit dihedral angle and mesh unit side length ratio, and generating a high-quality tetrahedral mesh; Calculating physical property parameters based on the material information, respectively calculating mass density, elastic modulus, Poisson's ratio and friction coefficient of the scene function layer through a material-physical mapping network, and establishing a corresponding relationship between the material information and the physical property parameters; The physical property parameters are distributed to each mesh unit of the high-quality tetrahedral mesh to generate a three-dimensional physical mesh model with physical properties.

6. The method according to claim 1, characterized in that Calling the lighting rendering module to perform lighting calculation and shadow mapping on each of the three-dimensional mesh models based on the hierarchical relationship of the scene function layer includes: The hierarchical relationship includes the spatial position relationship and occlusion relationship between the functional layers; Based on the hierarchical relationship, a hierarchical illumination propagation tree is constructed, wherein the nodes of the hierarchical illumination propagation tree represent scene function layers, the connecting edges between the nodes represent illumination propagation paths, and an illumination attenuation weight is assigned to each illumination propagation path according to the spatial position relationship and the occlusion relationship; According to the structure of the hierarchical illumination propagation tree, direct illumination and indirect illumination of each of the three-dimensional grid models are calculated to generate an illumination interaction matrix, wherein the illumination interaction matrix records the light energy transfer relationship between different functional layers; A depth map of each of the three-dimensional mesh models is generated based on the illumination interaction matrix, and a shadow mapping matrix is ​​calculated in combination with the illumination attenuation weights, and the illumination distribution of each functional layer is dynamically updated and optimized according to the shadow mapping matrix.

7. A modular engine-based metaverse scene rapid construction system, used to implement the method as described in any one of claims 1 to 6, characterized in that: include: The first unit is used to obtain a user's creation instruction for a metaverse scene, wherein the creation instruction includes scene type information and scene environment parameters; The second unit is used to call a corresponding scene rendering module from a preset modular engine library according to the scene type information, wherein the scene rendering module includes a physical engine module, a lighting rendering module and an environmental special effect module; The third unit is used for extracting features and stratifying the scene environment parameters based on a deep learning scene intelligent analysis model, parsing the scene environment parameters into multiple scene function layers, each of which includes corresponding geometric information, material information, interaction rules and environmental constraints; a progressive hierarchical relationship is formed between the scene function layers, and the environmental constraints of the upper scene function layer will optimize and adjust the rendering results of the lower scene function layer; The fourth unit is used to call the physical engine module, construct a three-dimensional mesh model according to the geometric information and material information of each scene function layer, and implant corresponding physical property parameters in each three-dimensional mesh model; A fifth unit is used to call the lighting rendering module to perform lighting calculation and shadow mapping on each of the three-dimensional mesh models based on the hierarchical relationship of the scene function layers; The sixth unit is used to call the environmental special effects module and add corresponding environmental special effects to each of the scene function layers according to the scene environment parameters.

8. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Interaction structure and interaction method based on VR (virtual reality) role

    CN111667560A

  • Sluice informatization system construction method based on digital twin technology

    CN118332832A