Real-time Rendering Optimization Method and System in Metaverse Scene Building Engine

By building a spatial octree structure, calculating occlusion coefficients, simplifying geometric data and dynamically merging material instances in the metaverse scene construction engine, the problems of low rendering efficiency, insufficient occlusion culling and inflexible material management in the existing technology are solved, and more efficient rendering performance and smoother user experience are achieved.

CN119941956BActive Publication Date: 2025-06-13HANGZHOU MOXI TECH DEV CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510429171.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-06-13
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

The existing technology has problems such as inefficiency in real-time rendering optimization in the metaverse scene construction engine, inadequate occlusion culling, and lack of flexibility in material instance management, resulting in poor user experience and rendering performance.

Method used

By building a spatial octree structure and traversing from top to bottom, nodes that are not in the view cone are eliminated; occlusion coefficients are calculated based on the occlusion information and occlusion culling is performed; detail hierarchy level is determined based on the distance from the center of the enclosure box to the viewpoint and the scene object geometry data is simplified; material instances are dynamically merged to generate a material instance cache pool.

Benefits of technology

It improves rendering efficiency, optimizes the rendering process, improves the real-time and smoothness of scene rendering, and reduces the rendering burden while ensuring visual effects, improving overall rendering performance and resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941956B_ABST
    Figure CN119941956B_ABST
Patent Text Reader

Abstract

The present invention provides a real-time rendering optimization method and system in a metaverse scene construction engine, relating to the technical field of the metaverse. The method and system include constructing a spatial octree structure of scene objects, reducing rendering load based on frustum culling, occlusion culling, and level of detail simplification techniques; and performing dynamic material merging using an instance cache pool based on material similarity. The present invention effectively reduces the rendering calculation amount, improves the rendering efficiency and smoothness of the metaverse scene, and enables complex scenes to achieve high-quality real-time rendering with limited hardware resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the metaverse technology, and in particular to a real-time rendering optimization method and system in a metaverse scene construction engine. Background Art

[0002] In recent years, with the rapid development of virtual reality and augmented reality technologies, the concept of the metaverse has gradually become a hot topic. As a virtual shared space, the metaverse can provide users with immersive experiences, attracting more and more developers and enterprises to invest resources in scene construction and content creation. To achieve high-quality real-time rendering, developers need efficient rendering optimization methods to ensure that various objects can be smoothly displayed in complex three-dimensional scenes.

[0003] Existing technologies have some defects and deficiencies in real-time rendering optimization in metaverse scene construction engines. First, traditional rendering methods often cannot effectively manage a large number of scene objects when dealing with complex scenes, resulting in low rendering efficiency and affecting the user experience. Second, existing occlusion culling technologies usually rely on simple bounding box judgments and lack in-depth analysis of occlusion information, easily leading to unnecessary rendering overhead. Finally, the management and merging process of material instances often lack flexibility and cannot be dynamically adjusted according to the changes in the actual scene, resulting in waste of resources and degradation of rendering performance.

[0004] Therefore, in view of these problems, it is particularly important to develop an efficient real-time rendering optimization method to improve the performance of the metaverse scene construction engine and the user experience. Summary of the Invention

[0005] Embodiments of the present invention provide a real-time rendering optimization method and system in a metaverse scene construction engine, which can solve the problems in the prior art.

[0006] In a first aspect of embodiments of the present invention, a real-time rendering optimization method in a metaverse scene construction engine is provided, including:

[0007] Obtain three-dimensional scene data to be rendered in the metaverse scene construction engine, where the three-dimensional scene data includes geometric data, material data, and lighting data of scene objects;

[0008] Based on the three-dimensional scene data, construct a spatial octree structure of scene objects, and record the bounding box information and occlusion information of the scene objects corresponding to each node in the spatial octree structure;

[0009] Traverse each node recursively from top to bottom in the spatial octree structure according to the user's viewpoint position information, determine whether the bounding box of each node is within the viewing frustum, and directly prune the nodes outside the viewing frustum; for the nodes within the viewing frustum, calculate the occlusion coefficient based on the occlusion information of the node, and perform occlusion culling when the occlusion coefficient is greater than the preset occlusion threshold; and determine the level of detail hierarchy of the node according to the distance from the center of the node's bounding box to the viewpoint, and generate the geometric data of the scene object after level of detail simplification.

[0010] Divide the geometric data of the scene object into multiple rendering batches, and each rendering batch contains scene objects with the same material properties.

[0011] Establish a material instance cache pool based on material similarity for the scene objects in each rendering batch, dynamically merge the material instances according to the degree of change of the material properties, and merge multiple materials with material property similarity higher than the set similarity threshold into one material instance.

[0012] Based on the three-dimensional scene data, construct a spatial octree structure of the scene object, and record the bounding box information and occlusion information of the scene object corresponding to each node in the spatial octree structure, including:

[0013] Calculate the overall boundary range of the scene based on the geometric data of the scene object, determine the minimum vertex coordinates and maximum vertex coordinates of the scene bounding box, construct the root node space of the octree, and map the scene object to the root node space according to its position information.

[0014] Traverse the scene objects in the root node space. When the number of scene objects contained in a node exceeds the preset segmentation threshold, evenly divide the space of the node into 8 sub-node spaces, and reassign the scene objects in the node to the corresponding sub-node spaces according to the position information; recursively execute the above space division process until the number of scene objects in all nodes does not exceed the preset segmentation threshold.

[0015] Traverse all nodes of the octree from top to bottom: for each node, calculate the axis-aligned bounding box information of the node, including the minimum vertex coordinates and the maximum vertex coordinates; calculate the surface area occlusion ratio of the node: obtain it by calculating the ratio of the area of the node surface occluded by other scene objects to the total surface area of the node; project the node onto six orthogonal direction planes to form a depth occlusion map, and record the maximum depth value at each projection pixel position.

[0016] For the nodes within the viewing frustum, calculate the occlusion coefficient based on the occlusion information of the node, and when the occlusion coefficient is greater than the preset occlusion threshold, perform occlusion culling, including:

[0017] For the nodes within the viewing frustum, calculate their comprehensive occlusion coefficient:

[0018] Based on the depth occlusion maps in six orthogonal projection directions, calculate the visibility of each projection plane in combination with the current view, and calculate the viewing distance attenuation factor by obtaining the distance from the node to the viewpoint; weighted-combine the visibility and the viewing distance attenuation factor according to a preset weight coefficient to obtain the comprehensive occlusion coefficient of the node.

[0019] Construct an adaptive occlusion threshold adjustment mechanism: Based on the target frame rate set by the system, obtain the actual frame rate of the current rendering, and calculate the ratio of the two; multiply the ratio by a preset adjustment coefficient as the dynamic adjustment factor, and multiply it by the base occlusion threshold to obtain the dynamic occlusion threshold of the current frame.

[0020] Perform occlusion judgment on each node in the node set to be processed one by one: Compare the comprehensive occlusion coefficient of each node with the currently calculated dynamic occlusion threshold. When the comprehensive occlusion coefficient of the node is greater than the dynamic occlusion threshold, mark the node as the occluded state, add it to the occlusion culling list, and perform occlusion culling.

[0021] Determine the level of detail (LOD) of the node according to the distance from the center of the bounding box of the node to the viewpoint, and generate the geometric data of the scene object after LOD simplification, including:

[0022] Calculate the distance from the center of the bounding box to the current viewpoint, and determine the level of detail of the node according to the ratio of this distance to the reference size of the node. Specifically: Take the logarithm of the ratio with base 2 and round down to obtain the level of detail.

[0023] For the nodes with the determined level of detail, calculate their simplification rate: Use the power of 2 to the level of detail of the node as the denominator to obtain the target simplification ratio; Based on the target simplification ratio, calculate the folding cost of the mesh edges, where the folding cost includes a weighted combination of the geometric error metric value and the attribute error metric value.

[0024] Perform priority sorting according to the folding cost of the edges, and gradually perform the edge folding operation according to the preset simplification ratio: First fold the edge with the smallest folding cost of the mesh edge, update the positions and attribute information of the adjacent vertices, and continuously perform edge folding until the target simplification rate is reached to generate the geometric data of the scene object after LOD simplification.

[0025] Establish a material instance cache pool based on material similarity for the scene objects in each rendering batch, dynamically merge the material instances according to the degree of change of the material attributes, and merge multiple materials with material attribute similarity higher than the set similarity threshold into one material instance, including:

[0026] Obtain the material data of the scene objects in the rendering batch, extract features from the material data, and construct a feature vector space; establish a material descriptor index tree based on the feature vector space, and the material descriptor index tree is used for the rapid retrieval and matching of materials;

[0027] Calculate the similarity between materials based on the material descriptor index tree: for non-texture attributes, directly calculate the Euclidean distance of the attribute values; for texture attributes, obtain the texture similarity by calculating the normalized difference of the texture pixel values; perform weighted combination on the similarities of each attribute to obtain the comprehensive similarity value between material pairs;

[0028] Divide the materials into multiple groups based on the comprehensive similarity value, construct a reference relationship graph of material instances within each material group, and record the association information of each material instance being referenced by the scene objects; perform hierarchical clustering on each material group again to generate an optimal merging sequence;

[0029] Perform material instance merging according to the optimal merging sequence: calculate the merged material parameters based on the attribute values of each instance within the material group, the non-texture parameters are obtained by weighted average, and the texture parameters are generated by mixed sampling; update the material reference relationship of the scene objects, and point the reference of the original material instance to the newly merged material instance.

[0030] Divide the materials into multiple groups based on the comprehensive similarity value, construct a reference relationship graph of material instances within each material group, and record the association information of each material instance being referenced by the scene objects; perform hierarchical clustering on each material group again to generate an optimal merging sequence including:

[0031] Construct a bipartite graph structure of the material group, the bipartite graph structure includes a set of material instance vertices and a set of scene object vertices, and establish a reference relationship edge set between the set of material instance vertices and the set of scene object vertices;

[0032] Calculate the reference weight in the bipartite graph structure, calculate the area ratio of a single material instance being referenced by a single scene object according to the usage area of each material instance on the surface of the scene object, and use the area ratio as the weight value of the reference relationship;

[0033] Construct a hierarchical clustering tree based on the bipartite graph structure, calculate the distance metric value for each pair of material clusters, and the distance metric value is obtained by calculating the average value of the distances between all material pairs in the two material clusters;

[0034] Take the material similarity, reference relationship overlap degree, and rendering cost as evaluation indicators, perform weighted combination on the three evaluation indicators to obtain the combined evaluation score between material clusters;

[0035] Insert the material cluster pairs with combined evaluation scores higher than the preset evaluation threshold into the priority queue in descending order of scores; extract the material cluster pair with the highest score from the priority queue as the optimal merging object.

[0036] In the second aspect of the embodiments of the present invention, a real-time rendering optimization system in a metaverse scene building engine is provided, including:

[0037] A first unit for obtaining three-dimensional scene data to be rendered in a metaverse scene building engine, where the three-dimensional scene data includes geometric data, material data, and lighting data of scene objects;

[0038] A second unit for constructing a spatial octree structure of scene objects based on the three-dimensional scene data, and recording the bounding box information and occlusion information of the scene objects corresponding to each node in each node of the spatial octree structure;

[0039] A third unit for recursively traversing each node from top to bottom in the spatial octree structure according to the user's viewpoint position information, determining whether the bounding box of each node is within the viewing frustum, and directly pruning the nodes outside the viewing frustum; for the nodes within the viewing frustum, calculating the occlusion coefficient based on the occlusion information of the nodes, and performing occlusion culling when the occlusion coefficient is greater than the preset occlusion threshold; and determining the level of detail of the nodes according to the distance from the center of the bounding box of the nodes to the viewpoint, and generating the geometric data of the scene objects after level-of-detail simplification;

[0040] A fourth unit for dividing the geometric data of the scene objects into multiple rendering batches, and each rendering batch contains scene objects with the same material attributes;

[0041] A fifth unit for establishing a material instance cache pool based on material similarity for the scene objects in each rendering batch, dynamically merging the material instances according to the change degree of the material attributes, and merging multiple materials with material attribute similarity higher than the set similarity threshold into one material instance.

[0042] In the third aspect of the embodiments of the present invention

[0043] Provide an electronic device, including:

[0044] A processor;

[0045] A memory for storing instructions executable by the processor;

[0046] Wherein, the processor is configured to call the instructions stored in the memory to execute the foregoing method.

[0047] In the fourth aspect of the embodiments of the present invention,

[0048] Provided is a computer-readable storage medium storing computer program instructions, which, when executed by a processor, implement the foregoing method.

[0049] The beneficial effects of this application are as follows:

[0050] The beneficial effects of the present invention are mainly reflected in the following three aspects

[0051] By constructing a spatial octree structure and performing a top-down recursive traversal, nodes outside the viewing frustum can be effectively pruned, thereby reducing the number of scene objects to be rendered and improving the rendering efficiency; calculating the occlusion coefficient based on occlusion information and performing occlusion culling can further optimize the rendering process, avoid invalid rendering calculations, and enhance the real-time performance and smoothness of scene rendering; by simplifying the level of detail of the geometric data of scene objects and dynamically merging material instances, the rendering burden can be reduced while ensuring the visual effect, and the overall rendering performance and resource utilization rate can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 It is a schematic flowchart of the real-time rendering optimization method in the metaverse scene construction engine according to an embodiment of the present invention;

[0053] Figure 2 It is a schematic diagram for comparing the occlusion culling efficiency of different methods;

[0054] Figure 3 It is a tabular diagram for comparing the occlusion culling efficiency of different methods;

[0055] Figure 4 It is a schematic diagram of the material instance merging system interface. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0056] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are only some of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0057] The technical solutions of the present invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0058] Figure 1 It is a schematic flowchart of the real-time rendering optimization method in the metaverse scene construction engine according to an embodiment of the present invention, as Figure 1 shown, the method includes:

[0059] Obtain the three-dimensional scene data to be rendered in the metaverse scene construction engine, where the three-dimensional scene data includes geometric data, material data, and lighting data of scene objects;

[0060] Based on the three-dimensional scene data, construct a spatial octree structure of scene objects, and record the bounding box information and occlusion information of the scene objects corresponding to each node in the spatial octree structure;

[0061] According to the user's viewpoint position information, recursively traverse each node from top to bottom in the spatial octree structure, determine whether the bounding box of each node is within the viewing frustum, and directly prune the nodes outside the viewing frustum; for the nodes within the viewing frustum, calculate the occlusion coefficient based on the occlusion information of the node, and perform occlusion culling when the occlusion coefficient is greater than the preset occlusion threshold; and determine the level of detail of the node according to the distance from the center of the node's bounding box to the viewpoint, and generate the geometric data of the scene object after level-of-detail simplification;

[0062] Divide the geometric data of the scene objects into multiple rendering batches, where each rendering batch contains scene objects with the same material attributes;

[0063] Establish a material instance cache pool based on material similarity for the scene objects in each rendering batch, dynamically merge the material instances according to the degree of change of the material attributes, and merge multiple materials with material attribute similarity higher than the set similarity threshold into one material instance.

[0064] In an optional implementation manner, constructing a spatial octree structure of scene objects based on the three-dimensional scene data, and recording the bounding box information and occlusion information of the scene objects corresponding to each node in the spatial octree structure includes:

[0065] Calculate the overall boundary range of the scene based on the geometric data of the scene objects, determine the minimum vertex coordinates and maximum vertex coordinates of the scene bounding box, construct the root node space of the octree, and map the scene objects to the root node space according to their position information;

[0066] Traverse the scene objects in the root node space. When the number of scene objects contained in a node exceeds the preset segmentation threshold, evenly divide the space of the node into 8 sub-node spaces, and reassign the scene objects in the node to the corresponding sub-node spaces according to their position information; recursively execute the above space division process until the number of scene objects in all nodes does not exceed the preset segmentation threshold;

[0067] Traverse all the nodes of the octree from top to bottom: For each node, calculate the axis-aligned bounding box information of the node, including the minimum vertex coordinates and the maximum vertex coordinates; calculate the surface area occlusion ratio of the node: obtain it by calculating the ratio of the area of the node's surface occluded by other scene objects to the total surface area of the node; project the node onto six orthogonal direction planes to form a depth occlusion map, and record the maximum depth value at each projected pixel position.

[0068] Specifically, traverse all the objects in the scene to obtain the vertex coordinate data of each object. For a scene containing 1000 3D models, each model may contain hundreds to thousands of vertices. By comparing the coordinate values of all vertices, determine the minimum vertex coordinates and the maximum vertex coordinates of the entire scene. For example, in an urban scene, the minimum vertex coordinates may be (-500, -500, 0), and the maximum vertex coordinates may be (500, 500, 100), representing a spatial range of 1000×1000×100 cubic meters. Based on these two coordinate values, construct an axis-aligned bounding box as the root node space of the octree, and map all the scene objects into this root node space according to their position information.

[0069] Set the preset segmentation threshold to 10, that is, when the number of scene objects contained in a node exceeds 10, evenly divide the space of this node into 8 sub-node spaces. The division method is to divide the space of the current node along the midpoints of the x, y, and z axes, forming 8 sub-spaces of equal size. For example, for the root node space from (-500, -500, 0) to (500, 500, 100), the first sub-node space obtained after division is from (-500, -500, 0) to (0, 0, 50), the second sub-node space is from (0, -500, 0) to (500, 0, 50), and so on.

[0070] Reassign the scene objects in the current node to the corresponding sub-node spaces according to their position information. For each scene object, calculate the intersection relationship between its bounding box and each sub-node space. If the bounding box of the scene object is completely located within a certain sub-node space, assign this object to this sub-node; if the bounding box of the scene object intersects with multiple sub-node spaces, assign this object to all the intersecting sub-nodes. For example, a building model located at coordinates from (-10, -10, 30) to (10, 10, 60) will be assigned to the relevant sub-nodes that contain this coordinate range.

[0071] Recursively execute the above spatial partitioning process until the number of scene objects in all nodes does not exceed the preset segmentation threshold of 10. In practical applications, it may be necessary to set a maximum recursion depth (such as 8 levels) to prevent over-subdivision. For dense areas containing a large number of small objects, the octree may reach deeper levels, while for open areas, the depth of the octree may be shallower.

[0072] After completing the spatial partitioning, traverse all the nodes of the octree from top to bottom, and calculate and record the bounding box information and occlusion information of each node. For each node, first calculate its axis-aligned bounding box information, including the minimum vertex coordinates and the maximum vertex coordinates. This can be achieved by merging the bounding boxes of all the scene objects contained in this node. For example, if a node contains 3 scene objects, and their bounding boxes are from (-10, -10, 0) to (10, 10, 20), from (-5, -15, 0) to (15, 5, 15), and from (-8, -5, 0) to (5, 12, 18) respectively, then the bounding box of this node is from (-10, -15, 0) to (15, 12, 20).

[0073] The surface area occlusion ratio refers to the ratio of the area of the node's surface occluded by other scene objects to the total surface area of the node. The calculation method is as follows: First, determine the six surfaces (front, back, left, right, top, bottom) of the node, and each surface is a rectangle; then for each surface, check whether other objects in the scene occlude this surface; finally, calculate the proportion of the occluded area to the total surface area. For example, if the total surface area of a node is 200 square meters, and the area occluded by other objects is 120 square meters, then its surface area occlusion ratio is 0.6 or 60%.

[0074] Project the node onto six orthogonal direction planes (i.e., along the positive x-axis, negative x-axis, positive y-axis, negative y-axis, positive z-axis, negative z-axis) to form a depth occlusion map. The specific method is as follows: For each projection direction, create a two-dimensional grid (such as 256×256 pixels) to represent the projection plane; project the bounding box of the node onto this plane to determine the pixel range covered by the projection; for each covered pixel, record the depth value of the farthest scene object in the direction from the viewpoint to this pixel. For example, looking from directly above (positive z-axis), a node covers 20×20 pixels on the projection plane. For each pixel position, record the maximum z coordinate value at this position to form a 20×20 depth map.

[0075] Figure 2 It is a schematic diagram for comparing the occlusion culling efficiency of different methods; Figure 3 It is a tabular graph for comparing the occlusion culling efficiency of different methods; Figure 2 and Figure 3It shows the comparison of occlusion culling efficiency of different methods under various scene complexities. The horizontal axis in the figure represents the number of scene objects (ranging from 1K to 500K), and the vertical axis represents the occlusion culling efficiency percentage (0% - 60%). The four methods are identified by different shapes: this solution (circle), traditional octree (square), uniform grid (triangle), and KD - tree (diamond). Judging from the data trend, this solution has the highest occlusion culling efficiency in scenes of various complexities, increasing from approximately 32% at 1K objects to approximately 62% at 500K objects. In contrast, the efficiency of the traditional octree increases from approximately 25% to approximately 45%; the KD - tree increases from approximately 26% to approximately 47%; the uniform grid has the lowest efficiency, increasing from approximately 20% to approximately 33%. The occlusion culling efficiency of all methods increases with the increase of scene complexity, but this solution always maintains a leading advantage of 10 - 20 percentage points, indicating its more significant performance advantage in large - scale scenes.

[0076] The spatial octree structure constructed by the above - mentioned method not only records the spatial distribution information of scene objects, but also contains detailed bounding box information and occlusion information, providing efficient data - structure support for subsequent scene rendering, visibility judgment, and spatial query.

[0077] In an optional implementation manner, for the nodes within the frustum, calculate the occlusion coefficient based on the occlusion information of the node. When the occlusion coefficient is greater than the preset occlusion threshold, performing occlusion culling includes:

[0078] For the nodes within the frustum, calculate their comprehensive occlusion coefficient:

[0079] Based on the depth occlusion maps in six orthogonal projection directions, calculate the visibility of each projection plane in combination with the current view angle, and calculate the viewing - distance attenuation factor by obtaining the distance from the node to the viewpoint; weight - combine the visibility and the viewing - distance attenuation factor according to the preset weight coefficient to obtain the comprehensive occlusion coefficient of the node;

[0080] Construct an adaptive occlusion - threshold adjustment mechanism: Based on the target frame rate set by the system, obtain the actual frame rate of the current rendering, and calculate the ratio of the two; multiply the ratio by the preset adjustment coefficient as the dynamic adjustment factor, and multiply it by the base occlusion threshold to obtain the dynamic occlusion threshold of the current frame;

[0081] Perform occlusion judgment on each node in the set of nodes to be processed one by one: Compare the comprehensive occlusion coefficient of each node with the currently calculated dynamic occlusion threshold. When the comprehensive occlusion coefficient of the node is greater than the dynamic occlusion threshold, mark the node as the occluded state, add it to the occlusion - culling list, and perform occlusion culling.

[0082] During the three-dimensional scene rendering process, it is necessary to perform frustum culling on the nodes in the scene and retain the set of nodes within the frustum. For these nodes, the present invention realizes efficient occlusion culling by calculating the comprehensive occlusion coefficient and comparing it with the dynamically adjusted occlusion threshold.

[0083] In one embodiment, for the nodes within the frustum, the system first constructs depth occlusion maps in six orthogonal projection directions. These six directions respectively correspond to the positive X, negative X, positive Y, negative Y, positive Z, and negative Z directions of the object coordinate system. The depth occlusion map in each direction records the depth information of each point in the scene when observed from that direction. The resolution of the depth occlusion map can be set to 512×512 pixels, and each pixel stores a 32-bit floating-point depth value.

[0084] When calculating the comprehensive occlusion coefficient of a node, it is first necessary to obtain the visibility information of the node under the current view. For each node, the system projects the eight vertices of its bounding box onto the six depth occlusion maps and checks whether these vertices are occluded by other objects. Specifically, when implementing, the world coordinates of the node vertices are converted to the coordinate system of the corresponding depth map, and the depth value of this point in the depth map is obtained and compared with the depth value stored in the depth map. If the depth value of the node vertex is greater than the value stored in the depth map, it means that this vertex is occluded by other objects.

[0085] For each projection direction, the system calculates the proportion of the occluded vertices among the eight vertices of the node to obtain the occlusion rate in that direction. For example, if a certain node has 6 vertices occluded in the depth map in the positive X direction, the occlusion rate in that direction is 75%.

[0086] The system calculates the weight of each projection plane according to the angle between the current camera view and the six projection directions. The smaller the angle, the greater the weight. Specifically, the weight can be calculated by the power function of the cosine value of the angle, such as weight = the cube of the cosine value. For example, if the angle between the current view and the positive Z direction is 30 degrees, the weight in that direction is approximately 0.65.

[0087] The system calculates the distance from the node to the viewpoint to obtain the viewing distance attenuation factor. The viewing distance attenuation factor decreases as the distance increases and can be calculated in the following way: when the distance from the center point of the node to the viewpoint is 100 meters, the attenuation factor is 0.8; when the distance is 200 meters, the attenuation factor is 0.6; when the distance is 500 meters, the attenuation factor is 0.3.

[0088] The system performs weighted averaging on the occlusion rates in the six directions according to the corresponding weights to obtain the basic occlusion coefficient of the node. Then, the basic occlusion coefficient is multiplied by the viewing distance attenuation factor to obtain the comprehensive occlusion coefficient of the node. For example, if the weighted average occlusion rate of a certain node in the six directions is 80% and the viewing distance attenuation factor is 0.7, then its comprehensive occlusion coefficient is 56%.

[0089] To adapt to different hardware performances and scene complexities, the present invention constructs an adaptive occlusion threshold adjustment mechanism. The system first sets a target frame rate, such as 60 frames per second. During the rendering process, the system obtains the current actual frame rate, such as 45 frames per second, and calculates the ratio of the two, 45 / 60 = 0.75.

[0090] The system multiplies this ratio by a preset adjustment coefficient to obtain a dynamic adjustment factor. The preset adjustment coefficient can be set to 1.5, so the dynamic adjustment factor is 0.75×1.5 = 1.125. Multiply this factor by the base occlusion threshold to obtain the dynamic occlusion threshold for the current frame. If the base occlusion threshold is set to 50%, then the dynamic occlusion threshold is 50%×1.125 = 56.25%.

[0091] When the actual frame rate is lower than the target frame rate, the dynamic occlusion threshold will increase accordingly, causing more nodes to be culled, thereby improving the rendering efficiency; conversely, when the actual frame rate is higher than the target frame rate, the dynamic occlusion threshold will decrease, reducing the number of culled nodes and improving the rendering quality.

[0092] The system performs an occlusion judgment on each node in the set of nodes to be processed. Compare the comprehensive occlusion coefficient of the node with the currently calculated dynamic occlusion threshold. When the comprehensive occlusion coefficient of the node is greater than the dynamic occlusion threshold, mark the node as the occluded state and add it to the occlusion culling list. For example, if the comprehensive occlusion coefficient of a certain node is 65% and the current dynamic occlusion threshold is 56.25%, then the node is marked as the occluded state.

[0093] The system performs a culling operation on the nodes in the occlusion culling list, and these nodes will not participate in the subsequent rendering process, thereby reducing the rendering load. For the nodes that are not culled, the system continues to execute the normal rendering process.

[0094] In practical applications, the system can also set personalized occlusion threshold adjustment coefficients according to the importance of the nodes. For example, for key objects in the scene, the occlusion threshold adjustment coefficient can be set to 0.8 so that it is not easily culled; for secondary objects, the adjustment coefficient can be set to 1.2 so that it is more easily culled.

[0095] Through the above technical solutions, the present invention realizes adaptive occlusion culling based on a multi-directional depth occlusion map, can dynamically adjust the culling strategy according to the system performance, and improves the rendering efficiency while ensuring the rendering quality.

[0096] In an alternative embodiment, the level-of-detail hierarchy of the node is determined according to the distance from the center of the node's bounding box to the viewpoint, and generating the geometric data of the scene object after level-of-detail simplification includes:

[0097] Calculate the distance from the center of the bounding box to the current viewing point. Based on the ratio of this distance to the reference size of the node, determine the level of detail (LOD) of the node. Specifically: Take the base-2 logarithm of the ratio and round down to obtain the LOD level.

[0098] For the nodes whose LOD levels are determined, calculate their simplification rates: Use the power of 2 to the node's LOD level as the denominator to obtain the target simplification ratio. Based on the target simplification ratio, calculate the folding cost of the mesh edges, where the folding cost includes a weighted combination of the geometric error metric value and the attribute error metric value.

[0099] Perform priority sorting according to the folding cost of the edges and gradually execute the edge folding operation according to the preset simplification ratio: First, fold the edge with the minimum folding cost of the mesh edges, update the positions and attribute information of the adjacent vertices, and continuously execute the edge folding until the target simplification rate is reached to generate the geometric data of the scene object after LOD simplification.

[0100] The bounding box can adopt an axis-aligned bounding box (AABB), which is represented by the minimum point coordinates (xmin, ymin, zmin) and the maximum point coordinates (xmax, ymax, zmax). The calculation of the center point coordinates of the bounding box is as follows:

[0101] The x coordinate of the center point = (xmin + xmax) / 2;

[0102] The y coordinate of the center point = (ymin + ymax) / 2;

[0103] The z coordinate of the center point = (zmin + zmax) / 2;

[0104] Assume the current viewing point coordinates are (viewX, viewY, viewZ), then the distance is calculated as the Euclidean distance between two points. For example, for the center point of the bounding box (centerX, centerY, centerZ), the distance value distance is calculated as the square root of the sum of the squares of the differences in each coordinate between the two points.

[0105] The reference size can be selected as the diagonal length of the bounding box, and the calculation method is the distance between the maximum point and the minimum point of the bounding box. For example, for the diagonal length baseSize of the bounding box, it can be expressed as the square root of the sum of the squares of the differences in the coordinates between the maximum point and the minimum point.

[0106] Based on the distance and the reference size, calculate the ratio ratio = distance / baseSize. This ratio represents the proportion of the viewing distance relative to the object size and is used to determine the appropriate level of simplification.

[0107] To determine the level of detail (LOD) level, take the base-2 logarithm of the ratio and round down. The LOD level lod = floor(log2(ratio)). When ratio is less than 1, lod is negative, indicating that higher precision is required; when ratio is greater than or equal to 1, lod is non-negative, indicating that simplification can be performed. For example, when ratio = 5, log2(5) is approximately 2.32, and rounding down gives lod = 2.

[0108] According to the determined LOD level, calculate the target simplification rate simplificationRate = 1 / (2^lod). For example, when lod = 2, simplificationRate = 1 / 4, indicating that 25% of the geometric information of the original model is retained. When lod is negative, simplificationRate is greater than 1, and no simplification is performed at this time.

[0109] Calculate the folding cost for each edge of the model. The folding cost comprehensively considers geometric error and attribute error. The geometric error metric reflects the impact of edge folding on the shape of the model and can be evaluated by calculating the volume change or quadratic error matrix before and after edge folding. The attribute error metric considers changes in attributes such as texture coordinates, normals, and colors.

[0110] Take a triangular mesh model as an example. This model contains 1000 vertices and 1800 triangles. For the edge (v1, v2), where the coordinates of v1 are (1.2, 3.4, 5.6) and the coordinates of v2 are (1.3, 3.5, 5.5), calculate its folding cost. Assume that the geometric error metric of this edge is 0.015, indicating the impact on the surrounding geometry after folding; the texture coordinate error is 0.008, and the normal error is 0.012. Let the geometric error weight be 0.7 and the attribute error weight be 0.3, then the total folding cost is 0.015×0.7 + (0.008 + 0.012)×0.3 = 0.0155.

[0111] Based on the calculated folding costs, construct a priority queue and sort all edges in ascending order of folding cost. For example, if the cost of edge e1 is 0.0155, the cost of edge e2 is 0.0187, and the cost of edge e3 is 0.0142, then the sorting result is e3, e1, e2.

[0112] According to the target simplification rate, determine the number of vertices to be retained. For example, for the original number of vertices 1000 and a simplification rate of 0.25, the target number of vertices is 250.

[0113] When performing the edge collapse operation, each time the edge with the minimum cost is taken out from the priority queue for collapse. For example, first collapse edge e3, and merge its two vertices into a new vertex. The position of the new vertex can be the midpoint of the original two vertices, or the optimal position determined according to the principle of minimizing error. Suppose edge e3 connects vertex v3(2.1, 1.5, 3.2) and v4(2.2, 1.6, 3.1), and the position of the new vertex after collapse is (2.15, 1.55, 3.15).

[0114] After the collapse operation, update the collapse costs of the affected edges and adjust the priority queue. Continue to perform edge collapse until the target number of vertices or the simplification rate is reached. For example, for a target number of 250 vertices, 750 edge collapse operations need to be performed.

[0115] Generate the simplified geometric data, including vertex coordinates, indices, and attribute information such as texture coordinates and normals. For example, the simplified model contains 250 vertices and approximately 450 triangles, and the data volume is reduced by approximately 75%.

[0116] In practical applications, multiple levels of detail models can be pre-generated according to different viewing distances, stored in memory, and an appropriate model can be selected for rendering according to the detail level calculated in real time. It is also possible to calculate the simplified model in real time, especially for dynamically changing scenes.

[0117] Through the above method, the model complexity can be dynamically adjusted according to the viewing distance. Simplified models are used for distant objects, and high-precision models are maintained for nearby objects, thereby improving the rendering efficiency while ensuring visual quality. Tests show that this method can increase the frame rate by 30% - 50% while maintaining acceptable visual quality.

[0118] In an alternative embodiment, a material instance cache pool based on material similarity is established for the scene objects in each rendering batch, and the material instances are dynamically merged according to the degree of change of the material attributes. Merging multiple materials with material attribute similarity higher than the set similarity threshold into one material instance includes:

[0119] Obtain the material data of the scene objects in the rendering batch, extract features from the material data, and construct a feature vector space; establish a material descriptor index tree based on the feature vector space, and the material descriptor index tree is used for fast retrieval and matching of materials;

[0120] Calculate the similarity between materials based on the material descriptor index tree: for non-texture attributes, directly calculate the Euclidean distance of the attribute values; for texture attributes, obtain the texture similarity by calculating the normalized difference of the texture pixel values; combine the similarities of each attribute with weights to obtain the comprehensive similarity value between material pairs;

[0121] Divide the materials into multiple groups based on the comprehensive similarity value, construct a reference relationship graph of material instances within each material group, and record the association information of each material instance being referenced by scene objects; perform hierarchical clustering on each material group again to generate an optimal merging sequence;

[0122] Perform material instance merging according to the optimal merging sequence: calculate the merged material parameters based on the attribute values of each instance within the material group, where non-texture parameters are obtained through weighted averaging, and texture parameters are generated through mixed sampling; update the material reference relationship of the scene object, and point the reference of the original material instance to the newly merged material instance.

[0123] Figure 4 Figure [X] is a schematic diagram of the system interface for material instance merging. Among them, material data includes non-texture attributes (such as diffuse coefficient, specular intensity, metallicity, roughness, etc.) and texture attributes (such as diffuse texture map, normal map, specular map, etc.). Feature extraction is performed on each material to construct a feature vector space. During the feature extraction process, for non-texture attributes, the numerical values are directly used as feature components; for texture attributes, the main features of the texture are extracted through downsampling and principal component analysis. For example, a 1024×1024 diffuse texture map is downsampled to 64×64, and then the mean, variance, and histogram features of the RGB channels are extracted as the feature representation of the texture.

[0124] Based on the extracted feature vectors, construct a material descriptor index tree. This index tree adopts a KD-tree structure, where each node represents a material instance, and the distance between nodes represents the similarity between materials. The construction process of the KD-tree is as follows: First, sort all material instances according to the first-dimensional feature, and select the median as the root node; then recursively divide according to the next-dimensional feature in the left and right subtrees until all material instances are added to the tree. For example, for a scene containing 100 material instances, the depth of the constructed KD-tree is about 7 layers, and the query efficiency can reach the O(log n) level.

[0125] The similarity calculation between materials is divided into two parts: non-texture attribute similarity and texture attribute similarity. For non-texture attributes, the normalized Euclidean distance is used to calculate the similarity. For example, if the diffuse coefficient of material A is (0.8, 0.2, 0.3) and the diffuse coefficient of material B is (0.7, 0.3, 0.3), then the distance between them in this attribute is 0.14, and the similarity is 1 - 0.14 = 0.86. For texture attributes, the similarity is obtained by calculating the normalized difference of texture pixel values. Specifically, when implemented, 100 corresponding pixels are randomly sampled from the two textures, and their average color difference is calculated, and then normalized to the [0, 1] interval to obtain the similarity value. For example, if the average color difference after sampling of two similar metal textures is 15%, then the texture similarity is 0.85.

[0126] The weight distribution is determined according to the material type and rendering importance. For example, for metal materials, the metallicity weight is 0.4, the roughness weight is 0.3, the diffuse texture weight is 0.2, and the normal map weight is 0.1. Suppose the similarities of two metal materials in these attributes are 0.95, 0.9, 0.8, and 0.7 respectively. Then the comprehensive similarity is 0.95×0.4 + 0.9×0.3 + 0.8×0.2 + 0.7×0.1 = 0.88.

[0127] Based on the comprehensive similarity value, the materials are divided into multiple groups using the threshold method. Set the similarity threshold to 0.85. Then the materials with a comprehensive similarity greater than 0.85 are divided into the same group. Within each material group, a reference relationship graph of material instances is constructed to record the association information of each material instance being referenced by scene objects. The reference relationship graph is stored using an adjacency list, and each material instance node contains a list of scene objects that reference the material.

[0128] The hierarchical clustering process is as follows: Initially, each material instance is an independent cluster, and the similarity between all pairs of clusters is calculated; the two clusters with the highest similarity are merged, and the inter-cluster similarity is updated; repeat the above steps until the preset number of clusters is reached or the similarity is lower than the threshold. For example, a group containing 8 similar materials may eventually be merged into 3 material instances through hierarchical clustering, and the merging sequence is [(M1,M3),(M2,M5),(M4,M7),(M6,M8),((M1,M3),(M4,M7)),((M2,M5),(M6,M8))].

[0129] For non-texture parameters, the merged value is obtained through weighted averaging. The weights can be determined based on the frequency of material references. For example, material A is referenced 10 times and material B is referenced 5 times. The merged diffuse coefficient is (A.diffuse×10 + B.diffuse×5) / 15. For texture parameters, a new texture is generated through mixed sampling. Specifically, when implementing, a new texture with the same resolution as the original texture is created. For each pixel position, the pixel values at the corresponding positions of the original textures are mixed according to the weights. For example, two 512×512 diffuse texture maps are merged with a weight ratio of 6:4. Then each pixel value in the new texture is the weighted average of the pixel values at the corresponding positions of the original textures.

[0130] After completing the material merging, update the material reference relationship of the scene objects, and point the references of the original material instances to the newly merged material instances. At the same time, maintain the material instance cache pool, regularly clean up the un-referenced material instances, and trigger the re-evaluation and merging process when the scene changes.

[0131] Through the above method, in a scene containing 1000 scene objects and using 800 material instances, after material merging, the number of material instances can be reduced to about 300, the rendering batches are reduced by about 60%, the rendering performance is improved by about 35%, while keeping the reduction of visual quality controlled within an acceptable range (PSNR value greater than 32 dB).

[0132] In an alternative embodiment, the materials are divided into multiple groups based on the comprehensive similarity value. A reference relationship graph of material instances is constructed within each material group, and the association information of each material instance being referenced by scene objects is recorded; Hierarchical clustering is performed on each material group again, and the optimal merging sequence generated includes:

[0133] Construct a bipartite graph structure of the material group, where the bipartite graph structure includes a set of material instance vertices and a set of scene object vertices, and a reference relationship edge set is established between the set of material instance vertices and the set of scene object vertices;

[0134] Calculate the reference weight in the bipartite graph structure. According to the usage area of each material instance on the surface of the scene object, calculate the area ratio of a single material instance being referenced by a single scene object, and use the area ratio as the weight value of this reference relationship;

[0135] Based on the bipartite graph structure, construct a hierarchical clustering tree, and calculate the distance metric value for each pair of material clusters. The distance metric value is obtained by calculating the average value of the distances between all pairs of materials in the two material clusters;

[0136] Take the material similarity, reference relationship overlap degree, and rendering cost as evaluation indicators, perform weighted combination on the three evaluation indicators to obtain the merging evaluation score between material clusters;

[0137] Insert the pairs of material clusters with a merging evaluation score higher than the preset evaluation threshold into the priority queue in descending order of the score; Extract the pair of material clusters with the highest score from the priority queue as the optimal merging object.

[0138] The comprehensive similarity value can be calculated by comparing features such as the color, texture, glossiness, and roughness of the materials. For example, for two materials A and B, extract their feature vectors, calculate the Euclidean distance between the feature vectors, and normalize the distance value to the range of 0 - 1 as the similarity value. When the similarity value is greater than 0.8, these two materials can be classified into the same group.

[0139] Within each material group, construct a reference relationship graph of material instances. The reference relationship graph records the association information of each material instance being referenced by scene objects. For example, material M1 is referenced by objects O1 and O2, and material M2 is referenced by objects O2 and O3. These reference relationships form a network structure.

[0140] Next, perform hierarchical clustering on each material group to generate an optimal merging sequence. The specific steps are as follows:

[0141] Construct a bipartite graph structure for the material group. This bipartite graph contains two vertex sets: the material instance vertex set and the scene object vertex set. For example, the material instance set contains {M1, M2, M3}, and the scene object set contains {O1, O2, O3}. Establish a set of reference relationship edges between the two sets. If material M1 is referenced by scene object O1, then an edge is established between M1 and O1.

[0142] According to the usage area of each material instance on the surface of the scene object, calculate the proportion of the area where a single material instance is referenced by a single scene object, and use this area proportion as the weight value of the reference relationship. For example, if material M1 covers 80% of the surface area of object O1, then the reference weight from M1 to O1 is 0.8. If the surface area of object O1 is 100 square units, then the actual usage area of M1 on O1 is 80 square units.

[0143] Calculate the distance metric value for each pair of material clusters. This distance metric value is obtained by calculating the average of the distances between all pairs of materials in the two material clusters. For example, cluster C1 contains materials {M1, M2}, and cluster C2 contains materials {M3, M4}, then the distance between C1 and C2 is (distance(M1, M3) + distance(M1, M4) + distance(M2, M3) + distance(M2, M4)) / 4.

[0144] When evaluating the merging of material clusters, consider three key metrics: material similarity, reference relationship overlap, and rendering overhead. Material similarity reflects the degree of closeness between two materials in terms of visual characteristics; reference relationship overlap indicates the degree to which two materials are referenced by the same scene object; rendering overhead considers the impact on rendering performance after merging.

[0145] Perform a weighted combination of these three evaluation metrics to obtain the merging evaluation score between pairs of material clusters. For example, the weight of material similarity can be set to 0.5, the weight of reference relationship overlap to 0.3, and the weight of rendering overhead to 0.2. Suppose the material similarity between material clusters C1 and C2 is 0.9, the reference relationship overlap is 0.7, and the rendering overhead reduction index after merging is 0.8, then the merging evaluation score is 0.5×0.9 + 0.3×0.7 + 0.2×0.8 = 0.83.

[0146] Insert the pairs of material clusters with a merging evaluation score higher than the preset evaluation threshold into the priority queue in descending order of the score. For example, set the preset evaluation threshold to 0.7, then all pairs of material clusters with a score greater than 0.7 will be inserted into the priority queue. Suppose there are three pairs of material clusters with scores of 0.83, 0.76, and 0.72 respectively, and they will all be inserted into the priority queue and sorted from high to low according to the score.

[0147] Extract the material cluster pair with the highest score from the priority queue as the optimal merging object. In the above example, the material cluster pair with a score of 0.83 will be merged first. After merging, it is necessary to update the bipartite graph structure and related reference relationships, then recalculate the merging evaluation scores between the remaining material cluster pairs, and update the priority queue.

[0148] In practical applications, an iterative termination condition can be set, for example, stop merging when the priority queue is empty or the highest score is lower than a certain threshold. In this way, while maintaining visual quality, the number of materials can be minimized, and the rendering efficiency can be improved.

[0149] Taking a specific scenario as an example, assume there is a 3D scene with 100 objects and 50 materials. After the initial grouping, 5 material groups are formed, with each group containing approximately 10 materials. In the first material group, through bipartite graph analysis, it is found that the merging evaluation score of materials M1 and M2 is 0.92, which is much higher than the preset threshold of 0.7. Therefore, they are merged. After merging, the number of materials in the scene is reduced to 49, and the visual quality hardly changes. After the complete merging process, the final number of materials may be reduced to about 30, the rendering time is reduced by about 40%, while maintaining an acceptable visual quality.

[0150] Through the above material merging optimization method, the number of materials in the 3D scene can be effectively reduced, the rendering overhead can be reduced, and the application performance can be improved, which is especially suitable for resource-constrained environments such as real-time rendering and mobile devices.

[0151] The real-time rendering optimization system in the metaverse scene construction engine of the embodiment of the present invention includes:

[0152] The first unit is used to obtain the three-dimensional scene data to be rendered in the metaverse scene construction engine, and the three-dimensional scene data includes the geometric data, material data, and lighting data of the scene objects;

[0153] The second unit is used to construct a spatial octree structure of the scene objects based on the three-dimensional scene data, and record the bounding box information and occlusion information of the scene objects corresponding to each node in each node of the spatial octree structure;

[0154] The third unit is used to recursively traverse each node from top to bottom in the spatial octree structure according to the user's viewpoint position information, determine whether the bounding box of each node is within the viewing frustum, and directly prune the nodes outside the viewing frustum; for the nodes within the viewing frustum, calculate the occlusion coefficient based on the occlusion information of the node, and perform occlusion culling when the occlusion coefficient is greater than the preset occlusion threshold; and determine the level of detail of the node according to the distance from the center of the bounding box of the node to the viewpoint, and generate the geometric data of the scene object after level-of-detail simplification;

[0155] A fourth unit, configured to divide the scene object geometry data into multiple rendering batches, each rendering batch including scene objects having the same material attributes;

[0156] A fifth unit, configured to establish a material instance cache pool based on material similarity for the scene objects in each rendering batch, dynamically merge the material instances according to the degree of change of the material attributes, and merge multiple materials with material attribute similarity higher than a set similarity threshold into one material instance.

[0157] In a third aspect of the embodiments of the present invention,

[0158] Provided is an electronic device, including:

[0159] A processor;

[0160] A memory for storing instructions executable by the processor;

[0161] Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.

[0162] In a fourth aspect of the embodiments of the present invention,

[0163] Provided is a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.

[0164] The present invention may be a method, an apparatus, a system, and / or a computer program product. The computer program product may include a computer-readable storage medium, on which computer-readable program instructions for executing various aspects of the present invention are loaded.

[0165] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A real-time rendering optimization method in a metaverse scene building engine, characterized in that: include: Obtaining three-dimensional scene data to be rendered in the Metaverse scene building engine, wherein the three-dimensional scene data includes geometric data, material data, and lighting data of scene objects; Based on the three-dimensional scene data, construct a spatial octree structure of the scene object, and record, in each node of the spatial octree structure, bounding box information and occlusion information of the scene object corresponding to the node; According to the user's viewpoint position information, recursively traverse each node from top to bottom in the spatial octree structure, determine whether the bounding box of each node is within the viewing cone, and directly prune the nodes outside the viewing cone; for the nodes within the viewing cone, calculate the occlusion coefficient based on the occlusion information of the node, and perform occlusion culling when the occlusion coefficient is greater than a preset occlusion threshold; and determining the level of detail of the node according to the distance from the center of the bounding box of the node to the viewpoint, generating scene object geometric data after the level of detail is simplified; dividing the scene object geometric data into a plurality of rendering batches, each rendering batch containing scene objects with the same material attributes; A material instance cache pool based on material similarity is established for each scene object in each rendering batch. Material instances are dynamically merged according to the degree of change of material properties. Multiple materials with material property similarity higher than the set similarity threshold are merged into one material instance, including: Acquire material data of scene objects in a rendering batch, perform feature extraction on the material data, and construct a feature vector space; establish a material descriptor index tree based on the feature vector space, and the material descriptor index tree is used for rapid retrieval and matching of materials; Calculating the similarity between materials based on the material descriptor index tree: for non-texture attributes, directly calculating the Euclidean distance of the attribute value; for texture attributes, obtaining the texture similarity by calculating the normalized difference of the texture pixel value; performing weighted combination of the similarities of the various attributes to obtain the comprehensive similarity value between the material pairs; Divide the materials into multiple groups based on the comprehensive similarity value, construct a reference relationship graph of material instances in each material group, and record the association information of each material instance referenced by the scene object; perform hierarchical clustering on each material group again to generate an optimal merge sequence; Material instance merging is performed according to the optimal merging sequence: merged material parameters are calculated based on the attribute values ​​of each instance in the material group, non-texture parameters are obtained by weighted averaging, and texture parameters are generated by mixed sampling; the material reference relationship of the scene object is updated, and the reference of the original material instance is pointed to the merged new material instance.

2. The method according to claim 1, characterized in that Based on the three-dimensional scene data, constructing a spatial octree structure of the scene object, and recording the bounding box information and occlusion information of the scene object corresponding to each node of the spatial octree structure includes: Calculating the overall boundary range of the scene based on the geometric data of the scene object, determining the minimum vertex coordinates and the maximum vertex coordinates of the scene bounding box, constructing the root node space of the octree, and mapping the scene object to the root node space according to its position information; Traversing the scene objects in the root node space, when the number of scene objects contained in the node exceeds the preset segmentation threshold, evenly dividing the node space into 8 sub-node spaces, and reallocating the scene objects in the node to the corresponding sub-node spaces according to the position information; recursively executing the above space division process until the number of scene objects in all nodes does not exceed the preset segmentation threshold; Traverse all nodes of the octree from top to bottom: for each node, calculate the axial bounding box information of the node, including the minimum vertex coordinates and the maximum vertex coordinates; calculate the surface area occlusion ratio of the node: obtained by calculating the ratio of the area of ​​the node surface occluded by other scene objects to the total surface area of ​​the node; project the node onto six orthogonal planes to form a depth occlusion map, and record the maximum depth value of each projected pixel position.

3. The method according to claim 2, characterized in that For a node within the view frustum, the occlusion coefficient is calculated based on the occlusion information of the node. When the occlusion coefficient is greater than a preset occlusion threshold, occlusion culling is performed, including: For the nodes within the viewing cone, calculate their comprehensive occlusion coefficient: Based on the depth occlusion map of six orthogonal projection directions, the visibility of each projection surface is calculated in combination with the current viewing angle, and the distance from the node to the viewpoint is calculated to obtain the view distance attenuation factor; the visibility and the view distance attenuation factor are weighted and combined according to a preset weight coefficient to obtain a comprehensive occlusion coefficient of the node; Build an adaptive occlusion threshold adjustment mechanism: Based on the target frame rate set by the system, obtain the actual frame rate of the current rendering and calculate the ratio of the two; multiply the ratio by the preset adjustment coefficient as the dynamic adjustment factor, and multiply it by the basic occlusion threshold to obtain the dynamic occlusion threshold of the current frame; Perform occlusion judgment on each node in the processing node set: compare the comprehensive occlusion coefficient of each node with the currently calculated dynamic occlusion threshold. When the comprehensive occlusion coefficient of the node is greater than the dynamic occlusion threshold, mark the node as occluded, add it to the occlusion culling list, and perform occlusion culling.

4. The method according to claim 1, characterized in that: The node's level of detail is determined based on the distance from the center of the node's bounding box to the viewpoint, and the generated scene object geometry data after level of detail simplification includes: Calculate the distance from the center of the bounding box to the current viewpoint, and determine the node's level of detail based on the ratio of the distance to the node's base size. Specifically, take the base 2 logarithm of the ratio and round it down to get the level of detail. For nodes with a certain level of detail, calculate their simplification rate: use 2 raised to the power of the node level as the denominator to get the target simplification ratio; based on the target simplification ratio, calculate the folding cost of the mesh edge, where the folding cost includes a weighted combination of the geometric error metric and the attribute error metric; Prioritize edges according to their folding costs, and perform edge folding operations step by step according to the preset simplification ratio: first fold the edges with the smallest folding costs, update the positions and attribute information of adjacent vertices, and continue to perform edge folding until the target simplification rate is reached, generating scene object geometry data after detail level simplification.

5. The method according to claim 1, characterized in that Dividing the materials into multiple groups based on the comprehensive similarity value, constructing a reference relationship graph of the material instances in each material group, and recording the associated information of each material instance being referenced by the scene object; performing hierarchical clustering on each material group again to generate an optimal merge sequence includes: Constructing a bipartite graph structure of a material group, the bipartite graph structure comprising a material instance vertex set and a scene object vertex set, and establishing a reference relationship edge set between the material instance vertex set and the scene object vertex set; Calculating the reference weight in the bipartite graph structure, calculating the area ratio of a single material instance referenced by a single scene object according to the used area of ​​each material instance on the surface of the scene object, and using the area ratio as the weight value of the reference relationship; Building a hierarchical clustering tree based on the bipartite graph structure, and calculating a distance metric value for each pair of material clusters, wherein the distance metric value is obtained by calculating the average distance between all material pairs in the two material clusters; The material similarity, reference relationship overlap and rendering cost are used as evaluation indicators, and the three evaluation indicators are weighted and combined to obtain the combined evaluation score between material cluster pairs. The material cluster pairs whose merging evaluation scores are higher than a preset evaluation threshold are inserted into a priority queue in descending order of scores; and the material cluster pair with the highest score is extracted from the priority queue as the optimal merging object.

6. A real-time rendering optimization system in a metaverse scene building engine, used to implement the method as claimed in any one of claims 1 to 5, characterized in that: include: The first unit is used to obtain three-dimensional scene data to be rendered in the Metaverse scene building engine, wherein the three-dimensional scene data includes geometric data, material data and lighting data of scene objects; A second unit is used to construct a spatial octree structure of the scene object based on the three-dimensional scene data, and record, in each node of the spatial octree structure, bounding box information and occlusion information of the scene object corresponding to the node; The third unit is used to recursively traverse each node from top to bottom in the spatial octree structure according to the user viewpoint position information, determine whether the bounding box of each node is within the viewing cone, and directly prune the nodes outside the viewing cone; for the nodes within the viewing cone, calculate the occlusion coefficient based on the occlusion information of the node, and perform occlusion culling when the occlusion coefficient is greater than a preset occlusion threshold; And determine the detail level of the node according to the distance from the center of the bounding box of the node to the viewpoint, and generate the geometric data of the scene object after the detail level is simplified; A fourth unit, configured to divide the scene object geometry data into a plurality of rendering batches, each rendering batch containing scene objects having the same material attributes; The fifth unit is used to establish a material instance cache pool based on material similarity for scene objects in each rendering batch, dynamically merge material instances according to the degree of change of material properties, and merge multiple materials whose material property similarity is higher than a set similarity threshold into one material instance.

7. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described in any one of claims 1 to 5.

8. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Rendering optimization method and device

    CN106504185A

  • Three-dimensional rendering method and device, equipment and medium

    CN117237502A