Urban digital twinning scene static LOD processing method
By screening the model in the visual vertebra and calculating the visual similarity, and dynamically adjusting the LOD level, the problem of visual jump in traditional LOD technology is solved, and more efficient rendering quality and efficiency are achieved.
Patent Information
- Application Number
- CN202510866029.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-26
AI Technical Summary
When objects move rapidly or viewing angles change, traditional LOD technology causes models to switch instantly between discrete levels, resulting in obvious visual jumps (popping) problems.
By obtaining the visual vertebra corresponding to the current viewpoint position and view angle, calculate the perceived screen pixel threshold and visual similarity of the model to be rendered, dynamically adjust the LOD level of the model, and use geometric deformation or transparency mixing algorithms to make a smooth transition to avoid visual jumps.
While maintaining user visual fidelity, it significantly reduces rendering overhead, improves rendering quality and efficiency, and avoids unnecessary model processing.
Smart Images

Figure CN120355830A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of model rendering processing, and particularly to a method for static LOD processing in an urban digital twin scene. Background Art
[0002] Traditional LOD (Level of Detail) technology balances rendering performance and picture quality by dynamically adjusting model details.
[0003] However, due to its discrete level switching and the design of static error thresholds, the LOD system triggers level switching by presetting the SSE (Screen-Space Error) threshold. When an object moves quickly or the viewing angle changes suddenly, the error threshold of SSE is disconnected from the actual visual perception, causing the model to switch instantaneously between discrete levels and resulting in obvious visual popping.
[0004] The above content is only used to assist in understanding the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention
[0005] The main purpose of this application is to provide a method for static LOD processing in an urban digital twin scene, aiming to solve the technical problem of visual popping during model rendering in the current three-dimensional scene.
[0006] To achieve the above purpose, this application proposes a method for static LOD processing in an urban digital twin scene, and the method includes: Obtain the frustum corresponding to the current viewpoint position and viewing angle, and determine the model to be rendered located within the frustum in the three-dimensional scene; Calculate the perceivable screen pixel threshold corresponding to the model to be rendered at the current viewpoint position; Obtain the visual similarity between the model to be rendered and the original model; If the visual similarity is less than the preset similarity, determine the target LOD of the model to be rendered corresponding to the perceivable screen pixel threshold.
[0007] In an embodiment, the step of determining the target LOD of the model to be rendered corresponding to the perceivable screen pixel threshold if the visual similarity is less than the preset similarity includes: If the visual similarity is less than the preset similarity, determine the LOD sequence with a screen space error less than or equal to the perceivable screen pixel threshold; Set the minimum LOD level of the LOD sequence as the target LOD.
[0008] In one embodiment, before the step of obtaining a frustum corresponding to the current viewpoint position and viewing angle and determining a model to be rendered located within the frustum in a three-dimensional scene, the static LOD processing method for the urban digital twin scene further includes: Determine the LOD corresponding to each of the original models in the three-dimensional scene, and construct a quadtree or octree according to the LOD. The leaf nodes of the quadtree or the octree include multiple LOD levels corresponding to each of the original models; Determine the screen space error corresponding to each LOD level at each viewing distance.
[0009] In one embodiment, the step of calculating the perceivable screen pixel threshold corresponding to the model to be rendered at the current viewpoint position includes: Obtain the screen physical size and pixel density of the acquisition device, and the viewing distance between the acquisition device and the model to be rendered; Calculate the product of the screen physical size and the pixel density, and the quotient between the product and the viewing distance; Set the product of the quotient, the minimum resolvable angle of the human eye, and the proportional coefficient between the angle and the radian as the perceivable screen pixel threshold.
[0010] In one embodiment, the step of obtaining the visual similarity between the model to be rendered and the original model includes: Determine the screen space geometric error and the screen texture error between the model to be rendered and the original model; Based on a preset weight value, fuse the screen space geometric error and the screen texture error to obtain a screen space visual error; Calculate the visual similarity according to the natural exponential function between the attenuation coefficient and the screen space visual error.
[0011] In one embodiment, after the step of, if the visual similarity is less than a preset similarity, determining the target LOD of the model to be rendered corresponding to the perceivable screen pixel threshold, the static LOD processing method for the urban digital twin scene further includes: If the target LOD is different from the LOD selected in the previous frame, perform a smooth process on the target LOD based on a geometric deformation or transparency blending algorithm to avoid visual jumps; Send the model to be rendered to the rendering pipeline so that the rendering pipeline performs a rendering process on the model to be rendered.
[0012] In one embodiment, the step of obtaining a frustum corresponding to the current viewpoint position and viewing angle and determining a model to be rendered located within the frustum in a three-dimensional scene includes: Calculate six plane equations based on the viewpoint position and the viewing angle, and form the viewing frustum based on the six plane equations; Determine the intersection detection result between the viewing frustum and the three-dimensional scene corresponding to the original model, and set the model that intersects with the plane of the viewing frustum and is inside the viewing frustum as the model to be rendered.
[0013] In one embodiment, after the step of obtaining the visual similarity between the model to be rendered and the original model, the static LOD processing method for the urban digital twin scene further includes: If the visual similarity is greater than or equal to the preset similarity, jump to execute the step of obtaining the viewing frustum corresponding to the current viewpoint position and viewing angle, and determining the model to be rendered located inside the viewing frustum in the three-dimensional scene.
[0014] In addition, to achieve the above object, the present application also proposes a level of detail processing device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the static LOD processing method for the urban digital twin scene as described above.
[0015] In addition, to achieve the above object, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, it implements the steps of the static LOD processing method for the urban digital twin scene as described above.
[0016] One or more technical solutions proposed by the present application have at least the following technical effects: First, obtain the viewing frustum corresponding to the current viewpoint position and viewing angle, so as to screen out the model to be rendered in the visible part inside the viewing frustum in the three-dimensional scene, and avoid unnecessary model processing. Subsequently, calculate the perceivable screen pixel threshold of the model to be rendered at the current viewpoint position to evaluate the visual saliency of the model on the screen. Then, obtain the visual similarity between the model to be rendered and the original model, and when the similarity is lower than the preset value, adaptively determine the target LOD level according to the perceivable screen pixel threshold. Based on this, through dynamic perception threshold and visual similarity evaluation, the LOD decision is changed from discrete switching based on preset rules to continuous adaptation based on actual perception, intelligently adjusting the level of detail, while maintaining the visual fidelity perceived by the user to avoid visual jumps, significantly reducing the rendering overhead and improving the rendering quality. Description of the Drawings
[0017] The drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0018] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0019] Figure 1 It is a schematic flowchart provided for the first embodiment of a static LOD processing method for an urban digital twin scenario in the present application; Figure 2 It is a schematic flowchart provided for the second embodiment of a static LOD processing method for an urban digital twin scenario in the present application; Figure 3 It is a schematic overall working flowchart obtained by combining the embodiments of the present application; Figure 4 It is a schematic diagram of the device structure of the hardware operating environment involved in a static LOD processing method for an urban digital twin scenario in the embodiments of the present application.
[0020] The realization of the purpose, functional features and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. Detailed Embodiments
[0021] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.
[0022] Traditional LOD (Level of Detail) technology balances rendering performance and picture quality by dynamically adjusting model details.
[0023] However, due to its discretized level switching and the design of the static error threshold, the LOD system triggers level switching by presetting the SSE (Screen-Space Error) threshold. When the object moves rapidly or the viewing angle changes suddenly, the error threshold of the SSE is disconnected from the actual visual perception, causing the model to switch instantaneously between discrete levels, resulting in obvious visual popping.
[0024] The main solution of the embodiments of the present application is: obtaining the frustum corresponding to the current viewpoint position and viewing angle, and determining the model to be rendered located in the frustum in the three-dimensional scene; Calculating the perceivable screen pixel threshold corresponding to the model to be rendered at the current viewpoint position; Obtaining the visual similarity between the model to be rendered and the original model; If the visual similarity is less than the preset similarity, determining the target LOD of the model to be rendered corresponding to the perceivable screen pixel threshold.
[0025] Specifically, through dynamic perception threshold and visual similarity evaluation, the LOD decision is changed from discrete switching based on preset rules to continuous adaptation based on actual perception, intelligently adjusting the level of detail, significantly reducing the rendering overhead while maintaining the visual fidelity perceived by the user and improving the rendering quality.
[0026] It should be noted that the execution entity of this embodiment can be a computing service device with data processing, network communication, and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device, a level-of-detail processing device, etc. that can implement the above functions. Hereinafter, taking the level-of-detail processing device as an example, this embodiment and the following embodiments will be described.
[0027] To better understand the technical solution of this application, the following will be described in detail in combination with the specification drawings and specific implementation manners.
[0028] The embodiment of this application provides a method for processing static LOD of an urban digital twin scene. Refer to Figure 1 , Figure 1 which is a schematic flowchart of the first embodiment of a method for processing static LOD of an urban digital twin scene in this application.
[0029] In this embodiment, the method for processing static LOD of an urban digital twin scene includes steps S10 to S40: Step S10, obtain the viewing frustum corresponding to the current viewpoint position and viewing angle, and determine the models to be rendered in the three-dimensional scene that are located within the viewing frustum.
[0030] It should be noted that in three-dimensional computer graphics, a frustum-shaped spatial region defined from the observer (viewpoint) according to the set viewing angle (such as horizontal / vertical field of view angle) and near / far clipping planes is the viewing frustum. After calculating the viewing frustum, the spatial index (such as quadtree / octree) can be quickly traversed to mark which nodes are within the viewing frustum (need to be rendered) and which are outside the viewing frustum (can be ignored or only the lowest LOD is retained). It can be understood that only the objects within this region may be rendered onto the screen. Therefore, during the rendering process of a specific frame, after the viewing frustum culling test, the three-dimensional models in the three-dimensional scene that are determined to be located inside the viewing frustum and need to be drawn onto the screen are the models to be rendered. Among them, each model in the three-dimensional scene, such as buildings, roads, etc., corresponds to a separate model to be rendered.
[0031] The rendering engine of the level-of-detail processing device first obtains the current viewpoint position (usually three-dimensional spatial coordinates) and viewing angle (including the viewing direction, horizontal / vertical field of view angle) of the observer in the virtual scene. Subsequently, according to the viewpoint position, viewing direction, field of view angle, and the pre-set near clipping plane distance and far clipping plane distance, six plane equations (top, bottom, left, right, near, far) are calculated, and these six planes together enclose a frustum. It can be understood that the frustum can be determined by conventional calculation methods, and this application will not elaborate on it here. Among them, when the viewpoint position and viewing angle of the observer change, the frustum needs to be re-determined.
[0032] After obtaining the frustum, it is necessary to determine the models to be rendered located within the frustum. As an optional implementation manner for determining the models to be rendered, the spatial index (quad-tree / octree) can be quickly traversed, and the nodes within the frustum (to be rendered) and those outside the frustum (can be ignored or only the lowest LOD is retained) are marked, that is, the models corresponding to the nodes within the frustum are used as the models to be rendered.
[0033] Optionally, in another optional implementation manner for determining the models to be rendered, a frustum culling test can be performed on all candidate 3D models in the scene. Specifically, it includes performing an intersection detection between the bounding volume of the model and the six planes of the frustum. If the bounding volume is completely outside a certain frustum plane, the model is invisible and is culled. If the bounding volume intersects with the frustum or is completely within the frustum, the model or its intersecting part is marked as the model to be rendered. Thus, the intersection detection result between the frustum and the 3D scene corresponding to the original model can be determined, and the models that intersect with the frustum plane and are within the frustum are set as the models to be rendered. Finally, the set of models that pass the test is output as the models to be rendered for this frame.
[0034] Exemplarily, in a 3D urban scene, the user character is located at the coordinates (0, 10, 0), facing directly forward (such as the negative Z-axis direction), the field of view angle is 90 degrees, the near clipping plane distance is 1 meter, and the far clipping plane distance is 100 meters. The rendering engine calculates the corresponding frustum and detects that a house (its AABB is within the frustum) and a tree (completely within the frustum) in the scene are within the frustum, while another tree behind the user character is culled. The house and the tree are the models to be rendered for the current frame.
[0035] Step S20, calculate the perceivable screen pixel threshold corresponding to the model to be rendered at the current viewpoint position.
[0036] The Perceivable Screen Pixel (PSP) hereinafter refers to the minimum critical value at which the pixel size occupied by a feature (such as a vertex, an edge, or a tiny area) on the surface of a 3D model after being projected onto the screen (imaging plane) at the current viewpoint position and viewing angle reaches the minimum value that the human eye can just clearly distinguish its details.
[0037] In this embodiment, for the selected model to be rendered, the spatial distance (Distance), i.e., the viewing distance, from the surface feature points (or the model center / vertices of the bounding box) of the model to the current viewpoint position can be calculated. At the same time, based on the physical size of the screen and the field of view (FOV) of the current acquisition device and by applying the perspective projection principle, the pixel size corresponding to a unit spatial length on the screen at this distance, i.e., the pixel density, can be calculated. The PSP is calculated through the pixel density, the physical size of the screen, and the viewing distance.
[0038] As an alternative implementation for determining the PSP, the physical size of the screen and the pixel density of the acquisition device, as well as the viewing distance between the acquisition device and the model to be rendered, can be obtained. Subsequently, the product of the physical size of the screen and the similarity density is calculated, and then the quotient between the product and the viewing distance is calculated. Finally, in combination with the quotient, the minimum resolution angle of the human eye, and the product of the angle and the radian coefficient, the PSP is calculated. Therefore, the calculation formula for the PSP is as follows:
[0039] where θ min is the minimum resolution angle of the human eye, usually taken as 0.016°, and π / 180 is the coefficient of the angle-to-radian ratio. It can be understood that the minimum resolution angle of the human eye is a radian coefficient, so parameter conversion is required.
[0040] This embodiment provides a quantization standard by calculating the PSP for LOD decision-making, accurately selects the coarsest LOD level that meets the visual fidelity when an upgrade is needed, and intelligently adjusts the LOD based on the PSP, thereby reasonably simplifying the distant model and improving the rendering efficiency of the model.
[0041] Step S30: Obtain the visual similarity between the model to be rendered and the original model.
[0042] In this embodiment, the above-mentioned original model is the version with the highest geometric detail and texture resolution in the 3D model, representing the highest visual fidelity of the model. Visual similarity (SIM, Similarity) is used to measure the degree of closeness in visual appearance between the current model instance used for rendering (usually a model at a certain LOD level) and its corresponding original model. The similarity is usually in the range of [0, 1] or [0%, 100%], and 1 (or 100%) means exactly the same. Among them, each model to be rendered corresponds to an original model. For example, if the model to be rendered is a tree, its original model is also a tree in the corresponding scene.
[0043] It can be understood that the PSP is the core basis for the LOD upgrade decision of the model to be rendered, while the visual similarity is the prerequisite for triggering the LOD upgrade. When the visual quality of the current model is insufficient (similarity < preset value), it is necessary to start the fine LOD selection based on visual perception. Therefore, after calculating the PSP, it is also necessary to obtain the visual similarity between the models.
[0044] In this embodiment, the visual similarity between the two models can be directly obtained from the local database, or the similarity between the two can be calculated in real time.
[0045] As an alternative implementation, during the model preprocessing stage (offline process), for each generated LOD-level model, calculate its visual similarity with the original model so that the visual similarity can be directly obtained from the database later. Common algorithms include geometric error metrics: such as through the Hausdorff distance (calculating the maximum and minimum distances between two sets of mesh surface points), root mean square vertex error (RMS), etc. Image space metrics: Render the original model and the LOD model at multiple predefined viewpoints, compare the differences in the rendered result images. Common methods include structural similarity index (SSIM), peak signal-to-noise ratio (PSNR), or perception-driven difference map analysis, and obtain an average similarity score by integrating the results from multiple perspectives.
[0046] Exemplarily, the model to be rendered currently uses the LOD1-level model. In the model resource library, the visual similarity between the LOD1 model and its original model (LOD0) has been calculated by the SSIM algorithm (rendering and comparing at 8 standard viewpoints) during the preprocessing stage and stored as 0.92. The rendering engine directly reads the pre-stored similarity value of 0.92.
[0047] Step S40, if the visual similarity is less than the preset similarity, determine the target LOD of the model to be rendered corresponding to the perceivable screen pixel threshold.
[0048] In this embodiment, the preset similarity is a threshold value set by the user or the system (e.g., 0.9 or 90%), representing the lowest acceptable visual quality similarity. When the visual similarity between models is lower than this value, it is considered that the visual quality is insufficient and a switch to a higher-detail LOD needs to be considered. Among them, the target LOD refers to the level of detail finally selected for the model to be rendered and to be used in this rendering. The higher the LOD level, such as LOD0, the richer the model details, the more vertices / faces, and the greater the rendering overhead. While the lower the level, such as LOD3, the more simplified the model and the smaller the rendering overhead.
[0049] Therefore, when the visual similarity is less than the preset similarity, it indicates that the visual quality of the current LOD model is insufficient and there is a large gap from the original model. At this time, it is necessary to determine the target LOD according to the PSP calculated in step S20. Among them, the PSP can correspond to one LOD or multiple LODs.
[0050] As an alternative implementation, when the PSP corresponds to multiple LODs, in the process of determining the target LOD, it is possible to traverse the LOD levels that are higher (more detailed) than the current LOD level at the current viewing distance, and filter out the LOD sequence whose screen space error is less than or equal to the perceivable screen threshold according to the screen space error corresponding to each level. Subsequently, find the LOD level with the coarsest detail, that is, the largest LOD level number and the smallest error but still meeting the conditions, from the selected LOD sequence as the target LOD. It can be understood that the projection of the combined error of the coarsest detail on the screen is less than the pixel threshold distinguishable by the human eye, indicating that the vision is already good enough. Selecting the coarsest level (highest LOD number) that meets this condition can save the most rendering resources and improve the rendering efficiency while ensuring the visual quality. Among them, if there is no higher-level LOD that meets the conditions, the original model (LOD0) is selected as the target LOD. The screen space error refers to the pixel size occupied by projecting the geometric error of the model at a specific LOD level (the maximum distance from the model surface point to the original model surface or a certain statistical distance) onto the imaging plane under the current viewing point and screen settings. It predicts the maximum geometric distortion degree that may be observed on the screen when rendering with this LOD level, and it can be a manually set parameter or a parameter calculated in real time based on an engine such as 3Dtiles.
[0051] Therefore, if the visual similarity is less than the preset similarity, it is necessary to determine the LOD sequence whose screen space error is less than or equal to the PSP, and then set the minimum LOD level of the LOD sequence as the target LOD. Among them, the minimum LOD level is the LOD with the largest number in the LOD sequence. For example, in LOD0-LOD6, LOD6 is the minimum LOD level.
[0052] Optionally, if the visual similarity is greater than or equal to the preset similarity, it indicates that the currently used LOD model has reached an acceptable standard in terms of visual quality and there is no need to switch to a higher-level LOD (because a higher level will not bring a significant visual improvement). At this time, there is no need to perform excessive processing on the currently to-be-rendered model, and the processing action in step S10 can be jumped to and executed to process other to-be-rendered models. It can be understood that when asynchronously executing the processing of multiple to-be-rendered models in step S10, the change of the viewing angle can be waited for and the action in step S10 can be re-executed at this time.
[0053] This embodiment provides a method for static LOD processing in an urban digital twin scenario. By combining the spatial relationship of the current view point with the inherent visual quality attribute of the model, that is, calculating the visual similarity, and introducing the human eye visual perception threshold as the key decision-making basis, the intelligence and adaptability of LOD selection are realized. When it is detected that the visual quality of the current model is insufficient, the lowest-detail LOD level that meets the visual fidelity requirements is accurately selected based on the actual visible influence of the model details in the current screen space, effectively avoiding unnecessary detail rendering in visually imperceptible areas or using a model with too low details in visually sensitive areas, resulting in image defects. Thus, on the premise of ensuring the user's visual experience, LOD is dynamically selected based on the visual similarity and the perceivable screen pixel threshold, realizing the dynamic optimal allocation of rendering computing resources and significantly improving the real-time rendering efficiency and smoothness of complex three-dimensional scenes.
[0054] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar content as that in the above first embodiment can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 2 , step S30 further includes steps S31 to S33: Step S31, determining the screen space geometric error and the screen texture error between the to-be-rendered model and the original model.
[0055] Step S32, fusing the screen space geometric error and the screen texture error based on a preset weight value to obtain a screen space visual error.
[0056] Step S33, calculating the visual similarity according to the natural exponential function between the attenuation coefficient and the screen space visual error.
[0057] In this embodiment, the three-dimensional space can be projected onto the screen space, and then the similarity between the two can be calculated through the screen-space geometric error (SGE) and the screen-space texture error (STE) between the models. In order to calculate the visual similarity through two different error parameters, thereby improving the accuracy of similarity calculation. Among them, the methods for calculating the screen-space geometric error and the screen-space texture error are existing technical methods, which will not be elaborated in this application. After obtaining two different types of errors, the errors can be subjected to mapping processing, including weighted fusion of the screen-space geometric error and the texture mapping error based on different weight ratios, so as to obtain the screen-space visual error (SVE), and then improve the accuracy of similarity calculation through different weight ratios. Finally, the calculation is performed based on a preset similarity algorithm.
[0058] Exemplarily, taking a Python code function as an example to illustrate the calculation process of visual similarity. def calculate_SIM(M_high, M_low, viewpoint): # 1. Projection from three-dimensional space to screen space SGE = project_to_screenspace(M_high, M_low, viewpoint) # Screen-space geometric error, M_high is the original model, M_low is the model to be rendered, and viewpoint is the viewing point STE = texture_error_map(M_high, M_low) # Screen-space texture error # 2. Perceptual weighted fusion (the human eye has different sensitivities to geometry / texture) SVE = α * SGE + β * STE # α = 0.6, β = 0.4 (calibrated by experiments) # 3. Visual similarity calculation SIM = exp(-k * SVE) # k is the attenuation coefficient, calibrated by PSP, and SIM is the visual similarity return SIM.
[0059] This embodiment provides a method for static LOD processing in an urban digital twin scenario. By calculating the actual visual similarity through two types of errors, namely the screen-space geometric error and the screen-space texture error, the accuracy of visual similarity calculation is improved, thereby improving the accuracy when making LOD decisions based on visual similarity subsequently.
[0060] Based on the second embodiment of the present application, in the third embodiment of the present application, the same or similar content as that in the above first embodiment can be referred to the above introduction and will not be elaborated hereinafter. On this basis, it is also necessary to calculate the screen space error in the preprocessing stage so as to select the LOD sequence based on the screen space error. Therefore, before step S10, the static LOD processing method for the urban digital twin scene further includes steps S01 to S02: Step S01, determine the LOD corresponding to each of the original models in the three-dimensional scene, and construct a quadtree or octree according to the LOD.
[0061] In this embodiment, in addition to selecting the LOD for the model to be rendered in the real-time rendering stage, it is also necessary to prepare the model parameters in the offline preprocessing stage, including generating a series of LOD models with different levels of detail based on each model (such as buildings, roads, etc.) in the original three-dimensional urban scene. After constructing the LOD models with different levels of detail, it is necessary to construct a spatial index to organize the entire scene into a quadtree (2D scene) or octree (3D scene) structure. Among them, the quadtree recursively divides the two-dimensional space into four quadrants, and the octree recursively divides the three-dimensional space into eight octants. Each node represents a spatial region, and the leaf node stores the objects located within that region. The end node in the tree data structure that is no longer further divided. In this embodiment, the leaf node associates the three-dimensional model within its spatial region and all the pre-generated LOD level data, that is, each leaf node contains the model within that region and its multiple LOD levels.
[0062] Specifically, the model of the detail level processing device traverses all the original models in the three-dimensional scene. For each original model, it loads all the pre-generated LOD level model data (such as LOD0 (original), LOD1, LOD2, LOD3). Each LOD level includes a simplified mesh, texture, and associated attributes. Subsequently, based on the spatial range (bounding box) of the entire scene, a quadtree (suitable for models mainly distributed on the ground) or an octree (suitable for scenes that require fine three-dimensional space division, such as flight or underwater scenes) is selected for construction.
[0063] After constructing a quadtree or octree, starting from the root node (representing the entire scene space), according to the set partitioning rules (such as the number of models in the node exceeding the threshold, or the node depth not reaching the upper limit), the spatial region of the current node is recursively divided into four (for quadtree) or eight (for octree) sub-regions (child nodes). Each model and all its LOD-level reference or index data are assigned to the leaf nodes that intersect with its bounding volume (such as AABB). A model may be assigned to multiple leaf nodes. Finally, the information stored in each leaf node includes: the spatial region represented by the node, and all the models located in this region and the access information of all their LOD levels.
[0064] Example: In a three-dimensional city scene, an octree needs to be constructed for the scene. The root node represents the entire city space (such as 10000m x 10000m x 1000m). After multiple levels of recursive partitioning, a leaf node in the city center area may contain all LOD levels (LOD0 fine model, LOD1 medium model, LOD2 simplified model, LOD3 low model) of the "city hall building" model. Another leaf node in the suburb contains all LOD levels of the "windmill" model.
[0065] Step S02, determine the screen space error corresponding to each LOD level at each viewing distance.
[0066] The viewing distance refers to the spatial straight-line distance from the observer (camera viewpoint) to a certain point on the surface of the three-dimensional model (usually taking the model center or the bounding box center). Each LOD level corresponds to a screen space error at different viewing distances. Therefore, in this embodiment, it is necessary to pre-calculate the reference value of the screen space error or visual similarity of each LOD level (relative to a finer LOD than the previous level) at different standard viewing distances. This reference value can be interpolated according to the actual viewing distance in the real-time stage.
[0067] When generating LOD by model simplification, the geometric error GE_lod of this LOD relative to the original model has usually been calculated. Therefore, the formula for the screen space error of a certain LOD level at the viewing distance d_i is: SSE_lod(d_i) = GE_lod * PixelPerUnit(d_i), where PixelPerUnit() is used to calculate the pixel density, and the calculation method is similar to PSP. This formula means that when observing the model at the distance d_i, the maximum geometric deviation that may be generated on the screen by the details of the model surface using this LOD level is approximately SSE_lod(d_i) pixels.
[0068] Exemplarily, for the LOD2 level of the "city hall building", its geometric error GE_lod2 = 0.5 meters is known. The set screen parameters are H_px (screen height) = 1080, FOV_v (vertical field of view) = 60° ≈ 1.0472 rad, and the sampling viewing distances: d_i = [10, 20, 50, 100, 200, 500] meters. Then calculate SSE_lod2(d_i): When d_i = 50m, PixelPerUnit(50) = 1080 / (2 * 50 * tan(1.0472 / 2)) ≈ 1080 / (100 * 0.5773) ≈ 18.70 px / m, and SSE_lod2(50) = 0.5 * 18.70 ≈ 9.35 pixels.
[0069] When d_i = 100m: PixelPerUnit(100) = 18.70 / 2 = 9.35 px / m (since the distance is doubled and the density is halved), SSE_lod2(100) = 0.5 * 9.35 ≈ 4.68 pixels.
[0070] The final stored SSE table for LOD2 is: {10: ≈ 37.4 px, 20: ≈ 18.7 px, 50: ≈ 9.35 px, 100: ≈ 4.68 px, 200: ≈ 2.34 px, 500: ≈ 0.94 px}.
[0071] It should be noted that the above parameters are only for explanation and are not a limitation to this application.
[0072] This embodiment provides a method for processing static LOD in a city digital twin scenario. By constructing a spatial index structure (quadtree / octree) and associating all LOD levels of the model to the leaf nodes, the process of "determining the model to be rendered according to the frustum" in the subsequent step S10 is significantly accelerated. Only by traversing the tree nodes intersecting with the frustum can the candidate model set be quickly obtained, thereby improving the rendering efficiency. At the same time, by pre-calculating and storing the screen space error of each LOD level at different viewing distances, when it is necessary to determine the target LOD, the estimated screen space error SSE_lod(d) of this LOD at the current distance can be directly obtained by quickly looking up the table or interpolation according to the actual distance d from the current model to the viewpoint, avoiding the real-time projection calculation of geometric error to screen space for each model in each frame, greatly reducing the calculation overhead, and improving the real-time performance of the overall rendering pipeline.
[0073] Based on the first embodiment of the present application, in the fourth embodiment of the present application, the same or similar content as that in the above first embodiment can be referred to the above introduction and will not be repeated hereinafter. On this basis, after step S40, in order to avoid jumps, boundary cases also need to be processed. If the LOD level selected for the current frame is different from that of the previous frame (especially when the viewpoint moves slowly, it may cause LOD switching), geometric morphing or alpha blending techniques are used for smooth transition to avoid visual jumps. Therefore, if the target LOD is different from the LOD selected for the previous frame, the target LOD is smoothly processed based on the geometric morphing or alpha blending algorithm to avoid visual jumps. It can be understood that the process of smooth processing is a known technique and will not be elaborated in the present application.
[0074] Immediately after obtaining the target LOD, the model of the selected LOD level also needs to be submitted to the rendering pipeline for drawing, that is, the model to be rendered is sent to the rendering pipeline so that the rendering pipeline can perform rendering processing on the model to be rendered.
[0075] Exemplarily, to help understand the implementation process of the static LOD processing method for the urban digital twin scene obtained by combining the above embodiments, please refer to Figure 3 , Figure 3 A schematic diagram of the overall work flow of a static LOD processing method for an urban digital twin scene is provided. Specifically: The viewpoint system of the device performs frustum culling through the scene manager by updating the viewpoint position or direction, and then submits the list of visible objects, that is, the models to be rendered in the frustum. Then for each visible object, first calculate the PSP threshold of the object (the model to be rendered), then calculate the visual similarity SIM through the visual similarity metric model, and then make a judgment based on the SIM value. If it is greater than or equal to the preset threshold of 0.95, the current LOD is maintained. If SIM is less than 0.95, a higher-precision LOD is requested from the LOD model library based on PSP. Then, after providing a new LOD model, smooth transition rendering processing is performed based on the rendering engine to complete the processing of the level of detail of the models in a three-dimensional scene, which can effectively improve the rendering efficiency of the models in a dynamic urban scene.
[0076] It should be noted that the above examples are only for understanding the present application and do not constitute a limitation on the static LOD processing method for the urban digital twin scene of the present application. Based on this technical concept, more forms of simple transformations are within the protection scope of the present application.
[0077] The present application provides a level-of-detail processing device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the static LOD processing method for the urban digital twin scenario in the above first embodiment.
[0078] Reference is made below to Figure 4 , which shows a schematic structural diagram of a level-of-detail processing device suitable for implementing the embodiments of the present application. The level-of-detail processing device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (PADs), portable multimedia players (PMPs), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 4 The level-of-detail processing device shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present application.
[0079] As Figure 4As shown in the figure, the level-of-detail processing device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1002 or the program loaded from the storage device 1003 into the random access memory (RAM) 1004. In the random access memory 1004, various programs and data required for the operation of the level-of-detail processing device are also stored. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the input / output interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the level-of-detail processing device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows a level-of-detail processing device having various systems, it should be understood that it is not required to implement or have all the shown systems. More or fewer systems may be implemented or had alternatively.
[0080] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program may be downloaded and installed from the network through the communication device, or installed from the storage device 1003, or installed from the read-only memory 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the methods of the embodiments disclosed in the present application are executed.
[0081] The level-of-detail processing device provided by the present application adopts the static LOD processing method for the urban digital twin scene in the above-mentioned embodiment, and can solve the technical problem of visual jump during the model rendering of the current 3D scene. Compared with the prior art, the beneficial effects of the level-of-detail processing device provided by the present application are the same as those of the static LOD processing method for the urban digital twin scene provided by the above-mentioned embodiment, and other technical features in the level-of-detail processing device are the same as those disclosed in the method of the previous embodiment, and will not be elaborated here.
[0082] It should be understood that each part disclosed in this application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.
[0083] As described above, the above are only specific embodiments of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all of them should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
[0084] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the static LOD processing method for urban digital twin scenarios in the above embodiments.
[0085] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memories, read-only memories, erasable programmable read-only memories (EPROMs), optical fibers, portable compact disk read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or combined with an instruction execution system, device, or device. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, radio frequency (RF), etc., or any suitable combination of the above.
[0086] The above computer-readable storage medium can be included in the level-of-detail processing device; it can also exist separately without being assembled into the level-of-detail processing device.
[0087] The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed by the level-of-detail processing device, the level-of-detail processing device is caused to: Obtain a frustum corresponding to the current viewpoint position and viewing angle, and determine the models to be rendered located within the frustum in the three-dimensional scene; Calculate the perceivable screen pixel threshold corresponding to the model to be rendered at the current viewpoint position; Obtain the visual similarity between the model to be rendered and the original model; If the visual similarity is less than the preset similarity, determine the target LOD of the model to be rendered corresponding to the perceivable screen pixel threshold.
[0088] Computer program code for performing the operations of the present application can be written in one or more programming languages or combinations thereof. The above programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0089] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0090] The modules involved in the embodiments of the present application can be implemented in software or in hardware. In some cases, the name of the module does not constitute a limitation on the unit itself.
[0091] The readable storage medium provided by this application is a computer-readable storage medium. The computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned static LOD processing method for urban digital twin scenarios, and can solve the technical problem of visual jump during model rendering in current 3D scenarios. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by this application are the same as those of the static LOD processing method for urban digital twin scenarios provided by the above embodiments, and will not be elaborated here.
[0092] The above are only partial embodiments of this application, and do not limit the patent scope of this application accordingly. Any equivalent structural transformation made under the technical concept of this application by using the content of the specification and drawings of this application, or any direct / indirect application in other related technical fields, is included in the patent protection scope of this application.
Claims
1. A static LOD processing method for urban digital twin scenarios, characterized in that, The static LOD processing method for the urban digital twin scenario includes: Obtain the frustum corresponding to the current viewpoint position and viewing angle, and determine the models to be rendered located within the frustum in the three-dimensional scene; Calculate the perceivable screen pixel threshold corresponding to the models to be rendered at the current viewpoint position; Obtain the visual similarity between the models to be rendered and the original models; If the visual similarity is less than the preset similarity, determine the target LOD of the models to be rendered corresponding to the perceivable screen pixel threshold.
2. The static LOD processing method for urban digital twin scenarios according to claim 1, characterized in that, The step of, if the visual similarity is less than the preset similarity, determining the target LOD of the models to be rendered corresponding to the perceivable screen pixel threshold includes: If the visual similarity is less than the preset similarity, determine the LOD sequence with a screen space error less than or equal to the perceivable screen pixel threshold; Set the minimum LOD level of the LOD sequence as the target LOD.
3. The static LOD processing method for an urban digital twin scenario according to claim 2, wherein Before the step of obtaining the frustum corresponding to the current viewpoint position and viewing angle, and determining the models to be rendered located within the frustum in the three-dimensional scene, the static LOD processing method for the urban digital twin scenario further includes: Determine the LOD corresponding to each of the original models in the three-dimensional scene, and construct a quadtree or octree based on the LOD. The leaf nodes of the quadtree or octree contain multiple LOD levels corresponding to each of the original models; Determine the screen space error corresponding to each LOD level at each viewing distance.
4. A static LOD processing method for urban digital twin scenarios according to claim 1, characterized in that The step of calculating the perceivable screen pixel threshold corresponding to the models to be rendered at the current viewpoint position includes: Obtain the screen physical size and pixel density of the acquisition device, and the viewing distance between the acquisition device and the models to be rendered; Calculate the product of the screen physical size and the pixel density, and the quotient between the product and the viewing distance; Set the product of the quotient, the minimum resolvable angle of the human eye, and the proportional coefficient between the angle and the radian as the perceivable screen pixel threshold.
5. A static LOD processing method for urban digital twin scenarios according to claim 1, characterized in that The step of obtaining the visual similarity between the models to be rendered and the original models includes: Determine the screen space geometric error and screen texture error between the models to be rendered and the original models; Fuse the screen space geometric error and the screen texture error based on a preset weight value to obtain a screen space visual error; Calculate the visual similarity according to the natural exponential function between the attenuation coefficient and the screen space visual error.
6. The static LOD processing method for urban digital twin scenarios according to claim 1, characterized in that, After the step of, if the visual similarity is less than the preset similarity, determining the target LOD of the models to be rendered corresponding to the perceivable screen pixel threshold, the static LOD processing method for the urban digital twin scenario further includes: If the target LOD is different from the LOD selected in the previous frame, perform a smooth process on the target LOD based on the geometric deformation or transparency blending algorithm to avoid visual jumps; Send the models to be rendered to the rendering pipeline so that the rendering pipeline performs rendering processing on the models to be rendered.
7. A static LOD processing method for urban digital twin scenarios according to claim 1, characterized in that The step of obtaining the frustum corresponding to the current viewpoint position and viewing angle, and determining the models to be rendered located within the frustum in the three-dimensional scene includes: Calculate six plane equations according to the viewpoint position and the viewing angle, and form the viewing frustum based on the six plane equations; Determine the intersection detection result between the viewing frustum and the three-dimensional scene corresponding to the original model, and set the model that intersects with the plane of the viewing frustum and is inside the viewing frustum as the model to be rendered.
8. The static LOD processing method for an urban digital twin scenario according to claim 1, wherein, After the step of obtaining the visual similarity between the model to be rendered and the original model, the static LOD processing method for the urban digital twin scene further includes: If the visual similarity is greater than or equal to the preset similarity, jump to execute the step of obtaining the viewing frustum corresponding to the current viewpoint position and viewing angle, and determining the model to be rendered located inside the viewing frustum in the three-dimensional scene.
Citation Information
Patent Citations
Three-dimensional rendering method and device, equipment and medium
CN117237502A
Urban digital twinning scene LOD processing method
CN119228975A
Three-dimensional model LOD continuous display method and device and electronic equipment
CN119579761A