Environment visualization simulation method based on wave collapse mechanism

By using an environmental visualization simulation method based on wave collapse mechanism, the problem of uneven resource allocation in 3D environment generation is solved, achieving efficient and stable visualization of complex scenes. It is suitable for high-performance environments such as interactive visualization and virtual simulation.

CN122023675APending Publication Date: 2026-05-12NANJING YUTIAN ZHIYUN SIMULATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANJING YUTIAN ZHIYUN SIMULATION TECH CO LTD
Filing Date
2026-04-13
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing methods for generating 3D environments struggle to achieve efficient and stable visual quality under limited computing power, especially in complex scenes where resource allocation is uneven, leading to an unstable visual experience.

Method used

An environmental visualization simulation method based on wave collapse mechanism is adopted. By obtaining the generation modal probability distribution of three-dimensional visual scene units, the visual information density and saliency factor are calculated. By using the line-of-sight correlation sensitivity coefficient and the shaping sorting function, progressive state convergence and semantic consistency propagation are achieved.

Benefits of technology

While ensuring visual quality, it reduces computational overhead and improves the real-time generation efficiency and stability of complex scenes, making it particularly suitable for high-performance environments such as interactive visualization and virtual simulation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122023675A_ABST
    Figure CN122023675A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of environment visualization, and discloses an environment visualization simulation method based on a wave collapse mechanism, which comprises the following steps: firstly, acquiring a three-dimensional visual scene unit in a current vision field range, and initializing the three-dimensional visual scene unit to generate modal probability distribution; then, visual information density features, visual saliency factors and generation uncertainty indexes are calculated, and a line-of-sight correlation sensitivity coefficient is obtained based on logarithm space regression to determine a visual convergence hysteresis threshold and a frame-level state establishment quota. And carrying out dynamic sorting on the importance of the scene units through a shaping sorting function, executing progressive state convergence and semantic association propagation on key units, and finally outputting an environment visualization result conforming to the target visual quality. According to the method, the rendering and generation calculation burden can be remarkably reduced, the real-time generation efficiency and visual stability of a complex scene are improved, and the method is suitable for the fields of virtual simulation, game rendering, digital twinning, intelligent visualization systems and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of environmental visualization technology, and more specifically, to an environmental visualization simulation method based on wave collapse mechanism. Background Technology

[0002] With the rapid development of technologies such as virtual simulation, digital twins, and intelligent visualization, the demand for real-time generation and dynamic presentation of large-scale 3D environments is constantly growing. Especially in application scenarios such as urban simulation, interactive virtual spaces, and intelligent driving simulation, systems need to continuously generate structurally complex, semantically rich, and visually stable 3D visual environments under limited computing power. To address this, an increasing number of rendering and generation frameworks are incorporating mechanisms such as probabilistic models, sparse representations, and semantic inference to improve the visualization quality of dynamic environments. However, faced with high-resolution viewpoints, large scene scales, and real-time interaction requirements, traditional visualization methods still struggle to achieve a balance between efficiency and visual quality.

[0003] Existing 3D environment generation methods largely rely on fixed rules or static priorities for rendering resource allocation, making it difficult to reflect real-time changes in the importance of each 3D scene unit within the viewport. Furthermore, traditional LOD (Level of Detail) based or local heuristic techniques are prone to slow visual convergence, inconsistent local semantics, and uneven resource allocation when handling highly complex scenes. When scenes exhibit significant differences in texture density, uneven semantic distribution, or frequent viewpoint movements, existing methods cannot effectively quantify generation uncertainty or dynamically adjust the rendering order based on visual saliency, leading to decreased overall generation efficiency and unstable visual experience. Summary of the Invention

[0004] This invention provides an environmental visualization simulation method based on wave collapse mechanism, which solves the technical problems mentioned in the background art.

[0005] This invention provides an environmental visualization simulation method based on wave collapse mechanism, including: Obtain the three-dimensional visual scene units within the current field of view, and initialize the generation modal probability distribution of each of the three-dimensional visual scene units; Calculate the visual information density features and visual saliency factor of the three-dimensional visual scene unit. Based on the mapping relationship between the visual saliency factor and the generation uncertainty index derived from the generation modality probability distribution, solve the line-of-sight sensitivity coefficient and the visual convergence hysteresis threshold. Based on the visual convergence hysteresis threshold, a frame-level state establishment quota for the current simulation frame is set. An integer sorting function is constructed using the gaze-related sensitivity coefficient to prioritize high visual saliency factors. In accordance with the order determined by the integer sorting function, progressive state convergence calculation and semantic association propagation are performed on the three-dimensional visual scene units within the frame-level state establishment quota to output a visualized environment scene.

[0006] The beneficial effects of this invention are as follows: by introducing multimodal indicators such as visual information density, visual saliency, and generation uncertainty, and combining them with the line-of-sight sensitivity coefficient and the shaping and sorting mechanism, dynamic, controllable, and progressive convergence and semantic consistency propagation of 3D visual scene units are achieved. This invention can reduce unnecessary computational overhead while ensuring visual quality, allowing rendering resources to be concentrated on key areas of the view domain, significantly improving the real-time generation efficiency and stability of complex scenes. It is particularly suitable for high-performance environments such as interactive visualization, virtual simulation, and digital twins. Attached Figure Description

[0007] Figure 1 This is a flowchart of an environmental visualization simulation method based on wave collapse mechanism according to the present invention; Figure 2 This is a schematic diagram of a three-dimensional visual scene unit within the current field of view of the present invention. Detailed Implementation

[0008] The subject matter described herein will now be discussed with reference to exemplary embodiments. It should be understood that these embodiments are discussed only to enable those skilled in the art to better understand and implement the subject matter described herein, and changes may be made to the function and arrangement of the elements discussed without departing from the scope of this specification. Various processes or components may be omitted, substituted, or added as needed in the examples. Furthermore, features described in some examples may be combined in other examples.

[0009] like Figure 1 As shown, an environmental visualization simulation method based on wave collapse mechanism includes: Obtain the three-dimensional visual scene units within the current field of view, and initialize the generation modal probability distribution of each of the three-dimensional visual scene units; Calculate the visual information density features and visual saliency factor of the three-dimensional visual scene unit. Based on the mapping relationship between the visual saliency factor and the generation uncertainty index derived from the generation modality probability distribution, solve the line-of-sight sensitivity coefficient and the visual convergence hysteresis threshold. Based on the visual convergence hysteresis threshold, a frame-level state establishment quota for the current simulation frame is set. An integer sorting function is constructed using the gaze-related sensitivity coefficient to prioritize high visual saliency factors. In accordance with the order determined by the integer sorting function, progressive state convergence calculation and semantic association propagation are performed on the three-dimensional visual scene units within the frame-level state establishment quota to output a visualized environment scene.

[0010] Preferably, the step of acquiring the three-dimensional visual scene units within the current field of view includes: Traverse the hierarchical nodes based on the 3D tile data structure and calculate the screen space error of each hierarchical node using the following formula: ; in, Indicates time Tile node Screen space error, This represents the camera constants related to viewport height and field of view. Represents tile nodes Geometric error parameters, Represents the virtual camera to tile node The Euclidean distance of the enclosing volume; Determine whether the screen space error is greater than the preset maximum allowable screen space error. If it is, refine the child nodes of the tile node; otherwise, determine the tile node and its associated semantic objects as the three-dimensional visual scene unit.

[0011] The current field of view is the visible space within the current simulation frame, defined by the virtual camera's viewing direction, the near clipping plane, the far clipping plane, and the display window.

[0012] The hierarchical node τ is a discrete node identifier organized based on a three-dimensional tile data structure.

[0013] The virtual viewpoint position refers to the position and viewing attitude of the virtual camera in the 3D scene within the current simulation frame. It can be obtained directly through the camera state interface output by the rendering engine, the view matrix resolution result, or the viewpoint cache of the interactive controller.

[0014] Screen space error is a metric used to measure the pixel-level geometric error of a tile node τ after it is projected onto the screen at the current time t. The larger the metric, the less geometric precision the current level of nodes has, and the more refinement is needed.

[0015] Time t is the time identifier or frame number identifier corresponding to the current simulation frame.

[0016] The camera constant is a projection scaling factor that converts geometric errors in world space into screen pixel errors.

[0017] The viewport height is the number of pixels in the vertical direction of the current rendering window or the current rendering target.

[0018] The field of view is the observation angle parameter of the virtual camera under the current projection model.

[0019] The geometric error parameter is the maximum allowable geometric deviation of a tile node τ relative to a higher-precision geometric representation at the current level. It can be obtained from the geometric error field in the 3D tile metadata, offline simplified logs, or error files recorded during the model preprocessing stage.

[0020] Euclidean distance is the spatial distance between the virtual camera and the reference bounding volume of the tile node τ at the current time t. It is used to characterize the attenuation effect of observation distance on screen error. The greater the distance, the smaller the impact of the same geometric error on the screen.

[0021] The preset maximum allowable error threshold is the upper limit of screen space error that determines whether hierarchical nodes should be further refined. A value of 12 pixels is preferred. Too small a threshold will cause distant nodes to be over-refined, increasing the loading burden, while too large a threshold will cause distortion of foreground contours and loss of detail.

[0022] The node to be rendered is a hierarchical node whose screen space error does not exceed the preset maximum allowable error threshold and is retained in the current frame to participate in rendering and semantic calculation.

[0023] Index mapping is the association structure between the node to be rendered and the underlying semantic object.

[0024] The underlying semantic object is the basic object unit with clear semantic labels in the scene data, such as buildings, roads, green spaces, facilities, or water fragments.

[0025] A semantic object set is a collection of all underlying semantic objects that are associated with a given node through an index mapping.

[0026] The 3D visual scene unit is the basic operating unit that actually performs probability initialization, saliency calculation, sorting and scheduling, and state convergence within the current field of view.

[0027] In detail, the camera constant is calculated as follows: Read the vertical pixel count and vertical field of view of the current viewport. First, take half of the vertical pixel count, then divide it by the tangent of half the vertical field of view to obtain the scaling factor that maps the world length to the screen pixel length for the current viewport. For example, when the viewport height is 1080 and the vertical field of view is 60 degrees, κcam can be approximately 935.3. When the system has multiple viewports, κcam should be calculated and used separately for each viewport.

[0028] In detail, the geometric error parameters are determined as follows: First, the geometric error metadata field corresponding to the hierarchical node in the 3D tile data is read. If the original data does not provide this field, then during the offline preprocessing stage, the maximum vertex deviation or maximum surface distance of the node's current geometric representation relative to its high-precision source model is used as the geometric error parameter, and the result is written back to the node metadata. For example, the corresponding geometric error parameter for a simplified building mesh can be calculated using the maximum surface distance between the original model and the simplified model.

[0029] In detail, the Euclidean distance is calculated as follows: The Euclidean distance from the virtual camera's optical center to the nearest point in the bounding box of the tile node is uniformly used as the calculation benchmark, instead of using both the bounding box and the bounding box center as references. When the camera is outside the bounding box, the distance from the optical center to the nearest point in the bounding box is calculated directly. When the camera enters the bounding box, the distance calculation benchmark is preferentially switched to the nearest spatial distance from the optical center to the actual geometric surface contained within the node or the bounding box of deeper child nodes. Only when the aforementioned internal geometric information is missing is the lower limit of the distance fixed at 0.1 meters or an equivalent small positive value in the scene's basic units to avoid the denominator approaching zero. For example, when the camera is close to the surface of a tall building, the shortest distance to the nearest face of the bounding box should be taken, rather than the distance to the center of the bounding box.

[0030] In detail, the preset maximum allowable error threshold is set by jointly determining the threshold based on the terminal resolution, target frame rate, and scene complexity. For desktop real-time simulation scenes, 12 pixels is preferred, with a preferred range of 8 to 16 pixels. For mobile devices or very large scenes, this can be relaxed to 16 to 24 pixels. If the system detects that the real-time frame rate is more than 10 times lower than the target frame rate, the threshold can be temporarily increased by 2 to 4 pixels to reduce the refinement pressure.

[0031] In detail, the termination method for recursive refinement of hierarchical nodes is as follows: refinement terminates when the screen space error of a node is no greater than the preset maximum allowable error threshold; refinement terminates when the node is already a leaf node; and refinement terminates when the preset maximum hierarchical depth is reached. When child node data has not yet been fully loaded, the parent node is retained as a temporary node to be rendered. For example, if the error of a foreground road node exceeds the threshold, but its next-level child nodes have not yet been fully loaded from disk or network, the parent node should be used to maintain the current frame output until the child nodes are available.

[0032] In detail, the index mapping between the nodes to be rendered and the underlying semantic objects is established as follows: During the offline preprocessing stage, a forward mapping is established using the node's unique identifier as the key and a list of semantic object identifiers as the value; simultaneously, a reverse mapping is established using the semantic object identifier as the key and a list of associated node identifiers as the value. Mapping determination prioritizes a joint assessment of three rules: spatial intersection area ratio, bounding volume containment relationship, and semantic label consistency. For example, when a node covers both the road surface and the curb, both object identifiers should be written into the node's mapping list.

[0033] In detail, the formation of a 3D visual scene unit is as follows: When a node to be rendered is mapped to multiple underlying semantic objects, a composite 3D visual scene unit is formed with the node as the geometric boundary and an ordered list of object identifiers as the semantic content. When a underlying semantic object spans multiple nodes to be rendered, the object is split into multiple local instances of the nodes according to the node boundaries, and a probability state is maintained for each local instance. For example, a road that spans multiple tile nodes should form independent local road units within each node so that subsequent saliency calculations and quota scheduling can be performed at the node granularity.

[0034] Preferably, the initialization of the generation modal probability distribution of each of the three-dimensional visual scene units includes: Based on the prior weights in the preset wave collapse mode library, the three-dimensional visual scene unit is calculated using the following formula. Select the Initial probabilities of semantic patterns: ; in, Indicates time The three-dimensional visual scene unit In the first The probability components of a semantic pattern Indicates the first The frequency weight of each semantic pattern in the pattern library This indicates the total number of candidate semantic patterns in the pattern library. Indicates the first Frequency weights of semantic patterns; will be determined by The resulting probability vector is determined as the generated mode probability distribution.

[0035] The wave collapse pattern library is a pre-built collection of candidate semantic patterns, along with a unified data repository containing the weights, constraints, and instance resources corresponding to each pattern. Ideally, the library should contain 6 to 12 basic urban scene semantic patterns. Too few patterns reduce the scene's expressive power, while too many patterns significantly increase the computational burden of probability updates and compatibility propagation.

[0036] Candidate semantic patterns are discrete semantic state options that a 3D visual scene unit can choose during the initialization phase. Each pattern corresponds to a candidate generation result that can be assigned a probability, such as semantic types like roads, buildings, vegetation, squares, or water bodies.

[0037] Sample statistical frequency refers to the frequency of occurrence of each candidate semantic pattern obtained from historical samples, labeled data, or statistical corpora. It can be obtained by counting patterns in a city 3D sample library, labeled scene database, or historical simulation results.

[0038] Prior frequency weights are the prior importance assigned to each candidate semantic pattern by the designer based on domain experience, business objectives, or rule knowledge when sufficient sample statistical frequencies are lacking. Ideally, these weights should be positive numbers between 0.01 and 1.00. This is because prior frequency weights need to reflect pattern preferences without allowing a few patterns to prematurely monopolize all probabilities during the initialization phase.

[0039] The initial probability value is the probability that each candidate semantic pattern is assigned to a 3D visual scene unit at the initialization time.

[0040] The generated modal probability distribution is a complete probabilistic representation of all candidate semantic modes for a given 3D visual scene unit.

[0041] A probability vector is a vector structure composed of multiple probability components arranged in a fixed pattern.

[0042] The 3D visual scene unit index i is a discrete number used to identify a specific 3D visual scene unit within the current view.

[0043] The semantic pattern index s is a discrete number used to identify a specific candidate semantic pattern in the pattern library.

[0044] The probability component is the probability value of a 3D visual scene unit i being in the s-th semantic mode at time t.

[0045] The frequency weight is the weight value corresponding to the s-th semantic pattern in the initialization formula.

[0046] The total number of patterns, M, is the total number of all candidate semantic patterns in the pattern library.

[0047] The summation pattern index r is an auxiliary number used to traverse all candidate semantic patterns and complete the weight accumulation.

[0048] The frequency weight is the weight value corresponding to the r-th semantic pattern when normalized and summed.

[0049] In detail, the wave collapse pattern library is constructed as follows: First, high-frequency semantic structures are extracted from the target scene samples. Then, each type of semantic structure is organized into discrete pattern entries, and each pattern entry is written with a pattern identifier, semantic name, sample statistical frequency, prior frequency weight, output geometric resource identifier, and adjacency constraint. For example, for urban road scenes, basic entries such as road patterns, sidewalk patterns, green belt patterns, building frontage patterns, and parking area patterns can be constructed.

[0050] In detail, the method for choosing or fusing the weights of sample statistical frequency and prior frequency is as follows: when the number of valid samples for a certain pattern in the sample library reaches the preset minimum sample size, the sample statistical frequency is used as the primary weight source. When the number of valid samples is insufficient, the final initial weight is the linear fusion result of sample statistical frequency accounting for 0.7 and prior frequency accounting for 0.3. For example, when there are insufficient samples of newly added characteristic landscape patterns, artificial priors can be appropriately retained to avoid them being completely suppressed in the initialization stage.

[0051] In detail, the additive normalization process is as follows: First, all weights are non-negative, and outlier weights less than zero are reset to zero. Then, all valid weights are summed. If the sum is greater than zero, the initial probability value of each mode is calculated by dividing its respective weight by the total weight. If the sum is equal to zero, all candidate semantic modes are set to an equal probability distribution. For example, when a new scene unit has no reliable prior knowledge, all modes can be initialized to a uniform distribution.

[0052] In detail, the method of using the same or differentiated initial probability distribution for different 3D visual scene units is as follows: first, the units are grouped according to the region type to which the 3D visual scene unit belongs, the underlying semantic object label, and the upper-layer tile category, and then a corresponding initial probability template is configured for each group. For example, the road node group prioritizes increasing the probability of patterns related to roads, sidewalks, and traffic facilities, while the waterfront node group prioritizes increasing the probability of patterns related to water bodies, embankments, and vegetation.

[0053] In detail, the storage and normalization verification method for the generated modal probability distribution is as follows: For each 3D visual scene unit, maintain a floating-point vector of length M equal to the total number of modes, and fix the vector indices to correspond one-to-one with the mode library order. After each initialization or update, recalculate the sum of all probability components. If the deviation of the sum from 1 exceeds 0.0001, normalization is performed again. For example, when accumulated errors occur after compatibility propagation, this verification step can be used to restore a valid probability distribution.

[0054] Preferably, calculating the visual information density features and visual saliency factors of the three-dimensional visual scene unit includes: The three-dimensional visual scene unit is calculated using the following formula. The visual information density features mentioned above: ; in, Indicates time The aforementioned visual information density features, This represents the total number of intervals in the feature histogram. Indicates the first The probability that each feature interval lies within the projection region; The visual saliency factor is calculated using the following formula: ; ; in, This represents the visual saliency factor. Indicates the weight of the center cone of the field of view. Representation unit The coordinates of the center of the screen projection. Represents the coordinates of the screen's geometric center. Indicates the width parameter of the center cone. Represents the entropy sensitivity coefficient. This represents the set of cells within the current field of view.

[0055] The screen projection area is the two-dimensional pixel area covered by the 3D visual scene unit after it is projected onto the current screen. This area defines the sampling range of rendering feature data and the calculation range of visual information density.

[0056] Rendering feature data consists of pixel-level feature information related to visual complexity, collected from the screen projection area, such as color changes, depth changes, normal changes, or material changes. It can be obtained through a deferred rendering buffer, a screen-space sampling pass, or a separate feature extraction rendering pass.

[0057] A feature distribution histogram is a distribution structure formed by statistically analyzing the rendered feature data within a screen projection area according to preset intervals.

[0058] The visual information density feature is a complexity index calculated using the Shannon entropy of the feature distribution histogram. The larger the index, the richer, more complex, and more worthy of priority attention the rendered features within the projection region are.

[0059] Shannon entropy is a statistical entropy result calculated based on the probability of each feature interval. It is used to measure the degree of dispersion and uncertainty of feature distribution within the screen projection area.

[0060] The total number of feature intervals is the number of discrete intervals used when constructing the feature distribution histogram. A value of 32 is preferred. Too few intervals will result in the loss of detailed variations, while too many intervals will introduce noise and reduce statistical stability.

[0061] The feature interval index is the number of a statistical interval in the feature distribution histogram.

[0062] The feature interval probability is the proportion of samples of 3D visual scene unit i in the b-th feature interval at time t.

[0063] The projection center coordinates are the two-dimensional center position coordinates of the three-dimensional visual scene unit i in the current screen space.

[0064] The screen geometric center coordinates are the two-dimensional geometric center coordinates of the current display window or the current rendering viewport.

[0065] The field-of-view cone weight is a center preference weight derived from the distance attenuation relationship between the unit's projection center and the screen's geometric center. The larger this weight, the closer the unit is to the center of vision, and the more priority should be given to ensuring visual quality there.

[0066] The center cone width parameter is a scale parameter that controls the rate at which the weight of the center cone of the viewport decays. It is preferably 0.25 times the number of pixels on the shorter side of the screen. This is because if σ is too small, the center preference will be too sharp, while if σ is too large, it will weaken the ability of the viewport center to distinguish priorities.

[0067] The entropy sensitivity coefficient η is a weighting coefficient that controls the strength of the influence of visual information density features on the visual saliency factor. A value of 1.0 is preferred. This is because a small η will weaken the priority of complex regions, while a large η will cause locally highly complex regions to excessively occupy processing resources.

[0068] The weighted result is the original significance value before normalization obtained by multiplying the visual information density feature by the weight of the visual field center cone after exponential amplification.

[0069] The visual saliency factor is the relative saliency weight obtained by normalizing each 3D visual scene unit within the current field of view. The visual saliency factor reflects the importance of a unit's contribution to visual quality under the current observation conditions.

[0070] The set of units within the current field of view is the set of all 3D visual scene units that enter the significance calculation, regression fitting, ranking, and quota statistics at the current time t.

[0071] In detail, the method for collecting rendering feature data is as follows: In the rendering pipeline, three basic sampling channels—color, depth, and normal—are enabled for the 3D visible scene units within the current viewport. Valid pixel samples are extracted from the screen projection area of ​​each unit. When the projection area is large, a 2x2 step downsampling is used. When the projection area is small, full pixel sampling is used. For example, downsampling can be used for foreground building facades to reduce statistical overhead, while small traffic signs should maintain full pixel sampling to avoid information loss.

[0072] In detail, the feature types and interval division methods of the feature distribution histogram are as follows: at least one of the three types of features—brightness gradient magnitude, depth gradient magnitude, and normal variation—is selected as the statistical object. Before the statistics are performed, a local spatial difference operator is introduced to weight the structured features of adjacent pixels, and then normalize them to the zero-to-one interval for equal-width binning, thereby distinguishing regular structures from disordered noise.

[0073] In detail, the probability of a feature interval is calculated as follows: The number of valid pixels falling into the b-th feature interval within the screen projection area of ​​the 3D visual scene unit i is counted, and then divided by the total number of valid pixels to obtain the basic probability of that interval. Before participating in subsequent logarithmic calculations, a preset small positive number is added to the basic probabilities of all intervals for smoothing and re-normalization to obtain the final probability of that interval. Invalid pixels include pixels outside the projection area, cropped pixels, and pixels with undefined depth. For example, if a region has 200 valid pixels, and 40 of them fall into the 5th feature interval, the corresponding feature interval probability is 0.2.

[0074] In detail, the calculation method for the projection center coordinates is as follows: The two-dimensional centroid coordinates of the effective projected pixel set are preferentially used as the projection center coordinates. When the effective projected pixels are too few, the center of the projection bounding box is used as the backoff value. For example, when the projection boundary of an irregular tree crown is significantly concave or convex, the centroid of the effective pixels should be used as the center, rather than simply taking the center of the bounding box.

[0075] In detail, the setting method for the center cone width parameter σ is as follows: First, read the number of pixels on the shorter side of the current viewport, then set σ to 0.25 times the number of pixels on the shorter side. When stronger center focus is needed, it can be lowered to 0.20 times the number of pixels on the shorter side; when a smoother center attenuation is needed, it can be increased to 0.35 times the number of pixels on the shorter side. For example, in a 1920x1080 resolution screen, σ can be preferentially set to 270.

[0076] In detail, the entropy sensitivity coefficient η is set as follows: First, compare the consistency between the visual saliency ranking and the manually focused areas under different η values ​​on the development set, and then select the η with the best consistency as the running parameter. For general scenarios, a value of 1.0 is preferred; for scenarios with significant texture differences, a value of 1.2 to 1.8 can be used; and for scenarios with weak texture differences, a value of 0.5 to 0.8 can be used. For example, in complex commercial street scenarios, η can be appropriately increased to strengthen the priority of highly complex areas.

[0077] In detail, the numerical stabilization method for the visual saliency factor is as follows: First, the original weighted result of each unit is calculated. Then, it is locally normalized by dividing it by the maximum weighted result of all units in the current field of view, so that the numerical distribution remains within a fixed range, thereby avoiding the magnitude collapse phenomenon caused by a large increase in the number of scene units. If the maximum value is extremely small, the visual saliency factor of all units in the current field of view is set to equal probability.

[0078] In detail, the handling method when the unit set in the current field of view is empty or the number of samples is too small is as follows: If the unit set in the current field of view is empty, skip the visual saliency factor normalization and set the quota of this frame to 0. When the number of units in the unit set in the current field of view is less than 3, sort directly by the original weighted result without performing global normalization comparison.

[0079] Preferably, the generation uncertainty index derived from the generation mode probability distribution includes: Based on the Shannon entropy definition, the three-dimensional visual scene unit is calculated using the following formula. The generation uncertainty index: ; in, Indicates time The aforementioned generation uncertainty index, This indicates the total number of candidate semantic patterns. In the generation mode probability distribution, the first... The probability components of a semantic pattern.

[0080] The generation uncertainty index Hi(t) is an entropy-type uncertainty measure calculated from the generation mode probability distribution. A larger index indicates that the unit is more difficult to determine among multiple candidate semantic modes. A smaller index indicates that the unit's generation state is closer to convergence.

[0081] In detail, the natural logarithm term is handled when the probability component is zero as follows: first, the lower limit of all probability components is truncated to 0.000001, and then the natural logarithm is calculated.

[0082] In detail, the normalization verification method for generating modal probability distributions is as follows: Before calculating the generation uncertainty index, sum all probability components of the current element. If the deviation of the sum from 1 exceeds 0.0001, then divide each probability component by the sum. If the sum is less than 0.000001, then revert to the previous valid probability distribution or uniform distribution.

[0083] In detail, the unified range processing method for comparing generated uncertainty indices across scenarios is as follows: When comparing scenarios with a total number M of different patterns, the generated uncertainty index is divided by the maximum entropy benchmark corresponding to the total number M of patterns, so that the result falls within the range of 0 to 1. If the sorting is only within the same pattern library, the original generated uncertainty index can be used directly. For example, no additional normalization is needed within the same city pattern library, but a unified range is recommended when comparing across city or task pattern libraries.

[0084] Preferably, the step of calculating the gaze-related sensitivity coefficient and the visual convergence hysteresis threshold based on the mapping relationship between the visual saliency factor and the uncertainty index includes: For the three-dimensional visual scene unit within the field of view, let the intermediate variable... , The line-of-sight correlation sensitivity coefficient is calculated using the least squares method with the following formula: ; in, This represents the gaze-related sensitivity coefficient. This indicates the generation uncertainty index. This represents the visual saliency factor. and To prevent small positive numbers with undefined logarithms, and These are the means of the corresponding intermediate variables. This represents the set of units within the current field of view; the visual convergence hysteresis threshold is calculated using the following formula: ; in, This represents the visual convergence hysteresis threshold.

[0085] The intermediate variable is the logarithmic space variable obtained by adding a small positive number εH to the generation uncertainty index and then taking the natural logarithm.

[0086] The intermediate variable is the logarithmic space variable obtained by adding a small positive number εA to the visual saliency factor and then taking the natural logarithm.

[0087] The gaze-related sensitivity coefficient α(t) is obtained from the least squares regression slope between intermediate variables xi and yi. It represents the degree of coupling sensitivity between the generation uncertainty index and the visual saliency factor, and directly affects the subsequent visual convergence hysteresis threshold and the shaping ranking function.

[0088] The small positive number εH is a numerical protection term added to avoid undefined uncertainty indices when taking the natural logarithm. It is preferably 0.000001. This is because this value must be small enough to avoid altering the original uncertainty distribution, while also being large enough to avoid logarithmic divergence caused by zero or minimum values.

[0089] The small positive number εA is a protective measure added to prevent the visual saliency factor from being undefined or too small in the denominator of the natural logarithm and ranking formulas. It is preferably 0.000001. This value is necessary to ensure stability in logarithmic and division operations without significantly distorting the original saliency ranking relationship.

[0090] The mean of the intermediate variable is the average value of all intermediate variable xi samples within the current field of view.

[0091] The mean of the intermediate variable is the average value of all intermediate variable yi samples within the current field of view.

[0092] The visual convergence hysteresis threshold is a frame-level proportional threshold obtained by further transforming the gaze-related sensitivity coefficient. The visual convergence hysteresis threshold is used to determine how many proportions of the three-dimensional visual scene units in the current frame need to perform state convergence, thereby mapping gaze sensitivity to discrete processing quotas.

[0093] In detail, the values ​​of the tiny positive numbers εH and εA are determined as follows: both are uniformly set to 0.000001. When the system uses single-precision floating-point and the minimum probability of the scenario fluctuates significantly, this can be relaxed to 0.00001. When the system uses double-precision floating-point and the probability distribution is stable, it can be reduced to 0.00000001. For example, in a single-precision implementation on a mobile device, to avoid repeated logarithmic underflow, 0.00001 can be used preferentially.

[0094] In detail, the selection method for effective regression samples is as follows: only 3D visual scene units with a projected area of ​​no less than sixteen effective pixels within the current field of view, both the visual saliency factor and the generation uncertainty index being valid positive values, and the generation uncertainty index being greater than the preset lower bound threshold for high-level non-convergence judgment, are retained as regression samples. By increasing this lower bound threshold, units that are close to convergence are eliminated in advance to prevent them from generating extreme shifts in the logarithmic space that would affect the fitting slope. When the number of effective samples is less than five, the valid gaze-related sensitivity coefficient of the previous frame is directly used or the default value of zero is used.

[0095] In detail, the least squares method handles outliers as follows: when the variance of the intermediate variable xi is less than a preset threshold, causing the denominator to approach zero, the gaze-related sensitivity coefficient is directly set to the valid value of the previous frame. When outliers exist, samples that are more than three times the absolute deviation of the median are removed before regression. For example, if an extremely bright billboard at a certain instant causes an abnormal amplification of the visual saliency factor, the outlier sample can be removed before fitting.

[0096] In detail, the inter-frame smoothing method for the gaze-related sensitivity coefficient is as follows: After obtaining a new regression result in each frame, the final gaze-related sensitivity coefficient is generated using an exponential smoothing method with the current frame result accounting for 0.3 and the previous frame result accounting for 0.7. If there is no valid result in the previous frame, the result of the current frame is used directly.

[0097] In detail, the boundary limitation method for the visual convergence hysteresis threshold is as follows: First, calculate the original visual convergence hysteresis threshold, and then truncate the result to the interval between 0 and 1. When the truncated value is 0 and there are still units in the current field of view, a fallback can be performed at the minimum processing ratio of 0.05. For example, when extreme parameters cause the original calculated result to be 1.2, it should be truncated to 1. When the original result is negative, it should be truncated to 0, and whether to enable the minimum processing ratio should be determined according to the system policy.

[0098] In detail, the protection method when the line-of-sight sensitivity coefficient plus one approaches zero is as follows: a smooth transition function is introduced. When the absolute value of the coefficient plus one is in the critical region of a set small interval, the arctangent function is used to continuously and non-linearly map the mutation result, smoothly transitioning it to the set safety lower bound, thus avoiding the phenomenon of quota numerical jump.

[0099] Preferably, the step of setting the frame-level state of the current simulation frame based on the visual convergence hysteresis threshold to establish the quota includes: The quota is established by calculating the frame-level state using the following formula: ; in, This indicates that the frame-level state establishes the quota. This represents the visual convergence hysteresis threshold. This indicates the total number of the three-dimensional visual scene units within the current field of view. This represents the floor function operator.

[0100] The frame-level state establishment quota is the upper limit on the number of 3D visual scene units allowed to perform state convergence calculations in the current simulation frame. It links the visual convergence hysteresis threshold with the number of units in the current viewport, enabling the allocation of processing resources per frame.

[0101] In detail, the total number of 3D visual scene units within the current viewport is counted as follows: After completing hierarchical node filtering and semantic mapping, duplicates are first deduplicated using the unique identifier of each 3D visual scene unit, and then the number of deduplicated units is counted as the number of units in the current viewport. When a low-level semantic object is split into multiple node local instances, it should be counted separately for each local instance. For example, if a cross-node road is split into three local units, the number of units in the current viewport should be counted as 3, not 1.

[0102] In detail, the boundary handling method for frame-level state establishment quota is as follows: When there are no 3D visible scene units in the current viewport, the frame-level state establishment quota is set to 0. When there are units in the current viewport and the rounding result is less than 1, the frame-level state establishment quota is set to at least 1. When the rounding result is greater than the number of units in the current viewport, the frame-level state establishment quota is set to equal the number of units in the current viewport.

[0103] Preferably, the step of constructing a shaping ranking function that prioritizes high visual saliency factors using the gaze-related sensitivity coefficient includes: The integer sorting function value of each of the three-dimensional visual scene units is calculated using the following formula: ; in, Indicates time The integer sorting function value, This indicates the generation uncertainty index. This represents the visual saliency factor. This represents the gaze-related sensitivity coefficient. This indicates a non-zero correction term to prevent the denominator from being zero; in accordance with The numerical values ​​are sorted in ascending order for all the three-dimensional visual scene units to obtain a determined processing sequence.

[0104] The sorting function value is a sorting criterion obtained by combining the uncertainty index, visual saliency factor, and line-of-sight sensitivity coefficient. As a relative dimensional index, this value is only used for ascending order comparison of units within the same simulation frame; the smaller the value, the higher the priority in the current frame. No absolute value comparisons of this function value are performed between different simulation frames.

[0105] The processing sequence is an ordered list formed by sorting all 3D visual scene units within the current viewport in ascending order according to the integer sorting function value.

[0106] In detail, the numerical stabilization method for the integer ranking function value is as follows: The calculation is performed using the visual saliency factor, εA, and line-of-sight sensitivity coefficient, all of which have undergone boundary protection. When the visual saliency factor is too small, causing the denominator to approach zero, the result of adding a small positive number εA to the visual saliency factor is used directly. When the line-of-sight sensitivity coefficient plus one approaches zero, safety protection is implemented first, and then the integer ranking function value is calculated. For example, when the visual saliency factor in the edge region is extremely small but the generation uncertainty index is large, adding εA can prevent the ranking value from being infinitely amplified.

[0107] In detail, the stable sorting method when sorting values ​​are tied is as follows: When the difference in the integer sorting function values ​​of two or more 3D visual scene units is less than 0.000001, the visual saliency factor is compared first, and the unit with the larger visual saliency factor is ranked first. If the visual saliency factors are still the same, the generation uncertainty index is compared, and the unit with the larger generation uncertainty index is ranked first. If they are still the same, they are sorted in ascending order according to the unit's unique identifier.

[0108] In detail, the connection between the processing sequence and the frame-level state-established quota is as follows: First, a complete processing sequence is generated for all 3D visual scene units in the current field of view. Then, the units with the previous frame-level state-established quota are extracted from the head of the processing sequence as the execution set for this frame. The remaining units are retained for re-evaluation in the next frame.

[0109] In detail, the safety protection method for the power factor is as follows: when the gaze-related sensitivity coefficient plus one is less than or equal to 0, the result is not directly used as the denominator. Instead, the gaze-related sensitivity coefficient is increased to a value of one before the reciprocal is calculated.

[0110] Preferably, the step of performing progressive state convergence calculation and semantic association propagation on the 3D visual scene units within the frame-level state establishment quota according to the order determined by the integer sorting function, and outputting a visualized environment scene, includes: For sorted first The three-dimensional visual scene unit of the position The branchless asymptotic state convergence calculation is performed using the following formula: ; in, Indicates the first The probability components of the pattern Indicates the sequence number Decreasing annealing temperature parameters Indicates the total number of patterns; The semantic association is calculated and the message is propagated and the adjacent units are updated using the following formula. probability distribution: ; ; in, Indicates the dissemination of a message. This represents the safety factor used to prevent the probability from reaching zero. Indicating semantic relations The semantic compatibility matrix coefficients are as follows; The pattern with the highest probability in each of the three-dimensional visual scene units is mapped to a three-dimensional geometric entity for rendering, and a visualized environment scene is output.

[0111] Processing progress refers to the degree of convergence that the current simulation frame has reached within the frame-level state-established quota.

[0112] The temperature parameter is an annealing-type parameter used to control the sharpening intensity of the probability distribution. The mapping relationship is adjusted according to the processing progress, so that the earlier the sorted units correspond to the higher the temperature parameter to maintain their smoothness, and the later the sorted units correspond to the lower temperature parameter to speed up convergence. The overall value range is between 0.35 and 1.00.

[0113] The sorting number n is the position number of a certain 3D visual scene unit in the current frame processing sequence.

[0114] The compatibility propagation message is semantic compatibility constraint information sent by a source unit to its neighboring units along the topological relationships of the city's semantic space. The calculation method involves calculating the local propagation message vector along a single topological relationship edge. When a target unit has multiple neighboring source units, the local propagation message vectors calculated according to this formula are multiplied one by one according to the candidate semantic pattern components. Then, the total number of neighboring source units of the target unit is used to take the square root of the multiplication result, thereby eliminating the degree bias of network topology nodes and obtaining the final comprehensive propagation message.

[0115] Neighboring units are the identifiers of target units that are directly connected to the current source unit in the topological relationship of the urban semantic space and need to be affected by semantic propagation.

[0116] The safety factor is a lower bound protection quantity added to prevent the probability terms in the compatibility propagation message from reaching zero during propagation. In addition operations, this safety factor is equivalent to multiplying a one-dimensional vector of the same length as the number of candidate semantic patterns by its corresponding scalar, ensuring that all components are accumulated element-wise along the same dimension. A value of one-thousandth is preferred.

[0117] The topological relationships in urban semantic space are a network of relationships such as adjacency, connection, inclusion, street proximity, or adjacency between semantic objects in a scene.

[0118] A semantic compatibility matrix is ​​a set of matrix parameters that describes the degree of compatibility between different semantic patterns under a specific semantic relation. Preferably, it is an M-by-M matrix constructed according to relation type, with matrix elements preferably being non-negative numbers between 0.05 and 1.00. This is because the compatibility matrix must reflect the differences in preferences between patterns while avoiding large areas of zero values ​​that could lead to propagation degradation.

[0119] The semantic compatibility matrix coefficients are the compatibility coefficient terms corresponding to the source semantic pattern under the condition of semantic relations.

[0120] The semantic relation 'r' is a relation number used to distinguish different types of topological relations, such as adjacent, connected, contained, adjacent to, or covering.

[0121] The probability distribution vector is the probability state vector held by neighboring unit j for all candidate semantic modes. The probability distribution vector is updated and renormalized into a new generative modality probability distribution after receiving the compatibility propagation message.

[0122] The semantic pattern with the highest probability is the candidate pattern with the highest probability value for a certain 3D visual scene unit at the current moment.

[0123] Three-dimensional geometric entities are geometric models, procedural geometry, or composite geometry resources that correspond to the semantic patterns with the highest probability and can be directly instantiated into the scene.

[0124] In detail, the temperature parameter decays with processing progress as follows: First, the sorting sequence number of the current unit in the execution set of this frame is converted to a processing progress between zero and one. Then, the temperature parameter gradually decreases from the higher digit to the lower digit as the processing sequence proceeds. When faster convergence in the later stages is required, an exponential decay with a slower initial decay followed by a faster decay is used. For example, when the current frame quota is ten, the first unit uses a temperature parameter close to 1.00, and the tenth unit uses a temperature parameter close to 0.35.

[0125] In detail, the renormalization method after power-law sharpening is as follows: First, sharpen each probability component of the 3D visual scene unit i by dividing 1 by a power of the temperature parameter, then sum all the sharpened probability components. If the sum is greater than 0, then each probability component is divided by the sum. If the sum is less than 0.000001, then revert to the probability distribution before sharpening or directly retain the one-bit effective distribution of the highest probability mode. For example, when all probabilities are extremely small, a lower limit protection should be applied before summing.

[0126] In detail, the construction method of urban semantic spatial topology is as follows: In the offline stage, spatial adjacency analysis is performed on all underlying semantic objects, and relationship edges are established according to five categories of rules: contact, connection, containment, intersection, and nearest neighbor based on distance thresholds. These relationship edges are then mapped to the 3D visual scene unit layer. For example, roads and sidewalks can be established as adjacencies through boundary contact relationships, and buildings and roads can be established as nearest neighbors based on street-facing distance thresholds.

[0127] In detail, the semantic compatibility matrix is ​​constructed as follows: An M-by-M compatibility matrix is ​​created for each semantic relation. Each element in the matrix represents the compatibility strength between the source and target patterns under that relation. Matrix elements are preferably non-negative numbers between 0.05 and 1.00. For example, under the adjacency relation, the compatibility value between roads and sidewalks can be set higher, while the compatibility value between roads and rooftops should be set lower.

[0128] In detail, the semantic compatibility matrix coefficients are a complete row of compatibility coefficient vectors corresponding to the source semantic patterns s under the semantic relation r. Each component in this vector corresponds to the compatibility level of the target unit when selecting each candidate pattern. For example, when the source pattern is a road and the relation is adjacent, this vector can give a higher compatibility value for sidewalks, curbs, and green belts, and a lower compatibility value for water bodies and roofs.

[0129] In detail, the aggregation method for the compatibility propagation message mi→j is as follows: when an adjacent unit j simultaneously receives propagation messages from multiple source units, the message vectors generated by each source unit are first multiplied by their components to obtain a composite message, and then the composite message is normalized or truncated to a lower bound. When enhanced robustness is required, it can also be changed to accumulation in the logarithmic field followed by inverse transformation. For example, when a plaza unit is adjacent to both a road and a building, the probability distribution can be corrected by combining the messages from both sides.

[0130] In detail, the probability distribution vector is updated as follows: the original probability distribution vector of the adjacent unit j is multiplied by the integrated compatibility propagation message component by component to obtain the updated temporary vector. Then, the temporary vector is summed and normalized to restore the sum of all probability components to 1. If the summation result is too small, a safety factor δ is added to each component before normalization. For example, if all components of a target unit are suppressed to extremely low levels after propagation, δ should be added first before restoring a valid distribution.

[0131] In detail, the method for mapping the semantic pattern with the highest probability to a 3D geometric entity is as follows: First, a mapping table from pattern identifiers to geometric resource identifiers is established. Then, the corresponding 3D geometric entity is instantiated based on the spatial location, orientation, and scale parameters of the current unit. For example, when the semantic pattern with the highest probability is "road," road patches, lane lines, and auxiliary components can be retrieved from the resource library and trimmed and arranged according to the unit boundary.

[0132] like Figure 2 As shown, Figure 2 This demonstrates the spatial geometry principles and screen spatial error (SSE) calculation logic for acquiring 3D visual scene units within the current viewport. The left side of the image presents a hierarchical data structure of 3D tiles organized using an octree. Nodes covered by the view frustum and determined to require loading are bolded and marked as 3D visual scene units within the current viewport, representing the working set. The filtering results are shown in the middle of the image. The imaging model of the virtual camera is displayed, and the vertical field of view is marked. And from the camera optical center to the tile node Euclidean distance of the center of the enclosing volume The display diagram on the right side of the image maps the geometric relationships of three-dimensional space to two-dimensional pixel space. The shaded area intuitively represents the screen space error generated by that tile node from the current viewpoint. The mathematical formula below the image. This corresponds to its spatial diagram. Specifically, screen error is directly proportional to the inherent geometric error of the tile itself, and inversely proportional to the viewing distance; the system is based on this... The comparison result between the value and the preset threshold determines whether to refine the hierarchical nodes on the left side of the graph or directly output them as visual units, thereby realizing scene streaming loading with adaptive view.

[0133] The embodiments of this example have been described above. However, this example is not limited to the specific implementation methods described above. The specific implementation methods described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms based on the guidance of this example, and all of them are within the protection scope of this example.

Claims

1. An environmental visualization simulation method based on wave collapse mechanism, characterized in that, include: Obtain the three-dimensional visual scene units within the current field of view, and initialize the generation modal probability distribution of each of the three-dimensional visual scene units; Calculate the visual information density features and visual saliency factor of the three-dimensional visual scene unit. Based on the mapping relationship between the visual saliency factor and the generation uncertainty index derived from the generation modality probability distribution, solve the line-of-sight sensitivity coefficient and the visual convergence hysteresis threshold. Based on the visual convergence hysteresis threshold, a frame-level state establishment quota for the current simulation frame is set. An integer sorting function is constructed using the gaze-related sensitivity coefficient to prioritize high visual saliency factors. In accordance with the order determined by the integer sorting function, progressive state convergence calculation and semantic association propagation are performed on the three-dimensional visual scene units within the frame-level state establishment quota to output a visualized environment scene.

2. The environmental visualization simulation method based on wave collapse mechanism according to claim 1, characterized in that, Retrieve the 3D visual scene units within the current view area, including: Traverse the hierarchical nodes based on the 3D tile data structure, and calculate the screen space error of each hierarchical node according to the virtual viewpoint position; recursively refine the hierarchical nodes whose screen space error exceeds the preset maximum allowable error threshold, and take the hierarchical nodes whose screen space error does not exceed the preset maximum allowable error threshold as the nodes to be rendered at the current moment; establish an index mapping between the nodes to be rendered and the underlying semantic objects, and take the set of semantic objects obtained by the mapping as the 3D visual scene unit.

3. The environmental visualization simulation method based on wave collapse mechanism according to claim 1, characterized in that, Initialize the generation modal probability distribution of each of the three-dimensional visual scene units, including: Call the preset wave collapse mode library to obtain the sample statistical frequency or prior frequency weight of each candidate semantic mode in the wave collapse mode library; perform additive normalization on the prior frequency weight of all candidate semantic modes to calculate the initial probability value of each candidate semantic mode; use the vector set formed by the initial probability values ​​as the generation modal probability distribution of each of the three-dimensional visual scene units at the initial time.

4. The environmental visualization simulation method based on wave collapse mechanism according to claim 1, characterized in that, Calculating the visual information density features and visual saliency factors of the three-dimensional visual scene unit includes: Rendering feature data is collected within the screen projection area of ​​the 3D visual scene unit, a feature distribution histogram is constructed, and the Shannon entropy of the feature distribution histogram is calculated. The calculated Shannon entropy value is used as the visual information density feature. The Euclidean distance between the projection center coordinates of the 3D visual scene unit in screen space and the geometric center coordinates of the screen is calculated. Based on the Euclidean distance, the central cone weight of the field of view with a central decay distribution is calculated. The visual information density feature and the central cone weight of the field of view are weighted and combined using an exponential function, and the weighted result of all units within the current field of view is normalized. The normalized value is used as the visual saliency factor.

5. The environmental visualization simulation method based on wave collapse mechanism according to claim 1, characterized in that, The generation uncertainty index derived from the generation mode probability distribution includes: The probability components of each candidate semantic mode in the generated modality probability distribution are traversed; the product of each probability component and its corresponding natural logarithm is calculated, and all product results are summed; the negative value of the summation result is used as the generation uncertainty index of the three-dimensional visual scene unit.

6. The environmental visualization simulation method based on wave collapse mechanism according to claim 4, characterized in that, Based on the mapping relationship between the visual saliency factor and the uncertainty index, the line-of-sight sensitivity coefficient and the visual convergence hysteresis threshold are calculated, including: Calculate the natural logarithm of the generation uncertainty index and the visual saliency factor for each of the three-dimensional visual scene units, and use the least squares method to perform regression fitting on the logarithmic linear relationship between the generation uncertainty index and the natural logarithm of the visual saliency factor. Use the value two as the base and the opposite of the reciprocal of the sum of the visual saliency factor and the value one as the exponent. Perform a power operation and use the result of the power operation as the visual convergence hysteresis threshold.

7. The environmental visualization simulation method based on wave collapse mechanism according to claim 1, characterized in that, Based on the visual convergence hysteresis threshold, the frame-level state of the current simulation frame is set to establish the quota, including: The total number of the three-dimensional visual scene units within the current field of view is counted; the product of the visual convergence hysteresis threshold and the total number is calculated; the product is rounded up, and the integer value is used as the frame-level state establishment quota for the current simulation frame to complete state convergence.

8. The environmental visualization simulation method based on wave collapse mechanism according to claim 1, characterized in that, Constructing an integer ranking function that prioritizes high visual saliency factors using the aforementioned gaze-related sensitivity coefficients includes: Calculate the reciprocal of the sum of the gaze-related sensitivity coefficient and the numerical value one, and use this reciprocal as the exponent; add the visual saliency factor of the three-dimensional visual scene unit to a preset non-zero correction term as the base, and perform an exponentiation operation; calculate the ratio of the generation uncertainty index of the three-dimensional visual scene unit to the result of the exponentiation operation, and use this ratio as the integer sorting function value; arrange all three-dimensional visual scene units within the current field of view in ascending order of the integer sorting function values, and generate a processing sequence that prioritizes processing high visual saliency factors.

9. The environmental visualization simulation method based on wave collapse mechanism according to claim 1, characterized in that, According to the order determined by the integer sorting function, progressive state convergence calculation and semantic association propagation are performed on the three-dimensional visual scene units within the frame-level state establishment quota, and a visualized environment scene is output, including: For each of the three-dimensional visual scene units within the frame-level state-established quota, the corresponding generated modal probability distribution is subjected to power-law sharpening and renormalization using a temperature parameter that decays with the processing progress, achieving progressive state convergence. Based on the topological relationship of the urban semantic space, a predefined semantic compatibility matrix is ​​queried, a compatibility propagation message containing a safety factor is calculated, and the generated modal probability distribution of adjacent units is updated using this compatibility propagation message, wherein the safety factor is used to ensure that the probability value is always greater than zero. The semantic mode with the highest probability among each of the three-dimensional visual scene units is selected for geometric instantiation rendering to generate a visualized environment scene.