Dynamic rendering method and system for a game context based on a virtual game
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- AHA ENTERTAIN (SHANGHAI) CO LTD
- Filing Date
- 2026-05-15
- Publication Date
- 2026-08-07
AI Technical Summary
[0003]虚拟游戏运行时的状态数据通常涵盖场景拓扑结构、角色交互时序及环境物理响应等多源异构模态,现有方法往往仅对这类数据进行简单的线性叠加或基于阈值的硬性触发,未进行跨模态语义对齐与有效聚合,使得生成的背景渲染数据存在语义冲突;同时,现有渲染方法在处理背景数据时,通常将场景的固有物理属性与动态的氛围扰动耦合在一起,缺乏有效的正交分解机制与背景物理约束关系把控,这种耦合导致物理结构与氛围特征相互干扰,使得生成的背景环境基础特征图与氛围扰动特征图的精准性大幅降低,导致了游戏背景画面的精准性较低
[0016] (1) During the operation of the virtual game, multi-source heterogeneous state data of the virtual game is acquired, cross-modal semantic alignment is performed on the multi-source heterogeneous state data, and the background rendering data of the virtual game at the current moment is constructed by combining the aggregation mechanism; the background rendering data is input into the pre-trained background style decoupling network, the background style decoupling network performs orthogonal decomposition on the background rendering data, and generates mutually independent background environment basic feature maps and atmosphere perturbation feature maps by combining background constraint relationships during the decomposition process. The background rendering data is introduced to further control the background constraint relationships and improve the accuracy of background environment basic feature maps and atmosphere perturbation feature maps.
Smart Images

Figure CN122530403A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of virtual games, and in particular to a dynamic rendering method and system for game backgrounds based on virtual games. Background Technology
[0002] With the continuous development of virtual game technology and the increasing demand of players for immersion, dynamic rendering technology of game backgrounds has become a key research direction in the field of game graphics. In complex virtual game environments, the background not only needs to present a static geometric spatial structure, but also needs to undergo real-time and dynamic visual evolution with changes in game logic, character status and environmental physical conditions.
[0003] The state data during virtual game runtime typically encompasses multi-source heterogeneous modalities such as scene topology, character interaction timing, and environmental physical responses. Existing methods often only perform simple linear superposition or hard triggering based on thresholds on this type of data, without cross-modal semantic alignment and effective aggregation, resulting in semantic conflicts in the generated background rendering data. At the same time, when processing background data, existing rendering methods usually couple the inherent physical properties of the scene with dynamic atmospheric perturbations, lacking an effective orthogonal decomposition mechanism and control over background physical constraints. This coupling causes mutual interference between physical structure and atmospheric features, significantly reducing the accuracy of the generated background environment basic feature map and atmospheric perturbation feature map, resulting in low accuracy of the game background image. Summary of the Invention
[0004] The purpose of this invention is to overcome the shortcomings of the prior art. This invention provides a dynamic rendering method and system for game backgrounds based on virtual games.
[0005] This invention provides a dynamic rendering method for game backgrounds based on virtual games, including:
[0006] During the operation of the virtual game, multi-source heterogeneous state data of the virtual game runtime is acquired, cross-modal semantic alignment is performed on the multi-source heterogeneous state data, and the background rendering data of the virtual game at the current moment is constructed by combining the aggregation mechanism.
[0007] The background rendering data is input into a pre-trained background style decoupling network. The background style decoupling network performs orthogonal decomposition on the background rendering data and generates mutually independent background environment basic feature maps and atmosphere perturbation feature maps by combining background constraint relationships during the decomposition process.
[0008] Obtain a preset dynamic style primitive library and combine it with the atmosphere perturbation feature map to determine the corresponding target style content. Then, perform multi-scale feature fusion between the target style content and the background environment basic feature map to generate the corresponding background potential content.
[0009] The pose parameters of the current virtual camera are collected, the latent background content is input into the preset neural network, and the corresponding latent space ray stepping operation is triggered in combination with the pose parameters of the current virtual camera. In this way, multi-frequency sparse sampling is performed during the stepping process, and finally the game background screen with parallax effect and dynamic depth of field is output.
[0010] This invention provides a dynamic rendering system for game backgrounds based on virtual games, which is applied to the aforementioned dynamic rendering method for game backgrounds based on virtual games; the dynamic rendering system for game backgrounds based on virtual games includes:
[0011] The virtual game module is used to acquire multi-source heterogeneous state data during the operation of the virtual game, perform cross-modal semantic alignment on the multi-source heterogeneous state data, and construct the background rendering data of the virtual game at the current moment by combining the aggregation mechanism.
[0012] The background rendering module is used to input background rendering data into a pre-trained background style decoupling network. The background style decoupling network performs orthogonal decomposition on the background rendering data and generates mutually independent background environment basic feature maps and atmosphere perturbation feature maps by combining background constraint relationships during the decomposition process.
[0013] The background latent content module is used to obtain a preset dynamic style primitive library, and combine it with the atmosphere perturbation feature map to determine the corresponding target style content. The target style content is then fused with the background environment basic feature map at multiple scales to generate the corresponding background latent content.
[0014] The game background image module is used to collect the pose parameters of the current virtual camera, input the potential background content into the preset neural network, and combine the pose parameters of the current virtual camera to trigger the corresponding latent space ray stepping operation, thereby performing multi-frequency sparse sampling during the stepping process, and finally outputting a game background image with parallax effect and dynamic depth of field.
[0015] Compared with the prior art, the beneficial effects of the present invention are:
[0016] (1) During the operation of the virtual game, multi-source heterogeneous state data of the virtual game is acquired, cross-modal semantic alignment is performed on the multi-source heterogeneous state data, and the background rendering data of the virtual game at the current moment is constructed by combining the aggregation mechanism; the background rendering data is input into the pre-trained background style decoupling network, the background style decoupling network performs orthogonal decomposition on the background rendering data, and generates mutually independent background environment basic feature maps and atmosphere perturbation feature maps by combining background constraint relationships during the decomposition process. The background rendering data is introduced to further control the background constraint relationships and improve the accuracy of background environment basic feature maps and atmosphere perturbation feature maps.
[0017] (2) Obtain the preset dynamic style primitive library and determine the corresponding target style content by combining the atmosphere perturbation feature map. Perform multi-scale feature fusion between the target style content and the background environment basic feature map to generate the corresponding background potential content. Collect the pose parameters of the current virtual camera, input the background potential content into the preset neural network, and trigger the corresponding latent space ray stepping operation by combining the pose parameters of the current virtual camera. Perform multi-frequency sparse sampling during the stepping process, and finally output the game background screen with parallax effect and dynamic depth of field. Further control the background potential content, fully consider the pose parameters of the current virtual camera and the background potential content, and improve the accuracy of the game background screen. Attached Figure Description
[0018] Figure 1 This is a flowchart illustrating the dynamic rendering method for game backgrounds based on virtual games in an embodiment of the present invention.
[0019] Figure 2 This is a flowchart illustrating step S11 in the dynamic rendering method for game backgrounds based on virtual games in an embodiment of the present invention.
[0020] Figure 3 This is a flowchart illustrating step S12 in the dynamic rendering method for game backgrounds based on virtual games in an embodiment of the present invention.
[0021] Figure 4 This is a flowchart illustrating step S13 in the dynamic rendering method for game backgrounds based on virtual games in an embodiment of the present invention.
[0022] Figure 5 This is a flowchart illustrating step S14 in the dynamic rendering method for game backgrounds based on virtual games in an embodiment of the present invention.
[0023] Figure 6 This is a schematic diagram of the structural composition of a dynamic rendering system for game backgrounds based on virtual games in an embodiment of the present invention. Detailed Implementation
[0024] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.
[0025] Please see Figures 1 to 6 A dynamic rendering method for game backgrounds based on virtual games, applied to virtual game scenes; the dynamic rendering method for game backgrounds based on virtual games includes:
[0026] Step S11: During the operation of the virtual game, acquire multi-source heterogeneous state data of the virtual game, perform cross-modal semantic alignment on the multi-source heterogeneous state data, and construct the background rendering data of the virtual game at the current moment by combining the aggregation mechanism;
[0027] Step S12: Input the background rendering data into the pre-trained background style decoupling network. The background style decoupling network performs orthogonal decomposition on the background rendering data and generates mutually independent background environment basic feature maps and atmosphere perturbation feature maps by combining background constraint relationships during the decomposition process.
[0028] Step S13: Obtain the preset dynamic style primitive library, and determine the corresponding target style content by combining the atmosphere perturbation feature map. Perform multi-scale feature fusion between the target style content and the background environment basic feature map to generate the corresponding background potential content.
[0029] Step S14: Collect the pose parameters of the current virtual camera, input the background latent content into the preset neural network, and combine the pose parameters of the current virtual camera to trigger the corresponding latent space ray stepping operation, thereby performing multi-frequency sparse sampling during the stepping process, and finally outputting a game background screen with parallax effect and dynamic depth of field.
[0030] refer to Figure 2 In step S11, the specific steps are as follows:
[0031] S111: Real-time monitoring of the virtual game's operation process. By deploying data probes at the bottom layer of the game engine, multi-source heterogeneous state data of the virtual game's operation is obtained in real time. This multi-source heterogeneous state data covers scene topology, character interaction timing, and environmental physical response content.
[0032] S112: Input multi-source heterogeneous state data into the data processing network and perform cross-modal semantic alignment based on the spatiotemporal attention mechanism. After alignment, introduce a dynamic aggregation mechanism to adaptively aggregate the semantically aligned heterogeneous features along the spatial and temporal dimensions to construct background rendering data representing the virtual game at the current moment.
[0033] In the embodiments of this application, the operation process of the virtual game is monitored in real time. Multi-source heterogeneous state data of the virtual game is obtained in real time through data probes deployed at the bottom layer of the game engine. The multi-source heterogeneous state data covers scene topology, character interaction timing and environmental physical response content.
[0034] At this time, during the operation of the virtual game, non-intrusive data probes are deployed in the underlying logic execution flow and memory management area of the game engine. These probes are attached to the main game thread in the form of engine subsystem hooks and perform asynchronous data sampling with a preset frame rate or event triggering mechanism. This allows for real-time interception and acquisition of the original multi-source heterogeneous state data stream during runtime without interfering with the core logic operation of the game. The data probes include memory reflection probes, event callback probes, and physical frame interception probes, which correspond to the mirror reading of static structure data, the subscription and capture of logical event streams, and the extraction of intra-frame snapshots of physical simulation results, respectively.
[0035] The intercepted multi-source heterogeneous state data is structured and analyzed according to its data modality and physical meaning, and divided into three orthogonal data dimensions: scene topology data, character interaction time series data, and environmental physical response content. Among them, the scene topology data represents the static spatial geometric features and connectivity of the virtual game scene, and is analyzed as a set of scene graph nodes and a spatial voxel-occupied grid. The character interaction time series data represents the dynamic causal logic and state evolution of entities in the game, and is analyzed as a discrete event sequence with timestamps and a continuous attribute change vector. The environmental physical response content represents the dynamic physical field changes of the virtual environment, and is analyzed as a grid-based fluid velocity field, soft body deformation gradient, and collision impulse distribution matrix.
[0036] To address the differences in sampling frequency and modal heterogeneity among the three dimensions of parsed data, a timestamp alignment mechanism based on game engine logical frames is adopted. Using the engine's logical frame number as the global time reference, the nearest neighbor interpolation strategy is used to complete the frames of low-frequency sampled scene topology data, while temporal sliding window mean pooling is used to reduce the frequency and compress the high-frequency sampled environmental physics response content, aligning the three to the same logical time axis. Tensor encapsulation is performed according to the structural characteristics of each dimension of data. The scene topology is encapsulated as a three-dimensional sparse tensor, the character interaction time sequence is encapsulated as a two-dimensional temporal feature matrix, and the environmental physics response is encapsulated as a two-dimensional multi-channel field tensor. Finally, a standardized set of multi-source heterogeneous state data tensors is output for subsequent cross-modal semantic alignment network calls.
[0037] Specifically, during the operation of this sandbox survival game, memory reflection probes and event callback probes are injected into the underlying physical and logical subsystems of the engine, respectively. The memory reflection probes asynchronously read the terrain height map and voxel occlusion data in the engine memory at a frequency of 30 times per second. The event callback probes subscribe to the "character being hit" and "sanity value change" event streams inside the engine. When the game character encounters an unspeakable monster in the ruins, the probes can intercept the character's hit sequence and sanity value decay event in real time without increasing the burden on the game's main thread, providing the underlying data source for the subsequent alienation rendering of the Cthulhu-style background.
[0038] The data intercepted by the probes was structured and analyzed: the dungeon's passageway direction, wall thickness, and occlusion relationships were analyzed into voxel-occupied grids in the scene topology data; the duration the player held the torch, the operation of activating the Forbidden Codex, and the numerical fluctuations of sanity decrease were analyzed into continuous attribute change vectors with timestamps in the character interaction time sequence data; the direction of torch swaying caused by wind in the dungeon, the ripple diffusion radius of the water surface stepped on by the character, and the feedback force of colliding with walls were analyzed into the fluid velocity field and collision impulse distribution matrix in the environmental physics response content.
[0039] To address the frequency differences in heterogeneous data within the Cthulhu dungeon scene of the survival game, this paper aligns the following: "Scene Topology: Dungeon Spatial Structure" rarely changes during gameplay and is considered low-frequency data. Nearest neighbor interpolation is used to extend it to each logical frame. "Character Interaction: Continuously Decreasing Sanity" is logical frame synchronous data, maintaining the original frame rate. "Environmental Physics Response: High-Frequency Flickering of Torch Flames and Changes in Airflow" is high-frequency physics frame data. Sliding window mean pooling is used to down-compress it to align with the logical frames. The dungeon voxel structure is encapsulated as a 3D sparse tensor, sanity and operation sequences are encapsulated as a 2D temporal feature matrix, and wind and water ripple physics fields are encapsulated as 2D multi-channel field tensors. After standardization and encapsulation, these data are ready to be input into the subsequent spatiotemporal attention network to drive the dynamic generation of fog and tentacles in the Cthulhu background.
[0040] Furthermore, multi-source heterogeneous state data is input into a data processing network and cross-modal semantic alignment is performed based on a spatiotemporal attention mechanism. After alignment, a dynamic aggregation mechanism is introduced to adaptively aggregate the semantically aligned heterogeneous features along the spatial and temporal dimensions to construct background rendering data representing the virtual game at the current moment. This approach takes into account the overall consideration of the semantically aligned heterogeneous features along the spatial and temporal dimensions, ensuring the accuracy of the background rendering data representing the virtual game at the current moment.
[0041] At this point, the multi-source heterogeneous state data tensor set, standardized and encapsulated by S111, is input into the feature embedding layer of the data processing network. The scene topology sparse tensor, the role interaction temporal matrix, and the environment physical field tensor are mapped to a unified-dimensional implicit feature space to obtain the corresponding topology label sequence, interaction label sequence, and physical label sequence. A dual-stream cross-attention module based on a spatiotemporal attention mechanism is constructed. Using the time axis as the semantic anchor, the interaction label sequence is used as the dominant query feature and cross-attention operations are performed with the topology label sequence and the physical label sequence, respectively. This forces the interaction features to retrieve and associate spatial structure features and physical response features with causal logic in the time dimension. At the same time, a spatiotemporal masking mechanism is introduced to shield invalid zero-value regions, thereby eliminating modal barriers between multi-source data and outputting a cross-modal aligned feature set that is mutually corroborated at the semantic level and strictly aligned in the spatiotemporal dimension.
[0042] After obtaining the cross-modal alignment feature set, a dynamic aggregation mechanism is introduced to perform adaptive aggregation along the spatial dimension. Based on the spatial gradient distribution and local occupancy density in the scene topology features, a spatial adaptive weight matrix is calculated. This weight matrix exhibits a high response value in the "spatial gradient abrupt change region, i.e., the junction of the topological structure edge and occlusion" and a low response value in the flat open region. The spatial adaptive weight matrix is then multiplied and weighted element-wise with the alignment features that incorporate the physical response, so that the physical field perturbation and interaction alienation features are non-uniformly distributed and aggregated according to the complexity of the spatial structure. This compresses and aggregates the discrete multi-point spatial features into a local background spatial aggregation feature map that highly matches the current scene spatial morphology.
[0043] Based on spatial aggregation, adaptive aggregation along the time dimension is further performed; the historical background rendering latent state of the current logical frame is extracted, and combined with the local background spatial aggregation features of the current frame, the time-adaptive forget gate and update gate are calculated through a gated loop mechanism; when the character interaction timing features indicate that a drastic state change has occurred (such as instantaneous high damage), the update gate opening is increased, forcibly increasing the weight ratio of the current frame's spatial aggregation features to achieve rapid response switching of the background atmosphere; when the interaction timing features indicate that it is in a stable exploration period, the forget gate opening is increased, making the historical latent state dominate the weight to achieve a smooth temporal transition of the background atmosphere. The features aggregated in the time dimension are decoded and reconstructed by a multilayer perceptron, outputting background rendering data that represents the virtual game at the current moment, possessing both transient responsiveness and temporal coherence.
[0044] Specifically, in this sandbox survival game, topological markers representing dungeon structure, interactive markers representing a sudden drop in player sanity, and physical markers representing a chaotic wind field caused by torches are input into the data processing network. Through cross-attention operations, the interactive markers representing a "sudden drop in sanity" are used as query features. The network precisely retrieves and focuses on the "dark topological area: topological marker" in the current dungeon corner and the "turbulent area: physical marker" where the wind is blowing towards the player. This makes these three originally independent heterogeneous data semantically related, namely, "when the player suffers a mental shock, he notices the anomaly and unnatural cold wind in the dark corner." This completes cross-modal semantic alignment and outputs a set of aligned features containing Cthulhu-style horror-related semantics.
[0045] In the Cthulhu-style dungeon scene, high-response spatial adaptive weights are calculated based on the corners of the dungeon walls and the edges of the door frames. These high-response weights are then applied to the aligned "turbulent cold wind" and "sudden drop in sanity" features, causing the atmospheric features representing indescribable terror to preferentially concentrate in the dark corners of the dungeon and the gaps in the door frames, while the feature responses in the center of the empty corridor are adaptively suppressed. As a result, the originally uniformly diffused alienation features are aggregated into a local background spatial aggregation feature map that spreads close to the walls and dark corners, which is consistent with the visual logic of "terror breeds in shadowy corners" in the Cthulhu background.
[0046] When a player suddenly looks directly at a statue of the ancient god Cthulhu in a sandbox survival game, the interaction timing feature detects a "sudden and precipitous drop in sanity." The time-adaptive update door opening instantly widens, causing "local background spatial aggregation features" representing intense fear in the current frame, such as the rapid expansion of dark corner shadows and the distortion of wall textures, to directly overwrite the historical stable state with extremely high weight, achieving a dramatic and sudden rendering of the game background plunging into darkness. When the player continues to slowly explore the dark corridor, the sanity fluctuates slowly, the time-adaptive forgetting door takes precedence, and the spread of blood streaks on the walls and the gathering of fog in the background evolve smoothly and gradually according to the historical state, avoiding temporal flickering in the background rendering. The final output is smooth background rendering data for the current moment on the timeline that matches the rhythm of horror.
[0047] refer to Figure 3 In step S12, the specific steps are as follows:
[0048] S121: Label the pre-trained background style decoupling network and load the background rendering data. At this time, determine the corresponding dual-stream orthogonal decomposition module based on the detection of the background style decoupling network. The dual-stream orthogonal decomposition module performs feature mapping and orthogonal constraint decomposition on the background rendering data.
[0049] S122: During the decomposition process, the background rendering data is further combined with the background constraints of the virtual game, and the contrastive learning loss factor is combined to generate mutually independent background environment basic feature maps and atmosphere perturbation feature maps. The background environment basic feature map represents the inherent physical properties of the scene; the atmosphere perturbation feature map represents dynamic emotions and environmental perturbations.
[0050] In the embodiments of this application, a pre-trained background style decoupling network is labeled, and background rendering data is loaded. At this time, the corresponding dual-stream orthogonal decomposition module is determined based on the detection of the background style decoupling network. The dual-stream orthogonal decomposition module performs feature mapping and orthogonal constraint decomposition on the background rendering data, which is compatible with the overall consideration of the background style decoupling network and ensures the accuracy of the corresponding dual-stream orthogonal decomposition module.
[0051] At this time, in the rendering pipeline of the virtual game, in response to the background rendering data generation instruction at the current moment, the pre-trained background style decoupling network is marked and called. At the same time, the background rendering data output from the previous step, which represents the global state of the game at the current moment, is loaded into the input tensor buffer of the network. The background style decoupling network has completed parameter solidification through a two-branch contrastive learning strategy during the pre-training stage. It contains a feature encoder and a two-stream decoder. The feature encoder is used to perform dimensionality reduction and high-dimensional semantic extraction on the loaded background rendering data to generate shared initial background latent features, providing a unified feature input source for subsequent two-stream orthogonal decomposition.
[0052] The feature encoder of the background style decoupling network performs dimensionality detection and semantic distribution parsing on the initial background latent features to determine the dynamic configuration parameters of the dual-stream orthogonal decomposition module adapted to the current data. The dual-stream orthogonal decomposition module includes a basic feature branch and an atmosphere feature branch. The network adaptively adjusts the channel attention weights and receptive field size of the dual-stream branches according to the ratio of high-frequency spatial gradient energy to low-frequency color variance in the initial background latent features. If high-frequency gradient energy is detected to be dominant, a large receptive field and high weights are configured for the basic feature branch to strengthen geometric edge parsing. If low-frequency color variance is detected to be dominant, a large receptive field is configured for the atmosphere feature branch to strengthen global atmosphere mapping, thereby ensuring that the decoupling network can perform optimal path decomposition operations for background data in different states.
[0053] The initial background latent features are input into the configured dual-stream orthogonal decomposition module. The basic feature branch and the atmosphere feature branch perform spatial mapping transformation on the initial background latent features through independent learnable mapping matrices to initially generate basic background environment mapping features and atmosphere perturbation mapping features. During the parallel process of mapping transformation, an orthogonal constraint decomposition mechanism is forcibly introduced to construct an orthogonal penalty term in the network's latent space, calculate the inner product matrix between the basic mapping features and the atmosphere mapping features, and minimize the Frobenius norm of this inner product matrix through gradient inversion. This orthogonal constraint forces the basic feature branch to retain only the high-frequency spatial geometric shape and material albedo information, while the atmosphere feature branch retains only the low-frequency spatial illumination color temperature and volume perturbation information. Through maximum decoupling in the feature space, the final output is a mutually independent and non-interfering background environment basic feature map and atmosphere perturbation feature map.
[0054] Specifically, in the rendering thread of a sandbox survival game, when background rendering data containing the information that "the player's sanity is extremely low and they are in a dark corner of the dungeon" is generated, the system immediately marks and calls the pre-trained background style decoupling network to load the rendering data tensor containing complex state information into the network buffer. The feature encoder inside the network performs convolutional dimensionality reduction and semantic purification on the data to extract the initial background latent features containing the dungeon outline and the atmosphere of terror, preparing to provide a unified input for the subsequent separation of the dungeon's physical structure and the indescribable atmosphere.
[0055] The decoupled network detects the extracted initial background latent features and finds that the high-frequency spatial gradient energy is extremely high due to the presence of dungeon walls and door frames. At the same time, the low-frequency color variance also shows abnormal fluctuations due to the dark red filter caused by the sharp drop in sanity value. Based on this detection result, the network adaptively assigns convolutional kernels with large receptive fields and high channel weights to the basic feature branch to accurately extract the rough edges of the stone walls. At the same time, it configures dilated convolutions for the atmosphere feature branch to capture a large range of dark red mental pollution halo, ensuring that the hardness of the physical structure is not lost in the subsequent decomposition, while fully resolving the pervasive sense of the Cthulhu atmosphere.
[0056] The initial background latent features, intertwined with dungeons and horror, are input into the dual-stream branch for mapping transformation, while triggering an orthogonal constraint mechanism. The orthogonal constraint penalty term forces the calculation of the inner product between the basic mapping feature representing the "stone wall brick outline" and the atmosphere mapping feature representing the "dark red mist and tentacle phantoms," and forces the inner product to be minimized to zero. This process mathematically completely strips away the red light attached to the stone wall texture and the brick artifacts hidden in the dark red mist, so that the basic feature branch outputs a pure, light-free gray-white dungeon geometric structure map, while the atmosphere feature branch outputs a pure, entity-free dark red distorted field perturbation map, achieving an absolute orthogonal decomposition of physical entities and indescribable mental interference in the Cthulhu background.
[0057] Furthermore, during the decomposition process, the background rendering data is further combined with the background constraints of the virtual game, and the contrastive learning loss factor is used to generate mutually independent background environment basic feature maps and atmosphere perturbation feature maps. The background environment basic feature map represents the inherent physical properties of the scene; the atmosphere perturbation feature map represents dynamic emotions and environmental disturbances. The introduction of the atmosphere perturbation feature map to represent dynamic emotions and environmental disturbances, along with the introduction of background rendering data, further controls the background constraints and improves the accuracy of the background environment basic feature map and atmosphere perturbation feature map.
[0058] At this point, during the mapping of background rendering data through the dual-stream orthogonal decomposition module, the pre-set background constraint relationships within the virtual game engine are extracted. These background constraint relationships include scene geometric normal constraints, atmospheric scattering physical models, and material optical property boundaries. The background constraint relationships are then transformed into continuously differentiable soft constraint regularization terms and injected into the latent space mapping path of the orthogonal decomposition. Specifically, when the basic feature branch attempts to extract inherent physical properties, high-frequency noise that violates the topological continuity of the scene is suppressed through geometric normal constraints. When the atmosphere feature branch extracts environmental disturbances, the diffusion gradient and attenuation rate of low-frequency disturbances are constrained through the atmospheric scattering model. This utilizes prior physical knowledge to guide feature decoupling to converge along a reasonable direction, avoiding visual distortions that violate the underlying physical laws of the game from occurring in the decomposition results.
[0059] During the synchronous process of decomposing based on background constraints, a contrastive learning loss factor is introduced to increase the inter-class distance between basic features and atmosphere features in the feature latent space. A contrastive learning loss function is constructed with the basic feature map of the current batch as the query term and the corresponding atmosphere feature map as the negative sample term. By calculating the cosine similarity matrix between the query term and the negative sample term and maximizing the cross-entropy loss of this similarity matrix, the network is forced to suppress the mutual information sharing between basic features and atmosphere features during backpropagation. The contrastive learning loss factor and the aforementioned orthogonal constraint form a dual decoupling mechanism. The orthogonal constraint forces the dot product to be zero based on the orthogonality of the feature vector basis, while the contrastive learning loss factor pushes apart the cluster centers of the two types of features from the manifold structure of the semantic distribution, thereby maintaining extremely high decoupling robustness under the complex and variable input of dynamic rendering.
[0060] After dual-constraint decoding using background constraint relations and contrastive learning loss factors, the feature maps output by the dual-stream branches are semantically recalibrated. For the basic feature branches, the inherent physical properties of the scene are highlighted by spatial high-frequency enhancement operators, including geometric topological contours, rigid body deformation states, and surface basic albedo, to generate basic feature maps of the background environment.
[0061] At this point, for the original feature map output from the basic feature branch, the system first performs a spatial high-frequency enhancement operation to highlight the inherent physical properties of the scene. This high-frequency enhancement operation uses the Laplacian operator as the high-frequency extraction kernel, with a size of 3x3, a center coefficient of 4, and a neighborhood coefficient of -1. The system convolves the original feature map with the Laplacian operator to obtain the high-frequency component feature map. The system then performs element-wise weighted superposition of the original feature map and the high-frequency component feature map according to a weight ratio of 7:3, where the original feature map has a weight of 0.7 and the high-frequency component has a weight of 0.3. After superposition, the system uses a pixel-wise ReLU activation function to set negative values to zero to ensure that the output feature map is non-negative.
[0062] After the above high-frequency enhancement processing, in the feature map output by the basic feature branch, the pixel gradient magnitude corresponding to the geometric topological contour is increased by at least 30%, the edge response intensity corresponding to the rigid body deformation state is enhanced to more than 1.2 times the original value, and the low-frequency component corresponding to the surface basic albedo remains unchanged. Finally, a basic feature map of the background environment is generated. This feature map clearly presents the sharp edges and continuous surfaces of entities such as scene walls, ground, and door frames in the spatial dimension, and does not contain any color shift or light scattering information.
[0063] For the atmosphere feature branch, the low-frequency color shift, volumetric light scattering intensity and fluid disturbance field vector that represent dynamic emotions and environmental disturbances are filtered and enhanced through the channel attention mechanism to generate an atmosphere disturbance feature map. The two feature maps output are independent of each other at the data level and rigorously represent the inherent physical objectivity of the scene and the subjective emotions and dynamic disturbances that evolve with the game state at the semantic level.
[0064] At this point, the system performs two parallel operations on the input feature map: global average pooling and global max pooling. Global average pooling calculates the average pixel value of each feature channel across all spatial locations, while global max pooling calculates the maximum pixel value of each feature channel across all spatial locations. The results of the two pooling operations are output as two vectors, each with a length equal to the number of channels in the feature map. The system then inputs these two vectors into a shared two-layer fully connected network. The first fully connected layer of this network compresses the dimension of the input vector to one-quarter of the original number of channels and is followed by a ReLU activation function. The second fully connected layer restores the dimension to the original number of channels and is followed by a Sigmoid activation function. The output value of the Sigmoid function ranges from 0 to 1, with each channel corresponding to a weight coefficient. The system then averages the weight coefficients generated by the global average pooling path and the weight coefficients generated by the global max pooling path element-wise to obtain the final channel attention weight vector. In this vector, channels with weight coefficients greater than 0.6 are considered strongly correlated channels, and channels with weight coefficients less than 0.2 are considered weakly correlated channels. The system multiplies the original channel attention weight vector with the original feature map of the atmosphere feature branch channel by channel, that is, multiplying the two-dimensional matrix of each feature channel by its corresponding weight coefficient.
[0065] After multiplication, the feature amplitudes of strongly correlated channels can be preserved or enhanced, while the feature amplitudes of weakly correlated channels are suppressed to less than 20% of their original values. After the above channel attention filtering, in the feature map output by the atmosphere feature branch, the response intensity of the low-frequency color shift channel is maintained or enhanced, the contrast of the volumetric light scattering intensity channel is enhanced by at least one octave, and the spatial gradient of the fluid perturbation field vector is preserved in the edge regions and smoothed in the flat regions. The system finally generates an atmosphere perturbation feature map, which contains time-varying ambient color bias, fog concentration distribution, and non-physical visual distortion fields, but does not contain any geometric edge or entity contour information.
[0066] Specifically, in this Cthulhu dungeon scene, the engine's pre-defined "dungeon wall normal distribution constraints" and "enclosed space fog scattering constraints" are used as background constraints. During orthogonal decomposition, the wall normal constraints are injected as regularization terms into the basic feature branches, forcing the branches to strictly follow the normal orientation when extracting the physical contours of the stone bricks, thus avoiding the generation of false geometry suspended in mid-air. At the same time, the fog scattering constraints are injected into the atmosphere feature branches, restricting the low-frequency dark red fog feature representing "unspeakable mental pollution" to only diffuse along the surface normal direction of the dungeon floor and walls in a manner that conforms to the scattering rate of hydrodynamics, preventing the fog from penetrating the sealed walls and causing a violation of physical common sense, and ensuring that the spread of the alienated atmosphere conforms to the physical rules of sandbox games.
[0067] When decomposing dungeon scene data, the extracted "basic stone brick features" are used as query terms, and the "dark red mist atmosphere features" are used as negative sample terms to calculate the contrastive learning loss. When the network attempts to retain some red light and shadow information in the basic features, or to imply brick textures in the atmosphere features, the cosine similarity of the two types of features in the latent space increases. At this time, the contrastive learning loss factor produces a high penalty, forcibly pushing away the distribution center of the "stone brick physical feature cluster" and the "dark red mist emotional feature cluster" in the latent space through backpropagation. This ensures that even in dynamic situations where the player's sanity fluctuates drastically and the screen colors are severely distorted, the network can still clearly distinguish which is the physical entity and which is the emotional disturbance, thus preventing feature crosstalk between physical entities and mental illusions.
[0068] After dual constraints, the dungeon features are semantically recalibrated: the basic feature branch uses a spatial high-frequency enhancement operator to purify and generate a background environment basic feature map containing only the rough albedo of the dungeon stone walls, the rigid outline of the door frame, and the solid topology of the ground, presenting a cold gray-white geometric skeleton; the atmosphere feature branch amplifies low-frequency perturbations through channel attention to generate an atmosphere perturbation feature map representing "dynamic fear emotion" and "environmental physical turbulence". This map only contains a dark red global color bias vector that fluctuates with sanity value, an unnatural tentacle phantom deformation field distorted along the corner of the wall, and a suppressed dim light scattering intensity; the two maps are completely independent, the former defining "what" the dungeon is, and the latter defining "how" the dungeon "feels" at this moment, providing a pure decoupled foundation for subsequent style fusion.
[0069] refer to Figure 4 In step S13, the specific steps are as follows:
[0070] S131: Mark the preset dynamic style primitive library, use the atmosphere perturbation feature map as the query vector, and perform content similarity-based retrieval and adaptation in the dynamic style primitive library to determine the corresponding target style content;
[0071] S132: Cross-match the target style content with the background environment's basic feature map, thereby combining a multi-scale feature fusion mechanism during the matching process. Through cross-layer feature interweaving and channel attention mechanisms, feature recombination and detail presentation are carried out at multiple perceptual scales to generate background potential content containing high-dimensional style semantics and low-dimensional physical structure.
[0072] In the embodiments of this application, a preset dynamic style primitive library is marked, the atmosphere perturbation feature map is used as the query vector, and content similarity-based retrieval and adaptation are performed in the dynamic style primitive library to determine the corresponding target style content. This approach is compatible with the overall considerations in the dynamic style primitive library and ensures the accuracy of the corresponding target style content.
[0073] At this point, in response to the instruction to generate the atmosphere perturbation feature map, a preset dynamic style primitive library is marked and invoked. The dynamic style primitive library is a pre-constructed high-dimensional feature matrix containing various discrete and continuous style representation vectors. Each row of primitive vectors is obtained by extracting latent representations from massive multi-style background images through a pre-trained autoencoder and calculating cluster centers, and is accompanied by a mean vector and covariance matrix representing the statistical characteristics of the style. At the same time, the currently input atmosphere perturbation feature map is compressed and converted into a compact query vector with fixed dimensions through global average pooling and a multilayer perceptron mapping layer, so that its feature distribution space is strictly aligned with the latent space of the dynamic style primitive library, thus establishing a unified metric for subsequent cross-space similarity retrieval.
[0074] For the dynamic style primitive library, the system collects at least 10,000 game background image samples covering various art styles, including but not limited to realistic, cartoon, ink painting, gothic, and Cthulhu styles. The system uses a pre-trained autoencoder network to process these samples. The autoencoder consists of an encoder and a decoder. The encoder progressively downsamples and compresses the input image into a 128-dimensional latent feature vector, and the decoder reconstructs the original image size from the latent vector. During training, the autoencoder minimizes the mean square error between the reconstructed image and the original image.
[0075] After training, the system discards the decoder and retains only the encoder. All sample images are input into the encoder to obtain a set of 128-dimensional latent feature vectors. The system then uses K-means clustering to cluster this set, with a preset number of 512 clusters. After clustering convergence, the system calculates the center vector for each cluster, which is the arithmetic mean of all feature vectors within that cluster. Simultaneously, the system calculates a covariance matrix for each cluster, a 128x128 dimensional matrix used to characterize the distribution of feature vectors within that cluster. These 512 center vectors and their corresponding covariance matrices together constitute a dynamic style primitive library, where each center vector is called a style primitive.
[0076] The mapped query vector is matched with each primitive vector in the dynamic style primitive library based on content similarity. The metric distance between the query vector and each primitive vector is calculated, and the metric distance is measured using scaled dot product similarity or Mahalanobis distance based on covariance. Based on the calculated similarity score distribution, a temperature coefficient is introduced for scaling, and the similarity score is converted into an adaptive weight distribution through the Softmax function. In this distribution, primitives that are highly consistent with the semantics of the current atmosphere perturbation receive a weight close to 1, while primitives with semantic conflicts receive a weight close to 0. The style primitives in the primitive library are weighted and summed using this adaptive weight distribution to achieve a non-rigid soft-fit allocation, thereby combining the discrete primitive library into a continuous and accurate target style implicit representation that matches the current atmosphere state.
[0077] The implicit representation of the target style generated by soft adaptation is input into the style decoder associated with the dynamic style primitive library. The low-dimensional latent vectors are upsampled and reconstructed into a high-dimensional spatial feature map through a transposed convolutional network. At the same time, based on the adaptive weight distribution corresponding to the implicit representation of the target style, the mean vector and covariance matrix of each primitive involved in the combination are weighted and fused to extract the global statistical parameter set of the target style content. The target style content includes both style feature maps with spatial structural details and statistical parameters that control color distribution and texture response intensity. Together, they constitute the complete target style content that needs to be injected into the virtual game background at the current moment for subsequent multi-scale feature fusion.
[0078] Specifically, in the sandbox survival game, after generating an atmospheric perturbation feature map representing "dark red mist and low-sanity fear," the system calls a preset dynamic style primitive library. This primitive library pre-stores the mean and covariance of various Cthulhu-style feature vectors such as "flesh crawling style," "abyssal mist style," and "spirit world distortion style." At the same time, a multilayer perceptron is used to reduce the dimensionality of the complex atmospheric perturbation feature map to a one-dimensional compact query vector. This ensures that the mathematical space of the query vector is under the same metric as the primitive vectors such as "flesh crawling" and "abyssal mist" in the primitive library, thus making subsequent retrieval comparisons mathematically equivalent.
[0079] The query vector representing "dark red mist and low-sanity fear" was input into the primitive library for retrieval. The similarity between the query vector and each primitive in the library was calculated by scaling the dot product. The results showed that the query vector had a very high similarity to the "Abyssal Mist Style" primitive, a moderate similarity to the "Flesh and Blood Style" primitive, and a very low similarity to the "Spirit World Distortion Style" primitive. After the Softmax function was used to assign adaptive weights, "Abyssal Mist" received a weight of 0.6, "Flesh and Blood Style" received a weight of 0.4, and "Spirit World Distortion" received a weight close to 0. By weighted summation of these primitives according to their weights, a composite Cthulhu-style implicit representation was created, which had both a heavy and deep misty feel and a slight flesh and blood texture at the edges, perfectly matching the subtle atmosphere of the player's current sanity being on the verge of collapse.
[0080] An implicit representation input style decoder, composed of "Abyssal Mist" with a weight of 0.6 and "Flesh and Blood Creep" with a weight of 0.4, is upsampled and reconstructed to produce a spatial feature map with dark red tones and a viscous, creeping texture. Simultaneously, based on the weights of 0.6 and 0.4, the mean / covariance parameters of the cool and dark "Abyssal Mist" and the mean / covariance parameters of the scarlet "Flesh and Blood Creep" are weighted and fused to extract a set of color and texture response statistical parameters that are predominantly cool and dark with local scarlet shimmering. This spatial feature map and statistical parameters together constitute the "Abyssal Flesh and Blood Composite Target Style Content" required for the current dungeon, waiting to be fused with the dungeon's basic physical feature map to make the cold stone wall surface exhibit a viscous, dark red, alienated texture that fluctuates with sanity values.
[0081] Furthermore, the target style content is cross-matched with the background environment's basic feature map. In the matching process, a multi-scale feature fusion mechanism is combined. Through cross-layer feature interweaving and channel attention mechanism, feature reorganization and detail presentation are carried out at multiple perceptual scales to generate background latent content containing high-dimensional style semantics and low-dimensional physical structure.
[0082] At this point, the target style content and the background environment basic feature map are input into the feature interaction layer of the multi-scale feature fusion network. The background environment basic feature map is used as the query term, and the target style content is used as the key and value term. Feature matching operation based on cross-attention is performed. In this operation, the low-dimensional physical structure features perform spatial retrieval of the high-dimensional style semantic features through spatial location encoding, and determine the best anchor point for each style semantic information in the physical structure space. This aligns and anchors the diffuse style features without clear spatial affiliation to specific geometric contours and topological regions, realizing the initial pixel-level alignment and association between the physical rigid structure and the style flexible semantics in the feature space.
[0083] After cross-matching, a multi-scale feature fusion mechanism is introduced, inputting the background environment basic feature map and the matched style features into feature pyramid levels with different downsampling ratios. Cross-layer feature interleaving is performed at each level, with low-level high-resolution features rich in low-dimensional physical edge details and high-level low-resolution features containing high-dimensional style global semantics being bidirectionally skipped and fused element-wise. During the fusion process, high-frequency fidelity filtering is applied to low-level physical features and low-frequency smoothing diffusion is applied to high-level style features to address the differences in the receptive fields of features at different scales. This achieves the gradual penetration and feature recombination of physical structure and style semantics at multiple receptive scales, generating a multi-scale fused feature pyramid.
[0084] To address the redundancy and semantic conflicts in channel features caused by the superposition of physical and style features in the multi-scale fusion feature pyramid, a channel attention mechanism is introduced to dynamically calibrate the recombined features. Global average pooling and max pooling are performed on the multi-scale fusion features along the channel dimension to extract the statistical distribution characteristics of each channel, and normalized weight coefficients for each channel are generated through a shared multilayer perceptron network. The weight coefficients are then scaled channel-by-channel with the multi-scale fusion features to suppress channels with blurred physical structures caused by style-forced injection and enhance detail response channels that are highly consistent with style and physics. The calibrated multi-scale features are then upsampled, concatenated, and convolutionally decoded to generate background latent content that contains both high-dimensional style semantics and retains low-dimensional physical structure details.
[0085] Specifically, in the Cthulhu-style fantasy setting of a sandbox survival game, the basic feature map representing the background environment of the dungeon stone walls and door frames is used as the query item, and the target style content containing "abyssal mist and flesh crawling" is used as the key and value item for cross-matching. The high-frequency edge features of the dungeon walls are retrieved in the style features, and the texture of "flesh crawling" is accurately anchored to the physical surface and cracks of the stone walls. At the same time, the color features of "abyssal mist" are located to the depth structure area of the corridor, so as to avoid the mist texture from mistakenly covering and destroying the hard edges of the door frame, thus achieving precise spatial anchoring of the Cthulhu-style alienation style and the physical space of the dungeon.
[0086] The basic features of the dungeon and the matched Cthulhu-style features are fed into the feature pyramid, and cross-layer feature interweaving is performed. At the lowest high-resolution level, the microscopic sharp cracks of the dungeon brick walls (low-dimensional physical structure) are preserved, and a slight dark red style infiltration is injected. At the higher low-resolution level, the grand oppressive feeling of the entire corridor being shrouded in abyssal mist is preserved (high-dimensional style semantics), and the macroscopic perspective of the corridor is injected. Through bidirectional cross-layer connection, the macroscopic oppressive feeling of mist spreads downwards to guide the microscopic color tone, and the microscopic brick wall cracks support the macroscopic structure upwards. This makes the final reconstructed features have both the grand horror of the Cthulhu style and the microscopic granularity of the dungeon brick walls.
[0087] In this multi-scale feature blending dungeon and Cthulhu styles, some passages may have blurred physical boundaries of door frames due to excessive overlay of fog, while others may have lost their stone texture due to the forced injection of flesh and blood textures. Through a channel attention mechanism, pooling statistics are performed on the entire image and weights are calculated. The weights of channels that cause blurred boundaries and texture conflicts are automatically reduced, while the weights of channels that enhance the contrast between dark red luster and uneven brickwork are increased. After weight scaling and convolutional decoding, the final background latent content is generated. This content presents the grand abyss fog and the grotesque flesh and blood (high-dimensional style semantics) while still clearly preserving the hard texture of the dungeon stone walls and the sharp outline of the corridor door frames (low-dimensional physical structure), providing a perfect composite implicit expression for subsequent light step rendering.
[0088] refer to Figure 5 In step S14, the specific steps are as follows:
[0089] S141: When the user is in the dynamic acquisition area, multiple posture parameters of the user are determined based on the camera's acquisition of the user's actions, and the pose parameters of the current virtual camera are determined by combining the cross-matching of the spatial coordinate system of the virtual game.
[0090] S142: Map the background latent content to a preset implicit neural radiation network, and trigger the corresponding latent space ray stepping operation in combination with the pose parameters of the current virtual camera. During the stepping process, multi-frequency sparse sampling is performed based on the viewpoint-related adaptive step size strategy to reconstruct the ray attenuation and occlusion relationship through the separate sampling of high-frequency details and low-frequency contours. Finally, the corresponding game background image is output, which has parallax effect and dynamic depth of field.
[0091] In the embodiments of this application, when the user is in the dynamic acquisition area, multiple posture parameters of the user are determined based on the camera's acquisition of the user's actions, and the pose parameters of the current virtual camera are determined by combining the cross-matching of the spatial coordinate system of the virtual game. This takes into account the overall consideration of the cross-matching of the spatial coordinate system of the virtual game and ensures the accuracy of the pose parameters of the current virtual camera.
[0092] During the virtual game, the system monitors in real time whether the user is within a preset physical dynamic acquisition area. Once the user is confirmed to have entered the visual sensor coverage area, the camera array is triggered to capture continuous frame images of the user. The captured image sequence is input into a pre-constructed skeletal key point detection model. The pixel coordinates of a preset number of joints on the user's body are extracted on the two-dimensional image plane through spatial coordinate regression. The three-dimensional physical spatial coordinates of each joint are calculated based on the triangulation principle of multi-view vision. The three-dimensional physical spatial coordinates are then subjected to temporal smoothing and filtering noise reduction to determine multiple posture parameters that characterize the user's current physical posture. These posture parameters include the three-dimensional position vector of each joint and the rotation vector between adjacent joints.
[0093] The extracted physical space posture parameters are mapped and aligned with the world space coordinate system of the virtual game to construct an affine transformation matrix from the physical acquisition space to the virtual game space. Specifically, the head orientation vector and torso displacement vector, which represent the user's head and gaze direction, are extracted from the posture parameters and used as the control driving source for the virtual camera posture. The displacement vector in the physical space is normalized and scaled to match the unit scale of the virtual game space coordinate system, and the rotation vector in the physical space is cross-compared and rotated to match the Euler angles or quaternion expressions of the virtual game space coordinate system. This eliminates the mapping deviation caused by the inconsistency between the coordinate axis directions of the physical space and the virtual space, thereby restoring the virtual control intent synchronized with the user's physical actions in the virtual game space.
[0094] Based on the virtual control intent obtained from the above cross-matching, and combined with the basic global coordinates of the virtual character in the game world, the target position and target orientation of the current virtual camera are calculated. The target position and target orientation are then interpolated with the historical pose of the virtual camera in a time sequence, and a low-pass filtering mechanism is introduced to eliminate high-frequency jitter in the virtual camera's viewpoint caused by physical motion jitter. At the same time, based on the dwell time and rate of change of the gaze in the pose parameters, the focal length scaling factor of the virtual camera is dynamically adjusted. Finally, the current virtual camera pose parameters, which include three-dimensional spatial position coordinates, quaternion rotation posture, and focal length parameters, are generated for subsequent latent space ray stepping rendering.
[0095] Specifically, in this sandbox survival game, when a player is in a room that supports motion-sensing interaction, the array of depth cameras deployed in the room is triggered to continuously capture images of the player; the skeletal detection model extracts the coordinates of the player's shoulder, elbow, wrist, spine and other joints from the images, and calculates the three-dimensional coordinates of these joints in the physical room through triangulation; even if the player subconsciously leans back, shrinks, or nervously looks around while experiencing the terrifying atmosphere of the Cthulhu dungeon, the system can convert these into precise physical spatial posture parameters, providing a realistic physical motion driving source for the subsequent control of the virtual camera.
[0096] The system maps the player's displacement and rotation parameters as they "step back to the left and lower their head" in the physical room to the virtual world coordinate system of the Cthulhu Dungeon. Through affine transformation, the backward displacement in the physical room is normalized and scaled to the backward stride in the dungeon corridor. The head-turning action to the left in physical space is precisely converted into a quaternion rotation around the Y-axis and X-axis of the virtual character's spine in the virtual world. Thus, the player's shrinking and looking around in fear due to seeing terrifying illusions in reality is seamlessly and synchronously converted into the backward retreat and terrified backward look in the dark corridor of the dungeon from a virtual perspective, achieving a perfect mapping between physical fear and virtual perspective.
[0097] The system overlays the matched backward and panoramic intentions onto the virtual character's baseline coordinates in the dungeon, calculating the target position and orientation of the virtual camera. To prevent players from shaking violently due to body tremors caused by fright, the camera pose is smoothed through temporal interpolation and low-pass filtering to ensure that the backward and backward viewing processes are smooth and conform to inertia. At the same time, the system detects that the player's gaze lingers for a long time in the darkness at the end of the corridor and that the eyeballs shrink, resulting in changes in posture parameters. Based on this, the system dynamically reduces the camera focal length to simulate the physiological reaction of narrowing field of vision when humans are afraid. Finally, the system outputs virtual camera pose parameters that include position, smooth rotation, and fear zoom characteristics, providing a highly immersive observation perspective for subsequent ray stepping calculations.
[0098] Furthermore, the latent background content is mapped to a preset implicit neural radiation network, and the corresponding latent space ray stepping operation is triggered in combination with the pose parameters of the current virtual camera. During the stepping process, multi-frequency sparse sampling is performed based on a viewpoint-dependent adaptive step size strategy to reconstruct the light attenuation and occlusion relationship through the separate sampling of high-frequency details and low-frequency contours. Finally, the corresponding game background image is output. This game background image has parallax effect and dynamic depth of field, further controls the latent background content, fully considers the pose parameters of the current virtual camera and the latent background content, and improves the accuracy of the game background image.
[0099] At this point, the background latent content generated in the previous steps is used as the input tensor of the implicit neural radiation network. Through the three-dimensional hash grid encoder and multilayer perceptron inside the implicit neural radiation network, the two-dimensional multi-channel background latent content is deconstructed along the depth dimension and mapped into a continuous three-dimensional implicit radiation field distribution. At the same time, the pose parameters of the current virtual camera are extracted, and a set of world space rays penetrating the virtual game scene is generated based on the optical center position of the virtual camera and the pixel coordinates of the sensor plane. Driven by the set of rays, the latent space ray stepping operation is triggered. The origin and direction vector of the ray are input into the implicit neural radiation field, and discretization stepping is performed along the ray travel direction to query and obtain the volume density and color features of each three-dimensional sampling point on the ray path.
[0100] During the ray stepping process, a viewpoint-dependent adaptive step size strategy is introduced to perform multi-frequency sparse sampling. Based on the pose parameters of the current virtual camera, the focal depth information is extracted, and the ray sampling interval is divided into a high-frequency detail focal area and a low-frequency contour background area. For rays passing through the high-frequency detail focal area near the focal depth, the step size is adaptively reduced, and high-frequency dense sampling is performed to resolve subtle geometric undulations and texture features in the scene. For rays passing through the low-frequency contour background area far from the focal depth, the step size is adaptively increased, and low-frequency sparse sampling is performed to capture macroscopic spatial contours and global illumination attenuation. At the same time, the volume density gradient of adjacent sampling points is monitored in real time during the stepping process. When the density gradient is lower than a preset threshold, it is determined to be a homogeneous space and intermediate sampling points are skipped. When the density gradient changes abruptly, local high-frequency resampling is triggered, thereby achieving efficient allocation of computing resources and separation and capture of high- and low-frequency features.
[0101] Based on high-frequency detail features and low-frequency contour features obtained from multi-frequency sparse sampling, discrete volume rendering integral operations are performed to reconstruct the relationship between light attenuation and occlusion. The color features and volume density of each sampling point are weighted and accumulated along the ray direction. High-frequency dense sampling points provide sharp edge occlusion and high-frequency light and shadow transitions, while low-frequency sparse sampling points provide smooth light attenuation and diffuse reflection. At the same time, combined with the aperture parameters and depth of focus of the current virtual camera, Gaussian scattering kernel convolution is applied to the low-frequency contour features that deviate from the focal plane along the adjacent sampling points in the depth dimension during the integration process to simulate the defocus blur effect in physical optics. The integral results of all rays are mapped to the two-dimensional image plane to output a game background screen with accurate parallax effect and dynamic depth distribution.
[0102] In practice, the system takes the background latent content feature map generated in the previous steps as input. The feature map is 256 pixels wide and 256 pixels high with 32 channels. The system inputs this feature map into the two-dimensional feature extraction module in the implicit neural radiation field network. This module consists of three stacked convolutional layers. Each convolutional layer has a kernel size of 3x3, a stride of 1, padding of 1, and channels of 32, 64, and 128 respectively. After processing by this module, the system obtains a two-dimensional feature map with a size of 256x256 and 128 channels. The system then expands each spatial location in this two-dimensional feature map along the depth dimension.
[0103] The specific method is as follows: For each pixel coordinate on the feature map, the system extracts its 128-dimensional feature vector and combines it with the pixel's 3D coordinates in the virtual game world space, and inputs them together into the hash grid encoder. The hash grid encoder performs linear interpolation in an eight-level voxel grid based on the input 3D coordinates, extracts the 16-dimensional features stored in the corresponding voxel at each level, and concatenates the features from the eight levels to form a 128-dimensional encoding vector. This encoding vector is then input into a four-layer multilayer perceptron, with each layer having 256, 128, 64, and 32 nodes, respectively. Each layer is followed by a ReLU activation function, and the last layer outputs two values: volume density and a 32-dimensional color feature vector. By performing the above mapping on all pixel coordinates and their corresponding 3D spatial points on the background latent content feature map, the system deconstructs and reconstructs the original two-dimensional background latent content into a continuous 3D implicit radiation field distribution. Each 3D point in this radiation field has a scalar volume density value and a 32-dimensional color feature vector.
[0104] The system collects the pose parameters of the current virtual camera, including the camera's three-dimensional position coordinates in world space, rotational attitude represented by quaternions, horizontal field of view, vertical field of view, and depth of focus. Based on the camera's horizontal and vertical field of view, and combined with the resolution of the output image, the system calculates the ray direction vector corresponding to each pixel.
[0105] Specifically, for a point in the output image located at pixel coordinates (u, v), the system first normalizes it to a coordinate range of -1 to 1, then multiplies it by the focal length (calculated from the field of view), and then performs a rotation matrix transformation using a rotation quaternion to obtain the ray direction vector of that pixel in world space. The origin of the ray is the camera's position coordinates in world space; the system generates an independent ray for each pixel in the image, and the total number of rays is the image width multiplied by the image height. For example, for a 1920x1080 output image, the system generates approximately 2.07 million rays; each ray contains three attributes: origin coordinates, direction vector, and the associated pixel coordinates.
[0106] Before the ray stepping begins, the system partitions the travel range of each ray based on the current virtual camera's focal depth parameter. The system denotes the focal depth value as D_focus and sets a depth tolerance threshold Delta, which is D_focus multiplied by 0.1. The system divides the region on the ray from the camera's optical center to a depth of (D_focus minus Delta) into the foreground low-frequency region, the region from a depth of (D_focus minus Delta) to (D_focus plus Delta) into the focal high-frequency region, and the region with a depth greater than (D_focus plus Delta) into the background low-frequency region. The length ratios of the three partitions on the ray are not necessarily equal, but the center of the high-frequency region is always aligned with the focal depth position.
[0107] The system employs different sampling step size strategies based on the partition type where the ray is located. Within the high-frequency focusing region, the system sets the sampling step size to one-thousandth of the diagonal length of the scene bounding box. For example, if the bounding box of a virtual game scene is a cuboid 100 meters long, 100 meters wide, and 50 meters high, with a diagonal length of approximately 150 meters, then the sampling step size in the high-frequency region is 0.15 meters. The system continuously samples within the high-frequency region, with the number of sampling points being the length of the high-frequency region divided by the step size, and this number, after being rounded down, is at least 32.
[0108] Within the foreground low-frequency region, the system's sampling step size is set to one-thirtieth of the diagonal length of the scene bounding box, approximately 5 meters. The system collects a sampling point every 5 meters within this region. The number of sampling points is the length of the foreground low-frequency region divided by 5 meters and rounded down, but not less than 2 and not more than 10.
[0109] In the low-frequency background region, the system first calculates the distance between the start and end points of the region, and then divides the distance by 32 to obtain a reference step size. However, the final step size does not exceed one-twentieth of the diagonal length of the scene bounding box, i.e., no more than 7.5 meters. The system performs equal-interval sampling in this region with this reference step size, and the number of sampling points is fixed at 32.
[0110] In addition, the system monitors the volume density changes between adjacent sampling points in real time during the stepping process. Specifically, after each sampling step, the system calculates the absolute value of the difference between the volume density of the current sampling point and the volume density of the previous sampling point, and divides this difference by the step time to obtain the volume density gradient. If the volume density gradient of three consecutive sampling points is less than 0.01 per meter, the system determines that the current ray is in a homogeneous space and automatically skips up to five subsequent sampling steps, that is, directly skips a large step distance before resuming sampling. If the detected volume density gradient exceeds 0.5 per meter, the system determines that the current ray has crossed the material interface and triggers local resampling, that is, inserts two additional sampling points before and after the current sampling point position, and reduces the sampling step size of the inserted points to one-fifth of the original step size.
[0111] After completing multi-frequency sparse sampling, the system performs discrete volume rendering integral calculations on all sampling points along each ray. This operation processes each sampling point sequentially along the ray from near to far, using a standard volume rendering integral approximation: the red component of the final pixel color equals the sum of the red components of all sampling points multiplied by the transmittance, where transmittance is the exponentially decaying cumulative value of all volume densities from the ray's origin to the current sampling point; the green component is calculated in the same way as the blue component. This calculation formula includes a step size factor, where the contribution of each sampling point is proportional to its corresponding sampling step size.
[0112] During this integration process, the high-frequency dense sampling points, due to their short step size and drastic density gradient changes, mainly contribute to the boundary regions of abrupt density changes, producing sharp edge occlusion effects and clear light and shadow transitions; while the low-frequency sparse sampling points, due to their long step size and gradual density changes, contribute to smooth light attenuation and diffuse reflection accumulation.
[0113] While performing volumetric rendering integration, the system simulates dynamic depth-of-field effects by combining the aperture and depth-of-focus parameters of the current virtual camera. The aperture parameter is denoted as F_number, with a value range of 2.8 to 16; a smaller value indicates a larger aperture. The system first calculates the actual depth value of each sampling point along the ray direction from the camera's optical center. If the sampling point is located in the high-frequency region of focus (i.e., the depth value is within the range of D_focus plus or minus Delta), the system does not modify the contribution of that sampling point and outputs a clear image. If the sampling point is located in the low-frequency region of the foreground or background (i.e., the depth value deviates from the depth-of-focus by more than Delta), the system performs Gaussian scattering kernel convolution processing on the color contribution of that sampling point.
[0114] The specific implementation is as follows: The system calculates the size of the scattering kernel based on the distance of the sampling point from the focal depth. The deviation distance is equal to the absolute value of the difference between the sampling point depth and the focal depth. The unit of the scattering kernel size is pixels, and its calculation formula is the deviation distance divided by D_focus and then multiplied by a base kernel size, which is F_number multiplied by 0.5; the minimum value of the scattering kernel size is 0 pixels, and the maximum value is 16 pixels. The system weights the color contribution value of the sampling point within a circular area with a radius equal to the scattering kernel size around the projection position on the two-dimensional image plane according to a two-dimensional Gaussian distribution. The standard deviation of the Gaussian distribution is equal to one-third of the scattering kernel size; the system accumulates the allocated color contribution value to the target pixel and its surrounding neighboring pixels, rather than only accumulating it to the original projection position.
[0115] The Gaussian scattering kernel convolution described above is performed in real time during the volume rendering integration process, rather than as a post-processing filter. The integration results of all rays are accumulated into the two-dimensional image frame buffer according to their original pixel coordinates and Gaussian scattering contribution coordinates. In the final output game background image, objects near the depth of focus exhibit sharp edges and clear textures, while foreground and background objects deviating from the depth of focus exhibit different degrees of defocus blur. The degree of blur increases linearly with the deviation distance and is modulated by the aperture parameters, thereby achieving a physically believable dynamic depth of field effect.
[0116] Because the aforementioned ray stepping process independently calculates the color accumulation along the path in the 3D radiation field for each ray, and the ray direction is updated in real time as the virtual camera pose parameters change, when the position or orientation of the virtual camera changes, the projected position of the same 3D scene point on the image plane will shift accordingly, with the displacement of nearby objects being greater than that of distant objects. This difference in projection displacement caused by the change of viewpoint is the parallax effect. In each frame rendering, the system re-executes all the above steps based on the latest acquired virtual camera pose parameters. Therefore, the final output game background image has a complete motion parallax effect, that is, when the player moves the virtual camera, the foreground elements and background elements in the scene move relative to each other at different rates, enhancing the sense of immersion and 3D space.
[0117] Specifically, the background latent content containing the "physical structure of the dungeon stone walls" and the "semantic meaning of the abyss mist" is input into the implicit neural radiation network. The network maps it into a three-dimensional implicit radiation field filled with dark red glimmer and viscous fluid. When the system obtains the virtual camera pose parameters representing the player's "frightened retreat and look around", it emits tens of thousands of world space rays from the optical center of the virtual camera to the screen pixels. These rays enter the radiation field and perform ray walking, like projecting invisible probes in the indescribable dungeon space, preparing to sample the density of the viscous flesh walls and the dark red optical characteristics of the abyss mist along the way.
[0118] When the virtual camera focuses on the faintly visible Cthulhu statue at the end of the dungeon corridor due to the player's fright, the system executes an adaptive step size strategy based on the depth of focus. In the focal plane area close to the statue, the ray step size is drastically reduced, and high-frequency dense sampling is performed to accurately capture the writhing flesh texture and tiny cracks on the statue's surface (high-frequency details). However, in the corridor's front end and deep, dark corners far from the focus, the step size is greatly stretched, and low-frequency sparse sampling is performed, only calculating the macroscopic occlusion and light and shadow attenuation of the dense abyss fog (low-frequency outline). For the vast, purely dark space through which the ray passes, the system detects no change in density and skips it directly, avoiding wasting computing power. Thus, with limited computing power, it simultaneously reconstructs clear and terrifying alien details and oppressive, deep scene outlines.
[0119] The system performs volume rendering integration on the high-frequency flesh texture and low-frequency abyssal fog along the ray path; high-frequency dense sampling points strictly block the light behind, reconstructing the sharp and terrifying parallax outline of the statue's edge, creating a highly oppressive 3D parallax effect as the player's view moves; while low-frequency sparse sampling points accumulate the gradually decaying dark red halo behind the fog; while integrating, the system applies a Gaussian scattering kernel based on depth difference to the nearby corridor stone wall and the endless abyss outside the player's field of vision, making it present a strong physical defocus blur in the picture, rendering a dynamic depth-of-field picture on the screen where the alien statue at the focal point has clear and terrifying veins, and the foreground and background are swallowed by fog and blur, perfectly reproducing the physiological visual characteristics of a person's field of vision narrowing under extreme fear and focusing only on the source of terror.
[0120] Please see Figure 6 The dynamic rendering system for game backgrounds based on virtual games is applied to the aforementioned dynamic rendering method for game backgrounds based on virtual games; the dynamic rendering system for game backgrounds based on virtual games includes:
[0121] The virtual game module 21 is used to acquire multi-source heterogeneous state data during the operation of the virtual game, perform cross-modal semantic alignment on the multi-source heterogeneous state data, and construct the background rendering data of the virtual game at the current moment by combining the aggregation mechanism.
[0122] Background rendering module 22 is used to input background rendering data into a pre-trained background style decoupling network. The background style decoupling network performs orthogonal decomposition on the background rendering data and generates mutually independent background environment basic feature maps and atmosphere perturbation feature maps by combining background constraint relationships during the decomposition process.
[0123] Background latent content module 23 is used to obtain a preset dynamic style primitive library, and combine it with the atmosphere perturbation feature map to determine the corresponding target style content. The target style content is then fused with the background environment basic feature map at multiple scales to generate the corresponding background latent content.
[0124] The game background image module 24 is used to collect the pose parameters of the current virtual camera, input the potential background content into the preset neural network, and combine the pose parameters of the current virtual camera to trigger the corresponding latent space ray stepping operation, thereby performing multi-frequency sparse sampling during the stepping process, and finally outputting a game background image with parallax effect and dynamic depth of field.
[0125] It should be noted that although multiple modules are mentioned in the detailed description above, this division is not mandatory; in fact, according to the embodiments of this disclosure, the features and functions of two or more modules or described above can be embodied in one module; conversely, the features and functions of one module described above can be further divided into multiple modules to be embodied.
[0126] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the disclosure herein; this application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein; the specification and embodiments are to be considered exemplary only.
[0127] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A dynamic rendering method for game backgrounds based on virtual games, characterized in that, include: During the operation of the virtual game, multi-source heterogeneous state data of the virtual game runtime is acquired, cross-modal semantic alignment is performed on the multi-source heterogeneous state data, and the background rendering data of the virtual game at the current moment is constructed by combining the aggregation mechanism. The background rendering data is input into a pre-trained background style decoupling network. The background style decoupling network performs orthogonal decomposition on the background rendering data and generates mutually independent background environment basic feature maps and atmosphere perturbation feature maps by combining background constraint relationships during the decomposition process. Obtain a preset dynamic style primitive library and combine it with the atmosphere perturbation feature map to determine the corresponding target style content. Then, perform multi-scale feature fusion between the target style content and the background environment basic feature map to generate the corresponding background potential content. The pose parameters of the current virtual camera are collected, the latent background content is input into the preset neural network, and the corresponding latent space ray stepping operation is triggered in combination with the pose parameters of the current virtual camera. In this way, multi-frequency sparse sampling is performed during the stepping process, and finally the game background screen with parallax effect and dynamic depth of field is output.
2. The dynamic rendering method for game backgrounds based on virtual games according to claim 1, characterized in that, During the operation of the virtual game, multi-source heterogeneous state data of the virtual game runtime is acquired, cross-modal semantic alignment is performed on the multi-source heterogeneous state data, and background rendering data of the virtual game at the current moment is constructed by combining an aggregation mechanism, including: The system monitors the virtual game's operation in real time by using data probes deployed at the game engine's underlying layer to acquire multi-source heterogeneous state data during the game's runtime. This multi-source heterogeneous state data covers scene topology, character interaction timing, and environmental physical response.
3. The dynamic rendering method for game backgrounds based on virtual games according to claim 2, characterized in that, The process of acquiring multi-source heterogeneous state data during the operation of the virtual game, performing cross-modal semantic alignment on the multi-source heterogeneous state data, and constructing the background rendering data of the virtual game at the current moment using an aggregation mechanism, also includes: Multi-source heterogeneous state data is input into a data processing network and cross-modal semantic alignment is performed based on a spatiotemporal attention mechanism. After alignment, a dynamic aggregation mechanism is introduced to adaptively aggregate the semantically aligned heterogeneous features along the spatial and temporal dimensions to construct background rendering data representing the virtual game at the current moment.
4. The dynamic rendering method for game backgrounds based on virtual games according to claim 1, characterized in that, The process involves inputting background rendering data into a pre-trained background style decoupling network. This network performs orthogonal decomposition on the background rendering data and, during the decomposition process, combines background constraint relationships to generate mutually independent background environment basic feature maps and atmosphere perturbation feature maps. The pre-trained background style decoupling network is labeled, and the background rendering data is loaded. At this time, the corresponding dual-stream orthogonal decomposition module is determined based on the detection of the background style decoupling network. The dual-stream orthogonal decomposition module performs feature mapping and orthogonal constraint decomposition on the background rendering data.
5. The dynamic rendering method for game backgrounds based on virtual games according to claim 4, characterized in that, The step of inputting background rendering data into a pre-trained background style decoupling network, wherein the background style decoupling network performs orthogonal decomposition on the background rendering data, and generates mutually independent background environment basic feature maps and atmosphere perturbation feature maps by combining background constraint relationships during the decomposition process, further includes: During the decomposition process, the background rendering data is further combined with the background constraints of the virtual game, and simultaneously combined with the contrastive learning loss factor to generate mutually independent background environment basic feature maps and atmosphere perturbation feature maps. The background environment basic feature map represents the inherent physical properties of the scene; the atmosphere perturbation feature map represents dynamic emotions and environmental disturbances.
6. The dynamic rendering method for game backgrounds based on virtual games according to claim 1, characterized in that, The process involves acquiring a preset dynamic style primitive library, determining the corresponding target style content by combining it with an ambient perturbation feature map, and then performing multi-scale feature fusion between the target style content and the background environment basic feature map to generate the corresponding potential background content, including: The system labels a pre-defined dynamic style primitive library, uses the atmospheric perturbation feature map as the query vector, and performs content similarity-based retrieval and adaptation in the dynamic style primitive library to determine the corresponding target style content.
7. The dynamic rendering method for game backgrounds based on virtual games according to claim 6, characterized in that, The step of obtaining a preset dynamic style primitive library, determining the corresponding target style content by combining it with the ambient perturbation feature map, and performing multi-scale feature fusion of the target style content with the background environment basic feature map to generate the corresponding background latent content also includes: The target style content is cross-matched with the basic feature map of the background environment. In the matching process, a multi-scale feature fusion mechanism is combined. Through cross-layer feature interweaving and channel attention mechanism, feature recombination and detail presentation are carried out at multiple perceptual scales to generate background potential content containing high-dimensional style semantics and low-dimensional physical structure.
8. The dynamic rendering method for game backgrounds based on virtual games according to claim 1, characterized in that, The process involves acquiring the pose parameters of the current virtual camera, inputting the latent background content into a preset neural network, and triggering corresponding latent space ray stepping operations based on the pose parameters of the current virtual camera. This process performs multi-frequency sparse sampling during the stepping process, ultimately outputting a game background image with parallax effects and dynamic depth of field, including: When a user is in a dynamic acquisition area, multiple posture parameters of the user are determined based on the camera's capture of the user's movements, and the pose parameters of the current virtual camera are determined by cross-matching the spatial coordinate system of the virtual game.
9. The dynamic rendering method for game backgrounds based on virtual games according to claim 8, characterized in that, The process of acquiring the pose parameters of the current virtual camera, inputting the latent background content into a preset neural network, and triggering corresponding latent space ray stepping operations based on the pose parameters of the current virtual camera, thereby performing multi-frequency sparse sampling during the stepping process, and finally outputting a game background image with parallax effect and dynamic depth of field, also includes: The background latent content is mapped to a preset implicit neural radiation network, and the corresponding latent space ray stepping operation is triggered by the pose parameters of the current virtual camera. During the stepping process, multi-frequency sparse sampling is performed based on the viewpoint-related adaptive step size strategy to reconstruct the light attenuation and occlusion relationship through the separate sampling of high-frequency details and low-frequency contours. Finally, the corresponding game background image is output, which has parallax effect and dynamic depth of field.
10. A dynamic rendering system for game backgrounds based on virtual games, characterized in that, The dynamic rendering system for game backgrounds based on virtual games is applied to the dynamic rendering method for game backgrounds based on virtual games as described in any one of claims 1-9; The dynamic rendering system for game backgrounds based on virtual games includes: The virtual game module is used to acquire multi-source heterogeneous state data during the operation of the virtual game, perform cross-modal semantic alignment on the multi-source heterogeneous state data, and construct the background rendering data of the virtual game at the current moment by combining the aggregation mechanism. The background rendering module is used to input background rendering data into a pre-trained background style decoupling network. The background style decoupling network performs orthogonal decomposition on the background rendering data and generates mutually independent background environment basic feature maps and atmosphere perturbation feature maps by combining background constraint relationships during the decomposition process. The background latent content module is used to obtain a preset dynamic style primitive library, and combine it with the atmosphere perturbation feature map to determine the corresponding target style content. The target style content is then fused with the background environment basic feature map at multiple scales to generate the corresponding background latent content. The game background image module is used to collect the pose parameters of the current virtual camera, input the potential background content into the preset neural network, and combine the pose parameters of the current virtual camera to trigger the corresponding latent space ray stepping operation, thereby performing multi-frequency sparse sampling during the stepping process, and finally outputting a game background image with parallax effect and dynamic depth of field.