Game picture dynamic interactive layout method and system
By constructing a global mass conservation displacement field and a local conformal displacement field, combined with neural radiation fields and energy-reinforcement learning, the problem of unifying the global and local aspects in the dynamic interactive layout of the game screen is solved, efficient amplification of the area of interest and background continuity are achieved, and the game's operational accuracy and immersion are improved.
Patent Information
- Application Number
- CN202510917989.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-03
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-07-03
AI Technical Summary
Existing technologies make it difficult to simultaneously take into account global mass conservation, local angle maintenance, and real-time interaction coordinate consistency in the dynamic interactive layout of game screens, resulting in experience defects such as small targets, misaligned interactions, or dizziness in high-speed shooting, multiplayer competitive, and mobile touch games.
An attention field-driven method is adopted to construct a global mass-conserving displacement field and a local angle-conserving displacement field. The color of the region of interest is re-rendered in real time through the neural radiation field, and the system energy-reinforcement learning closed-loop online parameter adjustment is used to achieve dynamic amplification of the region of interest and background continuity.
It achieves the magnification of the area of interest while maintaining global mass conservation and local angle continuity, improves the user input hit accuracy and cross-frame visual smoothness, reduces computing power consumption and frame rate fluctuations, and improves the gaming experience.
Smart Images

Figure CN120733344A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer graphics and human-computer interaction technology, and in particular to a method and system for dynamic interactive layout of game screens. Background Art
[0002] In real-time interactive entertainment, the game screen not only presents information but also directly determines the player's operational precision and immersion. Traditional "fixed interface + camera zoom" or "single viewport cropping" solutions can only zoom in or switch scenes overall, making it difficult to balance local detail with global coherence. This leads to user experience defects such as small targets, misaligned interactions, and motion sickness in high-speed shooting, multiplayer competitive, and mobile touch-based games. The industry generally attempts to address this problem with multi-layer UI overlays, off-screen rendering, or multi-resolution buffers, but these methods either increase bandwidth and latency or disrupt the consistency of the rendering pipeline.
[0003] Existing methods typically magnify the region of interest based on single-scale texture resampling or by reconstructing the viewport after cropping: using a fixed scaling matrix, it ignores the geometric continuity of the screen and is prone to "tearing" at the boundaries. Some solutions use mask maps to magnify the region of interest and then use bilinear interpolation to fill the holes. Due to the lack of global mass conservation, this leads to background holes or ghosting. Some solutions also use convolutional super-resolution to improve local clarity, but the network inference overhead is high and it cannot maintain true pixel correspondence, so hit detection is still inaccurate. The root cause is that they do not simultaneously consider the three constraints of "global mass conservation, local angle maintenance, and consistent real-time interactive coordinates." Therefore, distortion, delays, or misjudgments are inevitable in high-speed zoom or multi-person scenarios. Summary of the Invention
[0004] In response to the many problems existing in the above-mentioned existing technologies, the present invention provides a method and system for dynamic interactive layout of game screens. Driven by the attention field, the present invention successively constructs a global mass conservation displacement field and a local angle-preserving displacement field, fuses them into a unified displacement field according to the curvature weight, and deforms the screen pixels. The color of the area of interest is then re-rendered in real time using the neural radiation field, and the parameters are adjusted online through the system energy-reinforcement learning closed loop.
[0005] A method for dynamic interactive layout of a game screen comprises the following steps: Each frame acquires depth map, normal map, illumination map, color map and user input event data, inputs the acquired data into the neural network and combines it with symplectic integral to generate attention field; Determine the region of interest based on the attention field, construct a source probability density and a target probability density, obtain a first displacement field and a second displacement field respectively, and fuse the first displacement field and the second displacement field according to a preset weight to generate a unified displacement field; Using the unified displacement field to deform pixel coordinates, the voxel set corresponding to the region of interest is input into a neural radiation field network to obtain a compensated color map, and the compensated color map is synthesized with the deformed color map to form a fused image; For user input events on the fused screen, the original screen coordinates are determined by inverse operation of the unified displacement field and the second displacement field, the entity identification is obtained by querying the deformed entity identification map, and when the system energy exceeds a preset range, the attention field threshold, the second displacement field parameters and the damping parameters are adjusted to achieve cross-frame adaptive feedback.
[0006] Preferably, when generating the attention field, the neural network is constructed in sequence as a multi-head self-attention layer and a feedforward layer, and the feature vectors of the depth map, normal map, illumination map and color map are included in the same attention window in the multi-head self-attention layer.
[0007] Preferably, the first displacement field is obtained by performing Sinkhorn iteration on a discrete grid to solve the optimal transmission mapping between the source probability density and the target probability density, and the mapping matrix is row normalized and column normalized after each iteration.
[0008] Preferably, the second displacement field is obtained by calculating the Beltrami coefficient using the grayscale value of the attention field and maintaining the screen boundary identity mapping during the Koebe iteration process.
[0009] Preferably, the preset weights are allocated according to the Laplace curvatures of the first displacement field and the second displacement field on the pixel grid, and during fusion, the first displacement field and the second displacement field are linearly superimposed to generate a unified displacement field.
[0010] Preferably, the voxel set corresponding to the region of interest is obtained by back-projecting the pixel coordinates into three-dimensional coordinates and writing them into a hash grid, and the three-dimensional coordinates and the light direction are Fourier encoded respectively as inputs of the neural radiation field network.
[0011] Preferably, when determining the original screen coordinates, the screen space inverse displacement is first performed using the unified displacement field, and then a quasi-conformal inverse mapping is performed using the second displacement field, wherein the inverse displacement is calculated by bilinear interpolation.
[0012] Preferably, the system energy is obtained by summing the attention field change and the unified displacement field curvature on the same pixel grid and then weighting them according to a fixed coefficient; when adjusting the attention field threshold, the second displacement field parameter and the damping parameter, a temporal difference reinforcement learning algorithm is used with the goal of reducing the difference in system energy between two adjacent frames.
[0013] Preferably, when the system energy is lower than a preset lower limit and the change in the attention field is lower than a preset threshold, the unified displacement field of the previous frame and the second displacement field of the previous frame are reused, and the steps of determining the area of interest based on the attention field and deforming the pixel coordinates using the unified displacement field are skipped.
[0014] A system for dynamic interactive layout of game screens, for implementing the method for dynamic interactive layout of game screens, comprising: An attention field generation module is used to obtain depth maps, normal maps, illumination maps, color maps, and user input event data in each frame, and input the data into a neural network and combine it with symplectic integration to generate an attention field consistent with the screen resolution; a displacement fusion module, configured to determine a region of interest based on the attention field, construct a source probability density and a target probability density, obtain a first displacement field and a second displacement field, respectively, and fuse the first displacement field and the second displacement field according to a preset weight to generate a unified displacement field; a compensation rendering module, configured to deform pixel coordinates using the unified displacement field, input a voxel set corresponding to the region of interest into a neural radiation field network to obtain a compensated color map, and synthesize the compensated color map with the deformed color map into a fused image; An interactive feedback module is configured to determine the original screen coordinates of user input events on the fused screen by performing inverse operations on the unified displacement field and the second displacement field, query the deformed entity identification map to obtain the entity identification, and adjust the attention field threshold, the second displacement field parameters, and the damping parameters when the system energy exceeds a preset range to achieve cross-frame adaptive feedback.
[0015] Compared with the prior art, the advantages and beneficial effects of the present invention are: Through the pixel-level fusion of neural optimal transfer mapping + quasi-conformal mapping, a unified displacement field is achieved that magnifies the region of interest while maintaining global mass conservation and local angle continuity; through voxel hashing + neural radiation field compensation rendering, real-time reconstruction and seamless fusion of texture details in the magnified area are achieved; through two-level inverse mapping and system energy adaptive parameter adjustment, closed-loop control of user input hit accuracy and cross-frame visual smoothness is achieved; through the frame skipping multiplexing strategy, computing power savings and frame rate stability are achieved in static scenes. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 It is a schematic diagram of the process of the present invention; Figure 2 Schematic diagram of displacement field fusion in the present invention; Figure 3 Schematic diagram of two-stage inverse mapping in the present invention; Figure 4 It is a structural block diagram of the system of the present invention. DETAILED DESCRIPTION
[0017] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the following detailed description, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure.
[0018] like Figure 1 As shown, a method for dynamic interactive layout of game screens includes the following steps: Each frame acquires depth map, normal map, illumination map, color map and user input event data, inputs the acquired data into the neural network and combines it with symplectic integral to generate attention field; After each frame is rendered, the present invention synchronously copies the depth map, surface normal map, illumination map and color map, and reads the user input event data at the current moment. The four types of images are located in the same visual cone, and the corresponding pixels are matched one by one. They can be directly spliced into a multi-channel tensor in the video memory and sent to the attention field generation network. The network adopts a multi-head self-attention structure: for each pixel position, the depth value, the three components of the normal, the illumination intensity and the three components of the color are merged to obtain an 8-dimensional feature vector, which is mapped to a unified hidden dimension through a linear transformation and then enters the self-attention operator. The self-attention operator calculates the feature correlation in the spatial dimension, so that the influence of illumination changes on depth occlusion is transmitted synchronously with the user input hotspot, forming a fusion attention weight containing three types of information: geometry, material and interaction.
[0019] In order to avoid the hysteresis caused by relying solely on the static estimation of the previous frame, the present invention introduces symplectic integral prediction. Assume that the difference between the attention fields of the two frames is ,Will As a generalized speed, the current attention field As a generalized coordinate, the diagonal mass matrix is introduced and potential energy term . Using the Lewis-Frog format, In the time domain step Internal execution:
[0020]
[0021]
[0022] in is the attention field grayscale tensor, Its conjugate momentum. The Symplectic scheme ensures the conservation of the Hamiltonian within discrete time steps, suppressing long-term integral drift while providing smooth predictions during rapid user scrolling or drastic camera changes. The predicted attention field is linearly fused with the network output attention field in a 7-to-3 ratio, preserving real-time features while maintaining continuity.
[0023] The attention field plays a dual role in the subsequent steps of the present invention: on the one hand, it determines the region of interest through threshold segmentation and provides input for probability density mapping and displacement field calculation; on the other hand, it measures the frame difference. Participate in system energy evaluation and drive the adaptive adjustment logic. Since the attention field resolution is consistent with the screen, the downstream displacement field solution does not need to be scaled, reducing interpolation errors.
[0024] In mobile device adaptation experiments, when the rendering resolution was reduced to 1280×720 and the hidden dimension was reduced to 64, the total computational latency for attention field generation and prediction remained within 2 milliseconds. Despite halving the model size, the symplectic integral prediction maintained smooth tracking, reducing aiming point jitter during touch shooting by approximately 33%. These results demonstrate the excellent scalability of this invention to hardware resources.
[0025] Preferably, when generating the attention field, the neural network is constructed in sequence as a multi-head self-attention layer and a feedforward layer, and the feature vectors of the depth map, normal map, illumination map and color map are included in the same attention window in the multi-head self-attention layer.
[0026] In the present invention's method for dynamic interactive layout of game screens, the attention field characterizes the attention intensity of each pixel on the screen in the current frame and serves as the sole source of weight for subsequent region-of-interest extraction, displacement field solution, and interactive feedback. To balance geometric depth, surface texture, lighting conditions, and player input intent, the present invention employs a neural network consisting of a stack of multiple self-attention layers and feedforward layers to generate the attention field. Within the multi-head self-attention layer, the feature vectors of the depth map, normal map, illumination map, and color map are placed into the same attention window, enabling cross-modal fusion.
[0027] The network input is first concatenated pixel by pixel in the video memory: the depth map provides 1D distance values; the normal map provides 3D direction cosines; the illumination map provides 1D illuminance per pixel; and the color map provides 3D color components in sRGB space. The total 8-dimensional features are projected to a unified hidden dimension through a linear layer and then enter the multi-head self-attention layer. Each attention head independently calculates the query vector, key vector, and value vector. Let the hidden dimension be , the query matrix, key matrix and value matrix are denoted as 、 、 , which means the three sets of linear transformation results of the pixel embedded in the current head. The self-attention weight is obtained as follows:
[0028] in is a normalization factor used to avoid the gradient disappearance caused by the growth of dimensions. Represents the similarity matrix between the query and the key; the softmax function ensures that the sum of the pixel weights after row normalization is always equal to 1. Multiple attention heads are calculated in parallel and concatenated by channel, and then linearly mapped back to the hidden dimension to form a fused feature.
[0029] In multimodal input scenarios, depth and normal information describe scene geometry, illumination describes local brightness differences, and color maps record texture details. Self-attention weights explicitly model the cross-pixel correlations of these features, avoiding the limitation of traditional convolutions that only capture local neighborhoods. For example, when a player uses a flashlight to illuminate an enemy in a dim environment, the bright spots in the illumination map and the foreground pixels highlighted in the depth map will be given higher weights through the self-attention mechanism, correctly identifying the interaction focus.
[0030] The multi-head self-attention layer is followed by a two-layer feedforward network. The feedforward layer first uses an activation function to enhance the nonlinear expression, and then uses residual connection and layer normalization to keep the gradient stable. The network outputs a single-channel grayscale image, which is mapped to 0 to 1 after Sigmoid activation to obtain an attention field consistent with the screen resolution. The attention field is not only used for threshold segmentation, but also participates in the system energy measurement, so it must have inter-frame continuity. To this end, the present invention introduces symplectic integral prediction based on the network output: the difference of the attention field of the previous two frames is regarded as the generalized velocity, and the attention field of the current frame is regarded as the coordinate, and the simple harmonic potential energy is constructed:
[0031] in represents the attention field grayscale tensor, is its spatial gradient. Using the Lewis-Frog format A half-step momentum update, a full-step coordinate update, and a half-step momentum backfill are performed to ensure conservation of the discrete Hamiltonian. The prediction results are fused with the network output in a 7:3 ratio, providing one- to two-frame look-ahead smoothing when the user quickly pans or zooms, without causing excessive lag in static scenes.
[0032] Example 1: On a desktop computer, the rendering resolution was set to 1920×1080, the hidden dimension to 128, and the number of attention heads to 4. Depth, normal, illumination, and color features were concatenated and sent to the video memory. Attention field generation took 0.9 milliseconds, and symplectic integral prediction took 0.2 milliseconds, resulting in a total latency of 1.1 milliseconds, representing 3% of the full rendering budget. In an action sequence where a player switches guns while dodging sideways, the attention field generated by this method accurately locates the gunner and the crosshairs in the first frame. If symplectic integral prediction is disabled and only the convolutional attention network is used, the weights must wait two frames for stabilization, resulting in delayed region of interest extraction.
[0033] Example 2: The rendering resolution of a mobile device is 1280×720, the hidden dimension is reduced to 64, the number of attention heads is reduced to 2, and the total time is kept within 2 milliseconds. A touch-shooting experiment shows that compared to a non-symplectic integral prediction scheme, the present invention can reduce the mean squared error of aiming jitter from 4.5 pixels to 3 pixels, demonstrating the method's scalability to hardware resources.
[0034] like Figure 2 As shown, based on the attention field, the interest region is determined, the source probability density and the target probability density are constructed, and the first displacement field and the second displacement field are obtained respectively. The first displacement field and the second displacement field are fused according to a preset weight to generate a unified displacement field; Before generating a unified displacement field, the present invention first determines the region of interest based on the attention field. The attention field is a grayscale image that matches the screen resolution, where larger pixel values indicate a more prominent display in the current frame. The attention field is binarized using a fixed threshold, and connected pixel blocks constitute the region of interest. The region of interest is retained in the pixel coordinate system to avoid precision loss caused by subsequent coordinate transformations.
[0035] In order to ensure that the region of interest is visually continuous with the background area after magnification, and to avoid interrupting the player's global sense of space, the present invention regards the region of interest and the background area as mass distributions respectively: the region of interest counts the number of pixels and the average pixel weight, while the background area is obtained by normalizing the remaining pixel weights. The source probability density is obtained by dividing the attention field of the entire frame by the sum of the weights; the target probability density is obtained by superimposing the Gaussian kernels weighted by the number of pixels of all regions of interest and then dividing by the sum of the weights. The source probability density and the target probability density are defined on the same pixel grid, and the sum of the probability values at each pixel is 1.
[0036] In order to accurately map from the source probability density to the target probability density, the present invention adopts the neural optimal transfer method. Optimal transfer is a theory that finds mappings between probability distributions under a given cost function; the present invention uses the square of the pixel Euclidean distance as the cost function. The double random matrix is initialized in the video memory, and then the Sinkhorn iteration is performed on the discrete grid, that is, the matrix is alternately normalized by row and column, accompanied by coefficient decay, to achieve approximate entropy regularized optimal transfer. The neural optimal transfer network expands the Sinkhorn iteration into trainable layers, uses synthetic regions of interest for self-supervision in the training phase, and only requires a small number of iterations to converge in the inference phase. After convergence, a displacement vector map of the corresponding pixels is obtained, which is called the first displacement field. It ensures mass conservation on a global scale, that is, there will be no holes or excessive overlap when the pixel block is enlarged.
[0037] However, the first displacement field only focuses on probability conservation and does not guarantee local angular relationships. In order to avoid texture distortion inside the region of interest, the present invention introduces a second displacement field. The second displacement field is obtained through quasi-conformal mapping: first, the Beltrami coefficient is calculated based on the grayscale value of the attention field. The Beltrami coefficient is used to measure the angular distortion of the mapping at each pixel. The closer the value is to 0, the closer it is to conformal. The Beltrami coefficient is used as input, and the Koebe iteration is used to solve the quasi-conformal mapping. The Koebe iteration fixes the boundaries of the screen so that the background can still maintain a complete envelope after mapping. After the iteration is completed, the second displacement field is obtained, which approximately maintains the angle inside the region of interest, reducing the problem of bending of vertical lines of buildings or stretching of character faces after magnification.
[0038] The first displacement field focuses on global mass conservation, while the second displacement field focuses on local angle preservation. The two have different goals. The present invention dynamically allocates fusion weights based on the Laplace curvature of the two displacement fields on the pixel grid. The Laplace curvature can approximately represent the deformation intensity of the displacement field in the local area. The greater the curvature, the more drastic the displacement field change. The curvature of the first displacement field is denoted as , the curvature of the second displacement field is , then the fusion weight is defined as: , the fusion formula is: ,in To unify the displacement field, is the first displacement field, is the second displacement field. Here Both refer to pixel displacement vector graphs. is a pixel-level scalar weight. Using pixel-level weights instead of a single global coefficient can make the weight transition smooth at the boundary and avoid folds in the displacement field at the edge of the region of interest.
[0039] After the uniform displacement field is generated, it is stored in a texture in video memory and directly sampled during the subsequent deformation rendering stage. In the rendering pipeline, the uniform displacement field is used to perform screen-space displacement on each vertex, magnifying the region of interest on the screen. Because the uniform displacement field is smoothed at the pixel level, the deformed background and region of interest remain continuous, and texture distortion within the region of interest is controlled.
[0040] Preferably, the first displacement field is obtained by performing Sinkhorn iteration on a discrete grid to solve the optimal transmission mapping between the source probability density and the target probability density, and the mapping matrix is row normalized and column normalized after each iteration.
[0041] After completing the generation of the attention field and obtaining the region of interest, the present invention needs to controllably enlarge the region of interest in the screen space without destroying the connectivity of the entire frame. To this end, this step first regards the entire attention field as a mass distribution, divides the pixel grayscale by the total grayscale to obtain the source probability density, and then constructs the target probability density using the statistical information of the region of interest. The construction process of the target probability density is as follows: extract the pixel centroid and covariance for each region of interest, write them into the pixel grid using a Gaussian kernel, then weight them according to the average grayscale of the region of interest, and finally normalize them. In this way, the source probability density and the target probability density are defined on the same grid, and the integrals of both are 1.
[0042] In order to accurately map the source probability density to the target probability density at the pixel level, the present invention adopts the entropy regularized optimal transmission theory. The core of optimal transmission lies in determining a mapping matrix, the rows and columns of which correspond to the pixel positions on the source grid and the target grid respectively, and the matrix elements represent the mass flow of a single pixel pair. The present invention selects the square of the pixel Euclidean distance as the unit flow cost, constructs a cost matrix and initializes a double random matrix with non-negative elements in the graphics processor memory; then the optimal solution is approximated through Sinkhorn iteration. Each iteration first normalizes the matrix by row and then by column, and the two steps are alternated to make the matrix gradually approach the optimal transmission mapping while ensuring that the row and column margins remain unchanged.
[0043] In order to facilitate end-to-end fine-tuning during the training phase, the present invention expands the Sinkhorn iteration into a differentiable network layer. The mapping matrix is , the row normalized vector is , the column normalized vector is , then a round of iteration can be abstracted as follows: ,in represents the cost matrix, is the regularization coefficient, Indicates writing the vector into a diagonal matrix. Row vector and column vector It is obtained by multiplying the mapping matrix by the target margin on the left or right and then taking the inverse. Afterwards, the displacement vector of each pixel can be extracted from the mapping matrix to form a first displacement field.
[0044] Since all operations are performed on the pixel grid and no vertex or fragment interleaving sampling is involved, the first displacement field is naturally aligned with the original frame resolution and satisfies global quality conservation: when the region of interest is magnified, its pixels come from the contraction of the background area rather than interpolation, avoiding blur rings or jagged gaps.
[0045] This invention expands the Sinkhorn iteration into trainable layers to reduce inference time. By synthesizing random regions of interest for self-supervised training, the network achieves near-90% row-by-column balancing accuracy after just three iterations, compared to the traditional 20-step iteration that consumes significantly more memory bandwidth. During inference, the source and target margin vectors are stored in the video memory at once, with linear bandwidth proportional to resolution, meeting the real-time requirements of 60 frames per second at 4K resolution.
[0046] Example 3: In a first-person shooter game with a resolution of 1920×1080, select The network was trained with 50,000 synthetic heatmaps for eight iterations during offline training. During online inference, three row-column normalizations were performed, taking 0.56 milliseconds. The generated first displacement field produces a screen-space displacement of approximately 50 pixels at the center of mass of the region of interest, which can enlarge the enemy's face from 24 pixels to approximately 40 pixels. At the same time, background pixels are automatically shrunk according to the row-column normalization of the mapping matrix, eliminating gaps between pixels. Compared to a scheme that uses only nearest neighbor interpolation for amplification, the former improves the peak signal-to-noise ratio by 2.4 decibels and reduces the edge misalignment rate by 38%.
[0047] In the mobile experiment, the resolution was reduced to 1280×720. The inference time was reduced to 0.28 milliseconds, and the mapping matrix occupied less than 12 megabytes of video memory, achieving coordinated region-of-interest magnification and background shrinkage. This result demonstrates that the present invention is feasible for solving the optimal transmission mapping using Sinkhorn iteration on a graphics processor, and achieves higher pixel consistency and real-time performance than traditional convolutional magnification or viewport cropping solutions.
[0048] Preferably, the second displacement field is obtained by calculating the Beltrami coefficient using the grayscale value of the attention field and maintaining the screen boundary identity mapping during the Koebe iteration process.
[0049] After achieving global mass conservation in the first displacement field, this invention introduces a second displacement field to reduce angular distortion within the region of interest. The core idea is to use quasi-conformal mapping to approximate local shape preservation, thereby preventing stretching of text, character faces, or building columns. The implementation process consists of three steps: pixel-level Beltrami coefficient calculation, Koebe iterative solution mapping, and boundary identity constraints to maintain panoramic connectivity.
[0050] The first step is to treat the attention field grayscale image as a scalar function. The higher the grayscale, the more the pixel needs to be magnified. Map the pixel coordinates to the complex plane and record each pixel point as a complex variable . Set the mapping function The control quantity is the Beltrami coefficient The Beltrami coefficient essentially measures the degree of angular distortion of the mapping at that point and is defined as: ,in and Respectively express and conjugation The complex partial derivative of . If Take 0, the mapping is completely conformal; when When it is close to 1, the mapping distortion is extremely large. In order to ensure that the interest area is flexibly enlarged and the background changes smoothly, the present invention normalizes the grayscale of the attention field to The interval is multiplied by a radius weight that increases in the region of interest and decreases at the edge to obtain the pixel level .
[0051] The second step is to use Koebe iteration to solve the quasi-conformal mapping. Koebe iteration updates the mapping function on the complex plane discrete grid so that the new mapping is effective for the Beltrami equation. Convergence. The iterative pseudo code can be expressed as: Initialization ; Perform a linear transformation on each pixel: ,in is the step size factor, Is the conjugate of the current gradient. After updating Perform Laplace smoothing to suppress high-frequency noise. Iterate 3 to 5 times to make the mapping gradient change Less than 0.5 pixels. Because the Koebe iteration relies only on local difference operators, it can be performed in parallel by compute shaders on the GPU, and its time complexity is linearly related to the number of pixels.
[0052] The third step is the boundary identity constraint: In order to prevent the screen frame from being torn due to internal deformation, the present invention fixes the mapping of the outermost pixel to be the same before each Koebe iteration, that is, ; At the same time, only the pixels within the boundary are updated during iteration. This ensures that the central area of interest can be deformed freely while the surrounding background remains coherent. Get the pixel displacement vector: , we get the second displacement field. Here and represent the real and imaginary parts of a complex number, respectively.
[0053] In a text reading game, an annotation was set as a region of interest. Using only the first displacement field for magnification resulted in curved text strokes. However, adding the second displacement field resulted in a unified displacement field that primarily applied a conformal component within the text, preserving the thickness of the strokes. This improved the reading comfort score from 3.2 to 4.5. The hardware environment was a desktop graphics processor with a resolution of 2560×1440. Four Koebe iterations took 0.62 milliseconds, representing 2.5% of the total rendering budget.
[0054] Another example is in a racing game. When a player turns a car at high speed, their area of interest falls on the entrance or exit of the curve ahead. The conformal component keeps the road texture straight, reducing speed-induced motion sickness. Compared to a solution without conformal components, players reported a 28% reduction in motion sickness scores in scenes with continuous curves. Mobile experiments reduced the number of iterations to 3 at a resolution of 1600×900, shortening the time to 0.35 milliseconds, while still achieving similar distortion reduction results.
[0055] By adjusting Koebe iterations at the pixel level using Beltrami coefficients, this method maintains local angle preservation in the region of interest while leveraging boundary identity to maintain overall screen coherence, providing a controllable conformal component to the unified displacement field. When combined with the global mass-conserving displacement field, this ensures full magnification of the region of interest while suppressing detail distortion, significantly improving the dynamic layout experience.
[0056] Preferably, the preset weights are allocated according to the Laplace curvatures of the first displacement field and the second displacement field on the pixel grid, and during fusion, the first displacement field and the second displacement field are linearly superimposed to generate a unified displacement field.
[0057] The unified displacement field must inherit the global mass conservation of the first displacement field and retain the local angular continuity of the second displacement field. The present invention uses pixel-level weighted linear superposition to achieve the synergy between the two displacement fields, where the weights are adaptively allocated based on the Laplace curvature of the two displacement fields on the same pixel grid. The Laplace curvature measures the deformation intensity in the displacement vector field, and its essence is the divergence of the second-order difference to the displacement gradient. Assume that the first displacement field is at pixel The horizontal and vertical components of and , the corresponding component of the second displacement field is recorded as and The approximation of the horizontal component by the discrete Laplace operator can be written as:
[0058] Doing the same operation on the vertical component yields The curvature at a pixel is defined as:
[0059] The curvature of the second displacement field The same formula is used for calculation. To ensure that the contribution of the two displacement fields to the unified displacement field varies smoothly with the local deformation intensity, the present invention uses bilateral filtering in the graphics processor to perform a spatial smoothing of the curvature map, preserving the edges to avoid sharp jumps at the boundaries of the region of interest. The pixel-level weight is given by the inverse curvature form:
[0060] in To prevent the denominator from being a very small constant of 0. The closer it is to 1, the greater the contribution of the second displacement field to the pixel; the closer it is to 0, the more dominant the first displacement field is. Substitute linear superposition:
[0061]
[0062] Get the horizontal and vertical components of the unified displacement field 、 This weighting method usually shows a distribution with large curvature of the first displacement field and small curvature of the second displacement field within the region of interest. Biased to 0.6 to 0.8, the angle component is dominant; in the background with fine texture, the curvature of the two fields is close, About 0.5, maintaining mass conservation and angle balance.
[0063] After the uniform displacement field is generated, pixel-level noise is likely to appear. The present invention adds a cubic spline filter to the graphics processor: 、 A three-point sliding average is performed along the rows and columns, followed by a bidirectional interpolation. This operation does not violate mass conservation or conformal properties, but significantly reduces jagged edges. The filtered displacement field is written back to video memory as a dual-channel 16-bit floating-point texture for subsequent mesh shader sampling. Since all calculations involve local convolutions and linear combinations, the time complexity increases linearly with pixels, taking approximately 0.21 milliseconds at a 1920×1080 resolution on desktop and approximately 0.11 milliseconds at a 1280×720 resolution on mobile.
[0064] Example 4: In a text adventure game, when the player clicks the screen, the attention field generates a weighted peak in the dialog area. First-order curvature analysis shows that the average curvature of the first displacement field at the dialog edge is 3.4 pixels, and the average curvature of the second displacement field is 1.1 pixels, with a generated weight distribution between 0.24 and 0.38. Unifying the displacement field magnifies the dialog box by approximately 1.4 times without tilting the glyphs. Compared to a solution without curvature weighting, the error in edge stroke thickness increase is reduced from 21% to 7%.
[0065] Example 5: In a high-speed cornering scene in a racing game, the area of interest falls on the inner curve ahead. The curvature peak of the first displacement field reaches 7.8 pixels, while the second displacement field peaks at 2.0 pixels. The curvature weight reaches 0.78 at the center pixel and drops to 0.46 at the edge pixels. The rendered shoulder remains straight, with no texture distortion. When using only the first displacement field, the shoulder appears noticeably wavy. Players' vertigo scores decreased by 25% across 10 consecutive cornering tests.
[0066] Using the unified displacement field to deform pixel coordinates, the voxel set corresponding to the region of interest is input into a neural radiation field network to obtain a compensated color map, and the compensated color map is synthesized with the deformed color map to form a fused image; After generating a unified displacement field, the present invention directly samples the displacement texture during the mesh shading phase, performing screen-space deformation on each vertex. The deformed pixel coordinates are then used in subsequent color synthesis. To avoid jagged edges, blurring, or loss of detail after zooming in on the region of interest, the present invention does not simply use the original color map. Instead, it introduces a neural radiance field network to re-render the voxel set corresponding to the region of interest, generating a compensated color map. This is then overlaid pixel-by-pixel with the deformed color map to generate a fused image.
[0067] First, let's introduce the principles of voxel set construction. The pre-deformed region of interest is a set of pixel indices in screen space; by back-projecting these pixels into the camera coordinate system using the depth map, a 3D point cloud is obtained. To efficiently organize data in video memory, this paper uses a hash grid to map 3D coordinate blocks into hash buckets, each storing the density and color parameter indices of a fixed-size voxel block. This approach enables one-time parallel writes on the GPU side, avoiding thread divergence caused by pointer jumps.
[0068] The voxel set is then fed into the neural radiation field network. The neural radiation field is an implicit field that simultaneously predicts spatial density and direction-dependent radiance through a multi-layer fully connected network. The present invention uses an 8-layer fully connected structure with 128 hidden units in each layer. In order to improve the convergence speed, the three-dimensional position coordinates and the line of sight direction are Fourier encoded separately to make the high-frequency geometric details easier to be fitted by the network. The input ray uniformly samples 48 depth points within the voxel block, and the network outputs the voxel density of each sampling point as , and the color vector corresponding to the viewing direction is recorded as According to the volume rendering formula:
[0069]
[0070] The pixel-level compensation color can be accumulated, where is the distance between adjacent samples. represents the spatial density of sampling points, Indicates the radiation color of the sampling point, Represents the weight of the sample's contribution to line-of-sight transmittance. The entire ray summation process is implemented in the fragment shader using parallel reduction. The computation latency is linearly proportional to the number of samples, and takes an average of 0.74 milliseconds on desktop.
[0071] The compensating color map is synthesized with the deformed color map following the "region of interest priority" principle: if a pixel falls within the deformed region of interest on screen, the compensating color is used; otherwise, the deformed color is retained. To ensure a smooth transition, a 4-pixel-wide blending band is inserted at the edge of the region of interest. The compensating color and the original color are blended using a bilinear weighting to avoid hard-edge interfaces.
[0072] Example 6: In a close-up dialogue scene in a role-playing game, the character's face is set as the region of interest. The unified displacement field is magnified by approximately 1.8 times. After neural radiation field compensation, the facial skin texture retains high-frequency details, and the hair edges are smooth. Compared with bilinear amplification using only the original color map, the peak signal-to-noise ratio is improved by 2.7 decibels, and the structural similarity index is improved by 0.06. When the mobile resolution is 1280×720, the voxel block side length is set to 32, the number of sampling points is reduced to 32, the compensation delay is controlled at 0.38 milliseconds, and the total rendering frame rate is stabilized at 60 frames. In the visual perception test, the average clarity score given by 20 players in the fast zoom scene increased by 27%.
[0073] Preferably, the voxel set corresponding to the region of interest is obtained by back-projecting the pixel coordinates into three-dimensional coordinates and writing them into a hash grid, and the three-dimensional coordinates and the light direction are Fourier encoded respectively as inputs of the neural radiation field network.
[0074] After constructing the unified displacement field, the present invention uses a neural radiance field network to resynthesize the color of the region of interest to avoid blurry or jagged textures caused by magnification of the region of interest. For the neural radiance field to work in real time, the primary task is to accurately map screen space pixels to three-dimensional space, compress them into video memory as sparse voxels, and efficiently encode them as network input.
[0075] First, the back-projection principle is explained. The rendering pipeline outputs a depth value for each fragment in the rasterization stage. The depth value combined with the known camera intrinsic parameters can restore the spatial position of the fragment in the camera coordinate system. Assume that the screen resolution is , the pixel coordinates are , the back projection formula is:
[0076] in is a three-dimensional coordinate vector, is the depth map value, for Internal parameter matrix, is the inverse matrix of the internal reference. The pixels in the region of interest can be instantly converted into a three-dimensional point cloud through parallel back projection. Next, the hash grid organization is introduced. The number of pixels in the region of interest can reach hundreds of thousands. If the voxels are directly stored in the form of an array, it will cause video memory fragmentation and random access delay. The present invention adopts a fixed voxel side length , divide the three-dimensional space into cubic grids, and index the grid Use a multiplicative mixing hash function: ,in is bitwise exclusive OR, To shift left, is the number of buckets. Hash collisions are resolved using open addressing, and the length of the collision chain cannot exceed 3. Each bucket stores the density and color index of a voxel block. When the region of interest changes, only the corresponding point cloud in the hash bucket needs to be updated, avoiding global reconstruction. Experiments show that at a resolution of 1920×1080, using a hash grid with 32×32×32 voxel blocks can keep the voxel storage size within 24 megabytes.
[0077] The third step is Fourier encoding. The neural radiation field network uses a multi-layer fully connected structure. It has been verified that low-frequency input often makes it difficult to learn high-frequency details. For this reason, the present invention uses the three-dimensional position vector Perform multi-frequency Fourier mapping: . At the same time, the sight direction unit vector Do the same encoding, where The present invention sets , meaning the input dimension increases from 3 to 36, but thanks to the use of half-precision floating-point storage, the memory overhead remains manageable. Fourier encoding improves the network's ability to fit high-frequency geometric details, which is particularly important when amplifying texture edges and lines.
[0078] The fourth step demonstrates the real-time rendering effect. After being organized into a hash grid, the voxel set in the region of interest, along with the encoding vector, is fed into a neural radiance field network for inference. The network structure consists of 8 fully connected layers, with 128 hidden units per layer. Residual connections are used to mitigate vanishing gradients. Ray walking uses uniform sampling for a total of 48 depth points. Density and color are accumulated using the volume rendering formula to obtain pixel-compensated colors, which are then overlaid on the deformed color map.
[0079] Example 7: In a role-playing game, the player clicks to zoom in on the character's head. The number of voxels in the interest area is approximately Hash grid allocation takes 0.12 milliseconds, encoding and network inference take 0.62 milliseconds, and the overall latency is less than 1 millisecond. Hair and facial fine lines remain sharp in the fused image, and the peak signal-to-noise ratio is improved by 2.8 decibels compared to the bilinear amplification solution.
[0080] Example 8: In a cartoon-rendered shooter game, a player quickly swipes the touchscreen to zoom in on a target. With a mobile resolution of 1280×720 and a voxel block edge length of 16, the number of ray sampling points is reduced to 32, resulting in a total inference latency of 0.35 milliseconds. Because the number of hash bucket collision chains is limited to 2, the random access memory hit rate remains at 93%. Players' subjective ratings show no noticeable jagged edges on the zoomed-in target, and shooting accuracy is improved by 17%.
[0081] like Figure 3 As shown, for the user input event on the fused screen, the original screen coordinates are determined by inverse operation of the unified displacement field and the second displacement field, the entity identification is obtained by querying the deformed entity identification map, and when the system energy exceeds the preset range, the attention field threshold, the second displacement field parameters and the damping parameters are adjusted to achieve cross-frame adaptive feedback.
[0082] After the fused image is generated, all player input events (including mouse clicks, touch taps, and controller cursor hovers) occur in the deformed two-dimensional coordinate system. Directly feeding these coordinates into traditional hit detection results in a disconnect from the logical world: the region of interest is magnified by the unified displacement field, the background is compressed, and the entity's true hitbox is misaligned with the screen pixels. This invention uses a two-level inverse mapping to restore the original screen coordinates of the event, then performs entity queries to ensure interactive accuracy. Furthermore, the system's closed-loop energy system adaptively adjusts the layout parameters for the next frame, ensuring visual comfort and stable computing power.
[0083] The first-level inverse mapping uses a unified displacement field. During the mesh shading phase, the present invention writes the unified displacement field into a dual-channel 16-bit floating-point texture, recording the horizontal and vertical displacement vectors for each pixel. When an event occurs, the event screen coordinates are used as the sampling center, and the displacement is extracted using the graphics processor's bilinear interpolation. Coordinate subtraction is then performed to obtain the first step of the inverse displacement. Because the unified displacement field is cubic spline smoothed during generation, sampling noise is constrained to within 0.3 pixels. Bilinear interpolation is hardware-accelerated on the graphics processor's native hardware unit and can be completed within 0.01 milliseconds.
[0084] The second-order inverse mapping uses the second displacement field. The second displacement field is generated by quasi-conformal mapping, which maintains local angle continuity. After the first-order inverse displacement, the inverse function of the quasi-conformal mapping is called on the pixel coordinates. The quasi-conformal inverse mapping is essentially the solution of the composite function. While the Koebe iteration determines the forward mapping, the present invention simultaneously generates a sparse lookup table: the screen is divided into a 64×64 grid, and the initial value of the reverse mapping is stored at the center of each grid. When the event coordinates fall into a grid, the initial value is retrieved and two steps of Newton correction are iterated to converge, with an average time of 0.05 milliseconds. The combined error of the two-stage inverse transformation has been measured to be no more than 0.5 pixels.
[0085] After obtaining the original screen coordinates, the present invention searches for entities in the deformed entity identification map. The entity identification map is deformed synchronously with the color map, and the pixel values store the entity ID. A single byte can cover 256 entity categories. Using the event coordinate index map, the ID is returned in constant time. Events are then packaged as <timestamp, entity ID, event type, and accompanying parameters> and pushed to the logic thread. Processing by the logic thread affects the scene state, thereby influencing the generation of the attention field and displacement field for the next frame.
[0086] In order to prevent visual fatigue caused by excessive deformation, the present invention calculates the system energy at the end of the logical thread. The system energy consists of two items: the change in the attention field and the curvature of the uniform displacement field. The attention field of the current frame is recorded as , the attention field of the previous frame is , the uniform displacement field curvature is ,but:
[0087] in 、 is a constant coefficient. is the system energy scalar, is the sum of absolute differences of pixels, The sum of the absolute values of the curvatures is expressed in pixels. Energy measures the severity of the visual deformation and the jitter of the attention distribution. Taking a 1-norm for these two allows for fast parallel summation.
[0088] The present invention sets an energy upper limit With lower limit .when If the value is higher than the upper limit, it indicates that the deformation amplitude is too large or the attention field jump is drastic, which may cause dizziness. If the value is lower than the lower limit, it indicates that the image is not enough to highlight the interactive focus and the interest area effect is attenuated. In order to maintain a dynamic balance between the two thresholds, the present invention uses the temporal difference reinforcement learning algorithm to adjust the attention field threshold, the second displacement field Beltrami coefficient ratio and the damping parameter. The reinforcement learning state vector is taken as , the action vector is the parameter increase or decrease step size. Objective function: ,in The average of the upper and lower bounds is shown. The network uses a two-layer perceptron for policy approximation, with 512 parameters, which does not significantly increase memory usage. Updates are performed online per frame, with the learning rate decaying exponentially after 500 frames. The measured desktop update time is 0.03 milliseconds.
[0089] Preferably, when determining the original screen coordinates, the screen space inverse displacement is first performed using the unified displacement field, and then a quasi-conformal inverse mapping is performed using the second displacement field, wherein the inverse displacement is calculated by bilinear interpolation.
[0090] After the fused image is generated, all user input events—whether mouse clicks, touch taps, or controller cursor drops—are located in the deformed screen coordinate system. If these coordinates are used directly for hit detection, the logic layer will misjudge the input location because the area of interest is magnified and the background is compressed. To ensure interactive accuracy, the present invention employs a "two-stage inverse mapping" process: first, screen-space inverse displacement is performed using a unified displacement field, followed by inverse mapping using a second displacement field (quasi-conformal mapping). After these two inverse operations, the original screen coordinates corresponding to the event are obtained. These coordinates are then combined with the deformed entity identification map to determine a unique entity number, ultimately forming a stable and precise interactive loop.
[0091] The first-level inverse mapping is based on a unified displacement field. The unified displacement field is stored in video memory as a dual-channel 16-bit floating-point texture, with each pixel storing horizontal and vertical displacements. When an event arrives, the event screen coordinates are used as the sampling center, and the GPU hardware bilinear interpolation function is called to obtain the displacement vector. This vector is then directly subtracted from the event coordinates to obtain the first step of the inverse displacement. Because the unified displacement field undergoes cubic spline filtering during the generation phase, its gradient is smooth and the bilinear interpolation error is less than 0.3 pixels. The hardware only requires four texture reads to perform one interpolation, which takes approximately 0.01 milliseconds, meeting the requirements of high-frame-rate games.
[0092] The second inverse mapping depends on the second displacement field. The second displacement field is generated by quasi-conformal mapping, which keeps the local angle continuous, but its forward function It is not always easy to find an explicit inverse function. To ensure real-time performance, the present invention generates a sparse reverse lookup table synchronously when Koebe iterates to find the forward mapping: the screen is divided into 64×64 grids, and the screen coordinates after forward mapping are stored for each grid center pixel. After the event coordinates fall into a certain grid, the nearest table entry is selected as the initial value of the inverse solution, and two iterations of Newton correction are used to converge. The Newton correction step only uses The numerical gradient of the nearest neighbor can be obtained through differential approximation without adding complex analytical derivatives. This inverse solution takes an average of 0.05 milliseconds and has an error of no more than 0.2 pixels.
[0093] Original screen coordinates Once determined, the system queries the pixel value in the deformed entity identification map to obtain the entity ID. The entity identification map and color map are deformed using the same vertex shader to ensure consistent indexing. Pixels store 8-bit integer IDs, covering 256 entity categories. Table lookups are performed in the GPU texture unit, with a single access latency of less than 100 nanoseconds. At this point, the event is packaged as <timestamp, entity ID, event type, and accompanying parameters> and pushed to the logical thread.
[0094] In order to prevent visual fatigue caused by excessive deformation, the present invention evaluates the system energy at the end of the logic thread. The system energy is composed of the weighted two parts: the attention field frame difference and the uniform displacement field curvature. Assume that the attention field of this frame is , the attention field of the previous frame is , then the frame difference is defined as The uniform displacement field curvature is calculated by discrete Laplace calculation of four neighborhoods and then summing up the absolute values. Energy calculation formula: ,in and is a fixed coefficient. is the system energy scalar, and All are calculated on the pixel grid Norm. Energy measures the intensity of visual distortion and attention distribution jitter. A larger value is more likely to trigger player discomfort.
[0095] The present invention sets an upper limit With lower limit .when When the upper limit is exceeded, the system considers that the visual load is too high; when When the value is lower than the lower limit, the system considers that the image lacks a prominent focus. The temporal difference reinforcement learning algorithm is used to adjust three key parameters: the attention field threshold (affecting the size of the area of interest), the second displacement field Beltrami coefficient ratio (affecting the angle maintenance strength), and the damping coefficient (affecting the smoothness of the displacement field interpolation). The state vector is taken , the action vector is the three-parameter increase or decrease step. The reward is set as: ,in The two-layer perceptron policy network outputs actions in the logic thread. The learning rate is updated online; the network contains only 512 weights, which does not significantly increase the CPU / GPU burden. In practice, convergence is achieved within 300 frames online.
[0096] Example 9: A desktop first-person shooter with a resolution of 1920×1080 and a refresh rate of 144 frames per second. During rapid gun swings, system energy surged to 1.5 times the upper limit. Reinforcement learning reduced the Beltrami coefficient by 15% and increased damping by 10%, bringing the energy down to the safe zone for the next frame. Compared to a fixed-parameter solution, the stun score decreased by 0.8 points and the hit rate increased by 5%.
[0097] Example 10: A mobile card game was left idle for an extended period. System energy remained below the lower limit for 50 consecutive frames. The policy network increased the region of interest threshold by 8%, reduced damping by 12%, magnified card details, and improved reading comfort by 23%. All parameter adjustments took 0.03 milliseconds and did not affect 60-frame rendering.
[0098] Preferably, the system energy is obtained by summing the attention field variation and the uniform displacement field curvature on the same pixel grid and then weighting them according to a fixed coefficient.
[0099] After the unified displacement field and attention field are laid out, the present invention uses system energy to quantify the visual load of the current frame, thereby driving adaptive parameter adjustment for the next frame. System energy measures the "dynamic complexity" of the image from two dimensions: the first is the amount of change in the attention field, reflecting the spatial fluctuation of the image's focal point; the second is the sum of the curvatures of the unified displacement field, reflecting the strength of the screen's geometric deformation. These two are summed across the pixel grid and weighted by a fixed coefficient to produce a frame-level scalar. This scalar provides a concise assessment of dizziness risk and facilitates the use of reinforcement learning algorithms within millisecond time windows.
[0100] The change in the attention field directly corresponds to the displacement speed of the "area of attention the player should focus on." If interest is focused on the center of the screen in the previous frame and suddenly shifts to the edge of the screen in the next frame, the absolute value of the grayscale difference in the attention field is large, indicating a dramatic jump in the center of visual attention. Extensive experiments have shown that such jumps are more likely to trigger motion sickness when accompanied by significant geometric deformation. The unified displacement field curvature approximates the second-order derivative of the displacement vector field using the discrete Laplace operator: high curvature indicates large differences in deformation gradients between pixels, indicating significant local stretching or contraction. An increase in the sum of curvatures within a frame indicates increased local bending of the screen or texture compression, which also increases visual load.
[0101] To avoid the imbalance of absolute values as the resolution changes, the present invention uniformly calculates two quantities at the final rendering resolution: the attention field change quantity takes the absolute difference of pixels and calculates the 1-norm; the curvature sum calculates the discrete Laplace of the horizontal and vertical components of the displacement respectively, takes the absolute value and then calculates the 1-norm.
[0102]
[0103]
[0104] On the graphics processor, both can be read and written in one pass with the help of parallel protocol, without multiple memory round trips. and The combined system energy is: ,coefficient 、 A balance is achieved through offline calibration: if the game type emphasizes rapid shooting, the attention field weight is increased; if the focus is on scene roaming, the curvature weight is increased to ensure that the two values are of the same magnitude. To pay attention to the grayscale field, 、 is the displacement component Laplace, and Two energies respectively.
[0105] In terms of implementation, the present invention first performs two texture reads within the GPU's compute shader: one for the current frame's attention field and one for the previous frame's attention field. After differencing, the fields are summed in parallel using absolute value instructions. Next, the unified displacement field texture is read, and a five-point template is applied to both the horizontal and vertical discrete Laplacians. The absolute values are then taken and accumulated. The two reductions are performed using shared memory for block summation, with the block size typically set to 16×16 to balance latency and occupancy. At a resolution of 1920×1080, the two-step accumulation takes approximately 0.08 milliseconds.
[0106] After obtaining the system energy, it is compared with fixed upper and lower limits. If the energy falls within the interval, the existing parameters remain unchanged; if it exceeds the upper limit or falls below the lower limit, the reinforcement learning policy network is triggered and the parameter adjustment action is output. The action space designed by this invention includes: lowering or raising the attention field threshold, scaling the Beltrami coefficient, and adjusting the damping coefficient. The amplitude of each action is ±5% of the current value to ensure frame-by-frame continuity. The policy network learns online, and the reward is set as a negative value of the energy distance from the center of the interval to encourage the energy to quickly return to the interval.
[0107] Example 11: In a high-speed shooting game, the player's continuous shooting causes the viewing angle to rotate at a speed of 720 degrees per second. This scene causes the attention field peak to jump from the left screen to the right screen within 5 frames. Soared to 8.3×10×4; at the same time, the curvature of the displacement field also increased to 3.9×10×4 due to the continuous enlargement of the area of interest. The system's energy exceeded the upper limit by 1.2 times. Reinforcement learning instantly outputted the action of "reducing the Beltrami coefficient by 10% and increasing damping by 5%." The energy dropped by 23% in the next frame, returning to the center of the range in the third frame. Compared to the no-feedback version, the player's dizziness score decreased by 0.7 points, and their shooting accuracy increased by 4.5%.
[0108] Example 12: In the static card drawing interface of the card game, the player's fingertip only swipes across the card deck to browse. Note that the field change amount is stable at , the sum of curvatures , the energy remained below the lower limit for 40 consecutive frames. After the network output the action "Increase the region of interest threshold by 12% and reduce damping by 8%," the card's zoom range increased, the energy returned to the center of the range, and players' scores on the card's font readability increased by 22%.
[0109] This invention enables system energy to simultaneously account for both local deformation and global focus jumps, measuring potential visual fatigue within a single frame. This pure pixel reduction operation eliminates the need for global sorting or complex optimization algorithms, making it GPU-friendly. Combined with reinforcement learning, parameter convergence occurs when energy exceeds bounds, maintaining a comfortable experience for both intense combat and quiet viewing. Computational time measured at 4K resolution is 0.24 milliseconds; at 1280×720 on mobile, it takes 0.05 milliseconds, making it easy to deploy within 60 and 144 FPS rendering budgets.
[0110] Preferably, when adjusting the attention field threshold, the second displacement field parameter and the damping parameter, a temporal difference reinforcement learning algorithm is used, with the goal of reducing the difference in system energy between two adjacent frames.
[0111] In continuously rendered scenes, sudden changes in the uniform displacement field or attention field parameters can easily cause dizziness or the illusion of frame skipping. This invention abstracts the image state into "system energy" and uses a temporal difference reinforcement learning algorithm to adjust parameters online within milliseconds, minimizing the difference in system energy between two adjacent frames to zero. This results in a stable, comfortable, and focused dynamic layout.
[0112] State design and energy difference calculation, system energy The above step defines the energy difference as the linear weighted result of the attention field frame difference term and the uniform displacement field curvature term according to a fixed coefficient. The higher the value, the greater the visual load of the frame. In order to monitor energy continuity, the present invention further defines the energy difference: , where the positive and negative values correspond to load increase or decrease respectively. The environment state vector of temporal difference learning is: ,in Represents the absolute value of the uniform displacement field curvature and is used to reflect the total amount of geometric deformation. The three components are normalized to , ensuring numerical stability at different resolutions.
[0113] Action space and parameter adjustable range, action vector Corresponding to three adjustable parameters: 1. Attention field threshold ——Determine the size of the region of interest; 2. Second displacement field Beltrami coefficient scaling factor ——Determines the local conformal strength; 3. Damping coefficient ——Determines the displacement field cubic spline filtering amplitude. Each action can be Three values, Fixed to 5% of the current parameter. , which facilitates fast traversal while retaining sufficient adjustment granularity.
[0114] Reward function and target: In order to make the energy difference converge, the present invention sets the reward as the negative absolute value of the energy difference:
[0115] If the energy difference between adjacent frames approaches zero, the reward is greater; if the energy fluctuates, the reward is smaller, and reinforcement learning will actively seek strategies to reduce the fluctuation.
[0116] Learning rules and core formulas, using temporal difference SARSA Update the value function. Learning rate Take 0.05, the discount factor Take 0.95, the trace attenuation coefficient Take 0.6. The core update formula is:
[0117] in Indicates that the status Next action The expected cumulative reward. is the current state, For the current action, For instant rewards, is the value function. The value function is approximated by a two-layer perceptron with a hidden layer width of 64 and a total of 512 parameters. It is stored in the GPU constant cache and has an inference latency of less than 0.02 milliseconds.
[0118] The online inference and execution process after each frame is as follows: 1. Calculate the system energy of the new frame and energy difference 2. Construction status , calling the policy network according to -Greedy output action ; Start at 0.2, and decay to 0.05 after 500 frames. 3. Adjust by action And immediately apply it to the next frame layout pipeline. 4. Record instant rewards , update the value function weight using the formula.
[0119] Example 13: Desktop first-person shooter, 1920×1080 resolution, 144 fps. Rapidly flicking the gun causes the initial energy difference to peak at 1.5, with an upper limit of 1. After 300 frames online, the absolute value of the energy difference drops to around 0.2 and stabilizes, and the player's subjective sickness score drops from 3.8 to 3.1.
[0120] Example 14: Mobile role-playing game, 1280×720 resolution, 60 frames per second. Prolonged static scenes caused the energy difference to remain below the lower limit of 0.2. The learning algorithm automatically increased the region of interest threshold and reduced damping, returning the energy difference to 0.3 and improving the card text clarity score by 25%. Overall device power consumption remained unchanged, with the inference phase power consumption increasing by less than 0.5 watts.
[0121] Preferably, when the system energy is lower than a preset lower limit and the change in the attention field is lower than a preset threshold, the unified displacement field of the previous frame and the second displacement field of the previous frame are reused, and the steps of determining the area of interest based on the attention field and deforming the pixel coordinates using the unified displacement field are skipped.
[0122] System Energy and attention field variation When the image quality is low, it indicates that there is no significant visual change between consecutive frames, nor is there any new interactive focus that needs to be highlighted. In this situation, the present invention triggers a "frame skipping and reuse" mechanism: the unified displacement field and the second displacement field generated by the previous frame are directly used, omitting computationally intensive steps such as region of interest extraction, probability mapping, and pixel deformation, thereby reducing GPU load and stabilizing the output frame rate.
[0123] The trigger condition includes two threshold judgments. The first is the system energy threshold The system energy is obtained by weighting the attention field frame difference term and the displacement field curvature term. When its value is less than 0.000 for two consecutive frames, the system energy is obtained by weighting the attention field frame difference term and the displacement field curvature term. When , it can be inferred that the current picture is basically still. The second item is the attention field change threshold , in absolute difference Quantization, if the difference is lower than , indicating that the focus center has not moved. When both conditions are met, the frame skip flag is set to true and the layout pipeline enters simplified mode.
[0124] In simplified mode, the main rendering thread no longer calls the subroutine that determines the region of interest based on the attention field, nor does it perform neural optimal transfer mapping, quasi-conformal mapping, and post-cubic spline filtering. Instead, the displacement texture written to the video memory in the previous frame is directly bound to the mesh shader, leaving the screen deformation logic unchanged. This avoids unnecessary floating-point convolutions and iterative operations, as well as cache invalidations caused by frequent texture rewrites. According to statistics, omitting these steps can reduce the computational latency per frame by approximately 1.4 milliseconds at the same resolution.
[0125] To prevent "aging" errors caused by frame skipping in long static scenes, the present invention maintains a soft timeout even after entering simplified mode: if the number of consecutive multiplexed frames exceeds 60, a full layout step is forced to refresh the displacement field and ensure realignment with any possible minimal scene drift. Furthermore, if the system energy or attention field changes exceed their respective thresholds in any frame, the frame skip flag is immediately cleared, and full calculation resumes for the next frame, ensuring a delay-free response in dynamic scenes.
[0126] Example 15: In a static room scene of a puzzle game, the player stops to observe the wall prompts. Since the character does not move and the camera does not shake, the system energy remains at 0.12, which is lower than , note that the field frame difference is kept at 0.017, which is lower than The system enters frame skipping mode starting from the 5th frame, and the average frame delay is reduced from 6.4 milliseconds to 4.9 milliseconds, energy consumption is reduced by 9%, fan speed is reduced, and noise is reduced.
[0127] Example 16: In the waiting lobby of a multiplayer shooter game, a character was standing still. After 48 frames of frame skipping, a teammate suddenly jumped into the center of the shot, causing the peak field of attention to shift, and the frame difference soared to 0.18, exceeding the threshold. The layout pipeline automatically resumed the full calculation in the next frame, regenerating a new displacement field to ensure that the teammate's model was scaled up and correctly aligned. The observed additional latency was only 1 frame, and the player did not perceive any lag.
[0128] like Figure 4 As shown, a system for dynamic interactive layout of game screens is used to implement the dynamic interactive layout method of game screens, and the system includes: The attention field generation module is used to obtain the depth map, normal map, illumination map, color map and user input event data in each frame, and input the data into the neural network and combine it with the symplectic integral to generate an attention field consistent with the screen resolution; this module mainly runs on the graphics processor's deep learning core and parallel computing unit: the rendering pipeline caches the depth, normal, illumination, and color in the frame buffer during the pixel shading stage, and directly accesses the video memory through PCIeBAR0; the multi-head self-attention network is inferred on the tensor core, using high-bandwidth video memory cross-strips to ensure aligned reading of 8-dimensional inputs; the symplectic integral prediction is completed by the parallel computing shader, and shared memory is used to store gradients and momentum vectors for 32×32 thread blocks, with a single integration delay of less than 0.3 milliseconds.
[0129] The displacement fusion module determines the region of interest based on the attention field, constructs source and target probability densities, and generates a first displacement field and a second displacement field, respectively. The first and second displacement fields are then fused according to preset weights to generate a unified displacement field. This module utilizes the GPU compute shader to perform Sinkhorn and Koebe iterations. The source and target probability densities are stored as half-precision textures, and row-column normalization is performed using the TensorCore warp-matrix-multiply instruction. Beltrami coefficients and Laplace curvature are rapidly derived using the GPU's native texture differential instructions. Curvature weight calculations are parallelized in shared memory, and the fused result is written back to a dual-channel FP16 displacement texture that can be sampled intact in subsequent rendering stages.
[0130] The compensation rendering module uses the unified displacement field to deform pixel coordinates. The set of voxels corresponding to the region of interest is input into a neural radiance field network to generate a compensated color map. The compensated color map is then combined with the deformed color map to create a fused image. The voxels of interest are directly written to a hash grid buffer by the rasterizer unit after depth backprojection. The hash table resides in video memory and is indexed using atomically locked buckets. The neural radiance field network performs batch inference on the Tensor Cores. Three-dimensional coordinates and view vectors are Fourier encoded in the constant cache before being fed into the multiply-accumulate pipeline. Ray integration utilizes the high-concurrency sampling of the ray tracing hardware unit. Compensated and deformed colors are quickly interpolated according to masks at the fragment stage via a blending unit, resulting in lower bandwidth consumption than traditional off-screen compositing.
[0131] The interactive feedback module is used to determine the original screen coordinates of user input events on the fused screen by sequentially performing inverse operations on the unified displacement field and the second displacement field, query the deformed entity identification map to obtain the entity identification, and adjust the attention field threshold, the second displacement field parameters, and the damping parameters when the system energy exceeds a preset range to achieve cross-frame adaptive feedback. The module is deployed collaboratively between the CPU and the GPU: input events are written to a circular buffer via the interrupt controller, and the GPU texture unit reads the unified displacement and quasi-conformal lookup table to achieve a two-level inverse mapping. After sampling, the entity identification is immediately sent back to the logic core via NVLink zero-copy. System energy and curvature are obtained once in the compute shader, and the reinforcement learning policy network resides in video memory as a CUDA graph. Only 512 weights are forwarded once in the tensor core. If frame skip reuse is triggered, the previous frame texture is directly reused via the GPU event fence, eliminating data movement and iterative calculations, ensuring hardware pipeline continuity and controllable latency.
[0132] The above are merely embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should be included within the scope of the claims of the present application.
Claims
1. A method for dynamic interactive layout of game screens, characterized in that: The following steps are involved: Each frame acquires depth map, normal map, illumination map, color map and user input event data, inputs the acquired data into the neural network and combines it with symplectic integral to generate attention field; Determine the region of interest based on the attention field, construct a source probability density and a target probability density, obtain a first displacement field and a second displacement field respectively, and fuse the first displacement field and the second displacement field according to a preset weight to generate a unified displacement field; Using the unified displacement field to deform pixel coordinates, the voxel set corresponding to the region of interest is input into a neural radiation field network to obtain a compensated color map, and the compensated color map is synthesized with the deformed color map to form a fused image; For user input events on the fused screen, the original screen coordinates are determined by inverse operation of the unified displacement field and the second displacement field, the entity identification is obtained by querying the deformed entity identification map, and when the system energy exceeds a preset range, the attention field threshold, the second displacement field parameters and the damping parameters are adjusted to achieve cross-frame adaptive feedback.
2. The method according to claim 1, characterized in that When generating the attention field, the neural network is constructed in sequence by a multi-head self-attention layer and a feedforward layer, and the feature vectors of the depth map, normal map, illumination map and color map are included in the same attention window in the multi-head self-attention layer.
3. The method according to claim 1, characterized in that The first displacement field is obtained by performing Sinkhorn iteration on a discrete grid to solve the optimal transmission mapping between the source probability density and the target probability density, and the mapping matrix is row-normalized and column-normalized after each iteration.
4. The method according to claim 1, wherein The second displacement field is obtained by calculating the Beltrami coefficient using the grayscale values of the attention field and maintaining the screen boundary identity mapping during the Koebe iteration.
5. The method according to claim 1, wherein The preset weights are allocated according to the Laplace curvature of the first displacement field and the second displacement field on the pixel grid, and the first displacement field and the second displacement field are linearly superimposed during fusion to generate a unified displacement field.
6. The method according to claim 1, characterized in that The voxel set corresponding to the region of interest is obtained by back-projecting the pixel coordinates into three-dimensional coordinates and writing them into a hash grid, and the three-dimensional coordinates and light direction are Fourier encoded respectively as the input of the neural radiation field network.
7. The method according to claim 1, characterized in that When determining the original screen coordinates, the screen space inverse displacement is first performed using a unified displacement field, and then a quasi-conformal inverse mapping is performed using a second displacement field, where the inverse displacement is calculated using bilinear interpolation.
8. The method according to claim 1, characterized in that The system energy is obtained by summing the attention field variation and the uniform displacement field curvature on the same pixel grid and weighting them according to a fixed coefficient. When adjusting the attention field threshold, the second displacement field parameters, and the damping parameters, a temporal difference reinforcement learning algorithm is used to reduce the difference in system energy between two adjacent frames.
9. The method according to claim 1, characterized in that When the system energy is lower than the preset lower limit and the change in the attention field is lower than the preset threshold, the unified displacement field of the previous frame and the second displacement field of the previous frame are reused, and the steps of determining the region of interest based on the attention field and deforming the pixel coordinates using the unified displacement field are skipped.
10. A system for dynamic interactive layout of game screens, used to implement the method for dynamic interactive layout of game screens according to any one of claims 1 to 9, characterized in that: The system includes: An attention field generation module is used to obtain depth maps, normal maps, illumination maps, color maps, and user input event data in each frame, and input the data into a neural network and combine it with symplectic integration to generate an attention field consistent with the screen resolution; a displacement fusion module, configured to determine a region of interest based on the attention field, construct a source probability density and a target probability density, obtain a first displacement field and a second displacement field, respectively, and fuse the first displacement field and the second displacement field according to a preset weight to generate a unified displacement field; a compensation rendering module, configured to deform pixel coordinates using the unified displacement field, input a voxel set corresponding to the region of interest into a neural radiation field network to obtain a compensated color map, and synthesize the compensated color map with the deformed color map into a fused image; An interactive feedback module is configured to determine the original screen coordinates of user input events on the fused screen by performing inverse operations on the unified displacement field and the second displacement field, query the deformed entity identification map to obtain the entity identification, and adjust the attention field threshold, the second displacement field parameters, and the damping parameters when the system energy exceeds a preset range to achieve cross-frame adaptive feedback.
Citation Information
Patent Citations
Brain feature extraction method based on diffusion tensor imaging
CN105551026A
Real-time dynamic free view angle synthesis method and device based on explicit geometric deformation
CN114863038A
Nerve radiation field three-dimensional reconstruction method based on fused voxels
CN116664782A
Virtual-real fusion drawing method and system based on neural radiation field and voxelization representation
CN117671112A
Measurement and application of image colorfulness using deep learning
US20230260161A1