Game animation character display method based on virtual reality technology
By optimizing dynamic rendering accuracy levels and real-time rendering strategies, combined with eye tracking and heat maps, the problems of device performance differences and multi-user interaction in virtual reality technology are solved, improving rendering efficiency and user experience.
Patent Information
- Application Number
- CN202510714736.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-09-12
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing virtual reality technology has problems in displaying game and animation characters, such as large differences in device performance, inflexible rendering strategies, and insufficient frame rate stability under multi-user interaction, resulting in a poor user experience.
Through dynamic rendering precision level adjustment, real-time rendering strategy optimization, cross-terminal posture synchronization and video memory management, combined with eye tracking and heat map optimization rendering resource allocation, rendering efficiency and stability are ensured.
It realizes dynamic adjustment of rendering parameters according to model complexity and observation distance, reduces resource waste, improves multi-user interaction experience and immersion, and ensures the consistency and smoothness of rendering effects.
Smart Images

Figure CN120635274A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of virtual reality technology, and more specifically, to a method for displaying game cartoon characters based on virtual reality technology. Background Art
[0002] With the rapid development of virtual reality technology, the display of game and animation characters is also experiencing new changes. Traditional display methods rely on two-dimensional images or simple three-dimensional models, limiting users to viewing game and animation characters from fixed perspectives and limited interactive methods. This lacks immersion and realism, and fails to meet users' demand for a comprehensive, multi-angle, and dynamic display of game and animation characters. With the advent of virtual reality technology, people have begun experimenting with using virtual reality devices to display game and animation characters. Through hardware such as head-mounted displays, users are able to create an immersive virtual environment and more intuitively experience the charm of game and animation characters.
[0003] However, existing methods for displaying game characters based on virtual reality technology still present several challenges. For one thing, the performance of virtual reality devices varies widely, with different devices exhibiting significant differences in their ability to render complex 3D models. While some high-end devices can effectively display high-precision 3D models, their high cost and demanding hardware requirements limit their widespread adoption. Meanwhile, some mid-range and low-end devices are prone to frame rate fluctuations and lag when rendering high-precision models, severely impacting the user experience. Furthermore, existing display methods employ relatively rigid rendering strategies and fail to fully consider the dynamic characteristics of game characters and user interactions within the virtual scene. For example, some complex models may render at high precision when viewed up close, but continue to render at high precision when viewed from a distance. This not only wastes computing resources but can also lead to inefficient rendering. Furthermore, when multiple users are simultaneously in the same virtual scene, existing methods struggle to achieve cross-device posture synchronization and stable frame rate control, leading to display inconsistencies between different devices.
[0004] In the process of implementing the embodiments of the present invention, there are at least the following problems or defects in the existing technology: First, it is impossible to dynamically adjust the rendering strategy according to the performance of the virtual reality device, resulting in large differences in display effects on different devices; second, the dynamic characteristics of game animation characters and user interaction behaviors are not fully considered, the rendering strategy is not flexible enough, and resource utilization is low; third, in multi-user scenarios, the cross-terminal posture synchronization and frame rate stabilization control capabilities are insufficient, affecting the multi-user interaction experience. Summary of the Invention
[0005] The present invention provides a method for displaying game cartoon characters based on virtual reality technology. The method is executed by a virtual reality character display system and includes: Acquiring three-dimensional model data of at least one game and animation character in a target display area through a user interaction terminal; For a game character, the VR content management platform determines the accuracy level of the character's dynamic rendering based on the 3D model data. The virtual reality content management platform determines a real-time rendering strategy based on the dynamic rendering accuracy level of at least one game animation character and terminal performance parameters. The real-time rendering strategy includes a model face reduction coefficient, a texture compression rate, and a skeletal animation update frequency. The virtual reality content management platform sends the real-time rendering strategy to the virtual reality supervision platform and the user interaction terminal, and generates rendering instructions and sends them to the virtual reality rendering engine platform, so that the rendering engine performs real-time rendering based on the rendering instructions. The virtual reality rendering engine platform is configured as the rendering core of the head-mounted display device; Based on the target display area and real-time rendering strategy, the virtual reality monitoring platform obtains the frame rate fluctuation data of the head-mounted display device according to the preset frame rate threshold, and determines the scene rendering stability based on the frame rate fluctuation data; In response to the scene rendering stability being lower than a preset stability threshold, the virtual reality supervision platform sends a rendering optimization instruction to the virtual reality content management platform based on the scene rendering stability, and adjusts the model data acquisition frequency and video memory cleanup cycle based on the scene rendering stability, and releases video memory resources according to the video memory cleanup cycle.
[0006] Furthermore, based on the 3D model data, determining the accuracy level of dynamic rendering of the game cartoon character includes: Extract model complexity features of game and animation characters based on 3D model data; Based on the model complexity characteristics, the visual detail weights of game and animation characters at multiple viewing distances are calculated; And based on the visual detail weight, determine the dynamic rendering accuracy level.
[0007] Furthermore, based on the model complexity characteristics, the weights of visual details of game characters at multiple viewing distances are calculated, including: Setting at least one distance evaluation interval; For a distance evaluation interval, based on the distance evaluation interval, model material properties, and model complexity characteristics, the detail weight prediction model is used to output the visual detail weight of the distance evaluation interval. The detail weight prediction model is a convolutional neural network model, and its training process includes: Acquire multiple first training samples and first labels, where the first training samples include a sample distance interval, a sample material attribute, and a sample model complexity, and the first label is a sample visual detail weight corresponding to the first training sample; Training an initial detail weight prediction model based on the plurality of first training samples and the first labels; And iteratively optimize the detail weight prediction model until the output error is lower than the preset loss threshold.
[0008] Furthermore, the input of the detail weight prediction model also includes the coordinates of the user's gaze focus, which are determined based on the projection position of the eye tracking data in the three-dimensional space; Training the initial detail weight prediction model includes: Based on the difference in rendering time between the focus area and the non-focus area in the historical rendering log, the first label is weighted and modified to generate a second label; And based on the plurality of first training samples and the second labels, an initial detail weight prediction model is trained.
[0009] Furthermore, based on the dynamic rendering accuracy level of at least one game cartoon character and the terminal performance parameters, determining the real-time rendering strategy includes: Calculate rendering efficiency coefficients at different precision levels based on the frame rate attenuation gradient of historical rendering records; And determine the real-time rendering strategy based on the dynamic rendering accuracy level, rendering efficiency coefficient and terminal performance parameters of at least one game animation character.
[0010] Furthermore, based on the dynamic rendering accuracy level of at least one game cartoon character and the terminal performance parameters, determining the real-time rendering strategy includes: Construct a performance optimization objective function with rendering strategy as a variable:
[0011] in, The time it takes to render the i-th game character. is the video memory occupancy value of the i-th game cartoon character, , is the weight coefficient determined by the terminal performance parameters, n is the total number of game anime characters, and i is the index of the game anime character; Based on the performance optimization objective function, the set of candidate rendering strategies is iterated by gradient descent until the convergence condition is met and the real-time rendering strategy is output.
[0012] Furthermore, the method further comprises: During the presentation, in response to the smoothness of the movements of the rendered game cartoon character falling below a preset threshold, identifying a high-load model component causing lag; Predict the game anime characters to be rendered that contain high-load model components and mark them as priority optimization targets; And based on the priority optimization object set, reconstruct the subsequent real-time rendering strategy.
[0013] Furthermore, obtaining the three-dimensional model data of at least one game cartoon character in the target display area through the user interaction terminal includes: Real-time collection of user hand motion vectors
[0014] in, is the coordinate value of the hand in three-dimensional space, is the hand rotation angle; Will Mapped to the skeleton nodes of the game animation characters, driving the model to generate interactive action data; The three-dimensional model data includes interactive action data.
[0015] Furthermore, the virtual reality monitoring platform obtains frame rate fluctuation data of the head-mounted display device according to a preset frame rate threshold based on the target display area and real-time rendering strategy, including: When multiple user interaction terminals are detected in the same virtual scene, a cross-terminal gesture synchronization channel is established; Synchronize the position matrix and rotation quaternion of the game characters in each terminal through the channel:
[0016] in, is the position matrix of the game cartoon characters in the kth terminal, N is the total number of terminals in the same virtual scene, is the spherical linear interpolation algorithm, is the network delay compensation factor, k is the index of the terminal, is the position matrix after synchronization, They are the rotation quaternions of the game characters in different terminals. is the rotation quaternion after synchronization; Calculates frame rate fluctuation data based on the synchronized position matrix and rotation quaternion.
[0017] Furthermore, the virtual reality supervision platform sends rendering optimization instructions to the virtual reality content management platform based on the scene rendering stability, including: Generate a 3D heat map based on the scene rendering stability. The heat map uses color gradients to identify the rendering load intensity of different areas in the virtual scene. The heat map is overlaid and displayed on a supervisory view layer of a head-mounted display device; the rendering optimization instruction includes region identification data of the heat map.
[0018] The above embodiments of the present invention have at least the following beneficial effects: 1. Through dynamic rendering accuracy grading and real-time strategy adjustment, the problem of uneven resource allocation when rendering multiple character models in virtual reality scenes is solved. It can automatically match the optimal rendering parameters according to model complexity, viewing distance and terminal performance, significantly reducing hardware load while ensuring visual effects, and avoiding frame rate drops or freezes caused by excessive rendering pressure.
[0019] 2. The introduction of a rendering optimization mechanism based on eye tracking and heat maps solves the problem of wasted rendering resources in the user's gaze and non-gaze areas in traditional VR systems. By intelligently identifying high-attention areas and dynamically allocating rendering resources, it not only improves the detail expression of the visual focus area, but also reduces the performance consumption of non-essential areas, achieving a more efficient human-computer interaction experience.
[0020] 3. Through dynamic memory cleanup and cross-terminal posture synchronization technology, the problem of insufficient rendering stability in multi-user collaborative virtual scenes is solved. It can monitor frame rate fluctuations in real time and automatically optimize memory management. At the same time, it ensures the synchronization of model movements and position data between multiple terminals, reduces delays and rendering asynchrony, and improves the smoothness and collaborative experience of large-scale virtual scenes. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The above and other objects, features and advantages of the exemplary embodiments of the present invention will become readily apparent by reading the following detailed description with reference to the accompanying drawings, in which several embodiments of the present invention are shown by way of example and not limitation, in which: Figure 1 This is a flowchart of a method for displaying game cartoon characters based on virtual reality technology provided by one embodiment of the present invention. DETAILED DESCRIPTION
[0022] The technical solutions of this application will be described clearly and completely below, in conjunction with the accompanying drawings. It should be understood that the described embodiments represent only a portion of the embodiments of this application, and not all of them. The components of this application, generally described and illustrated in the drawings herein, may be arranged and designed in a variety of different configurations. Therefore, the following detailed description of the embodiments of this application provided in the drawings is not intended to limit the scope of the claimed application, but rather merely represents selected embodiments of this application. All other embodiments derived by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application. It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings. Furthermore, in the description of this application, the terms "first," "second," etc., are used solely to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0023] In traditional virtual reality game character display systems, the conflict between mobile hardware computing power and the need for high-precision model rendering results in a breakdown in the mechanism for balancing scene rendering efficiency and image quality. Dynamic precision control lacks correlation with the user's focus, resulting in excessive consumption of model facets and texture resources in non-focus areas, and insufficient levels of detail in focus areas, leading to visual fragmentation. In multi-user collaborative scenarios, the inter-terminal posture synchronization algorithm fails to utilize rotational quaternion interpolation. The simple duplication of position matrices leads to deviations in skeletal animation data, triggering a surge in local rendering load. The performance monitoring system relies on macro-level hardware indicators and is unable to locate hotspots of video memory usage within three-dimensional space. Consequently, the video memory cleanup cycle settings are mismatched with the spatial load distribution of the scene.
[0024] For example, in a multi-person collaborative virtual comic exhibition scene, users wear head-mounted display devices to interact with cartoon characters in real time. When the user's gaze focus deviates from the current rendering area, the high-precision model of the non-focus area continues to occupy video memory resources, and the skeletal animation update frequency is not dynamically reduced as the line of sight shifts. When three users access the same virtual booth through different terminals, no cross-terminal posture synchronization channel is established between the terminals. Each terminal independently calculates the character position matrix and rotation quaternion, resulting in asynchronous posture presentation of the same cartoon character on the alien screen. The virtual scene supervision platform only monitors the global GPU usage rate and cannot identify video memory leaks in the booth hot zone model components. The periodic video memory cleanup operation mistakenly deletes high-priority model texture data.
[0025] If these issues are not addressed, the waste of rendering resources in non-focus areas will continue to reduce the effective rendering frame rate. The asynchrony of multi-terminal posture data will exacerbate cross-device interaction latency. The lack of spatial load hotspot identification will lead to a misalignment between graphics memory management strategies and real-time rendering requirements. This will lead to technical consequences such as character movement stuttering in virtual scenes, perspective tearing across multiple users, and the accumulation of fragmented graphics memory resources, ultimately severely compromising the coherence and authenticity of the immersive interactive experience.
[0026] When faced with the above problems, this application first addresses the issue of disconnection between dynamic precision control and user gaze focus, considers fusing eye tracking data with model complexity features, and predicts the weights of visual details in different distance intervals by training convolutional neural networks to achieve dynamic precision grading driven by gaze focus. For posture synchronization in multi-terminal collaborative scenarios, this application attempts to use the spherical linear interpolation algorithm of rotation quaternions to replace the traditional coordinate copying method, establish a cross-terminal posture synchronization channel, and ensure the posture consistency of skeletal animation data between different devices. In order to solve the defect of lack of spatial guidance for performance optimization, this application explores mapping frame rate fluctuation data into a three-dimensional heat map, and uses color gradients to identify rendering load distribution, providing a spatial dimension reference for the video memory cleanup strategy.
[0027] In this regard, the present application proposes a method to be executed by a virtual reality character display system, including: obtaining three-dimensional model data of at least one game and animation character in a target display area through a user interaction terminal; for a game and animation character, a virtual reality content management platform determines a dynamic rendering accuracy level of the game and animation character based on the three-dimensional model data; the virtual reality content management platform determines a real-time rendering strategy based on the dynamic rendering accuracy level of at least one game and animation character and terminal performance parameters, the real-time rendering strategy including a model face reduction coefficient, a texture compression rate, and a skeletal animation update frequency; the virtual reality content management platform sends the real-time rendering strategy to the virtual reality supervision platform and the user interaction terminal, and generates The rendering instruction is sent to the virtual reality rendering engine platform so that the rendering engine performs real-time rendering based on the rendering instruction. The virtual reality rendering engine platform is configured as the rendering core of the head-mounted display device; the virtual reality supervision platform obtains the frame rate fluctuation data of the head-mounted display device according to a preset frame rate threshold based on the target display area and the real-time rendering strategy, and determines the scene rendering stability based on the frame rate fluctuation data; and in response to the scene rendering stability being lower than the preset stability threshold, the virtual reality supervision platform sends a rendering optimization instruction to the virtual reality content management platform based on the scene rendering stability, and adjusts the model data acquisition frequency and the video memory cleanup cycle based on the scene rendering stability, and releases video memory resources according to the video memory cleanup cycle.
[0028] 3D model data refers to a data set containing the geometry, texture mapping, and motion information of game and animation characters. This can be achieved by exporting FBX files with skeletal animation using 3D modeling software, providing the basic data support for character rendering. Dynamic rendering accuracy refers to a rendering quality grading metric that dynamically adjusts based on model complexity and viewing distance. This can be achieved by using a convolutional neural network to predict the weight of visible detail at different distances. This is used to reduce rendering resource consumption in non-critical areas while ensuring visual quality. Terminal performance parameters refer to quantitative indicators of the number of GPU shader cores and video memory bandwidth of the headset. This can be achieved by accessing real-time data through the hardware query interface of the graphics API, and is used to constrain the computing power adaptation range of the rendering strategy. The model face reduction factor refers to a quantitative parameter for the reduction ratio of the number of triangles. This can be achieved by using an edge-folding algorithm to dynamically simplify the mesh topology, reducing GPU load during the geometry processing stage. Texture compression ratio refers to a numerical parameter for the proportional reduction of texture resolution. This can be achieved using the ASTC adaptive scalable texture compression format, balancing texture clarity and video memory usage. Among them, the skeletal animation update frequency refers to the control parameter of the number of skeletal matrix calculations per unit time, which can be implemented by the action sampling interval adjustment method based on quaternion interpolation, and is used to optimize the CPU overhead of skinning calculations. Among them, the scene rendering stability refers to the normalized measurement value of the inverse of the frame generation time variance, which can be calculated by counting the standard deviation of the rendering time of N consecutive frames through a sliding window, and is used to quantitatively evaluate the real-time performance status of the system. Among them, the video memory cleanup cycle refers to the time interval parameter for releasing abandoned resource buffers, which can be implemented by a dynamic memory management mechanism that combines reference counting with the least recently used strategy, and is used to prevent sudden performance degradation caused by video memory fragmentation.
[0029] The core innovation of this application lies in building a closed-loop control system for dynamic precision control and multi-platform collaborative optimization, driving the iterative optimization of real-time rendering strategies through quantifiable scene rendering stability indicators, and combining eye tracking data with multi-terminal posture synchronization technology to achieve spatially precise allocation of hardware resources while ensuring the visual expressiveness of high-precision models.
[0030] The working process and principle of this application are as follows: the virtual reality character display system obtains the three-dimensional model data of the game and animation characters in the target display area through the user interaction terminal. The virtual reality content management platform determines the dynamic rendering accuracy level of the game and animation characters based on the three-dimensional model data. The content management platform determines the real-time rendering strategy based on the dynamic rendering accuracy level and terminal performance parameters, including the model patch reduction coefficient, texture compression rate and skeletal animation update frequency. The content management platform sends the rendering strategy to the supervision platform and the user interaction terminal, and generates rendering instructions to send to the rendering engine platform to perform real-time rendering. Based on the target display area and the real-time rendering strategy, the virtual reality supervision platform obtains the frame rate fluctuation data of the head-mounted display device according to the preset frame rate threshold and calculates the scene rendering stability. When the stability is lower than the preset threshold, the supervision platform sends optimization instructions to the content management platform to adjust the model data acquisition frequency and the video memory cleanup cycle, and releases video memory resources on a cyclical basis.
[0031] As a preferred embodiment, the solution of this application is implemented as follows: The virtual reality character display system first obtains 3D model data of game and animation characters within a target display area through a user interaction terminal. The virtual reality content management platform receives the 3D model data, extracts model complexity features, calculates visible detail weights at different viewing distances, and determines the dynamic rendering accuracy level. Based on the dynamic rendering accuracy level and the terminal GPU performance parameters, the content management platform constructs a rendering performance optimization objective function and determines a real-time rendering strategy through gradient descent iteration. The rendering strategy includes the model patch reduction coefficient, texture compression ratio, and skeletal animation update frequency. The content management platform sends the rendering strategy to the virtual reality supervision platform and the user interaction terminal, and simultaneously generates rendering instructions and sends them to the virtual reality rendering engine platform. The rendering engine platform, as the rendering core of the head-mounted display device, performs real-time rendering based on the rendering instructions. Based on the target display area and the real-time rendering strategy, the virtual reality supervision platform collects frame rate data from the head-mounted display device according to a preset frame rate threshold, calculates the frame rate fluctuation variance, and determines the scene rendering stability. When the stability falls below the preset threshold, the supervision platform generates a 3D spatial heat map to identify the rendering load distribution and sends rendering optimization instructions containing the heat map data to the content management platform. At the same time, the regulatory platform adjusts the model data acquisition frequency and memory cleanup cycle based on stability, and executes memory resource release operations according to the new cycle.
[0032] Through the above scheme, this application achieves a dynamic balance between rendering efficiency and image quality in virtual reality scenes. Through dynamic rendering accuracy levels and real-time rendering strategies, the system can adaptively adjust rendering parameters according to model characteristics and hardware capabilities to avoid waste of resources in non-focus areas. When multiple terminals collaborate, the cross-terminal posture synchronization channel ensures the consistency of skeletal animation data and reduces interaction delays. The three-dimensional spatial heat map provides an accurate spatial load distribution reference for video memory management, so that the video memory cleanup strategy is accurately matched with real-time rendering requirements. These technical means jointly improve the smoothness of character movements in virtual scenes, improve the consistency of multi-user perspectives, and optimize the efficiency of video memory resource utilization, thereby enhancing the user's immersive interactive experience.
[0033] In some of the above-mentioned solutions of this application, it is proposed to determine the dynamic rendering accuracy level based on three-dimensional model data to optimize the allocation of rendering resources. However, in this process, due to the lack of targeted analysis of the structural characteristics of the model itself and the failure to quantify the impact of different observation distances on the rendering detail requirements, the accuracy level division does not match the user's actual visual perception, making it difficult to dynamically balance the rendering load and visual effects.
[0034] In this regard, the present application further proposes to extract the model complexity characteristics of game cartoon characters based on three-dimensional model data; calculate the visual detail weights of game cartoon characters at multiple observation distances based on the model complexity characteristics; and determine the dynamic rendering accuracy level based on the visual detail weights.
[0035] Among them, the model complexity feature is realized by analyzing the vertex distribution density and the number of facets. The vertex distribution density is measured by the number of vertices per unit volume, and the number of facets is output by the triangle face statistics module. The calculation of the visual detail weight needs to be combined with the distance evaluation interval. For example, three distance evaluation intervals are set as within 1 meter, 1-5 meters, and above 5 meters. Each interval corresponds to a different detail retention coefficient, among which the detail retention coefficient of the close distance interval is 0.8-1.0, the medium distance interval is 0.5-0.8, and the long distance interval is 0.3-0.5. The weight calculation at different observation distances is realized by the convolutional neural network model. The input parameters include model complexity features and material reflectivity. The material reflectivity is obtained through the material property analysis module. The dynamic rendering accuracy level is divided according to the weight threshold. When the visual detail weight is greater than 0.7, it is classified as a high precision level, 0.4-0.7 is a medium precision level, and below 0.4 is a low precision level.
[0036] Specifically, during the processing of 3D model data, vertex density and facet count are extracted as model complexity features. Vertex density is determined by voxelizing the model space using the 3D meshing module. The number of vertices within each voxel is counted to generate a vertex density distribution map. Facet counts are accumulated by traversing the triangle facet indices within the model mesh data. During the calculation of visual detail weights, the distance evaluation interval is pre-divided into multiple levels, each corresponding to a weight calculation channel. The material attribute parsing module uses the model surface's diffuse reflectance, specular reflectance, and normal map intensity as material parameters. The convolutional neural network model utilizes a three-layer convolutional layer architecture: the first layer processes the vertex density distribution map, the second layer integrates the facet count with material parameters, and the third layer outputs weights for each distance interval. During the dynamic rendering precision level determination, weights are classified using a threshold comparison module. The classification results trigger the corresponding rendering parameter configuration. For example, the high precision level enables 8K textures and every-frame skeleton updates, the medium precision level enables 4K textures and every-other-frame skeleton updates, and the low precision level enables 2K textures and every-three-frame skeleton updates. Therefore, the combination of model complexity characteristics and observation distance enables dynamic matching of rendering resource allocation with visual perception requirements, avoiding resource waste caused by high-precision rendering in distant scenes, while ensuring the complete presentation of key details when observed at close range.
[0037] As a preferred embodiment, the solution of this application is specifically implemented as follows: Based on 3D model data, we extract model complexity characteristics for game and animation characters. Specifically, we analyze parameters such as the 3D model's vertex distribution, number of facets, and texture resolution to calculate a model complexity metric. For example, we can count the total number of vertices, facets, and materials in the model, and combine this with the model's bounding box volume to derive the geometric complexity per unit volume.
[0038] Furthermore, based on the model complexity characteristics, the visual detail weights of game and animation characters at multiple viewing distances are calculated. First, multiple distance evaluation intervals are set, such as close distance 0-2 meters, medium distance 2-5 meters, and long distance 5 meters or more. Then, for each distance interval, the visual detail weight is calculated using a pre-trained detail weight prediction model, combining the model complexity characteristics and material properties. The detail weight prediction model can adopt a convolutional neural network structure, with input including the distance interval, material properties, and model complexity characteristics, and outputting the corresponding visual detail weight.
[0039] Thus, based on the calculated visual detail weight, the dynamic rendering accuracy level is determined. Specifically, the visual detail weight can be mapped to a predefined accuracy level range, such as high, medium, and low. For example, when the visual detail weight at close range is higher than a threshold, the rendering accuracy level of the model when observed at close range is set to high; when the visual detail weight at medium range is in a medium range, the rendering accuracy level when observed at medium range is set to medium; when the visual detail weight at long distance is lower than a threshold, the rendering accuracy level when observed at long distance is set to low.
[0040] Through the above technical solution, this application realizes the dynamic rendering accuracy level division based on the structural characteristics of the model itself and the observation distance. By extracting the complexity characteristics of the model, it ensures that the accuracy evaluation is based on the geometric characteristics of the model itself. The visual detail weight is calculated in combination with the observation distance, and the necessity of detail preservation at different distances is quantified. Finally, the accuracy level is dynamically divided based on the visual detail weight, so that the rendering strategy can be adaptively adjusted according to the user's observation position. This method avoids the performance waste caused by over-rendering, while ensuring the detail presentation of the visual focus area, effectively balancing the rendering load and visual effect.
[0041] In some of the above-mentioned solutions of this application, a method of calculating the visual detail weight based on the model complexity characteristics is proposed to dynamically adjust the rendering accuracy. However, in this process, since the impact of the user's actual gaze focus on the perception of three-dimensional model details is not taken into account, the rendering resource consumption of the non-focus area does not match the visual requirements of the focus area, resulting in a waste of computing resources and an inability to accurately improve the rendering quality of the user's focus area.
[0042] In this regard, the present application further proposes to calculate the visual detail weights of game cartoon characters at multiple observation distances based on model complexity characteristics, including: setting at least one distance evaluation interval; for a distance evaluation interval, based on the distance evaluation interval, model material properties and model complexity characteristics, through a detail weight prediction model, outputting the visual detail weight of the distance evaluation interval, the detail weight prediction model is a convolutional neural network model, and its training process includes: obtaining multiple first training samples and first labels, the first training sample includes the sample distance interval, sample material properties and sample model complexity, and the first label is the sample visual detail weight corresponding to the first training sample; based on the multiple first training samples and the first label, training the initial detail weight prediction model; and iteratively optimizing the detail weight prediction model until the output error is lower than the preset loss threshold.
[0043] Among them, the distance evaluation interval can be set through spatial discretization, for example, divided into three intervals of 0-2 meters, 2-5 meters, and above 5 meters. The model material properties include surface reflectivity, roughness, and texture resolution parameters, and the model complexity characteristics include the number of vertices, triangle density, and the number of bone bindings. The input layer of the convolutional neural network model is set as a four-dimensional tensor structure. The first dimension corresponds to the distance interval index, the second dimension maps the spectral reflectance curve of the material attribute, the third dimension stores the model complexity feature matrix, and the fourth dimension retains the coordinates of the user's gaze focus.
[0044] During training, the weights of visual detail in samples were calibrated through human visual perception experiments. For example, the minimum threshold for discernible detail at various distances was collected from 50 test subjects in a VR environment. The loss function uses a weighted combination of mean squared error and perceptual similarity, with a default loss threshold of 0.15. Batch normalization layers were introduced during optimization to address the dynamic range of material properties, and an adaptive learning rate adjustment mechanism was used to accelerate model convergence.
[0045] Specifically, during the distance assessment interval setting phase, the model discretizes the three-dimensional space into multiple subregions, enabling it to independently calculate weights for different observation distances. For example, the 0-2 meter range corresponds to close-range observation mode, focusing on analyzing facial expression details; the range above 5 meters corresponds to long-range observation mode, which reduces the accuracy of rendering clothing wrinkles. The metallic and diffuse textures in the material properties are each encoded as a high-dimensional feature vector, and the spatial filtering effect of the convolution kernel is used to extract the correlation between the material and the geometric structure.
[0046] When constructing training samples, triangle face density data, part of the model's complexity features, is collected using a 3D scanner. For example, for anime character models, high-precision sampling of 200-500 facets per square millimeter is performed. During training, the convolutional neural network model extracts the nonlinear relationship between distance intervals and material properties through multi-layer convolution operations. For example, a 3×3 convolution kernel is used to perform local feature matching on material spectral curves. The model output layer uses a sigmoid activation function to constrain the predicted values to the range of 0-1, corresponding to the normalized value of the detail weight.
[0047] The optimized detail weight prediction model can dynamically generate weight values based on the real-time distance, material type, and model complexity. For example, when the metal material model is in the 2-5 meter distance range, the model automatically increases the weight value of the highlight reflective area, while reducing the weight distribution of the non-reflective area. This mechanism enables the rendering engine to prioritize visually sensitive areas based on the weight distribution, such as allocating 60% of rendering resources to model components with weight values higher than 0.8. By iteratively optimizing to a loss value below 0.15, the model's prediction error on the test set is controlled within ±5%, effectively balancing rendering quality and computational load.
[0048] As a preferred embodiment, the solution of this application is specifically implemented as follows: The distance evaluation intervals are set to 0-1m, 1-3m, 3-5m, and over 5m. For the 1-3m distance evaluation interval, a detail weight prediction model is used to output the visual detail weight for that distance evaluation interval based on this interval, the model's material properties, and its complexity. This detail weight prediction model uses a convolutional neural network architecture consisting of five convolutional layers and three fully connected layers.
[0049] The training process of the detail weight prediction model is as follows: First, multiple first training samples and first labels are obtained. The first training samples include the sample distance interval, sample material properties, and sample model complexity. The sample distance interval is represented by a normalized value between 0 and 1. The sample material properties include parameters such as the diffuse reflectance coefficient and the specular coefficient. The sample model complexity includes geometric features such as the number of vertices and the number of facets. The first label is the sample visual detail weight corresponding to the first training sample, obtained by manual annotation.
[0050] Based on the first training samples and the first labels, the initial detail weight prediction model is trained using stochastic gradient descent. The detail weight prediction model is iteratively optimized, and training is terminated when the output error falls below a preset loss threshold. The preset loss threshold can be set to 0.01.
[0051] Through the above technical solution, this application realizes the calculation of visual detail weights based on the dynamic prediction mechanism between partitions and the machine learning model. This method comprehensively considers the distance interval, material properties and model complexity characteristics, and can capture the visual significance differences of different materials and geometric structures at a specific distance, thereby generating visual detail weights that are more in line with actual perception needs. Compared with traditional static LOD technology, this method can allocate rendering resources more accurately, improve the rendering quality of the focus area, and reduce redundant calculations in the non-focus area, thereby improving rendering efficiency while ensuring visual effects.
[0052] In some of the above-mentioned solutions of the present application, since the dynamic changes of the user's actual gaze focus are not taken into consideration, the rendering resource allocation between the focus area and the non-focus area is unbalanced. The detail rendering of the focus area may be degraded due to insufficient resources, while the non-focus area will waste resources due to excessive rendering.
[0053] In this regard, the present application further proposes that the input of the detail weight prediction model also includes the coordinates of the user's gaze focus, which are determined based on the projection position of the eye tracking data in three-dimensional space; training the initial detail weight prediction model includes: based on the difference in rendering time between the focus area and the non-focus area in the historical rendering log, weighted correction is performed on the first label to generate a second label; based on multiple first training samples and the second label, the initial detail weight prediction model is trained.
[0054] The acquisition of the user's gaze focus coordinates involves capturing the user's eye movement data through an eye tracking device and mapping the data to a three-dimensional scene coordinate system. For example, a perspective projection matrix is used to convert two-dimensional eye movement data into three-dimensional space coordinates. During the model training phase, the rendering time differences recorded in the historical rendering log are quantified as regional weight correction coefficients. For example, the correction coefficient for the focus area is set to 1.2-1.5 times, and for the non-focus area to 0.6-0.8 times. The original labels are dynamically adjusted using a linear interpolation method. The distance evaluation intervals in the training samples and the gaze focus coordinates form a spatial correlation matrix, which is used to calculate the joint feature vector of the model input.
[0055] Specifically, during the model inference process, the coordinates of the user's gaze focus and the distance evaluation interval are jointly input into the convolutional neural network, and spatial attention features are extracted through multi-layer convolution kernels. During model training, the loss function is optimized based on the weighted second label, so that the prediction error weight of the focus area is increased by 30%-50%, forcing the model to prioritize learning the detail distribution rules of the gaze area. For example, during the training iteration process, when the rendering time of the focus area is more than 15% higher than that of the non-focus area, the loss weight of the corresponding sample is increased to 1.3 times the original value. Through this training mechanism, the visual detail weights output by the model form a gradient attenuation distribution around the focus, so that the rendering resource allocation ratio is positively correlated with the user's visual attention. In this way, the texture resolution of the focus area is increased by 20%-35% without increasing the complexity of the model, while the memory usage of the non-focus area is reduced by 18%-25%.
[0056] As a preferred embodiment, the solution of this application is specifically implemented as follows: The input to the detail weight prediction model includes the coordinates of the user's gaze focus. The coordinates of the user's gaze focus are determined by collecting eye movement data using an eye tracking device and projecting this data into three-dimensional space. For example, an infrared camera can be used to capture pupil movement trajectories and, combined with head posture information, calculate the intersection of the user's gaze in the virtual scene.
[0057] When training the initial detail weight prediction model, we first obtain historical rendering logs. This log records the rendering time data for different areas. Furthermore, we perform a weighted correction on the first label based on the difference in rendering time between the focus area and the non-focus area. Specifically, we set a weighting coefficient α. When the rendering time of the focus area is higher than that of the non-focus area, the corresponding first label is multiplied by (1+α); otherwise, it is multiplied by (1-α). This generates the second label.
[0058] Finally, the initial detail weight prediction model is trained using the multiple first training samples and the modified second labels. The training process can use a backpropagation algorithm to iteratively optimize the model parameters through multiple rounds until the prediction error is reduced to below a preset threshold.
[0059] Through the above technical solution, the present application realizes dynamic detail weight prediction based on the user's real-time gaze behavior. Due to the introduction of eye tracking data, the model can more accurately identify the scene area that the user is paying attention to, thereby optimizing the allocation of rendering resources. At the same time, through the weighted correction of the training labels, the model's ability to learn the detail weights of the focus area is enhanced. This method effectively balances the rendering quality of the focus area and the non-focus area, avoids the waste of resources caused by excessive rendering of the non-focus area, and ensures the quality of the detail rendering of the focus area. Ultimately, this technical solution improves the overall rendering efficiency and improves the user's visual experience.
[0060] In some of the above-mentioned solutions of this application, due to the nonlinear impact of different precision levels on rendering resource consumption, simply relying on current terminal parameters may cause the rendering strategy to be unable to accurately predict performance bottlenecks in complex scenarios, thereby causing the problem of aggravated frame rate attenuation.
[0061] In this regard, the present application further proposes to calculate the rendering efficiency coefficients of different accuracy levels based on the frame rate attenuation gradient of historical rendering records; and determine the real-time rendering strategy by combining the dynamic rendering accuracy level, rendering efficiency coefficient and terminal performance parameters of at least one game animation character.
[0062] The frame rate attenuation gradient is obtained by statistically analyzing the frame rate drop rates corresponding to different precision levels during the historical rendering process. For example, if the high precision level causes the frame rate to drop from 60FPS to 45FPS during three consecutive frame rendering cycles, the gradient value is -5FPS / frame. The rendering efficiency coefficient is configured as the output of the normalized function of the gradient value, which can be specifically expressed as
[0063] in is the absolute value of the frame rate attenuation gradient of the current accuracy level, It is the maximum gradient value in the historical records. When the gradient value of the high-precision level exceeds the preset threshold, its efficiency coefficient is given a higher resource sensitivity weight. For example, when the gradient value is -8FPS / frame, the efficiency coefficient is adjusted to 0.3, triggering the model patch reduction coefficient in the rendering strategy to be increased to 0.6. When the dynamic rendering accuracy level is combined with the terminal performance parameters, the video memory occupancy rate is quantified to a normalized value of 0-1, and weightedly fused with the efficiency coefficient to generate the texture compression rate parameter under the objective function constraint. The sliding window length of historical data is set to the most recent 30 rendering records to ensure that the efficiency coefficient can reflect the latest changes in hardware performance.
[0064] Specifically, in the process of determining the real-time rendering strategy, the historical frame rate data is first retrieved from the storage module, and the average attenuation gradient is calculated according to the accuracy level classification. For the dynamic rendering accuracy level in the current scene, the corresponding performance coefficient is matched, and multi-parameter fusion is performed in combination with the real-time memory occupancy rate and CPU load rate of the terminal GPU. For example, when the performance coefficient is 0.5 and the memory occupancy rate exceeds 70%, the skeletal animation update frequency is limited to 30Hz, and the model patch reduction coefficient is increased by 10%. In the weight allocation stage, the GPU model in the terminal performance parameters is mapped to the hardware capability index, which is used to dynamically adjust the influence ratio of the performance coefficient in strategy generation. For example, the hardware capability index of a high-end GPU is 1.2, which allows a higher texture compression rate threshold to be used at the same performance coefficient. By combining the historical performance attenuation trend with the real-time hardware status, the rendering strategy can reduce the resource allocation weight of the high-load accuracy level in advance. For example, when a complex physical simulation scene is predicted, the rendering time weight of the high-precision level is proactively increased. The frame rate is reduced from 0.6 to 0.4, thus avoiding sudden frame rate fluctuations. This process effectively overcomes the strategy lag caused by local parameter sampling in traditional methods, and enables real-time matching of rendering resource allocation with hardware performance fluctuation curves.
[0065] As a preferred embodiment, the solution of this application is specifically implemented as follows: Based on the frame rate attenuation gradient of historical rendering records, the rendering efficiency coefficients of different accuracy levels are calculated. Specifically, the rendering data within a certain period of time in the past (for example, 30 seconds) can be collected, and the rate of change of the frame rate over time can be calculated for each accuracy level. Furthermore, the rate of change data is fitted into a curve, and the slope of the curve is obtained as the frame rate attenuation gradient. In this way, a rendering efficiency coefficient between 0 and 1 can be assigned to each accuracy level. The larger the attenuation gradient, the smaller the efficiency coefficient.
[0066] Determine the real-time rendering strategy by combining the dynamic rendering accuracy level, rendering efficiency coefficient, and terminal performance parameters of at least one game character. For example, a weighted sum model can be constructed: Rendering score = A1 × accuracy level + A2 × rendering efficiency coefficient + A3 × CPU utilization + A4 × GPU utilization; A1 to A4 are weight coefficients that can be adjusted based on the specific scenario. A higher rendering score results in a more aggressive rendering strategy, such as a higher model face reduction coefficient. As a preferred implementation, multiple rendering score thresholds can be set to categorize rendering strategies into several levels.
[0067] Through the above technical solution, this application can predict performance bottlenecks based on historical rendering data, dynamically adjust the rendering strategy, and effectively suppress frame rate mutations caused by switching precision levels. At the same time, by introducing the rendering efficiency coefficient, the resource consumption patterns of different precision levels in a specific hardware environment are quantified, making the rendering strategy more accurate. In addition, the strategy generation mechanism that integrates historical performance data and real-time parameters improves the adaptability of the rendering strategy to complex scenes, effectively solving the problem of inaccurate rendering strategies caused by the lack of prediction of performance degradation trends in traditional methods.
[0068] In some of the above-mentioned schemes in this application, a method for determining real-time rendering strategies based on dynamic rendering accuracy levels and terminal performance parameters is proposed to optimize rendering efficiency. However, since rendering operations of different accuracy levels have a nonlinear relationship with the consumption of computing resources, it is difficult to accurately balance the contradiction between rendering time and video memory occupancy by simply superimposing parameters. As a result, the strategy selection lacks a mathematical optimization basis, and it is impossible to achieve optimal global performance under limited hardware resources.
[0069] In this regard, the present application further proposes to construct an efficiency optimization objective function with the rendering strategy as a variable, and perform gradient descent iteration on the set of candidate rendering strategies based on the objective function until the real-time rendering strategy is output when the convergence conditions are met.
[0070] Among them, the performance optimization objective function is designed to minimize the weighted sum of rendering time and video memory occupancy. Specifically, the weight coefficient is dynamically determined by the terminal performance parameters. For example, when the terminal video memory capacity is lower than the preset threshold, the weight coefficient of the video memory occupancy item is set to 2-3 times the weight of the time item, so as to give priority to avoiding video memory overflow. The set of candidate rendering strategies is generated by enumerating a combination of model patch reduction coefficients, texture compression rates, and skeletal animation update frequencies. Each strategy corresponds to a candidate solution for the objective function. During the gradient descent iteration process, the partial derivatives of the objective function with respect to the policy variables are calculated, and the policy parameters are updated along the negative gradient direction. When the difference in the objective function for three consecutive iterations is less than 0.5%, it is judged to be converged.
[0071] Specifically, the construction of the performance optimization objective function transforms the multi-parameter dynamic balance problem into a mathematical optimization problem. Each time a rendering strategy is generated, a historical performance database is first queried based on the terminal GPU model to obtain a baseline ratio of rendering time to video memory usage under typical load for that model, thereby initializing the weight coefficients. Subsequently, an initial set of 5-8 candidate strategies is generated by iterating over different combinations of patch reduction coefficients and texture compression rates. During the gradient descent process, the first-order derivative of the objective function with respect to the skeletal animation update frequency is calculated at each iteration, and the parameters are updated with a step size of 0.1, while monitoring whether the peak video memory usage exceeds the hardware limit. When the video memory usage is detected to be approaching the threshold, the weight coefficient is dynamically adjusted to increase the gradient component of the video memory item, forcing the optimization direction to shift towards the video memory safe zone. Through this process, a strategy solution that meets the terminal's real-time requirements can be obtained within 200-300 milliseconds, achieving a Pareto optimal balance between rendering time and video memory usage.
[0072] As a preferred embodiment, the solution of this application is specifically implemented as follows:
[0073] in, The time it takes to render the i-th game character. is the video memory occupancy value of the i-th game cartoon character, , is the weight coefficient determined by the terminal performance parameters, n is the total number of game anime characters, and i is the index of the game anime character; Based on the performance optimization objective function, the set of candidate rendering strategies is iterated by gradient descent until the convergence condition is met and the real-time rendering strategy is output.
[0074] For example, a stochastic gradient descent algorithm can be used for iterative optimization with a learning rate of 0.01. In each iteration, a candidate rendering strategy is randomly selected, the objective function value is calculated, and the parameters are updated. If the objective function value changes by less than 0.001 over 10 consecutive iterations, convergence is considered achieved, and the optimal real-time rendering strategy is output.
[0075] Through the above technical solution, the present application can realize the mathematical modeling and optimization of the rendering strategy. By constructing the performance optimization objective function, taking the rendering time and video memory occupancy as the core variables, a dynamic weight coefficient is introduced to adapt to the hardware characteristics of different terminals. The gradient descent iterative algorithm is used to optimize the candidate strategies, which can approach the global optimal solution. This method avoids relying on empirical strategy selection and ensures that a rendering strategy that meets the current hardware load capacity can be generated in different terminal environments. In addition, the dynamic adjustment mechanism of the weight coefficient in the objective function enables the system to automatically adjust the optimization direction according to real-time performance fluctuations to achieve the optimal allocation of rendering resources.
[0076] In some of the above-mentioned solutions of this application, a method for determining a real-time rendering strategy based on dynamic rendering accuracy level and terminal performance parameters is proposed to optimize rendering efficiency. However, during the display process, when the smoothness of the movements of the rendered game animation characters is partially stuck, the existing solution cannot quickly identify the high-load model components that specifically cause performance degradation, and it is difficult to adjust the rendering strategy in time to eliminate the stuck, resulting in a continuous deterioration of rendering efficiency.
[0077] In this regard, the present application further proposes that during the display process, in response to the movement smoothness of the rendered game cartoon character being lower than a preset threshold, the high-load model components that cause lag are identified; the game cartoon characters to be rendered that contain high-load model components are predicted and marked as priority optimization objects; and based on the set of priority optimization objects, the subsequent real-time rendering strategy is reconstructed.
[0078] Among them, the preset threshold can be set to the action smoothness critical value of 45 frames per second. When the value is detected to be lower than the threshold for 5 consecutive times within the real-time frame rate sampling period, the performance analysis thread is triggered. The high-load model component can obtain the video memory occupancy rate and the proportion of computing time of each model component through the GPU resource monitoring interface. For example, a model component with more than 20,000 faces is marked as a high-load component when the video memory occupancy rate exceeds 15%. The prediction of the priority optimization object set can be based on the component dependency graph of the models to be processed in the rendering queue. When it is detected that at least one high-load component is referenced by more than three models to be rendered, its associated models are automatically added to the optimization set. When reconstructing the rendering strategy, the gradient descent algorithm can be used to dynamically adjust the weight coefficient of the corresponding optimization object in the objective function, for example, the weight coefficient w2 of the video memory occupancy value Mi_i associated with the high-load component is increased to 0.7.
[0079] Specifically, when the motion smoothness monitoring module detects that the frame rate data consistently falls below a preset threshold, the performance diagnosis module immediately initiates component-level resource analysis. By comparing the computational time baselines of each component in historical rendering logs, for example, if the single-frame processing time of a skeletal animation component exceeds 8ms, the component is determined to be under high load. The prediction algorithm establishes a directed graph data structure based on the component reference relationships of each model in the rendering queue and performs topological sorting. Nodes of models to be rendered containing high-load components are marked as red, high-priority nodes. When refactoring the rendering strategy, the model patch reduction coefficient of marked objects is prioritized to 0.6, while the skeletal animation update frequency is reduced to 30Hz. Experimental data shows that this solution can reduce the rendering time of high-load components by 37% and reduce the overall system frame rate fluctuation variance to less than 12%. Through dynamic optimization of component granularity, key performance bottlenecks are prioritized while maintaining overall rendering quality, quickly eliminating lag.
[0080] As a preferred embodiment, the solution of this application is specifically implemented as follows: During the presentation, the smoothness of the action can be evaluated by monitoring the frame rate of the game's animated characters. When the action frame rate is detected to be lower than the preset threshold, for example, lower than 30 frames per second, the system triggers the jamming diagnosis process. First, by analyzing the time consumption data of each stage in the rendering pipeline, the high-load model components that cause performance bottlenecks are identified. These components may be high-facet geometry, complex skeletal animation, or large-scale texture maps.
[0081] Next, the system predicts the game characters to be rendered that contain these high-load components. This prediction process is based on the spatial distribution and movement trajectories of the characters in the scene, combined with information about the identified high-load components. The predicted character models are marked as priority optimization targets and added to a priority queue.
[0082] Finally, based on the prioritized set of objects, the system restructures subsequent real-time rendering strategies. For character models in the priority queue, more aggressive optimization measures can be taken, such as further reducing the number of model faces, compressing texture resolution, or reducing the number of keyframes in skeletal animation. Simultaneously, the system dynamically adjusts rendering resource allocation, allocating more GPU computing resources to these priority objects to ensure their rendering quality and smoothness are prioritized.
[0083] Through the above technical solution, the present application can achieve dynamic performance optimization during the display of game anime characters. By promptly identifying high-load model components that cause jamming, the system can quickly locate performance bottlenecks and avoid continued deterioration of rendering efficiency. Predicting characters to be rendered that contain high-load components and marking them as priority optimization objects enables the system to perform preventive optimization for potential problems in advance. Reconstructing the rendering strategy based on the priority optimization object set ensures that system resources can be concentrated on solving key performance problems, thereby effectively improving the overall rendering efficiency and the smoothness of the user experience. This dynamic optimization mechanism can not only cope with immediate performance degradation, but also prevent potential jamming problems, providing more stable and smooth visual effects for the display of game anime characters in a virtual reality environment.
[0084] In some of the above-mentioned solutions in this application, traditional three-dimensional model data acquisition only generates a static model based on a preset action library, which is unable to capture the spatial coordinates and rotational posture of the user's hand movements in real time, resulting in the user's interactive movements with the virtual character being stiff and lacking in natural continuity, affecting the immersive experience. At the same time, the static model data cannot adapt to the dynamic rendering strategy's optimization requirements for real-time interactive movements.
[0085] In this regard, the present application further proposes real-time collection of user hand motion vectors, where is the coordinate value of the hand in three-dimensional space and is the rotation angle of the hand; it will be mapped to the skeletal nodes of the game animation character to drive the model to generate interactive motion data; the three-dimensional model data includes the interactive motion data.
[0086] Among them, the hand motion vector is collected by an optical tracking device or an inertial measurement unit, and the coordinate value records the displacement of the hand in a three-dimensional rectangular coordinate system with millimeter-level accuracy. For example, a position sensor with a resolution of 0.1 mm is used. The rotation angle is expressed in the form of quaternion or Euler angle, and the sampling frequency is set to above 60Hz to meet real-time requirements. The bone node mapping adopts a bone weight matrix driving mechanism, inputs the original motion vector into the inverse kinematics solver, calculates the rotation matrix and translation vector of the corresponding bone node, for example, the Jacobian matrix iteration algorithm is used to realize the solution of joint angle and position, and ensures that the hand motion forms a 1:1 mapping relationship with the virtual character's bone motion. The interactive action data is generated by the bone animation interpolator, which converts the transformation parameters of the bone node into model vertex displacement data and encapsulates it into a time-stamp aligned action sequence frame. For example, each frame of action data contains the transformation parameters of 128 bone nodes for real-time call by the rendering engine.
[0087] Specifically, after the user's hand motion vectors are collected in real time, they are transmitted to the skeletal animation system based on a four-dimensional data structure. The three-dimensional coordinate values are converted into the global position of the skeleton root node through the coordinate transformation matrix, and the rotation angle is decomposed into local rotation parameters of multiple sub-joints through the quaternion spherical interpolation algorithm. During the skeletal driving process, the motion vector is matched with the preset bone hierarchy structure. For example, the rotation angle of the wrist joint drives the forearm bone first, and then affects the posture changes of the upper arm and shoulder bones through chain transmission. The interactive motion data generated is bound to the model vertex data and processed in parallel with the texture map and lighting parameters in the rendering pipeline, so that the dynamic rendering strategy can automatically adjust the mesh subdivision level according to the bone transformation frequency. For example, when the hand rotation angle change rate exceeds 30 degrees per second, the number of subdivision meshes in the corresponding model area is increased to 1.5 times the baseline value, while the number of meshes in the static area is reduced to 0.7 times the baseline value. This process reduces the action response time of the virtual character to less than 20 milliseconds by eliminating the action delay caused by manual keyframe setting. At the same time, the dynamically generated bone parameters are directly input into the rendering resource allocation module, reducing the video memory usage by 12%-18%.
[0088] As a preferred embodiment, the solution of this application is specifically implemented as follows: The user wears a hand interaction device equipped with a nine-axis inertial measurement unit, which captures the displacement vector and rotation quaternion of the palm in a three-dimensional coordinate system at a sampling rate of 200Hz. The displacement vector is filtered through a Kalman filter algorithm to eliminate sensor noise interference, and the rotation quaternion is converted to Euler angles to obtain the yaw, pitch, and roll angles. The motion capture data is encapsulated as a four-dimensional vector, where the first three dimensions represent the normalized spatial coordinate offset and the fourth dimension represents the normalized rotation angle value. The motion vector is mapped to the avatar's hand bone nodes using a bone weight matrix. This weight matrix dynamically adjusts the weight distribution ratio based on the relative position of the bone node and its parent node. When driving the model, the displacement and rotation of the bone nodes are smoothly transitioned using linear interpolation and spherical linear interpolation algorithms, respectively, to generate motion sequence data with timestamps. The final three-dimensional model data package embeds a motion sequence data block, which stores the motion parameters of each bone node within a unit time interval in a binary stream format.
[0089] Through the above technical solution, this application achieves continuous and accurate capture of the six-degree-of-freedom motion trajectory of the hand, and eliminates joint deformation distortion during motion transmission through a nonlinear mapping mechanism driven by bone weights. The deep fusion of motion sequence data and three-dimensional model data enables the rendering engine to directly parse motion parameters without performing additional data format conversion, reducing the calculation delay during dynamic rendering. This solves the problem of mechanical motion feedback caused by traditional static motion models, allowing the virtual character's limb movements to form a synchronous response with millimeter-level precision with the user's hand movements, while providing a low-latency motion data input interface for real-time dynamic rendering strategies.
[0090] In some of the above-mentioned solutions of this application, a method is proposed to obtain frame rate fluctuation data of a head-mounted display device through a virtual reality monitoring platform to evaluate the stability of scene rendering. However, when multiple user interaction terminals are in the same virtual scene, the existing technology uses a simple coordinate copying method for posture synchronization, resulting in the inability to accurately match the rotation posture data between terminals, causing the problem of increased rendering frame rate differences.
[0091] In this regard, the present application further proposes establishing a cross-terminal posture synchronization channel when multiple user interaction terminals are detected in the same virtual scene; synchronizing the position matrix and rotation quaternion of the game animation characters in each terminal through the channel; and calculating the frame rate fluctuation data based on the synchronized posture data.
[0092] The cross-terminal attitude synchronization channel is established through a private communication protocol that uses a time-division multiplexing mechanism to allocate transmission bandwidth. Position matrix synchronization uses a weighted average algorithm, and the weight coefficient is dynamically adjusted based on the distance between the terminal and the scene center point. For example, a weight coefficient of 0.7 is assigned when the distance exceeds a set threshold. Rotational quaternion synchronization uses an improved spherical linear interpolation algorithm. The interpolation step size is negatively correlated with the network delay between terminals. When the delay exceeds 50ms, the prediction compensation mechanism is automatically enabled. The network delay compensation factor α is set to a value range of 0.3-0.8. The specific value is calculated by real-time monitoring of the round-trip delay and packet loss rate. For example, when the round-trip delay is 30ms and the packet loss rate is less than 2%, α is set to 0.65.
[0093] Specifically, when multiple terminals enter the same virtual scene, independent communication resources are first allocated through a dedicated channel establishment module to avoid data conflicts caused by broadcast transmission. The synchronous calculation of the position matrix utilizes a distributed architecture. Each terminal uploads its local position data to a central node for weighted averaging, and the calculated results are then broadcast to all terminals. For example, when there are three terminals, the central node dynamically adjusts the weights based on the position deviations of each terminal, ensuring that the calculated results are closer to the actual positions of the majority of terminals. The synchronization of rotation quaternions is performed locally on the terminal. Interpolation algorithms are used to fuse the rotation data from other terminals with the local data. Timestamp alignment techniques are used during the fusion process to eliminate clock discrepancies. The frame rate fluctuation calculation module generates standardized motion trajectories based on the synchronized posture data. The frame rate fluctuation value is calculated by comparing the deviation between the actual trajectory output by the rendering engine and the standard trajectory. For example, if the posture synchronization error exceeds 0.05 radians, the system automatically triggers a frame rate compensation mechanism, reducing the texture resolution in non-critical areas to maintain overall rendering stability.
[0094] As a preferred embodiment, the solution of the present application is specifically implemented as follows: in a multi-person collaborative virtual exhibition scene, when three user interaction terminals are connected to the virtual art gallery scene at the same time, a cross-terminal posture synchronization channel is established through a dedicated data link. Each terminal uploads the position matrix and rotation quaternion of the local game cartoon character to the synchronization server, wherein the position matrix is processed by arithmetic averaging to generate a synchronization position, and the rotation quaternion is fused using a spherical linear interpolation algorithm. The network delay compensation factor is dynamically calculated based on the round-trip delay of each terminal, specifically the normalized result of the transmission delay ratio between each terminal and the server. The synchronized posture data is verified for the continuity of its motion trajectory by the physical simulation module of the rendering engine, and finally the frame rate monitoring module collects the rendering time of each frame to generate a fluctuation curve.
[0095] Through the above technical solutions, this application effectively solves the problem of reduced rendering stability caused by insufficient posture synchronization accuracy in multi-terminal collaborative scenarios. Through a dedicated synchronization channel and a mathematical fusion algorithm, the posture jump phenomenon caused by traditional coordinate replication is eliminated, making the transition of rotation movements natural and smooth. The network delay compensation mechanism dynamically corrects the time deviation of cross-terminal data transmission and ensures the real-time consistency of synchronized data. Based on precisely synchronized posture data, the frame rate fluctuation monitoring module can accurately reflect the stability state of multi-terminal collaborative rendering and provide reliable input for subsequent optimization strategies.
[0096] In some of the above-mentioned solutions of this application, a method is proposed to obtain frame rate fluctuation data of a head-mounted display device through a virtual reality monitoring platform to evaluate the stability of scene rendering. However, when multiple user interaction terminals are in the same virtual scene, the existing technology uses a simple coordinate copying method for posture synchronization, resulting in the inability to accurately match the rotation posture data between terminals, causing the problem of increased rendering frame rate differences.
[0097] In this regard, the present application further proposes establishing a cross-terminal posture synchronization channel when multiple user interaction terminals are detected in the same virtual scene; synchronizing the position matrix and rotation quaternion of the game animation characters in each terminal through the channel; and calculating the frame rate fluctuation data based on the synchronized posture data.
[0098] The cross-terminal attitude synchronization channel is established through a private communication protocol that uses a time-division multiplexing mechanism to allocate transmission bandwidth. Position matrix synchronization uses a weighted average algorithm, and the weight coefficient is dynamically adjusted based on the distance between the terminal and the scene center point. For example, a weight coefficient of 0.7 is assigned when the distance exceeds a set threshold. Rotational quaternion synchronization uses an improved spherical linear interpolation algorithm. The interpolation step size is negatively correlated with the network delay between terminals. When the delay exceeds 50ms, the prediction compensation mechanism is automatically enabled. The network delay compensation factor α is set to a value range of 0.3-0.8. The specific value is calculated by real-time monitoring of the round-trip delay and packet loss rate. For example, when the round-trip delay is 30ms and the packet loss rate is less than 2%, α is set to 0.65.
[0099] Specifically, when multiple terminals enter the same virtual scene, independent communication resources are first allocated through a dedicated channel establishment module to avoid data conflicts caused by broadcast transmission. The synchronous calculation of the position matrix utilizes a distributed architecture. Each terminal uploads its local position data to a central node for weighted averaging, and the calculated results are then broadcast to all terminals. For example, when there are three terminals, the central node dynamically adjusts the weights based on the position deviations of each terminal, ensuring that the calculated results are closer to the actual positions of the majority of terminals. The synchronization of rotation quaternions is performed locally on the terminal. Interpolation algorithms are used to fuse the rotation data from other terminals with the local data. Timestamp alignment techniques are used during the fusion process to eliminate clock discrepancies. The frame rate fluctuation calculation module generates standardized motion trajectories based on the synchronized posture data. The frame rate fluctuation value is calculated by comparing the deviation between the actual trajectory output by the rendering engine and the standard trajectory. For example, if the posture synchronization error exceeds 0.05 radians, the system automatically triggers a frame rate compensation mechanism, reducing the texture resolution in non-critical areas to maintain overall rendering stability.
[0100] As a preferred embodiment, the solution of the present application is specifically implemented as follows: in a multi-person collaborative virtual exhibition scene, when three user interaction terminals are connected to the virtual art gallery scene at the same time, a cross-terminal posture synchronization channel is established through a dedicated data link. Each terminal uploads the position matrix and rotation quaternion of the local game cartoon character to the synchronization server, wherein the position matrix is processed by arithmetic averaging to generate a synchronization position, and the rotation quaternion is fused using a spherical linear interpolation algorithm. The network delay compensation factor is dynamically calculated based on the round-trip delay of each terminal, specifically the normalized result of the transmission delay ratio between each terminal and the server. The synchronized posture data is verified for the continuity of its motion trajectory by the physical simulation module of the rendering engine, and finally the frame rate monitoring module collects the rendering time of each frame to generate a fluctuation curve.
[0101] Through the above technical solutions, this application effectively solves the problem of reduced rendering stability caused by insufficient posture synchronization accuracy in multi-terminal collaborative scenarios. Through a dedicated synchronization channel and a mathematical fusion algorithm, the posture jump phenomenon caused by traditional coordinate replication is eliminated, making the transition of rotation movements natural and smooth. The network delay compensation mechanism dynamically corrects the time deviation of cross-terminal data transmission and ensures the real-time consistency of synchronized data. Based on precisely synchronized posture data, the frame rate fluctuation monitoring module can accurately reflect the stability state of multi-terminal collaborative rendering and provide reliable input for subsequent optimization strategies.
[0102] The above are merely examples of the present application and are not intended to limit the scope of protection of the present application. Those skilled in the art will appreciate that various modifications and variations are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present application shall be included within the scope of protection of the present application.
Claims
1. A method for displaying game cartoon characters based on virtual reality technology, characterized in that: The method is executed by a virtual reality character display system, and includes: Acquiring three-dimensional model data of at least one game and animation character in a target display area through a user interaction terminal; For a game character, the VR content management platform determines the accuracy level of the character's dynamic rendering based on the 3D model data. The virtual reality content management platform determines a real-time rendering strategy based on the dynamic rendering accuracy level of at least one game animation character and terminal performance parameters. The real-time rendering strategy includes a model face reduction coefficient, a texture compression rate, and a skeletal animation update frequency. The virtual reality content management platform sends the real-time rendering strategy to the virtual reality supervision platform and the user interaction terminal, and generates rendering instructions and sends them to the virtual reality rendering engine platform, so that the rendering engine performs real-time rendering based on the rendering instructions. The virtual reality rendering engine platform is configured as the rendering core of the head-mounted display device; Based on the target display area and real-time rendering strategy, the virtual reality monitoring platform obtains the frame rate fluctuation data of the head-mounted display device according to the preset frame rate threshold, and determines the scene rendering stability based on the frame rate fluctuation data; In response to the scene rendering stability being lower than a preset stability threshold, the virtual reality supervision platform sends a rendering optimization instruction to the virtual reality content management platform based on the scene rendering stability, and adjusts the model data acquisition frequency and video memory cleanup cycle based on the scene rendering stability, and releases video memory resources according to the video memory cleanup cycle.
2. The method according to claim 1, characterized in that Based on the 3D model data, the accuracy level of dynamic rendering of game characters is determined as follows: Extract model complexity features of game and animation characters based on 3D model data; Based on the model complexity characteristics, the visual detail weights of game and animation characters at multiple viewing distances are calculated; And based on the visual detail weight, determine the dynamic rendering accuracy level.
3. The method according to claim 2, characterized in that Based on the model complexity characteristics, the weights of visual details of game and animation characters at multiple viewing distances are calculated, including: Setting at least one distance evaluation interval; For a distance evaluation interval, based on the distance evaluation interval, model material properties, and model complexity characteristics, the detail weight prediction model is used to output the visual detail weight of the distance evaluation interval. The detail weight prediction model is a convolutional neural network model, and its training process includes: Acquire multiple first training samples and first labels, where the first training samples include a sample distance interval, a sample material attribute, and a sample model complexity, and the first label is a sample visual detail weight corresponding to the first training sample; Training an initial detail weight prediction model based on the plurality of first training samples and the first labels; And iteratively optimize the detail weight prediction model until the output error is lower than the preset loss threshold.
4. The method according to claim 3, characterized in that The input of the detail weight prediction model also includes the coordinates of the user's gaze focus, which are determined based on the projection position of the eye tracking data in three-dimensional space; Training the initial detail weight prediction model includes: Based on the difference in rendering time between the focus area and the non-focus area in the historical rendering log, the first label is weighted and modified to generate a second label; And based on the plurality of first training samples and the second labels, an initial detail weight prediction model is trained.
5. The method according to claim 2, characterized in that Determining a real-time rendering strategy based on a dynamic rendering accuracy level of at least one game animated character and terminal performance parameters includes: Calculate rendering efficiency coefficients at different precision levels based on the frame rate attenuation gradient of historical rendering records; And determine the real-time rendering strategy based on the dynamic rendering accuracy level, rendering efficiency coefficient and terminal performance parameters of at least one game animation character.
6. The method according to claim 1, characterized in that Determining a real-time rendering strategy based on a dynamic rendering accuracy level of at least one game animated character and terminal performance parameters includes: Construct a performance optimization objective function with rendering strategy as a variable: ; in, The time it takes to render the i-th game character. is the video memory occupancy value of the i-th game cartoon character, , is the weight coefficient determined by the terminal performance parameters, n is the total number of game anime characters, and i is the index of the game anime character; Based on the performance optimization objective function, the set of candidate rendering strategies is iterated by gradient descent until the convergence condition is met and the real-time rendering strategy is output.
7. The method according to claim 6, characterized in that The method also includes: During the presentation, in response to the smoothness of the movements of the rendered game cartoon character falling below a preset threshold, identifying a high-load model component causing lag; Predict the game anime characters to be rendered that contain high-load model components and mark them as priority optimization targets; And based on the priority optimization object set, reconstruct the subsequent real-time rendering strategy.
8. The method according to claim 1, characterized in that Acquiring three-dimensional model data of at least one game and animation character in a target display area through a user interaction terminal includes: Real-time collection of user hand motion vectors ; in, is the coordinate value of the hand in three-dimensional space, is the hand rotation angle; Will Mapped to the skeleton nodes of the game animation characters, driving the model to generate interactive action data; The three-dimensional model data includes interactive action data.
9. The method according to claim 1, characterized in that Based on the target display area and real-time rendering strategy, the VR monitoring platform obtains frame rate fluctuation data of the head-mounted display device according to the preset frame rate threshold, including: When multiple user interaction terminals are detected in the same virtual scene, a cross-terminal gesture synchronization channel is established; Synchronize the position matrix and rotation quaternion of the game characters in each terminal through the channel: ; in, is the position matrix of the game cartoon characters in the kth terminal, N is the total number of terminals in the same virtual scene, is the spherical linear interpolation algorithm, is the network delay compensation factor, k is the index of the terminal, is the position matrix after synchronization, They are the rotation quaternions of the game characters in different terminals. is the rotation quaternion after synchronization; Calculates frame rate fluctuation data based on the synchronized position matrix and rotation quaternion.
10. The method according to claim 1, characterized in that The VR supervision platform sends rendering optimization instructions to the VR content management platform based on the scene rendering stability, including: Generate a 3D heat map based on the scene rendering stability. The heat map uses color gradients to identify the rendering load intensity of different areas in the virtual scene. The heat map is overlaid and displayed on a supervisory view layer of a head-mounted display device; the rendering optimization instruction includes region identification data of the heat map.
Citation Information
Cited By
Intelligent image quality optimization method and system applied to audio and video rendering
CN121078249A
Cross-platform application construction system and method
CN121187613A
A system and method for building cross-platform applications
CN121187613B
WebGL-based radar body rendering performance optimization system
CN121600151A
Mask art multi-dimensional visual state display method and mask art multi-dimensional visual state display system
CN121837502A