High-resolution rendering method and system of XR display screen
By performing distributed sampling and spatial geometric correction of the optical characteristics of the XR display, combined with local enhancement processing and multi-channel synchronization control, the problems of low rendering quality and computing efficiency of XR displays in the prior art are solved, and high-resolution rendering and resource optimization are achieved.
Patent Information
- Application Number
- CN202510608294.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-06-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The rendering method of existing XR displays cannot dynamically optimize users' real-time visual needs, resulting in waste of computing resources and poor display effects, and there are problems such as large spatial geometric registration errors and uneven rendering resolution.
By distribute the optical characteristics of the XR display, a four-dimensional rendering sampling matrix is established, and Latin hypercube sampling is performed to obtain multiple sets of rendering parameter combinations and high-resolution rendering image data. Then, spatial geometry correction processing is performed, local enhancement processing is performed, and rendering channel parameters of multiple display channels are calculated, and multi-channel synchronization control is performed.
It significantly improves the rendering quality and computing efficiency of XR display, reduces the system's resource consumption, and improves the registration accuracy and user experience of virtual and real scenes.
Smart Images

Figure CN120147499A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of XR display screens, and particularly to a high-resolution rendering method and system for XR display screens. Background Art
[0002] With the rapid development of virtual reality and augmented reality technologies, XR display screens, as key visual interaction devices, are facing dual challenges of rendering quality and computing efficiency. Traditional XR display rendering methods adopt fixed resolution and unified sampling strategies, and cannot be dynamically optimized according to the real-time visual needs of users, resulting in waste of computing resources and poor display effects.
[0003] Currently, there are problems such as large spatial geometric registration errors and uneven rendering resolutions in the fusion rendering of virtual and real scenes in XR displays. Due to the complex optical characteristics of XR displays, including multi-dimensional parameters such as field of view angle, spatial distortion, optical path difference, and fusion boundary, traditional rendering methods are difficult to accurately capture and model these characteristics, affecting the fusion effect of virtual and real scenes and the user experience. Summary of the Invention
[0004] This application provides a high-resolution rendering method and system for XR display screens, thereby improving the rendering quality and computing efficiency of XR displays, and at the same time reducing the resource consumption of the system.
[0005] In the first aspect of this application, a high-resolution rendering method for XR display screens is provided. The high-resolution rendering method for XR display screens includes: Performing distributed sampling on the optical characteristics of the XR display screen to obtain a four-dimensional rendering sampling matrix, and performing Latin hypercube sampling to obtain multiple sets of rendering parameter combinations and corresponding high-resolution rendering image data; Performing spatial geometric correction processing on the rendering parameter combinations and the high-resolution rendering image data to obtain spatial geometric correction parameters; Performing local enhancement processing on the fixation point coordinate sequence and the spatial geometric correction parameters to obtain local resolution enhancement data; Calculating the rendering channel parameters of multiple display channels according to the spatial geometric correction parameters and the local resolution enhancement data, and performing multi-channel synchronous control to obtain target stereoscopic image display parameters.
[0006] In the second aspect of this application, a high-resolution rendering system for XR display screens is provided. The high-resolution rendering system for XR display screens includes: A distributed sampling module, configured to perform distributed sampling on the optical characteristics of the XR display screen to obtain a four-dimensional rendering sampling matrix, and perform Latin hypercube sampling to obtain multiple sets of rendering parameter combinations and corresponding high-resolution rendering image data; A calibration processing module, configured to perform spatial geometric calibration processing on the combination of rendering parameters and the high-resolution rendering image data to obtain spatial geometric calibration parameters; An enhancement processing module, configured to perform local enhancement processing on the sequence of fixation point coordinates and the spatial geometric calibration parameters to obtain local resolution enhancement data; A synchronization control module, configured to calculate rendering channel parameters of multiple display channels according to the spatial geometric calibration parameters and the local resolution enhancement data, and perform multi-channel synchronization control to obtain target stereoscopic image display parameters.
[0007] Compared with the prior art, the present application has the following beneficial effects: By establishing a four-dimensional rendering sampling matrix and a Latin hypercube sampling strategy, the optical characteristic parameters of XR display are systematically collected and quantified. The spatial geometric calibration is performed using a dual-branch convolutional neural network structure, realizing the deep fusion of rendering parameter features and spatial features, and significantly improving the registration accuracy of virtual and real scenes. Based on the local enhancement processing method of residual convolution encoding and deconvolution decoding, high-resolution rendering of the fixation area and dynamic downsampling of the surrounding area are realized, optimizing the allocation efficiency of computing resources. A multi-channel rendering control strategy based on particle swarm optimization is designed. By adaptively adjusting rendering parameters and computing resource allocation, the collaborative working effect of each display channel is ensured. The sliding time window mechanism and the PID control algorithm are introduced to realize real-time adjustment of rendering parameters and multi-channel synchronization control, improving the stability and response speed of the XR display system. Through multi-dimensional parameter optimization and multi-channel collaborative control, the rendering quality and computing efficiency of XR display are significantly improved, while reducing the resource consumption of the system. Description of the Drawings
[0008] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0009] The structures, proportions, sizes, etc. shown in the drawings of this specification are only used to cooperate with the content disclosed in the specification for those who are familiar with this technology to understand and read, and are not used to limit the limited conditions under which the present invention can be implemented. Therefore, they do not have technical substance. Any modification of the structure, change of the proportional relationship, or adjustment of the size should still fall within the scope that can be covered by the technical content disclosed in the present invention without affecting the effects that the present invention can produce and the purposes that can be achieved.
[0010] Figure 1 It is a schematic flowchart of a high-resolution rendering method for an XR display screen provided by an embodiment of the present invention; Figure 2 It is a schematic block diagram of the structure of the high-resolution rendering system of the XR display screen provided by the embodiment of the present invention. Specific embodiments
[0011] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0012] The flowchart shown in the accompanying drawings is only an example, and does not necessarily include all contents and operations / steps, nor does it necessarily need to be executed in the described order. For example, some operations / steps can also be decomposed, combined or partially merged, so the actual execution order may change according to the actual situation.
[0013] It should also be understood that the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. As used in the specification of this application and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms.
[0014] It should be further understood that the term "and / or" used in the specification of this application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations. Please refer to Figure 1 , an embodiment of the high-resolution rendering method of the XR display screen in the embodiment of this application includes: Step 100: Perform distributed sampling on the optical characteristics of the XR display screen to obtain a four-dimensional rendering sampling matrix, and perform Latin hypercube sampling to obtain multiple sets of rendering parameter combinations and corresponding high-resolution rendering image data; It can be understood that the execution subject of this application can be the high-resolution rendering system of the XR display screen, or a terminal or a server, and specific limitations are not made here. This application takes the server as the execution subject as an example for description.
[0015] Specifically, the optical system of the XR display screen is modeled, and sampling division is performed for different spatial dimensions. In the horizontal direction, a uniform division method is adopted, and the entire field of view angle range is sliced at equal intervals to obtain multiple field of view angle sampling nodes. In the vertical direction, since the optical distortion effect is often more obvious at certain angles, an increasing interval method is used for division, so that the sampling is denser in the area with larger distortion, and the density of sampling points is reduced in the area with smaller distortion, obtaining spatial distortion sampling nodes. At the same time, considering the optical performance of the XR display screen at different depth levels, division is carried out along the Z-axis direction to obtain the optical path difference sampling layer. Since light is affected by refraction, reflection, etc. when passing through different media, resulting in differences in the optical path at different depths, the Z-axis is divided according to certain levels, so that the optical characteristics at different depths can be fully measured and calculated, thereby accurately characterizing the influence of the optical path difference. Grid division is performed on the XY plane to obtain the fusion boundary sampling grid. This grid division method helps to accurately sample the pixel boundaries and optical fusion regions of the XR display screen, ensuring the correct processing of the optical transition problem in the boundary region in multi-channel or multi-level display modes and improving the display quality. Spatial feature extraction is performed on the sampling points in the field of view angle sampling nodes, spatial distortion sampling nodes, optical path difference sampling layer, and fusion boundary sampling grid. For each sampling point, its three-dimensional coordinates, field of view angle, distortion coefficient, optical path difference, and boundary intensity value are extracted. These parameters jointly determine the optical performance of the XR display screen under different conditions. The three-dimensional coordinates are used to determine the position of each sampling point in space, and the field of view angle is used to describe the imaging characteristics of the point at different viewing angles; the distortion coefficient characterizes the deformation effect of the optical system on the image, the optical path difference reflects the temporal and spatial changes caused by different paths during light propagation, and the boundary intensity value is used to measure the optical fusion effect in the display edge region. These data jointly constitute a complete description of the optical characteristics of the XR display screen. The sampling data is subjected to feature matrix conversion to obtain a four-dimensional rendering sampling matrix. By arranging the data of all sampling points according to a certain dimension, each element in the matrix accurately reflects the optical characteristics at a specific position and viewing angle, constituting a complete high-dimensional data representation. Latin hypercube sampling is performed based on the four-dimensional rendering sampling matrix. Latin hypercube sampling is a statistically efficient sampling method that evenly distributes sampling points in a high-dimensional space, thereby ensuring the representativeness of different combinations of rendering parameters, obtaining multiple different combinations of rendering parameters, and each set of parameter combinations represents a possible display condition. Finally, multiple sets of rendering parameter combinations and corresponding high-resolution rendering image data are obtained.
[0016] The field-of-view angle dimension in the four-dimensional rendering sampling matrix is divided into intervals to ensure uniform coverage of optical characteristics within different angular ranges. The division process is optimized based on the optical characteristics of the XR display screen, such that sampling at different field-of-view angles covers the entire field-of-view range, while ensuring denser sampling in areas with drastic angular changes, thereby obtaining the field-of-view angle parameters. The spatial distortion dimension in the four-dimensional rendering sampling matrix is sampled in layers to describe the geometric deformation characteristics of the XR display screen at different positions. Since optical distortion is usually non-uniform and has a higher distortion degree in certain specific areas, the method of layer sampling effectively improves the modeling accuracy, that is, higher-density sampling is used in areas with larger distortion, while the number of sampling points is reduced in areas with smaller distortion to optimize the allocation of computing resources. In this way, the distortion coefficient parameters are obtained. At the same time, the optical path difference dimension in the four-dimensional rendering sampling matrix is depth-segmented to obtain optical parameters at different depth levels. The optical path difference is caused by the path change of light when propagating at different depth levels. The method of depth segmentation ensures that the optical characteristics at different depth levels can be fully sampled, enabling accurate modeling of the optical changes from the near view to the far view. Through this step, the depth level parameters are obtained, which are used for the rendering calculation of the XR display screen, enabling reasonable rendering of images in different depth-of-field ranges, thereby enhancing the stereoscopic vision effect and the user's sense of immersion. While modeling the optical distortion and depth information, the bandwidth of the fusion boundary dimension in the four-dimensional rendering sampling matrix is divided to ensure optimized optical fusion effects in the boundary area in the multi-channel or multi-screen splicing display mode. Through bandwidth division, the boundary transition zone width parameters are obtained, which are used for rendering calculation to ensure smooth optical transition in the boundary area, thereby avoiding visual breaks or color non-uniformity and enhancing the overall display consistency. Orthogonal combinations are made according to the field-of-view angle parameters, distortion coefficient parameters, depth level parameters, and boundary transition zone width parameters, and the possible rendering parameter space is traversed in a systematic manner to ensure that the selected parameter combinations represent the entire parameter range and effectively reduce the computational overhead to form multiple sets of rendering parameter combinations. High-resolution rendering image data corresponding to each set of rendering parameter combinations are collected under multiple environmental lighting conditions.
[0017] Step 200: Perform spatial geometric correction processing on the rendering parameter combinations and the high-resolution rendering image data to obtain spatial geometric correction parameters; Specifically, the rendering parameter combination is input into the first branch of the dual-branch convolutional neural network. This branch is used to process the input rendering parameters and convert them into feature vectors for geometric correction through feature mapping. The first branch consists of three fully connected layers, and after each fully connected layer, there is a batch normalization layer and a ReLU activation function to ensure that the model can converge stably during training and enhance the model's non-linear fitting ability. The introduction of the batch normalization layer effectively alleviates the problems of gradient vanishing and gradient explosion, while accelerating the training convergence. The ReLU activation function helps to improve the model's expressive ability, enabling it to capture the complex mapping relationships in the rendering parameters and generate highly abstract rendering parameter features. At the same time, the high-resolution rendered image data is input into the second branch of the dual-branch convolutional neural network. This branch extracts spatial feature information from the input image data to provide the geometric features of the image for subsequent correction calculations. The second branch consists of four convolutional blocks, and each convolutional block contains two 3×3 convolutional layers, a batch normalization layer, and a ReLU activation layer. Among them, the 3×3 convolutional layer is used twice in each convolutional block to ensure that the network can effectively capture local spatial structure features and gradually expand the receptive field, thereby extracting more global spatial information. Through the role of the batch normalization layer, while maintaining the stability of network training, the separability of features is improved, making it easier for subsequent fusion calculations to combine the rendering parameter features and the image spatial features. The introduction of the ReLU activation layer can increase the non-linear expressive ability of the model and prevent the gradient vanishing problem, enabling the model to effectively learn the distortion features, boundary features, and perspective information in the image during the deep learning process and generate a spatial feature map. The rendering parameter features and the spatial feature map are subjected to attention fusion processing to generate a fused feature map. During the attention fusion process, the correlation between the rendering parameter features and the spatial feature map is calculated, and the spatial features are dynamically weighted through attention weights to highlight the regional information most valuable for the correction task and suppress irrelevant interference information. Through this step, a fused feature map is obtained. The fused feature map is subjected to 3×3 convolutional processing to refine the features and remove redundant information, obtaining an intermediate feature map. The intermediate feature map is subjected to 3×3 convolutional processing to continue enhancing the feature expressive ability and improving the model's adaptability to complex geometric distortion situations, obtaining a preprocessed feature map. Based on the preprocessed feature map, 1×1 convolutional processing is performed to finally extract the spatial geometric correction parameters including the perspective transformation matrix, distortion correction coefficients, and registration offsets. The perspective transformation matrix is used to correct the geometric deformation of the image under different perspectives, so that the finally displayed image conforms to the perspective relationship in the real world. The distortion correction coefficients are used to correct geometric distortion problems such as barrel distortion, pincushion distortion, or radial distortion caused by the non-ideal imaging characteristics of the optical system, ensuring that the XR display screen can correctly present the picture.The registration offset is used to adjust the spatial alignment between different display areas, so that in the case of multi-channel rendering or tiled display, each part of the screen can achieve pixel-level precise alignment, thereby eliminating the misalignment phenomenon and finally achieving a high-precision XR display effect.
[0018] Step 300: Perform local enhancement processing on the sequence of fixation point coordinates and the spatial geometry correction parameters to obtain locally enhanced resolution data; It should be noted that the sequence of fixation point coordinates is input into six residual convolutional blocks for encoding. The role of the residual convolutional blocks is to extract the spatial distribution features of the fixation point coordinates, while ensuring the efficient transmission of information and the stable propagation of gradients in the deep network structure. Each residual convolutional block contains two 3×3 convolutional layers and adopts a skip connection structure, enabling the input features to maintain the original information while undergoing non-linear transformation, thereby preventing the problem of gradient disappearance and enhancing the training stability of the deep network. Through this encoding process, a fixation point feature map is obtained. This feature map represents the spatio-temporal distribution of the fixation points and captures the dynamic changes in the user's visually attended regions, so as to adjust the resolution allocation strategy for different regions in subsequent processing. At the same time, feature expansion processing is performed on the spatial geometric correction parameters to generate a correction parameter feature map. In the feature expansion processing, a series of fully connected layers are used to perform non-linear mapping on the correction parameters, and feature normalization techniques are used to ensure the consistency of the numerical range, making it suitable for subsequent deep learning calculations. The mapped correction parameter features are input into multiple convolutional layers for feature expansion, and multi-scale features of the spatial geometric correction parameters are extracted through local receptive fields, so as to perform more accurate matching with the fixation point feature map in the subsequent fusion process. After the expansion processing, a correction parameter feature map is obtained. Self-attention weight calculation is performed on each feature channel in the fixation point feature map to enhance the model's perception ability of the attended regions. The self-attention mechanism is used to calculate the importance weights of each channel, enabling the network to pay more attention to the key features in the fixation point region while suppressing the interference information in the irrelevant regions. The correlation between each channel and other channels is calculated, and the features are reconstructed through weighted operations to ensure that the information interaction between different feature channels maximally utilizes the spatial relationship, thereby enhancing the feature expression ability. After the self-attention calculation, an attention-weighted feature map is obtained. The attention-weighted feature map and the correction parameter feature map are concatenated in the channel dimension to obtain multi-scale features. The multi-scale features are upsampled, and the skip connection features corresponding to the encoding stages are fused at each scale level to restore high-resolution information. In this process, the method of transposed convolution or bilinear interpolation is used to gradually magnify the feature map, and the skip connection features from the encoding stage are combined at each layer to ensure that important spatial information is not lost during the upsampling process, generating a locally enhanced feature map. Spatial mapping transformation is performed on the locally enhanced feature map, and different resolution weights are assigned to different regions in the locally enhanced feature map according to the distribution of the attended regions to achieve precise resolution control. Within the region covered by the fixation point, the resolution is increased to A times the original resolution to ensure that the user's visually attended regions obtain clearer image details, while for the regions far from the fixation point, dynamic downsampling is used to compress their resolution to reduce the computational cost and optimize the resource utilization of the XR display screen.This step is completed by Gaussian weight mapping or an adaptive filtering strategy based on deep learning. Among them, the Gaussian weight mapping method generates a spatial weight map based on the fixation point coordinates and adjusts the resolution through pixel-by-pixel weighted operations. The adaptive filtering strategy, on the other hand, directly predicts the optimal resolution level for each pixel using a trained deep network to achieve the optimal computational resource allocation effect. Through spatial mapping transformation, locally enhanced resolution data is obtained.
[0019] Step 400: Calculate the rendering channel parameters of multiple display channels based on the spatial geometric correction parameters and the locally enhanced resolution data, and perform multi-channel synchronization control to obtain the target stereoscopic image display parameters.
[0020] Specifically, the rendering parameters of N display channels are organized into an N×M-dimensional state vector, where M represents the number of parameters required for each display channel. These parameters include the computational resource allocation weight, the rendering priority coefficient, and the cache policy parameters. By structuring and organizing the parameters of each display channel, a state representation matrix is formed. An initial particle swarm position matrix is constructed based on these channel parameters. Each particle in this matrix represents a possible parameter configuration scheme, and its initialization strategy adopts a uniform random distribution to ensure that the search space can be fully covered and sufficient solution candidates can be provided to improve the global search ability of the optimization algorithm. A multi-objective function is constructed for the spatial geometric correction parameters to construct a fitness function, which is used to measure the comprehensive performance of different particle swarm configuration schemes in terms of spatial correction optimization and rendering quality improvement. The fitness function combines multiple factors such as spatial geometric error, computational resource consumption, cache optimization strategy, and rendering accuracy in the fixation point area to ensure that the final optimization result maximally improves the rendering quality of the XR display under the condition of limited computational resources. To adapt the particle swarm optimization process to the dynamic changes of the local resolution enhancement data, the positions of the particle swarm are updated according to this enhancement data to obtain the particle swarm velocity matrix. The velocity update strategy is based on the inertia factor, the cognitive factor, and the social factor. The inertia factor controls the global exploration ability of the particle swarm search, the cognitive factor is used to enhance local search, and the social factor is used to guide the particles to converge to the optimal solution. With the completion of the calculation of the particle swarm velocity matrix, the positions of the particle swarm are iteratively updated. In each iteration, the fitness value of the current particle is calculated, and the local optimal solution and the global optimal solution are recorded to ensure that the search process gradually approaches the optimal configuration. When the number of iterations reaches B times, the particle swarm search is stopped, and the particle position with the highest fitness value is selected as the final optimization scheme. This optimized particle position represents the optimal display channel rendering parameter configuration scheme for subsequent computational resource allocation. On this basis, the computational resource allocation scheme for each display channel is determined according to the optimized particle position. To ensure the reasonable allocation of resources, the computational resources are hierarchically divided according to the rendering priority to obtain the resource allocation matrix. This hierarchical strategy combines the resolution requirements of the fixation point area, the spatial distortion correction requirements, and the computational overhead of multi-channel rendering to ensure that high-priority areas obtain higher computational resources, while low-priority areas adopt a lower computational overhead strategy to optimize the overall rendering performance. The resource allocation matrix is decoded and mapped according to the display channels to generate the rendering channel parameters. The decoding and mapping process involves converting the resource allocation matrix into the specific computational parameters of each channel, including the sampling density parameter, the update frequency parameter, and the cache policy parameter. Among them, the sampling density parameter determines the rendering accuracy distribution of each display channel, the update frequency parameter is used to dynamically adjust the rendering refresh rate to adapt to different user interaction scenarios, and the cache policy parameter affects the storage and transmission methods of multi-channel data, thereby optimizing the overall computational efficiency and data transmission bandwidth.Through this decoding process, it is ensured that the rendering system dynamically adjusts the rendering strategies of different display channels according to the current computing resource situation, so as to improve the rendering fluency and real-time performance of the XR display screen. The rendering channel parameters are synchronously controlled in multiple channels according to the sliding time window mechanism. By setting a sliding time window, the rendering progress of different channels is coordinated, so that the frame update rates of all channels are kept consistent, and the display delay problem between different channels is avoided. The sliding time window mechanism aligns the frame timestamps of each rendering channel and allows a small range of time errors within a certain time range to reduce the computational load while ensuring the visual synchronization effect. In this process, the prediction compensation algorithm is combined to optimize the dynamic adjustment strategy of the time window, so that the rendering system can still maintain a high rendering stability under high load conditions and ensure that the XR display screen provides a high-quality stereoscopic image display effect. The target stereoscopic image display parameters are obtained.
[0021] The rendering channel parameters are divided into time windows, and three controller gain coefficients are set for each display channel to form time window control parameters. During the time window division process, the rendering channel parameters are segmented along the time dimension, so that the rendering parameters in different time periods remain relatively stable, while allowing dynamic adjustment when necessary. The size of the time window is adaptively adjusted according to the frame rate requirements of the XR display screen, the user interaction response speed, and the multi-channel synchronization error, in order to ensure the best balance between low latency and high synchronization accuracy. And within each time window, the rendering state of each display channel is affected by computing resources, data transmission bandwidth, and optical distortion correction. Therefore, three controller gain coefficients are set for it. These coefficients are used to adjust spatial registration, resolution uniformity, and time synchronization respectively, so as to ensure that all channels always maintain a good cooperative state throughout the rendering process. The rendering errors of each display channel are calculated to obtain error feedback data, which includes spatial registration accuracy error, resolution uniformity error, and time synchronization error. Among them, the spatial registration accuracy error reflects the image geometric alignment error between different channels. The sources of this error include optical distortion, projection coordinate conversion error, and the accuracy limitations of the screen hardware itself, etc.; the resolution uniformity error describes the consistency of rendering quality between different channels, which is mainly caused by uneven computing resource allocation, unreasonable local enhancement strategies, or unbalanced rendering loads, etc.; the time synchronization error measures whether the frame updates of each channel are consistent, which is mainly caused by computing latency, data transmission latency, and system load fluctuations, etc. During the error calculation process, based on the current rendering channel state, it is compared with the ideal state, and statistical analysis methods are used to calculate the mean value, variance, and maximum error value of different error terms, in order to provide comprehensive error feedback data. Based on the error feedback data, the controller gain coefficients are dynamically adjusted to meet the optimization requirements in different rendering states. The adjustment of the controller gain coefficients adopts an adaptive adjustment strategy, that is, according to the change trend of the error, the gain parameters are updated in real time, so that the control intensity increases when the error is large, and the adjustment strength decreases when the error is small, in order to improve the stability of the system. The updated controller parameters are used for PID (Proportional-Integral-Derivative) control operations to optimize the synchronization of the rendering channels. During the PID control calculation process, the proportional term is used to quickly respond to the error change, the integral term is used to correct the long-term error, and the derivative term is used to suppress the situation where the error changes too fast, in order to ensure that the system can converge smoothly to the optimal synchronization state. The data output by the PID control is represented as a dynamic adjustment matrix, which contains the best rendering parameter adjustment amounts of each channel in different time windows. The controller output data is input into the rendering parameter scheduling module of each display channel, and the sampling density, update frequency, and cache policy are jointly optimized according to the computing resource allocation weights to generate a scheduling control matrix.The allocation of computing resources is optimized according to the rendering priorities of different channels, so that more computing resources are allocated in critical rendering areas (such as the user's fixation point area), while the allocation of computing resources is reduced in secondary areas (such as the peripheral vision) to improve the overall computing efficiency. The optimization of the sampling density is adjusted according to the error feedback data to minimize the resolution uniformity error, while ensuring that the local resolution enhancement strategy can effectively play its role; the optimization of the update frequency is adjusted in combination with the time synchronization error to make the frame rates of different channels as consistent as possible, thereby reducing the visual desynchronization problem between multiple channels; the optimization of the caching strategy is adjusted in combination with the spatial registration error to ensure that the data access speed matches the rendering calculation requirements and prevent latency problems caused by improper cache management. Through this joint optimization process, a scheduling control matrix is obtained, which is used to guide how each display channel reasonably allocates computing resources and adjusts the rendering strategy within different time windows. The scheduling control matrix is decoded to be converted into specific rendering parameters corresponding to each display channel, including rendering resolution parameters, refresh timing parameters, and cache configuration parameters. The rendering resolution parameters determine the final output image quality of each channel and are dynamically adjusted in combination with the local resolution enhancement data to ensure that the highest quality image can be provided in the user's fixation point area, while the resolution is appropriately reduced in the surrounding area to optimize the computing efficiency. The refresh timing parameters are used to ensure the frame synchronous update of all display channels to eliminate the flicker or misalignment problems caused by visual desynchronization, while the cache configuration parameters are used to optimize the data access and transmission strategies to ensure that the rendering calculation can be efficiently executed under limited hardware resources. Through the above steps, the target stereoscopic image display parameters are finally obtained.
[0022] In the embodiments of this application, by establishing a four-dimensional rendering sampling matrix and a Latin hypercube sampling strategy, the optical characteristic parameters of XR display are systematically collected and quantified. The dual-branch convolutional neural network structure is used for spatial geometric correction, realizing the deep fusion of rendering parameter features and spatial features, and significantly improving the registration accuracy of virtual and real scenes. Based on the local enhancement processing method of residual convolutional coding and deconvolution decoding, high-resolution rendering in the fixation area and dynamic downsampling in the surrounding area are realized, optimizing the allocation efficiency of computing resources. A multi-channel rendering control strategy based on particle swarm optimization is designed. By adaptively adjusting the rendering parameters and computing resource allocation, the collaborative working effect of each display channel is ensured. The sliding time window mechanism and the PID control algorithm are introduced to realize the real-time adjustment of rendering parameters and multi-channel synchronous control, improving the stability and response speed of the XR display system. Through multi-dimensional parameter optimization and multi-channel collaborative control, the rendering quality and computing efficiency of XR display are significantly improved, while the resource consumption of the system is reduced.
[0023] In a specific embodiment, the process of executing step 100 may specifically include the following steps: Perform uniform division processing on the horizontal direction of the XR display screen to obtain field-of-view angle sampling nodes; Perform incremental interval division on the vertical direction of the XR display screen to obtain spatial distortion sampling nodes; Perform depth-level division along the Z-axis direction of the XR display screen to obtain optical path difference sampling layers; Perform grid division on the XY plane to obtain fusion boundary sampling grids; Extract spatial features from the sampling points in the field-of-view angle sampling nodes, spatial distortion sampling nodes, optical path difference sampling layers, and fusion boundary sampling grids to obtain sampling data including three-dimensional coordinates, field-of-view angle, distortion coefficient, optical path difference, and boundary intensity values; Convert the sampling data into a feature matrix to obtain a four-dimensional rendering sampling matrix; Perform Latin hypercube sampling based on the four-dimensional rendering sampling matrix to obtain multiple groups of rendering parameter combinations and corresponding high-resolution rendering image data.
[0024] Specifically, uniformly divide the field-of-view angle in the horizontal direction to obtain field-of-view angle sampling nodes. Assume that the horizontal field-of-view angle range of the XR display screen is , then divide it at equal intervals to generate sampling points, and the angle value of each point is expressed as: where represents the position of the th field-of-view angle sampling point, and represents the angle increment between each field-of-view angle node. Uniformly cover the sampling points within the entire horizontal field of view, so that subsequent rendering calculations fully consider the optical characteristics at different horizontal angles. At the same time, adopt an incremental interval division method in the vertical direction to adapt to the non-uniform distribution characteristics of spatial distortion. Assume that the vertical field-of-view angle range of the XR display screen is , then adopt an exponential growth interval strategy to make the sampling denser in the area with larger distortion and sparser in the area with smaller distortion. Define the sampling angle as: where represents the th field-of-view angle sampling point in the vertical direction, is a parameter to control the incremental rate, and is the number of sampling points in the vertical direction. When , the sampling points near are dense, while those near The sampling points are relatively sparse. And depth levels are divided along the Z-axis direction of the XR display screen to consider the influence of optical path differences on the rendering results. Assume that the depth range in the Z-axis direction is , and the division is carried out according to logarithmic intervals to ensure higher sampling accuracy in the area close to the display screen, while appropriately reducing the sampling density in the area far from the display screen. Define the depth level sampling points as: where, represents the th depth sampling point, is the number of sampling layers in the depth direction. Higher optical path difference resolution is achieved in the closer depth layers to improve the rendering quality. At the same time, grid division is carried out on the XY plane to obtain the fusion boundary sampling grid. Assume that the XY plane size of the XR display screen is , then it is divided into cells in a uniform grid manner, and the center point coordinates of each grid are expressed as: where, respectively represent the center point coordinates of the th grid, are the step sizes of the grid division respectively. Spatial feature extraction is performed on the sampling points in the field of view angle sampling nodes, spatial distortion sampling nodes, optical path difference sampling layers, and fusion boundary sampling grids to obtain complete optical data. The spatial features of each sampling point include the three-dimensional coordinates , the field of view angle , the distortion coefficient , the optical path difference and the boundary strength . These parameters are calculated as follows respectively: where, is the distortion model, is the optical path calculation function, is the fusion boundary calculation function. These feature data are organized into a feature matrix and transformed to obtain a four-dimensional rendering sampling matrix, and each element of this matrix represents the complete optical information of a sampling point. After completing the construction of the four-dimensional rendering sampling matrix, in order to improve the calculation efficiency, the Latin hypercube sampling method is used for optimized sampling to ensure that the selected combination of rendering parameters has high representativeness. Assume that the dimensions of the four-dimensional rendering sampling matrix are , then define the sampling interval on each parameter dimension: Among them, represents different parameter dimensions. The Latin hypercube sampling method ensures uniform distribution of sampling points in each dimension while maximizing the independence between sampling points to improve sampling efficiency. The rendering parameter combinations based on sampling are used to generate high-resolution rendering image data, enabling the XR display to achieve optimal rendering under different optical conditions.
[0025] In a specific embodiment, the process of performing Latin hypercube sampling based on a four-dimensional rendering sampling matrix to obtain multiple sets of rendering parameter combinations and corresponding high-resolution rendering image data may specifically include the following steps: Divide the field of view angle dimension in the four-dimensional rendering sampling matrix to obtain the field of view angle parameters; Perform stratified sampling on the spatial distortion dimension in the four-dimensional rendering sampling matrix to obtain the distortion coefficient parameters; Perform depth segmentation on the optical path difference dimension in the four-dimensional rendering sampling matrix to obtain the depth level parameters; Perform bandwidth division on the fusion boundary dimension in the four-dimensional rendering sampling matrix to obtain the boundary transition zone width parameters; Orthogonally combine the field of view angle parameters, distortion coefficient parameters, depth level parameters, and boundary transition zone width parameters to obtain multiple sets of rendering parameter combinations; Perform image acquisition on multiple sets of rendering parameter combinations respectively under various environmental lighting conditions to obtain the high-resolution rendering image data corresponding to each set of rendering parameter combinations.
[0026] Specifically, divide the field of view angle dimension in the four-dimensional rendering sampling matrix. Assume that the horizontal and vertical field of view angle ranges of the XR display are and , then divide the field of view angle at equal intervals and to form horizontal field of view angle sampling points and vertical field of view angle sampling points. The calculation method for each angle point is as follows: Among them, and respectively represent the th and th field of view angle parameters in the horizontal and vertical directions, ensuring that the sampling points evenly cover the entire field of view range so that subsequent rendering calculations can fully consider the optical characteristics under different perspectives. To optimize the sampling of spatial distortion, perform stratified sampling on the spatial distortion dimension in the four-dimensional rendering sampling matrix. Assume that the value range of the spatial distortion coefficient is Adopt a non-uniform layering method to adapt to the distortion characteristics of different regions. Define the distortion coefficient sampling points as: where, represents the distortion coefficient of the th layer, is the number of distortion sampling layers, is the parameter for controlling the increasing rate. When , the sampling points near are denser, while the sampling points near are sparser, so as to provide higher sampling accuracy in the regions with larger distortion. This method can effectively optimize the distortion correction calculation and ensure higher resolution in the regions with significant optical distortion. At the same time, perform depth segmentation on the optical path difference dimension in the four-dimensional rendering sampling matrix to determine the parameters of different depth levels. Assume that the depth range of the XR display screen is , and use logarithmic interval division to ensure that the sampling in the near region is finer and the sampling in the far region is sparser. The specific calculation method is as follows: where, represents the depth sampling point of the th layer, is the number of levels in the depth direction. In this way, ensure that the optical path calculation can adapt to the changes of different depth levels. Especially in the XR scene, due to the significant change of depth of field, the optical characteristics of different depth levels may vary greatly. Therefore, this non-linear division method can better optimize the rendering accuracy. In addition, in the boundary region of the XR display screen, it is necessary to perform bandwidth division on the fusion boundary dimension in the four-dimensional rendering sampling matrix to ensure that the transition effect of different boundary regions is optimized. Assume that the range of the boundary transition bandwidth is , and perform division in a linearly increasing manner. The specific calculation is as follows: where, represents the width parameter of the th boundary transition band, is the number of boundary sampling layers. This division method ensures that the optical transition can be smoother at different splicing regions or display boundaries to avoid problems such as color difference or visual break. Orthogonally combine the field of view angle parameter, distortion coefficient parameter, depth level parameter and boundary transition band width parameter to construct multiple groups of rendering parameter combinations. Assume that the combination method of all parameters is , then adopt the method of orthogonal experimental design to reduce the computational complexity and ensure the uniform coverage of representative sampling points. The final number of orthogonal combinations is expressed as: Among them, represents the total number of final rendered parameter combinations, ensuring that different optical characteristics can be fully covered to optimize the rendering calculation. To improve the authenticity of the rendering results, image acquisition is performed on these rendered parameter combinations under various environmental lighting conditions to obtain high-resolution rendered image data. Assume that the environmental lighting conditions are represented by the brightness level as follows: Among them, represents the th light intensity level, is the number of sampling points for the lighting conditions. Rendering data acquisition is performed under different lighting conditions to ensure that the final rendered model can adapt to complex lighting environments and improve the realism of the display effect.
[0027] In a specific embodiment, the process of executing step 200 may specifically include the following steps: Input the rendered parameter combination into the first branch of the dual-branch convolutional neural network. The first branch contains three fully connected layers, and each fully connected layer is followed by a batch normalization layer and a ReLU activation function to obtain the rendered parameter features; Input the high-resolution rendered image data into the second branch of the dual-branch convolutional neural network. The second branch contains four convolutional blocks, and each convolutional block consists of two 3×3 convolutional layers, a batch normalization layer, and a ReLU activation layer to obtain the spatial feature map; Perform attention fusion processing on the rendered parameter features and the spatial feature map to obtain the fused feature map, and perform 3×3 convolutional processing on the fused feature map to obtain the intermediate feature map; Perform 3×3 convolutional processing on the intermediate feature map to obtain the preprocessed feature map, and perform 1×1 convolutional processing on the preprocessed feature map to obtain the spatial geometric correction parameters including the perspective transformation matrix, distortion correction coefficients, and registration offsets.
[0028] Specifically, a dual-branch convolutional neural network is constructed to process the rendered parameter features and the spatial feature information respectively. Input the rendered parameter combination into the first branch of the dual-branch convolutional neural network. This branch is mainly composed of three fully connected layers, and a batch normalization layer and a ReLU activation function are connected after each fully connected layer to ensure that the features can be stably propagated and prevent gradient disappearance. Assume that the input rendered parameter combination is the vector , where represents the number of rendered parameters, then the first-layer fully connected transformation is expressed as: Among them, is the weight matrix of the first layer, is the bias term, represents the ReLU activation function, is the output feature of the first layer. To ensure numerical stability, this output is batch-normalized to make the activation values more evenly distributed: where, and are the mean and standard deviation of batch normalization respectively. Similarly, the fully connected calculations for the second and third layers are: is input as the rendering parameter feature into the subsequent fusion module. At the same time, to extract the spatial features in the high-resolution rendered image data, this data is input into the second branch of the dual-branch convolutional neural network, which consists of four convolutional blocks, and each convolutional block is composed of two convolutional layers, a batch normalization layer, and a ReLU activation layer. Assume the input high-resolution rendered image data is , where and are the height and width of the image respectively, is the number of channels, then the calculation method of the first convolutional block is: where, is the convolution kernel, Conv represents the convolution operation, BN represents the batch normalization operation, is the output feature map of the first convolutional block. The subsequent three convolutional blocks are calculated in the same way: is input as the spatial feature map into the attention fusion module. In the attention fusion process, the rendering parameter feature is feature-aligned with the spatial feature map and the attention weights are calculated. Define the attention weight matrix as: where, is the attention mapping matrix, softmax ensures weight normalization. The final fused feature map is calculated as: Input the fused feature map into the convolutional layer for processing to obtain an intermediate feature map: Then perform a second convolutional processing to obtain a preprocessed feature map: Adopt the convolutional layer to map the preprocessed feature map to the final spatial geometric correction parameters: Among them, it includes a perspective transformation matrix , distortion correction coefficient k, and registration offset : This method combines rendering parameters and high-resolution image information to achieve high-precision geometric correction and ensure the rendering quality of the XR display system in various complex environments.
[0029] Among them, the spatial geometric correction parameters are also used for distortion compensation processing of high-resolution rendered image data, specifically including: performing a color space transformation on the high-resolution rendered image data, converting the RGB color space to the YUV color space to obtain a luminance component and a chrominance component; inputting the luminance component into a U-shaped encoder for feature extraction, where the U-shaped encoder optimizes the convolutional layer using the structural reparameterization technique. The encoder part includes five downsampling blocks, and the decoder part includes five upsampling blocks to obtain a multi-scale feature map; performing adaptive feature fusion on the multi-scale feature map, using a channel attention mechanism to weight and combine features at different scales to obtain a fused feature map; inputting the fused feature map into a style transfer network, where the style transfer network includes three residual blocks, and each residual block uses an instance normalization layer and a LeakyReLU activation function to obtain a style feature map; constructing a distortion compensation layer based on the style feature map, compensating and enhancing the local area of the image through adaptive weight adjustment to obtain a compensated luminance component; performing an inverse color space transformation on the compensated luminance component and the original chrominance component, converting the YUV color space back to the RGB color space to obtain a high-resolution rendered image after distortion compensation; combining the high-resolution rendered image after distortion compensation with the spatial geometric correction parameters to generate input data for subsequent local enhancement processing; performing a quality assessment on the input data, calculating the structural similarity index and the peak signal-to-noise ratio index, and dynamically adjusting the distortion compensation parameters according to the evaluation results to obtain optimized distortion compensation data.
[0030] In this embodiment, obtaining the spatial geometric correction parameters further includes: calculating the density distribution of the perspective transformation matrix in the spatial geometric correction parameters to obtain a rendering density distribution map of the local area; identifying the regional density peaks of the rendering density distribution map by calculating the density difference between each pixel point and its neighborhood points and recursively propagating the differences to obtain an initial set of density peak points; constructing an adaptive density radius based on the initial set of density peak points to dynamically adjust the influence range of each density peak point to obtain an optimized high-density area distribution; inputting the high-density area distribution into a self-adjusting probability clustering model, using the merging parameters to balance the density weights of each area to obtain regional clustering features; iteratively optimizing the regional clustering features, calculating the interaction intensity between the clustering center and the high-density points in each iteration, and stopping the iteration when the change rate of the interaction intensity for three consecutive iterations is less than 0.05 to obtain a stable clustering result; performing a partition mapping on the original spatial geometric correction parameters based on the stable clustering result, grouping the areas with similar features into the same rendering group to obtain hierarchical rendering parameters; adjusting the correction coefficients in partitions according to the hierarchical rendering parameters, calculating the perspective transformation matrix and the distortion correction coefficients separately for each rendering group to obtain locally optimized spatial geometric correction parameters; and assigning weights to the locally optimized spatial geometric correction parameters according to the density distribution of the rendering groups, assigning a larger correction weight to the high-density areas to obtain an adaptive correction parameter matrix.
[0031] In a specific embodiment, the process of executing step 300 may specifically include the following steps: Input the sequence of fixation point coordinates into six residual convolutional blocks for encoding. Each residual convolutional block contains two 3×3 convolutional layers and a skip connection structure to obtain a fixation point feature map, and perform feature expansion processing on the spatial geometric correction parameters to obtain a correction parameter feature map; Calculate the self-attention weights for each feature channel in the fixation point feature map to obtain an attention-weighted feature map, and perform channel dimension concatenation on the attention-weighted feature map and the correction parameter feature map to obtain a multi-scale feature; Perform upsampling processing on the multi-scale feature, and fuse the skip connection features corresponding to the encoding stage at each scale level to obtain a locally enhanced feature map; Perform a spatial mapping transformation on the locally enhanced feature map, assign different resolution weights to different regions in the feature map according to the fixation area distribution, increase the resolution of the fixation area to A times the original resolution, and perform dynamic downsampling on the surrounding areas to obtain locally enhanced resolution data.
[0032] Specifically, input the sequence of fixation point coordinates into six residual convolutional blocks for encoding. Each residual convolutional block consists of two It consists of a convolutional layer and a skip connection structure to ensure the effective propagation of information in the deep network and prevent the vanishing gradient. Assume that the sequence of fixation point coordinates is represented as where and are the height and width of the input feature map respectively, is the number of channels, then the calculation method of the first residual block is as follows: where are convolution kernels respectively, Conv represents the convolution operation, BN represents batch normalization, represents the ReLU activation function, and is the output of the first residual block. To ensure that features can be fully extracted, the subsequent five residual blocks are calculated in the same way: where , and finally the encoded fixation point feature map is obtained. At the same time, the feature expansion process is performed on the spatial geometric correction parameters to obtain the correction parameter feature map. Assume that the spatial geometric correction parameters include the perspective transformation matrix , the distortion correction coefficient and the registration offset , then these parameters are converted into the same feature space through the feature mapping network. The mapping transformation is performed on the perspective transformation matrix: where FC represents the fully connected layer, which is used to convert the matrix into a vector representation. Similarly, the mapping is performed on the distortion correction coefficient and the registration offset: The three groups of data are concatenated to obtain the correction parameter feature map : After obtaining the fixation point feature map and the correction parameter feature map, the self-attention weight is calculated for each channel in the fixation point feature map to enhance the feature expression ability of the key regions. Define the attention weight matrix : where and They are the weight matrix and bias term for attention calculation respectively. The final attention-weighted feature map is calculated as follows: The attention-weighted feature map is concatenated with the calibration parameter feature map in the channel dimension to obtain a multi-scale feature map: To optimize local resolution enhancement, the multi-scale feature map is upsampled, and the skip connection features corresponding to the encoding stage at each scale level are fused to restore high-resolution information. The upsampling operation is set to bilinear interpolation: And the skip connection features are fused at each scale level: The local enhanced feature map is obtained . A spatial mapping transformation is performed on the local enhanced feature map, and different resolution weights are assigned to different regions in the feature map according to the fixation area distribution. The weight mapping function is defined as: where is the current fixation point coordinate, controls the weight decay rate. The final local resolution enhancement data is calculated as follows: where is the resolution improvement multiple of the fixation area, serves as a dynamic downsampling factor to control the rendering precision of the surrounding area.
[0033] In a specific embodiment, the process of executing step 400 may specifically include the following steps: Organize the rendering parameters of N display channels into an N×M-dimensional state vector, where M is the number of parameters of each display channel, and the number of parameters includes the computational resource allocation weight, rendering priority coefficient, and cache policy parameter. Construct an initial particle swarm position matrix according to the number of parameters of the display channel; Construct a multi-objective function for the spatial geometry calibration parameters to obtain a fitness function, and update the velocity of the particle swarm position matrix according to the local resolution enhancement data to obtain a particle swarm velocity matrix; Perform position iterative update on the particle swarm velocity matrix, record the local optimal solution and the global optimal solution in each iteration, and stop the iteration when the number of iterations reaches B times to obtain the optimized particle positions; Determine the computational resource allocation scheme for each display channel according to the optimized particle positions, and hierarchically divide the computational resources according to the rendering priority to obtain a resource allocation matrix; Decode and map the resource allocation matrix according to the display channels to obtain the rendering channel parameters composed of the sampling density parameters, update frequency parameters, and cache policy parameters of each channel; Perform multi-channel synchronization control on the rendering channel parameters according to the sliding time window mechanism to obtain the target stereoscopic image display parameters.
[0034] Specifically, organize the rendering parameters of display channels into a -dimensional state vector, where represents the number of parameters of each display channel, including the computational resource allocation weight, rendering priority coefficient, and cache policy parameter. Set the state vector , where each element represents the -th parameter value of the -th channel, that is: Among them, represents the computational resource allocation weight of the -th channel, represents the rendering priority coefficient, represents the cache policy parameter, and so on. To optimize this state vector, construct an initial particle swarm position matrix , where is the number of particles in the particle swarm, and each particle represents a possible rendering parameter configuration scheme. The positions of the initial particle swarm are randomly initialized as: Among them, and are the minimum and maximum ranges of the -th parameter respectively, generates a random number between 0 and 1 to ensure that the parameters are evenly distributed within a reasonable range. During the particle swarm optimization process, construct a fitness function to evaluate the quality of different particle configuration schemes. This fitness function consists of multiple objective functions, including the spatial geometric correction error , computational resource utilization , and rendering uniformity . Set the total fitness function as: Among them, is the weight factor of the objective function, is calculated as follows: Among them, is the expected rendering parameter value. The computational resource utilization is defined as: Among them, is the computing resource consumption of the th channel, is the total computing resource limit. The rendering uniformity is calculated as follows: Among them, is the global mean value of the th parameter. Based on the fitness function, the particle velocity matrix and the position matrix are updated according to the particle swarm optimization rule. The velocity update formula is as follows: Among them, is the inertia factor, is the learning factor, is a random number, and are the local optimal solution and the global optimal solution respectively. The position update formula is: After each iteration, the local optimal solution and the global optimal solution are updated until the maximum number of iterations is reached. Then, the optimization stops, and the optimal particle position matrix is obtained. Based on the optimized particle positions, the computing resource allocation scheme for each display channel is determined, and the computing resources are hierarchically divided according to the rendering priority to form a resource allocation matrix : Among them, represents the rendering priority weight of the th channel, represents the total computing resources. The resource allocation matrix is decoded and mapped according to the display channels to generate rendering channel parameters, including the sampling density , the update frequency and the caching policy : To ensure the synchronous execution of all channels, a sliding time window mechanism is adopted for multi-channel synchronous control. The size of the time window is defined, and the update step size of each channel is adjusted within each window: Based on the sliding time window mechanism, the synchronous execution step sizes of all channels are adjusted to ensure that the multi-channels complete rendering within a consistent time interval, and finally, the target stereoscopic image display parameters are output .
[0035] In a specific embodiment, the process of performing the steps to perform multi-channel synchronization control on the rendering channel parameters according to the sliding time window mechanism to obtain the target stereoscopic image display parameters may specifically include the following steps: Perform time window partitioning on the rendering channel parameters, and set three controller gain coefficients for each display channel to obtain time window control parameters; Calculate the rendering error of each display channel to obtain error feedback data, where the error feedback data includes spatial registration accuracy error, resolution uniformity error, and time synchronization error; Dynamically adjust the controller gain coefficients according to the error feedback data to obtain updated controller parameters, and perform PID control operations on the updated controller parameters to obtain controller output data; Input the controller output data into the rendering parameter scheduling module of each display channel, and jointly optimize the sampling density, update frequency, and cache policy according to the computing resource allocation weight to obtain a scheduling control matrix; Perform decoding processing on the scheduling control matrix, and convert it into the rendering resolution parameters, refresh timing parameters, and cache configuration parameters corresponding to each display channel to obtain the target stereoscopic image display parameters.
[0036] Specifically, perform time window partitioning on the rendering channel parameters to ensure that different channels can be synchronously adjusted within a specified time. Assume that the time window size of the system is , and the total number of rendering channels is . If the rendering parameters of each display channel include computing resource allocation weight, sampling density, update frequency, and cache policy, then define the time window control parameter matrix , where each channel has three controller gain coefficients, which are , , , corresponding to spatial registration, resolution uniformity, and time synchronization optimization. The time window partitioning is expressed as: where, is the period of the rendering frame, ensuring that the adjustment within each time window can adapt to the frame update rate. After time window partitioning, calculate the rendering error of each display channel to form error feedback data , where the error feedback data of each channel includes the spatial registration accuracy error , the resolution uniformity error and the time synchronization error . The spatial registration accuracy error is calculated as follows: Among them, represents the th rendering parameter of the channel, represents the expected rendering parameter value, is the number of channel parameters. The resolution uniformity error is calculated as: Among them, is the parameter mean of all channels, which is used to measure the rendering consistency between different channels. The time synchronization error is calculated as follows: Among them, and are the refresh frequencies of channel and channel respectively, to ensure multi-channel synchronization. Based on the error feedback data, the controller gain coefficient is dynamically adjusted to optimize the response characteristics of different channels. The gain adjustment formula is set as: Among them, is the dynamic adjustment coefficient, which is used to control the adaptive change of the gain. The adjusted gain coefficient is used for PID (Proportional-Integral-Derivative) control calculation to generate the controller output data : Among them, is the integral gain, is the derivative gain, to ensure that the error adjustment has dynamic response ability. The controller output data is input into the rendering parameter scheduling module of each display channel, and the sampling density, update frequency, and cache policy are jointly optimized according to the calculation resource allocation weight to form a scheduling control matrix : The scheduling control matrix is decoded to obtain the specific rendering parameters of each channel, including the sampling density , refresh timing and cache configuration : To ensure that all channels can run synchronously, a sliding time window mechanism is adopted for multi-channel synchronization control, defining the time window size and calculating the synchronization offset of each channel: Adjust the rendering update rate of each channel by sliding a time window to ensure the synchronous output of the target stereoscopic image display parameters for multiple channels 。
[0037] The high-resolution rendering method of the XR display screen in the embodiment of the present application has been described above. Next, the high-resolution rendering system 10 of the XR display screen in the embodiment of the present application will be described. Please refer to Figure 2 In one embodiment, the high-resolution rendering system 10 of the XR display screen in the embodiment of the present application includes: A distributed sampling module 11, configured to perform distributed sampling on the optical characteristics of the XR display screen to obtain a four-dimensional rendering sampling matrix, and perform Latin hypercube sampling to obtain multiple sets of rendering parameter combinations and corresponding high-resolution rendering image data; A calibration processing module 12, configured to perform spatial geometric calibration processing on the rendering parameter combinations and high-resolution rendering image data to obtain spatial geometric calibration parameters; An enhancement processing module 13, configured to perform local enhancement processing on the fixation point coordinate sequence and spatial geometric calibration parameters to obtain local resolution enhancement data; A synchronization control module 14, configured to calculate the rendering channel parameters of multiple display channels according to the spatial geometric calibration parameters and local resolution enhancement data, and perform multi-channel synchronization control to obtain the target stereoscopic image display parameters.
[0038] Through the collaborative cooperation of the above-mentioned various components, by establishing a four-dimensional rendering sampling matrix and a Latin hypercube sampling strategy, the optical characteristic parameters of XR display are systematically collected and quantified. The spatial geometric calibration is performed using a dual-branch convolutional neural network structure, realizing the deep fusion of rendering parameter features and spatial features, and significantly improving the registration accuracy of virtual and real scenes. Based on the local enhancement processing method of residual convolution coding and deconvolution decoding, high-resolution rendering of the fixation area and dynamic downsampling of the surrounding area are realized, optimizing the allocation efficiency of computing resources. A multi-channel rendering control strategy based on particle swarm optimization is designed. By adaptively adjusting rendering parameters and computing resource allocation, the collaborative working effect of each display channel is ensured. The sliding time window mechanism and PID control algorithm are introduced to realize real-time adjustment of rendering parameters and multi-channel synchronization control, improving the stability and response speed of the XR display system. Through multi-dimensional parameter optimization and multi-channel collaborative control, the rendering quality and computing efficiency of XR display are significantly improved, while reducing the resource consumption of the system.
[0039] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the above-described systems, systems, and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0040] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing an electronic device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0041] As described above, the above embodiments are only used to illustrate the technical solutions of this application and are not intended to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of this application.
Claims
1. A high-resolution rendering method for an XR display screen, characterized in that: The method comprises: Distributed sampling is performed on the optical characteristics of the XR display to obtain a four-dimensional rendering sampling matrix, and Latin hypercube sampling is performed to obtain multiple sets of rendering parameter combinations and corresponding high-resolution rendering image data; Performing spatial geometric correction processing on the rendering parameter combination and the high-resolution rendering image data to obtain spatial geometric correction parameters; Performing local enhancement processing on the gaze point coordinate sequence and the spatial geometric correction parameters to obtain local resolution enhancement data; Rendering channel parameters of multiple display channels are calculated according to the spatial geometry correction parameters and the local resolution enhancement data, and multi-channel synchronization control is performed to obtain target stereoscopic image display parameters.
2. The high-resolution rendering method of an XR display screen according to claim 1, characterized in that: The distributed sampling of the optical characteristics of the XR display screen is performed to obtain a four-dimensional rendering sampling matrix, and Latin hypercube sampling is performed to obtain multiple sets of rendering parameter combinations and corresponding high-resolution rendering image data, including: The XR display screen is evenly divided in the horizontal direction to obtain the field of view angle sampling nodes; Divide the vertical direction of the XR display screen into increasing intervals to obtain spatial distortion sampling nodes; Perform depth level division along the Z-axis direction of the XR display screen to obtain the optical path difference sampling layer; Perform grid division on the XY plane to obtain a fused boundary sampling grid; Performing spatial feature extraction on the sampling points in the field of view angle sampling node, the spatial distortion sampling node, the optical path difference sampling layer and the fusion boundary sampling grid to obtain sampling data including three-dimensional coordinates, field of view angle, distortion coefficient, optical path difference and boundary intensity value; Performing feature matrix conversion on the sampling data to obtain a four-dimensional rendering sampling matrix; Latin hypercube sampling is performed based on the four-dimensional rendering sampling matrix to obtain multiple groups of rendering parameter combinations and corresponding high-resolution rendering image data.
3. The high-resolution rendering method of an XR display screen according to claim 2, characterized in that: The Latin hypercube sampling is performed based on the four-dimensional rendering sampling matrix to obtain multiple sets of rendering parameter combinations and corresponding high-resolution rendering image data, including: Dividing the field of view angle dimension in the four-dimensional rendering sampling matrix into intervals to obtain a field of view angle parameter; Performing layered sampling on the spatial distortion dimension in the four-dimensional rendering sampling matrix to obtain distortion coefficient parameters; Performing depth segmentation on the optical path difference dimension in the four-dimensional rendering sampling matrix to obtain a depth level parameter; Performing bandwidth division on the fusion boundary dimension in the four-dimensional rendering sampling matrix to obtain a boundary transition band width parameter; Performing orthogonal combination according to the field of view angle parameter, the distortion coefficient parameter, the depth level parameter and the boundary transition zone width parameter to obtain multiple groups of rendering parameter combinations; Image acquisition is performed on the multiple rendering parameter combinations under multiple ambient lighting conditions to obtain high-resolution rendering image data corresponding to each rendering parameter combination.
4. The high-resolution rendering method of an XR display screen according to claim 1, characterized in that: The performing spatial geometric correction processing on the rendering parameter combination and the high-resolution rendering image data to obtain spatial geometric correction parameters includes: Inputting the rendering parameter combination into a first branch of a two-branch convolutional neural network, wherein the first branch comprises three fully connected layers, each of which is followed by a batch normalization layer and a ReLU activation function, to obtain a rendering parameter feature; Inputting the high-resolution rendered image data into the second branch of the dual-branch convolutional neural network, wherein the second branch comprises four convolution blocks, each convolution block comprises two 3×3 convolution layers, a batch normalization layer, and a ReLU activation layer, to obtain a spatial feature map; Performing attention fusion processing on the rendering parameter feature and the spatial feature map to obtain a fused feature map, and performing 3×3 convolution processing on the fused feature map to obtain an intermediate feature map; A 3×3 convolution process is performed on the intermediate feature map to obtain a preprocessed feature map, and a 1×1 convolution process is performed on the preprocessed feature map to obtain spatial geometric correction parameters including a perspective transformation matrix, a distortion correction coefficient, and a registration offset.
5. The high-resolution rendering method of an XR display screen according to claim 1, characterized in that: The locally enhancing the gaze point coordinate sequence and the spatial geometry correction parameters to obtain local resolution enhancement data includes: Inputting the gaze point coordinate sequence into six residual convolution blocks for encoding, each residual convolution block includes two 3×3 convolution layers and a skip connection structure, obtaining a gaze point feature map, and performing feature expansion processing on the spatial geometry correction parameters to obtain a correction parameter feature map; Performing self-attention weight calculation on each feature channel in the fixation point feature map to obtain an attention-weighted feature map, and concatenating the attention-weighted feature map with the correction parameter feature map in channel dimension to obtain a multi-scale feature; Upsampling the multi-scale features, fusing the skip connection features of the corresponding encoding stage at each scale level to obtain a local enhanced feature map; The local enhanced feature map is spatially mapped and transformed, and different resolution weights are assigned to different areas in the feature map according to the distribution of the attention area. The resolution of the attention area is increased to A times the original resolution, and the surrounding area is dynamically downsampled to obtain local resolution enhanced data.
6. The high-resolution rendering method of an XR display screen according to claim 1, characterized in that: The step of calculating rendering channel parameters of multiple display channels according to the spatial geometry correction parameters and the local resolution enhancement data, and performing multi-channel synchronous control to obtain target stereoscopic image display parameters includes: Organizing the rendering parameters of N display channels into an N×M dimensional state vector, where M is the number of parameters of each display channel, the number of parameters including computing resource allocation weights, rendering priority coefficients, and cache strategy parameters, and constructing an initial particle swarm position matrix according to the number of parameters of the display channels; A multi-objective function is constructed for the spatial geometry correction parameters to obtain a fitness function, and a velocity update is performed on the particle swarm position matrix according to the local resolution enhancement data to obtain a particle swarm velocity matrix; Iteratively update the position of the particle swarm velocity matrix, record the local optimal solution and the global optimal solution in each iteration, stop the iteration when the number of iterations reaches B times, and obtain the optimized particle position; Determine a computing resource allocation scheme for each display channel according to the optimized particle positions, divide the computing resources into layers according to the rendering priorities, and obtain a resource allocation matrix; Decoding and mapping the resource allocation matrix according to display channels to obtain rendering channel parameters consisting of sampling density parameters, update frequency parameters and cache strategy parameters of each channel; The rendering channel parameters are subjected to multi-channel synchronous control according to a sliding time window mechanism to obtain target stereoscopic image display parameters.
7. The high-resolution rendering method of an XR display screen according to claim 6, characterized in that: The step of performing multi-channel synchronous control on the rendering channel parameters according to a sliding time window mechanism to obtain target stereoscopic image display parameters includes: Dividing the rendering channel parameters into time windows, and setting three controller gain coefficients for each display channel to obtain time window control parameters; Calculating the rendering error of each display channel to obtain error feedback data, wherein the error feedback data includes a spatial registration accuracy error, a resolution uniformity error, and a time synchronization error; Dynamically adjusting the controller gain coefficient according to the error feedback data to obtain updated controller parameters, and performing PID control operation on the updated controller parameters to obtain controller output data; The controller output data is input into the rendering parameter scheduling module of each display channel, and the sampling density, update frequency and cache strategy are jointly optimized according to the computing resource allocation weight to obtain a scheduling control matrix; The scheduling control matrix is decoded and converted into rendering resolution parameters, refresh timing parameters and cache configuration parameters corresponding to each display channel to obtain target stereoscopic image display parameters.
8. A high-resolution rendering system for an XR display, characterized in that: A high-resolution rendering system for an XR display screen according to any one of claims 1 to 7, wherein the high-resolution rendering system for the XR display screen comprises: The distributed sampling module is used to perform distributed sampling of the optical characteristics of the XR display screen to obtain a four-dimensional rendering sampling matrix, and perform Latin hypercube sampling to obtain multiple sets of rendering parameter combinations and corresponding high-resolution rendering image data; A correction processing module, used for performing spatial geometry correction processing on the rendering parameter combination and the high-resolution rendering image data to obtain spatial geometry correction parameters; An enhancement processing module, used for performing local enhancement processing on the gaze point coordinate sequence and the spatial geometric correction parameters to obtain local resolution enhancement data; A synchronization control module is used to calculate rendering channel parameters of multiple display channels according to the spatial geometry correction parameters and the local resolution enhancement data, and perform multi-channel synchronization control to obtain target stereoscopic image display parameters.
Citation Information
Cited By
Cache region correction method and device, electronic equipment and storage medium
CN120949986A
A cache area correction method and device, electronic equipment and storage medium
CN120949986B
Video synchronous acquisition method and device based on optical far image
CN121000964A
Method and system for batch generation and consistent output of document images based on browser instance pool
CN121598922A
Method and system for batch generating consistent output based on browser instance pool document image
CN121598922B