Virtual Exhibition Hall Generation Method and Device Based on Curved Screen

By establishing an accurate three-dimensional spatial coordinate system and multi-angle scanning mechanism on the arc-shaped LED display screen, combining multi-scale feature hierarchical network and graph convolution network, the visual distortion and discontinuity of subtitles on the arc-shaped screen are solved, and a high-precision and consistent subtitle display effect is achieved.

CN119722443BActive Publication Date: 2025-07-04SHENZHEN MEIYAD OPTOELECTRONICS CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510229522.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-07-04
Estimated Expiration
2045-02-28

AI Technical Summary

Technical Problem

The existing subtitle display technology is difficult to achieve refined display on curved LED display screens, resulting in distortion of the subtitle display effect as the viewing angle changes, and there is a problem of visual distortion and discontinuity of display between modules.

Method used

By establishing an accurate three-dimensional spatial coordinate system and multi-angle scanning mechanism, a multi-scale feature hierarchical network is used for feature extraction and fusion, combining self-attention mechanism and feature pyramid integration, a graph convolution network is introduced for module interactive analysis, a collaborative training framework for deep residual teacher network and lightweight student network is designed, multi-angle consistency loss training is carried out, and the display effect is ensured through gradient domain fusion of Poisson equations and ambient light compensation processing.

Benefits of technology

It improves the degree of refinement of subtitle display, ensures the consistency of display effects under different perspectives, eliminates mutations in the edge display and timing jitters, and achieves the rationality of subtitle layout and uniformity of brightness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119722443B_ABST
    Figure CN119722443B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of virtual exhibition halls, and discloses a method and device for generating a virtual exhibition hall based on a curved screen. The method includes: projecting the pre-input caption content in the virtual exhibition hall onto the surface of the display screen through a curved LED display screen, and obtaining a caption display area mapping matrix and an initial caption image; performing feature extraction and adaptive fusion operation on the display area to obtain a target mapping matrix and a hierarchical feature atlas; performing LED module interaction analysis on the target mapping matrix and the hierarchical feature atlas to obtain a display parameter matrix; inputting the display parameter matrix into a deep residual teacher network and a lightweight student network for multi-angle display consistency loss training to obtain a caption display model; jointly predicting the caption position and display brightness to obtain a target display parameter set, and outputting target caption display data, thereby effectively improving the refinement degree of caption display in the virtual exhibition hall and ensuring the consistency of display effects from different perspectives.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of virtual exhibition halls, and particularly to a method and device for generating a virtual exhibition hall based on a curved screen. Background Art

[0002] Due to its unique spatial curved surface characteristics, the curved LED display screen can bring an immersive visual experience to viewers and plays an important role in virtual exhibition halls. However, when displaying subtitle content on a curved LED display screen, due to the curvature change of the display screen surface, the subtitle display effect will be distorted with the change of the viewing angle, affecting the exhibition viewing experience.

[0003] The existing subtitle display technologies are mainly designed for flat display screens and are difficult to be directly applied to curved display scenarios. Traditional methods often use simple geometric transformations for subtitle projection and cannot effectively handle the visual distortion problems brought by the curved surface. At the same time, due to the modular design of the curved LED display screen, there are often discontinuities in the display effects between adjacent LED modules, resulting in visual defects such as breaks and jaggedness in subtitle display. Summary of the Invention

[0004] This application provides a method and device for generating a virtual exhibition hall based on a curved screen, thereby effectively improving the refinement degree of subtitle display in the virtual exhibition hall and ensuring the consistency of display effects from different perspectives.

[0005] In the first aspect of this application, a method for generating a virtual exhibition hall based on a curved screen is provided. The method for generating a virtual exhibition hall based on a curved screen includes:

[0006] Project the pre-input subtitle content in the virtual exhibition hall onto the surface of the display screen through the curved LED display screen to obtain a subtitle display area mapping matrix and an initial subtitle image;

[0007] Based on the subtitle display area mapping matrix and the initial subtitle image, perform feature extraction and adaptive fusion operations on the display area to obtain a target mapping matrix and a hierarchical feature atlas;

[0008] Perform LED module interaction analysis on the target mapping matrix and the hierarchical feature atlas to obtain a display parameter matrix;

[0009] Input the display parameter matrix into a deep residual teacher network and a lightweight student network for multi-angle display consistency loss training to obtain a subtitle display model;

[0010] Use the subtitle display model to jointly predict the subtitle position and display brightness to obtain a target display parameter set, and based on the target display parameter set, perform gradient domain fusion and temporal smoothing on the subtitle edge to output target subtitle display data.

[0011] The second aspect of the present application provides a virtual exhibition hall generation device based on an arc screen. The virtual exhibition hall generation device based on the arc screen includes:

[0012] A projection module, configured to project the pre-input caption content in the virtual exhibition hall onto the surface of the display screen through an arc-shaped LED display screen, and obtain a caption display area mapping matrix and an initial caption image;

[0013] A feature extraction module, configured to perform feature extraction and adaptive fusion operations on the display area based on the caption display area mapping matrix and the initial caption image to obtain a target mapping matrix and a hierarchical feature atlas;

[0014] An interaction analysis module, configured to perform LED module interaction analysis on the target mapping matrix and the hierarchical feature atlas to obtain a display parameter matrix;

[0015] A training module, configured to input the display parameter matrix into a deep residual teacher network and a lightweight student network for multi-angle display consistency loss training to obtain a caption display model;

[0016] An output module, configured to jointly predict the caption position and display brightness by using the caption display model to obtain a target display parameter set, and perform gradient domain fusion and temporal smoothing on the caption edge according to the target display parameter set, and output target caption display data.

[0017] Compared with the prior art, the present application has the following beneficial effects: By establishing an accurate three-dimensional space coordinate system and a multi-angle scanning mechanism, accurate modeling of the surface characteristics of the arc-shaped LED display screen is realized. A multi-scale feature hierarchical network is used for display area feature extraction and fusion, combined with self-attention mechanism and feature pyramid integration, effectively improving the refinement degree of caption display. The graph convolutional network is introduced for LED module interaction analysis, and by constructing the connection relationship matrix between modules, the problem of discontinuous display between adjacent LED modules is solved. A collaborative training framework of a deep residual teacher network and a lightweight student network is designed. Through knowledge distillation and multi-angle consistency loss constraints, the consistency of display effects under different perspectives is ensured. The joint optimization of caption position and display brightness is realized. Through the collaborative effect of the dual-path prediction branch, the rationality of caption layout and the uniformity of display brightness are ensured. A gradient domain fusion scheme based on the Poisson equation is proposed, combined with ambient light compensation and temporal smoothing processing, effectively eliminating the problems of sudden changes in caption edge display and temporal jitter. Description of the Drawings

[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0019] The structures, ratios, sizes, etc. shown in the drawings of this specification are only used to cooperate with the content disclosed in the specification for those familiar with this technology to understand and read, and are not used to limit the limiting conditions under which the present invention can be implemented. Therefore, they do not have a substantial technical meaning. Any modification of the structure, change in the proportional relationship, or adjustment of the size, without affecting the effects that the present invention can produce and the purposes that can be achieved, should still fall within the scope that can be covered by the technical content disclosed by the present invention.

[0020] Figure 1 is a schematic flowchart of a method for generating a virtual exhibition hall based on a curved screen provided by an embodiment of the present invention;

[0021] Figure 2 is a schematic block diagram of the structure of a device for generating a virtual exhibition hall based on a curved screen provided by an embodiment of the present invention. Detailed implementation manners

[0022] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0023] The flowchart shown in the drawings is only an example, and does not necessarily include all the content and operations / steps, nor does it necessarily need to be executed in the described order. For example, some operations / steps can also be decomposed, combined, or partially merged. Therefore, the actual execution order may be changed according to the actual situation.

[0024] It should also be understood that the terms used in this specification of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in this specification of the present application and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.

[0025] It should be further understood that the term "and / or" used in this specification of the present application and the appended claims refers to any combination and all possible combinations of one or more of the related listed items, and includes these combinations. Please refer to Figure 1, an embodiment of the virtual exhibition hall generation method based on a curved screen in the embodiments of the present application includes:

[0026] Step 100, project the pre-input caption content in the virtual exhibition hall onto the surface of the display screen through a curved LED display screen, and obtain a caption display area mapping matrix and an initial caption image;

[0027] It can be understood that the execution subject of the present application can be a virtual exhibition hall generation device based on a curved screen, or a terminal or a server, and specific limitations are not made here. In the embodiments of the present application, the server is used as the execution subject for illustration.

[0028] Specifically, by performing a zoning operation on the surface of the display screen, the surface of the curved LED display screen is divided into multiple grid points. Each grid point is assigned a unique identifier to form a grid point identification matrix. Use a laser rangefinder to scan the grid points from multiple angles. During the scanning process, the laser rangefinder emits laser beams within a preset angle range and measures the distance between the grid point and the rangefinder according to the time delay of the reflected signal to obtain accurate distance data. These distance data are organized into a grid point distance matrix, which corresponds one-to-one with the grid point identification matrix. Through multi-angle scanning, grid point data from different perspectives are obtained. Based on the grid point identification matrix and the distance matrix, calculate the specific coordinate values of each grid point in three-dimensional space, and establish a surface point cloud dataset of the curved LED display screen. Perform a surface fitting operation on the surface point cloud dataset, using a mathematical interpolation method or a least squares fitting method, to generate a display screen surface equation that can describe the shape of the display screen surface. Convert the pre-input caption content in the virtual exhibition hall into a font outline vector. The font outline vector is a vectorized description of the caption shape and can adapt to different resolution and deformation requirements. According to the display screen surface equation, perform projection calculation of the font outline on the surface of the display screen to determine the projection position of the caption on the curved display screen, and generate caption outline projection data. Perform rasterization processing on the caption outline projection data to map the continuous projection outline to the discrete LED module positions, generate a caption display area mapping matrix, and describe the correspondence between each LED module and the projection outline, so that the caption content can be accurately displayed on the screen. Perform font anti-aliasing processing based on the caption outline projection data, using technologies such as bilinear interpolation or sub-pixel rendering to eliminate the jagged effect at the font edges and improve the smoothness of the display. At the same time, perform rasterization processing to convert the vectorized font outline into a high-resolution font bitmap. Perform a resolution adaptive transformation on the high-resolution font bitmap to scale the font bitmap to the actual resolution of the display screen to ensure that the final image exactly matches the hardware resolution of the display screen, and generate an initial caption image.

[0029] Step 200: Based on the subtitle display area mapping matrix and the initial subtitle image, perform feature extraction and adaptive fusion operations on the display area to obtain the target mapping matrix and the hierarchical feature atlas;

[0030] Specifically, input the subtitle display area mapping matrix and the initial subtitle image into a multi-scale feature hierarchical network for multi-scale decomposition. The multi-scale hierarchical network splits the input data into n feature layers of different scales through hierarchical processing. Each feature layer corresponds to a specific resolution and spatial perception characteristic, and these feature layers together constitute a multi-dimensional description of the subtitle display area. Perform convolution operations and local feature extraction on the n feature layers of different scales respectively to generate n sets of original feature maps. The convolution operation can effectively capture complex edges, textures, and other significant area characteristics in the subtitle display area by extracting local spatial information and semantic features. At the same time, local feature extraction can refine the feature representation, making each original feature map have rich spatial hierarchical information and semantic hierarchical information. Construct a feature correlation graph based on the n sets of original feature maps to depict the mutual relationship between the feature layers. In the feature correlation graph, each feature node represents an original feature map, and the edges between the nodes represent the correlation of these feature maps in specific semantic or spatial dimensions. Generate a feature correlation matrix by calculating the similarity weights between the feature nodes. This matrix reflects the global similarity and correlation degree between different feature maps. Use the feature correlation matrix to perform weighted integration on the n sets of original feature maps, fuse the feature regions with stronger correlation, and generate a fused feature map. Perform self-attention operation on the fused feature map to obtain the attention weight matrix. The self-attention operation calculates the attention weight matrix within the global range and identifies the significant areas crucial for the display effect in the subtitle display area. Perform spatial re-weighting processing on the fused feature map based on the attention weight matrix, further enhance the high-weight areas, and at the same time suppress the influence of the low-weight areas to generate the target mapping matrix. Input the fused feature map and the attention weight matrix into a transposed convolution network for feature upsampling reconstruction. The transposed convolution network gradually restores the fused feature map from low resolution to high resolution through inverse convolution operations, and at the same time uses the attention weight matrix to dynamically adjust the reconstruction process to generate feature representations of different resolutions. Integrate the feature representations of different resolutions through a feature pyramid to generate a hierarchical feature atlas. The feature pyramid integration process can balance the detailed expression at high resolution and the global semantic expression at low resolution, forming a multi-level and multi-scale feature representation system.

[0031] Step 300: Perform LED module interaction analysis on the target mapping matrix and the hierarchical feature atlas to obtain the display parameter matrix;

[0032] It should be noted that for the target mapping matrix, through regional segmentation and by analyzing the geometric and display characteristics of the target mapping matrix, the entire display area is divided into multiple LED display modules with independent display capabilities, generating a module division matrix to clarify the display area range of each LED module and its positional relationship in the global mapping. Based on the module division matrix, a display module adjacency graph is constructed, where each LED display module is regarded as a node in the graph. To characterize the spatial relationship between modules, the Euclidean distance of the node geometric positions in the mapping matrix is used to calculate the spatial distance between each pair of module nodes, generating a spatial adjacency matrix. This adjacency matrix quantifies the spatial distribution and adjacency relationship between modules in numerical form. For each feature map in the hierarchical feature map set, perform regional aggregation operations, combine global features with module characteristics, extract and aggregate the feature regions corresponding to each LED display module to form a module feature matrix, which contains multi-dimensional feature information of each module, including features such as brightness, contrast, and the relationship with surrounding modules. Combining the spatial adjacency matrix, perform message passing and update on the module features. Through a graph neural network or a similar graph processing algorithm, introduce the spatial relationship between modules captured in the spatial adjacency matrix into the module feature update process to generate a node state matrix, which describes the semantic and spatial meaning of the features of each module in the global context. Perform edge attention operations on the node state matrix. By dynamically adjusting the weights of the connection edges between modules, identify the connection relationships between modules that have the greatest impact on the display effect, generating a module connection relationship matrix. The module connection relationship matrix is used to reflect the display cooperation requirements between different modules. For example, on a curved screen, adjacent modules need to synchronously adjust brightness or contrast to avoid visual discontinuity. After obtaining the module connection relationship matrix, combine the node state matrix and the module feature matrix to perform collaborative optimization on the brightness and contrast of the LED display modules, making the overall visual effect of the display area more uniform and natural. After the optimization is completed, use the module control parameters to densely sample the entire display area. By performing more refined parameter sampling and distribution fitting on the display units within each module, generate a display unit parameter distribution map. This distribution map refines the display parameters such as brightness and contrast of each display unit within the module, making the overall display effect reach a higher precision. Perform boundary constraint processing on the display unit parameter distribution map to ensure that the connection areas between modules comply with physical and visual rules during display transitions. Through boundary constraints, effectively reduce the visual fragmentation caused by uneven transitions between modules, and finally output a display parameter matrix containing the optimized parameters of all modules and display units.

[0033] Step 400: Input the display parameter matrix into a deep residual teacher network and a lightweight student network for multi-angle display consistency loss training to obtain a subtitle display model;

[0034] Specifically, the display parameter matrix is input into the deep residual teacher network. The deep residual teacher network performs multi-level feature transformation on the input display parameters. It extracts the spatial characteristics, brightness characteristics, and contrast characteristics in the display parameters layer by layer through residual blocks and generates a high-dimensional teacher network feature map. At the same time, to enhance the generalization ability of training, a data augmentation operation is performed on the display parameter matrix. The data augmentation methods include perspective change, brightness perturbation, contrast adjustment, etc., so as to generate a multi-perspective training dataset for simulating the performance of subtitle content under different viewing angles and display conditions on the curved screen. The multi-perspective training dataset is input into the lightweight student network for feature extraction and transformation. The lightweight student network realizes the efficient processing of the multi-perspective training data through a simplified convolutional structure. The generated student network feature map is lighter than the teacher network feature map but still retains the key features required for subtitle display. During this process, a feature alignment operation is performed. By comparing the similarity between the teacher network feature map and the student network feature map, the knowledge distillation loss is calculated. This loss value quantifies the degree of approximation of the student network to the learning results of the teacher network, ensuring that the student network can reproduce the display effect of the teacher network with lower computational complexity. At the same time, based on the multi-perspective training dataset, the display effect differences under different perspectives are calculated, and a multi-angle display consistency loss value is constructed. This loss value quantifies the difference degree of the display under different perspectives by analyzing the brightness uniformity, contrast consistency, and subtitle edge smoothness of the display effects from each perspective, providing an optimization target for network training. To comprehensively consider the knowledge distillation loss and the multi-angle display consistency loss, the two are weighted and combined to form an overall loss function. The lightweight student network is trained by backpropagation according to the overall loss function. Backpropagation updates the network parameters layer by layer, enabling the student network to gradually optimize its feature extraction ability and reduce the knowledge distillation loss and the display consistency loss. After multiple rounds of training, an optimized lightweight student network is obtained. Model pruning and quantization processing are performed on the optimized student network. The pruning operation reduces the model complexity by removing redundant neurons and connections in the network, while the quantization processing reduces the model volume and computational overhead by reducing the precision of parameter representation (such as from 32-bit floating-point numbers to 8-bit integers). A subtitle display model is generated.

[0035] Step 500: Jointly predict the subtitle position and display brightness using the subtitle display model to obtain a target display parameter set, and based on the target display parameter set, perform gradient domain fusion and temporal smoothing on the subtitle edge, and output the target subtitle display data.

[0036] Specifically, set the subtitle display model to the inference mode. By inputting the current subtitle content and display area information, the model generates initial prediction results. After decoding and analyzing these prediction results, they are separated into the output data of the position prediction branch and the brightness prediction branch, obtaining dual-path prediction features. The output data of the position prediction branch mainly describes the geometric layout characteristics of the subtitles, while the brightness prediction branch provides an estimation of the brightness distribution of each display unit. Perform coordinate transformation on the position prediction branch in the dual-path prediction features, mapping the subtitle positions predicted by the model to the physical coordinate system of the display screen. The coordinate transformation process needs to consider the geometric characteristics of the curved screen and the display area defined by the target mapping matrix, so that the generated subtitle layout coordinates can accurately adapt to the actual display environment of the curved screen. At the same time, for the output data of the brightness prediction branch, calculate the brightness distribution of the display units by performing brightness normalization operations. The purpose of the brightness normalization operation is to adjust the brightness value range to conform to the physical brightness range of the display screen, generating a brightness control sequence. The brightness control sequence is the result of distributed regulation of the brightness of each display unit. Calculate the position-brightness coupling coefficient based on the subtitle layout coordinates and the brightness control sequence. Through the calculation of the position-brightness coupling coefficient, the mutual influence between the subtitle position and the brightness distribution is quantified. The coupling relationship can reflect the compensation effect of brightness on the subtitle layout and the constraint effect of the subtitle position on the brightness distribution. Generate fusion weights using the coupling coefficient, and perform weighted combination on the subtitle layout coordinates and the brightness control sequence according to these weights to generate a set of target display parameters. Perform ambient light compensation operations on the set of target display parameters, and perform gradient domain fusion and temporal smoothing processing on the subtitle edges. During the gradient domain fusion process, calculate the gradient distribution of the subtitle edge region, and optimize the continuity of the edge transition region through a fusion algorithm to eliminate the jagged or broken problems existing in the subtitle edges. Gradient domain fusion makes the subtitle edge transition more natural and smooth by regulating the brightness gradient of adjacent pixels, enhancing the overall visual perception. At the same time, to cope with the movement or change of subtitles in dynamic display scenarios, perform temporal smoothing processing. Temporal smoothing filters the changes in the set of target display parameters in the time dimension, effectively reducing the jitter and jumping problems during subtitle updates, and ensuring the display stability of subtitles in a dynamic environment. Combining the influence of ambient light, perform ambient light compensation operations on the set of target display parameters. The ambient light compensation operation adjusts the brightness distribution in the set of target display parameters in real time by collecting the lighting conditions of the current display environment, so that the display effect of the subtitles can meet the viewing requirements under different light conditions. For example, increase the contrast and brightness of the subtitles in strong light environments, and reduce the brightness in weak light environments to avoid glare. The target subtitle display data after ambient light compensation has higher adaptability and visual comfort.

[0037] Perform local gradient calculation on the subtitle edge region according to the target display parameter set to generate a gradient value matrix, and at the same time extract the edge contour data. The gradient value matrix reflects the change amplitude of the brightness of subtitle edge pixels, while the edge contour data provides the geometric structure information of the subtitle edge. Construct a Poisson boundary equation based on the gradient value matrix and the edge contour data. The Poisson equation takes the gradient field as input, combines the boundary conditions, and generates an edge fusion function through numerical solution. This function is used to describe the pixel-level transition relationship between the subtitle edge and the background region. Use the edge fusion function to perform pixel-level gradual change processing on the subtitle edge region and the background region, and eliminate the jagged and broken phenomena at the edge by smoothing the gradient to generate edge transition region data. On this basis, calculate the regional fusion coefficient according to the distribution characteristics of the edge transition region data and the display region, and generate a regional weight matrix. The regional weight matrix is used to quantify the contribution ratio of different display regions to the overall display effect. To enhance the visibility of subtitle display under different lighting conditions, perform multi-point collection and spatial distribution analysis on the ambient light in the display region. Collect the light intensity data at multiple positions in the display region through an ambient light sensor, and generate an ambient light intensity map based on these data. The ambient light intensity map describes the lighting conditions of the environment where the display screen is located in a spatial distribution form, and can reflect the change of light intensity. According to the ambient light intensity map, perform block brightness compensation on the display region, and balance the influence of ambient light by adjusting the brightness of each display block to generate a light response matrix. The light response matrix ensures that the subtitle display will not be too dark in a strong light environment and will not be too bright in a weak light condition, improving the adaptability of the overall display effect. Perform joint optimization operation on the regional weight matrix and the light response matrix to construct a display compensation model. The display compensation model generates regional compensation parameters by integrating the results of regional weight and light response, quantifies the brightness adjustment requirements of each display unit under specific ambient light conditions, and ensures that the subtitle display can adapt to complex lighting environments. Use these compensation parameters to perform adaptive brightness adjustment on the subtitle display region to generate compensated display data, so as to achieve dynamic balance of brightness and improvement of display quality globally. Reconstruct the frame sequence of the compensated display data according to the display timing to construct a dynamic display sequence. The frame sequence reconstruction reorganizes the compensated data in the time dimension to ensure the natural content and brightness transition of each frame during the dynamic display of the subtitle. To eliminate the inter-frame jitter problem existing in the dynamic display sequence, perform time-domain filtering operation on the frame sequence. The time-domain filtering suppresses the pixel values with sudden changes by smoothing the inter-frame data to generate smooth sequence data. Input the smooth sequence data into the display buffer for frame synchronization control. The frame synchronization control coordinates the refresh timing and synchronization signal of the display screen to ensure that the subtitle display data can be loaded and output at a stable rate. During this process, reorganize the frame data according to the hardware characteristics of the display screen to make it adapt to the pixel distribution and refresh characteristics of the display screen, and finally generate the target subtitle display data.

[0038] In the embodiments of the present application, by establishing an accurate three-dimensional space coordinate system and a multi-angle scanning mechanism, an accurate modeling of the surface characteristics of the arc-shaped LED display screen is achieved. A multi-scale feature hierarchical network is used for feature extraction and fusion of the display area, combined with the self-attention mechanism and feature pyramid integration, effectively improving the refinement degree of subtitle display. A graph convolutional network is introduced for LED module interaction analysis. By constructing an inter-module connection relationship matrix, the problem of discontinuous display between adjacent LED modules is solved. A collaborative training framework of a deep residual teacher network and a lightweight student network is designed. Through knowledge distillation and multi-angle consistency loss constraints, the consistency of the display effects under different perspectives is ensured. The joint optimization of subtitle position and display brightness is realized. Through the collaborative effect of the dual-path prediction branches, the rationality of subtitle layout and the uniformity of display brightness are ensured. A gradient domain fusion scheme based on Poisson's equation is proposed, combined with ambient light compensation and temporal smoothing processing, effectively eliminating the problems of sudden changes in subtitle edge display and temporal jitter.

[0039] In a specific embodiment, the process of executing step 100 may specifically include the following steps:

[0040] Divide the surface of the arc-shaped LED display screen into multiple grid points, and assign a unique identifier to each grid point to obtain a grid point identification matrix;

[0041] Use a laser rangefinder to perform multi-angle scanning and distance data analysis on the grid points within a preset angle range to obtain a grid point distance matrix;

[0042] Based on the grid point identification matrix and the grid point distance matrix, calculate the three-dimensional space coordinate values of each grid point, establish a surface point cloud data set of the arc-shaped LED display screen, and perform a surface fitting operation on the surface point cloud data set to obtain a display screen surface equation;

[0043] Convert the subtitle content pre-input in the virtual exhibition hall into a font outline vector, calculate the projection position of the font outline on the surface according to the display screen surface equation to obtain subtitle outline projection data, and perform rasterization processing on the subtitle outline projection data to map the continuous projection outline to the discrete LED module positions to obtain a subtitle display area mapping matrix;

[0044] Based on the subtitle outline projection data, perform font anti-aliasing and rasterization processing to obtain a high-resolution font bitmap, and perform a resolution adaptive transformation on the high-resolution font bitmap to scale the font bitmap to the display screen resolution to obtain an initial subtitle image.

[0045] Specifically, divide the surface of the arc-shaped LED display screen into multiple grid points. Based on the geometric shape of the display screen, divide it into regular two-dimensional grids, where each grid point represents an independent physical position. Assume the width of the display screen is and the height is , the number of rows of the mesh division is , the number of columns is , then the total number of grid points is . Each grid point is uniquely identified by the row and column numbers to form a grid point identification matrix , where and respectively represent the numbers of the th row and the th column, satisfying , . For example, if , then the grid point identification matrix is a two-dimensional matrix of , numbered from (1,1) to (10,20). Use a laser rangefinder to perform multi-angle scanning and distance data analysis on the grid points within a preset angle range to obtain the distance data of the grid points. Assume that the laser rangefinder is located at the reference point , and the scanning angle for the grid point is , where is the horizontal angle and is the vertical angle, satisfying , where is the number of scanning angles. According to the ranging formula, the distance measured by the laser rangefinder satisfies:

[0046] ;

[0047] Through multi-angle scanning, a distance data matrix of each grid point is obtained, where each element represents the distance measured in the th scan. Based on the grid point identification matrix and the distance matrix, calculate the three-dimensional spatial coordinates of each grid point through trigonometric relationships. Assume that the unit vector in the laser direction is

[0048] , then the three-dimensional coordinates of the grid point are expressed as:

[0049] ;

[0050] ;

[0051] ;

[0052] Take the average value of all scanning angles for each grid point to reduce measurement errors and obtain the point cloud coordinates. The three-dimensional coordinates of all grid points form a point cloud data set Perform a surface fitting operation on the point cloud dataset to establish the surface equation of the display screen. Assume that the surface of the display screen is represented by a cubic polynomial, and its surface equation is:

[0053] ;

[0054] where, are the surface parameters determined by the least squares fitting method. During the fitting process, the coordinates of the point cloud data are used as inputs to solve for the parameter set that minimizes the error . After obtaining the surface equation of the display screen, convert the pre-entered caption content in the virtual exhibition hall into font outline vectors. Assume that the caption content is a text string, and use a font generation tool to represent it as vectorized contour data , where each point is the contour coordinate of the font. According to the surface equation , project the font outline onto the display screen surface to generate three-dimensional contour data:

[0055] ;

[0056] To adapt to the discrete structure of the LED display module, perform rasterization processing on the projected caption contour data to map the continuous three-dimensional contour to the discrete LED module positions. Let the center coordinate of each LED module be . The rasterization process is achieved by finding the nearest neighbor module, thereby generating a caption display area mapping matrix. Perform font anti-aliasing and rasterization processing on the caption contour projection data to generate a high-resolution font bitmap. Through the bilinear interpolation algorithm, reduce the jagged effect and ensure the smoothness of the font edges. The high-resolution font bitmap undergoes a resolution adaptive transformation to scale it to the actual resolution of the display screen, and finally obtain the initial caption image.

[0057] In a specific embodiment, the process of performing step 200 may specifically include the following steps:

[0058] Input the caption display area mapping matrix and the initial caption image into a multi-scale feature hierarchical network for multi-scale decomposition to obtain n different-scale feature layers, and perform convolution operations and local feature extraction on the n different-scale feature layers respectively to obtain n sets of original feature maps;

[0059] Construct a feature correlation graph based on the n sets of original feature maps, calculate the similarity weights between feature nodes to obtain a feature correlation matrix, and use the feature correlation matrix to perform weighted integration on the n sets of original feature maps to obtain a fused feature map;

[0060] Perform self-attention operation on the fused feature map to obtain an attention weight matrix, and perform spatial reweighting on the fused feature map based on the attention weight matrix to obtain a target mapping matrix;

[0061] Input the fused feature map and the attention weight matrix into a transposed convolution network for feature upsampling and reconstruction to obtain feature representations of different resolutions, and perform feature pyramid integration on the feature representations of different resolutions to obtain a hierarchical feature map set.

[0062] Specifically, map the subtitle display area mapping matrix and the initial subtitle image into a multi-scale feature hierarchical network for multi-scale decomposition. Assume that the size of the input mapping matrix and image is , that is, the height is and the width is . The multi-scale feature hierarchical network gradually reduces the size of the input data through downsampling operations. Let the size of the feature map of the th layer obtained by decomposition be . Then the number of decomposition scales determines the minimum resolution of the feature map. For example, for , the size of the last layer of feature map is . These multi-scale feature layers capture spatial information and semantic characteristics at different resolutions. After decomposition, perform convolution operations and local feature extraction on each feature layer to generate groups of original feature maps. Let the feature map of the th layer be , and the feature obtained after its convolution operation is:

[0063] ;

[0064] where Conv represents the convolution operation, is the weight of the convolution kernel of the th layer. The convolution operation extracts local spatial information, such as the gradient of the subtitle edge and the regional brightness distribution, while enhancing the feature expression ability. For each group of feature maps , extract significant features through pooling operations, such as the contrast characteristics between the subtitle contour and the background area. After obtaining groups of original feature maps, analyze the relationships between these features by constructing a feature correlation graph. The feature correlation graph uses the feature map of each layer as a node and describes their similarity with edges. Let the similarity between the feature maps of the th layer and the th layer be , then the similarity is calculated by cosine similarity or correlation:

[0065] ;

[0066] Among them, represents the inner product of two feature maps, and represents the norm of the feature map. By calculating the similarity between all feature maps, a feature correlation matrix is constructed, where each element represents the th layer and the th layer's feature correlation strength. When using the feature correlation matrix to weightedly integrate a group of original feature maps, the integrated feature map is calculated by the following formula:

[0067] ;

[0068] Among them, is the weight of the feature correlation matrix, representing the weighted ratio between feature maps. By weighted integration, redundant information in multi-scale features is eliminated, and the cross-scale feature expression ability is enhanced. Perform self-attention operation on the fused feature map to highlight the key region features. The self-attention mechanism calculates the global relationship matrix and , as well as the corresponding feature weight matrix , to obtain the attention weight matrix :

[0069] ;

[0070] Among them, are respectively the query, key, and value of the fused feature map, is the normalization coefficient of the feature dimension. The attention weight matrix performs spatial re-weighting on the fused feature map to generate the target mapping matrix:

[0071] ;

[0072] The target mapping matrix further improves the accurate expression ability of features in space, making the display effect of subtitles more prominent. Input the fused feature map and the attention weight matrix into the transposed convolutional network for feature upsampling reconstruction to generate feature representations of different resolutions. The transposed convolutional network gradually enlarges the feature map through inverse convolutional operations, and the size of the th layer feature map generated is:

[0073] ;

[0074] Among them and is the height and width of the minimum resolution layer. The upsampled multi-resolution feature maps are fused through feature pyramid integration, and the integration formula is:

[0075] ;

[0076] where, is the weight of each layer of feature map, is the layer's feature representation. The hierarchical feature map set generated by feature pyramid integration can retain both high-resolution details and low-resolution semantic features simultaneously.

[0077] In a specific embodiment, the process of executing step 300 may specifically include the following steps:

[0078] Perform region segmentation on the target mapping matrix, divide the display area into multiple LED display modules to obtain a module division matrix, and construct a display module adjacency graph based on the module division matrix. Take each LED display module as a graph node and calculate the spatial distance between nodes to obtain a spatial adjacency matrix;

[0079] Perform region aggregation operations on each feature map in the hierarchical feature map set to obtain a module feature matrix, and perform message passing and update on the node features based on the spatial adjacency matrix and the module feature matrix to obtain a node state matrix;

[0080] Perform edge attention operations on the node state matrix to obtain an inter-module connection relationship matrix, and perform collaborative optimization on the brightness and contrast of the LED display modules according to the inter-module connection relationship matrix to obtain module control parameters;

[0081] Use the module control parameters to densely sample the display area to obtain a display unit parameter distribution map, and perform boundary constraints on the display unit parameter distribution map to obtain a display parameter matrix.

[0082] Specifically, perform region segmentation on the target mapping matrix to divide the entire display area into multiple independent LED display modules. Assume the size of the target mapping matrix is , that is, the height is and the width is . By setting the row and column division parameters and , divide the display area into rectangular modules. The range of each module is defined as:

[0083] ;

[0084] where represents the row number, Indicates the column number. After the segmentation is completed, each module is assigned a unique identifier to form a module partition matrix , where represents the pixel point belongs to the module number. Based on the module partition matrix , a display module adjacency graph is constructed, with each module regarded as a node in the graph, and the relationship between adjacent modules is described by edges. Let the module and the module represent two adjacent modules respectively, and their spatial distance is calculated by the Euclidean distance of their geometric centers:

[0085] ;

[0086] where and are the central coordinates of the module and the module respectively. According to these distances, a spatial adjacency matrix is generated, where represents the adjacency relationship between the module and . The threshold method is used to limit the adjacency relationship, and only the module pairs with a distance less than a certain value are retained. The region aggregation operation is performed on each feature map in the hierarchical feature map set to obtain the module feature matrix. Assume that the feature map set contains layers of feature maps, and the size of each layer of feature map is , and the -th layer of feature map is , where is the number of channels. Through the region aggregation operation, the pixel values in the feature map are averaged and pooled according to the module partition to obtain the module feature matrix , and the specific formula is:

[0087] ;

[0088] where, represents the pixel set contained in the module , and is the number of elements in this set. The module feature matrix is combined with the spatial adjacency matrix, and the node features are passed and updated through the graph neural network to generate the node state matrix. The core of the node state update is to perform weighted aggregation of the features between modules based on the adjacency matrix , and the update formula is:

[0089] ;

[0090] where, is the module after the features representation module set of adjacent modules is a learnable weight matrix is an activation function (such as ReLU). After several iterations, the generated node state matrix reflects the global features of each module. Perform edge attention operation on the node state matrix, calculate the connection weights between modules, and generate the inter-module connection relationship matrix . The edge attention mechanism adjusts the adjacency weights by learning the feature relationships of each pair of adjacent modules:

[0091] ;

[0092] where represents feature concatenation is the learned weight vector. Based on cooperatively optimize the brightness and contrast of the module to obtain the module control parameter matrix. Use the module control parameter matrix to densely sample the display area, and generate the display unit parameter distribution map by mapping the module parameters to the pixel level . The brightness value of each pixel is interpolated according to the module control parameters, and at the same time, boundary constraint processing is performed in combination with the module boundary characteristics to adjust the parameter distribution in the edge transition region, and generate the final display parameter matrix.

[0093] In a specific embodiment, the process of executing step 400 may specifically include the following steps:

[0094] Input the display parameter matrix into the deep residual teacher network, perform multi-level feature transformation on the display parameters to obtain the teacher network feature map, and perform data augmentation on the display parameter matrix to generate a multi-view training dataset;

[0095] Input the multi-view training dataset into the lightweight student network for feature extraction and transformation to obtain the student network feature map, and perform feature alignment operation on the teacher network feature map and the student network feature map to obtain the knowledge distillation loss;

[0096] Calculate the display effect differences under different views based on the multi-view training dataset, construct the multi-angle display consistency loss value, and perform weighted combination of the knowledge distillation loss and the multi-angle display consistency loss value to obtain the overall loss function;

[0097] Perform backpropagation training on the lightweight student network according to the overall loss function, update the network parameters to obtain the optimized student network, and perform model pruning and quantization processing on the optimized student network to obtain the subtitle display model.

[0098] Specifically, the display parameter matrix is input into the deep residual teacher network. and represent the height and width of the display screen respectively, and [channel number of parameters (such as brightness, contrast, etc.)]. The deep residual teacher network performs multi-level feature transformation on the display parameters through multiple residual modules. Each residual module consists of convolution, batch normalization, and activation function, and its core formula is:

[0099] ;

[0100] where is the input feature map of the th layer, and are the convolution weight and bias respectively, is the activation function (such as ReLU), represents the convolution operation. By stacking residual modules layer by layer, the teacher network can extract the global and local features of the display parameters and generate a high-dimensional teacher network feature map , where is the final number of feature channels. To enhance the generalization ability of the training data, data augmentation is performed on the input display parameter matrix to generate a multi-view training dataset . Data augmentation includes operations such as rotation, flipping, brightness perturbation, and contrast adjustment. For example, if the initial brightness value of the display parameter matrix is , brightness perturbation is performed through a random gain :

[0101] ;

[0102] where is used to limit the brightness value range to [0, 1]. These augmentation operations generate a diverse dataset for training by simulating the display characteristics under different views and environmental conditions. The multi-view training dataset is input into the lightweight student network for feature extraction and transformation. The lightweight student network adopts efficient structures such as depthwise separable convolution to reduce the computational complexity while retaining the key features. Assuming that the feature map of the th layer of the student network is then its output is expressed as:

[0103] ;

[0104] where and are the weight matrices of depthwise convolution and pointwise convolution respectively. The feature map finally generated by the student network is , where , indicating that the dimension of the feature is compressed. To implement knowledge distillation, perform feature alignment operations on the feature maps of the teacher network and the student network feature map . Feature alignment minimizes the difference between the two, and defines the knowledge distillation loss as:

[0105] ;

[0106] where, and respectively represent the values of the teacher and student feature maps at position . At the same time, calculate the difference in the display effect under different perspectives based on the multi-perspective training dataset, and define the multi-angle display consistency loss value . Let the perspective set be , and at each perspective , the loss of the display effect is defined as the KL divergence between the target parameter distribution and the actual distribution:

[0107] ;

[0108] where, and respectively represent the target display distribution and the actual display distribution. By weighted combination of the knowledge distillation loss and the multi-angle display consistency loss, the overall loss function is defined as:

[0109] ;

[0110] where, and are weight hyperparameters used to balance the contributions of the two losses. After defining the overall loss function, use the backpropagation algorithm to train the lightweight student network and gradually update the network parameters. The formula for gradient update is:

[0111] ;

[0112] where, is the learning rate, is the value of the weight of the th layer at the th iteration. After training, perform model pruning and quantization on the optimized student network to reduce the computational complexity. The pruning operation reduces redundant parameters by removing unimportant neuron connections. Quantization compresses the weights and activation values from floating-point representation to low-precision representation (such as 8-bit integers) to reduce storage requirements and computational overhead. Finally, a subtitle display model is obtained.

[0113] In a specific embodiment, the process of executing step 500 may specifically include the following steps:

[0114] Set the subtitle display model to the inference mode, input the current subtitle content and display area information, obtain the initial prediction result, and perform decoding analysis on the initial prediction result to separate the output data of the position prediction branch and the brightness prediction branch, obtaining dual-channel prediction features;

[0115] Perform coordinate transformation on the position prediction branch in the dual-channel prediction features, map the predicted position to the physical coordinate system of the display screen to obtain the subtitle layout coordinates. At the same time, calculate the brightness distribution of the display unit based on the brightness prediction branch in the dual-channel prediction features, perform brightness normalization operation to obtain the brightness control sequence;

[0116] Calculate the position-brightness coupling coefficient according to the subtitle layout coordinates and the brightness control sequence to obtain the fusion weight, and use the fusion weight to perform weighted combination on the subtitle layout coordinates and the brightness control sequence to obtain the target display parameter set;

[0117] Perform ambient light compensation operation according to the target display parameter set, perform gradient domain fusion and temporal smoothing processing on the subtitle edge, and output the target subtitle display data.

[0118] Specifically, set the trained subtitle display model to the inference mode. The parameters of the model are frozen to ensure stable performance during inference and avoid unnecessary gradient calculations. Input the current subtitle content and the display area information into the model to generate the initial prediction result. Assume that the input subtitle content is vectorized text contour information, represented as a point set , and the display area information is a parameter set describing the geometry and resolution of the curved display screen. The model generates the initial prediction result , where contains the output data of the position prediction branch and the brightness prediction branch . When performing decoding analysis on the initial prediction result , extract the two-dimensional position information of the subtitle from the position prediction branch , represented as relative coordinates , where describes the position of the subtitle in the normalized display area. At the same time, extract the brightness distribution of the display unit from the brightness prediction branch , represented as a set of brightness values , where represents the normalized brightness value of the th subtitle pixel. Perform coordinate transformation on the output of the position prediction branch to map the normalized position to the physical coordinate system of the display screen. Assume that the physical resolution of the display screen is , the conversion formula is:

[0119] ;

[0120] Wherein, is the position coordinate in the physical coordinate system of the display screen, and are the screen width and height respectively. The converted coordinate set constitutes the subtitle layout coordinates, which are used to accurately describe the distribution of subtitle content on the physical screen. For the output of the brightness prediction branch, based on the brightness value set calculate the brightness distribution of the display unit, and at the same time perform brightness normalization operation to adapt to the brightness range of the display screen. Let the maximum and minimum brightness values of the display screen be and respectively, and the brightness normalization formula is:

[0121] ;

[0122] Wherein, is the normalized brightness value. The normalization operation ensures that the brightness value output by the model can be accurately mapped to the hardware range of the display screen, generating the brightness control sequence . After obtaining the subtitle layout coordinates and the brightness control sequence, calculate the coupling relationship between the two. The position-brightness coupling coefficient is defined by quantifying the influence of the brightness of each display unit on its position adjustment. Let the coupling coefficient of each display unit be , and its calculation formula is:

[0123] ;

[0124] Wherein, is the total number of display units, represents the relative brightness weight of the th unit. Using the coupling coefficient to perform weighted combination on the subtitle layout coordinates and the brightness control sequence to generate the target display parameter set . The position coordinates and brightness values in the target display parameters are calculated by the following formula:

[0125] ;

[0126] Wherein, is the physical coordinate after coupling adjustment, is the final brightness value. According to the target display parameter set , perform ambient light compensation operation to adapt to different lighting conditions. During the compensation process, collect the ambient light intensity around the display screen to generate the ambient light distribution map . The compensation formula is:

[0127] ;

[0128] Among them, is the compensation coefficient, is the brightness value after ambient light compensation. Perform gradient domain fusion on the subtitle edge area, and smooth the transition area by calculating the edge gradient field to ensure visual continuity between the subtitle and the background area. The fusion formula is:

[0129] ;

[0130] Among them, is the brightness value after gradient correction, and represent the Laplacian operator and the gradient operator respectively, is the brightness gradient. Smooth the display parameters in the time dimension to eliminate jumps and jitters during the dynamic change of the subtitle. Let the time step be , then the time series smoothing formula is:

[0131] ;

[0132] Among them, is the smoothing coefficient, represents the smoothed brightness value of the th frame. Input the parameters after ambient light compensation, gradient domain fusion and time series smoothing into the display buffer to ensure synchronization with the refresh of the display screen. Through the above steps, the generated target subtitle display data can present the subtitle content with the best effect under different lighting and viewing angles.

[0133] In a specific embodiment, the process of performing ambient light compensation operation according to the target display parameter set, performing gradient domain fusion and time series smoothing on the subtitle edge, and outputting the target subtitle display data may specifically include the following steps:

[0134] Perform local gradient calculation on the subtitle edge area according to the target display parameter set to obtain a gradient value matrix and edge contour data, and construct a Poisson boundary equation based on the gradient value matrix and edge contour data to obtain an edge fusion function;

[0135] Use the edge fusion function to perform pixel-level gradient processing on the subtitle edge area and the background area to obtain edge transition area data, and calculate the area fusion coefficient based on the edge transition area data and the display area distribution characteristics to obtain an area weight matrix;

[0136] Perform multi-point acquisition and spatial distribution analysis on the ambient light of the display area to obtain an ambient light intensity map, and perform block brightness compensation on the display area according to the ambient light intensity map to obtain a light response matrix;

[0137] Jointly optimize the regional weight matrix and the light response matrix to construct a display compensation model, obtain regional compensation parameters, and adaptively adjust the brightness of the subtitle display area according to the regional compensation parameters to obtain compensated display data;

[0138] Reconstruct the frame sequence of the compensated display data according to the display timing to obtain a dynamic display sequence, and perform a time-domain filtering operation on the dynamic display sequence to eliminate inter-frame jitter and obtain smooth sequence data;

[0139] Input the smooth sequence data into the display buffer for frame synchronization control, and perform data reorganization according to the refresh timing and synchronization signal of the display screen to obtain the target subtitle display data.

[0140] Specifically, based on the target display parameter set , perform local gradient calculation on the subtitle edge area to capture the brightness change information of the edge. Assume that the brightness distribution of the subtitle area is , then the horizontal and vertical components of the local gradient are respectively expressed as:

[0141] ;

[0142] Through the above calculation, generate a gradient value matrix for the subtitle edge area. Extract the edge contour data according to the gradient value matrix in combination with the high-contrast area in the pixel brightness distribution. This data describes the geometric shape and characteristics of the edge area. After obtaining the gradient value matrix and the edge contour data, use the Poisson equation to construct an edge fusion model to achieve a continuous brightness transition between the edge area and the background area. The basic form of the Poisson equation is:

[0143] ;

[0144] Among them, is the Laplace operator, represents the divergence of the gradient field. By numerically solving the above equation and combining the edge contour data as the boundary condition, obtain the edge fusion function , and the brightness value output by it realizes pixel-level gradual change within the edge area. Use the edge fusion function to fuse the subtitle edge area and the background area to generate edge transition area data . On this basis, combine the distribution characteristics of the display area (such as modular structure or geometric characteristics) to calculate the regional fusion coefficient , and generate the regional weight matrix . The fusion coefficient is defined as the normalized weight of the brightness difference between the edge area and the background area. For example:

[0145] ;

[0146] Region weight matrix represents the relative weight of each pixel during the fusion process. To adapt to different lighting environments, an ambient light compensation operation is performed on the display area. Distribution data of the ambient light around the display screen is obtained through a multi-point light intensity acquisition device to generate an ambient light intensity map . This intensity map reflects the influence of ambient light on each pixel. According to , block brightness compensation is performed on the display area to generate a light response matrix , and its calculation formula is:

[0147] ;

[0148] where is the light compensation coefficient, is the uncompensated target brightness. The region weight matrix and the light response matrix are jointly optimized to construct a display compensation model, thereby generating a region compensation parameter matrix . The formula for joint optimization is:

[0149] ;

[0150] Based on the region compensation parameters, adaptive brightness adjustment is performed on the subtitle display area to obtain compensated display data :

[0151] ;

[0152] The compensated display data is reconstructed into a frame sequence according to the display timing to ensure smooth display when the subtitle dynamically changes. Let the time step be , and the frame sequence reconstruction formula is:

[0153] ;

[0154] where is the smoothing coefficient, which is used to adjust the weight distribution between frames. To further eliminate inter-frame jitter, a time-domain filtering operation is performed on the frame sequence data, and high-frequency noise is suppressed through a low-pass filter to obtain smoothed sequence data . The filtering formula is:

[0155] ;

[0156] where is the filtering window size, which determines the intensity of time smoothing. The smoothed sequence data is input into the display buffer, through the refresh timing of the display screen and the synchronization signal are aligned to complete data recombination and generate target subtitle display data.

[0157] In this embodiment, it further includes the step of dynamically following the interactive subtitles between the user and the virtual exhibition hall: using a depth camera to collect the user's motion video data, annotating and extracting the human body joint points in each frame of the image, constructing a two-dimensional human body skeleton node map, and converting the two-dimensional human body skeleton node map into a three-dimensional space coordinate system to obtain the user's motion three-dimensional feature sequence; inputting the user's motion three-dimensional feature sequence into an improved Joint-UNet network, extracting the joint point features at different scales through multi-scale node pooling to obtain hierarchical motion features, and at the same time using skip connections to retain the detailed information to obtain a multi-level pose feature pyramid; constructing a self-regulating graph convolutional network based on the multi-level pose feature pyramid, automatically adjusting the spatial relationship and motion correlation degree between joint points through a learnable adjacency matrix, dynamically updating and propagating the node features to obtain an adaptive connection weight matrix; constructing a user motion trajectory prediction model according to the adaptive connection weight matrix, modeling the user's motion direction and speed, and combining historical trajectory data for motion prediction to obtain a predicted position sequence; projecting the predicted position sequence into the three-dimensional scene of the virtual exhibition hall, calculating the relative position relationship between the user and the display content, constructing an interactive mapping function based on the attention mechanism to obtain subtitle following policy parameters; calculating the target position and orientation of the subtitle in the three-dimensional space according to the subtitle following policy parameters, using the L1 and L2 fusion loss to smooth the following trajectory to eliminate following jitter and obtain a smooth motion trajectory; inputting the smooth motion trajectory into a motion compensation network to adjust the position, size and direction of the subtitle in real time to ensure that the subtitle is always at the best viewing angle of the user to obtain dynamically updated display parameters; performing display optimization processing on the dynamically updated display parameters, adjusting the spatial layout and display effect of the subtitle, and calling the step of "performing ambient light compensation operation according to the target display parameter set, performing gradient domain fusion and temporal smoothing processing on the subtitle edge, and outputting the target subtitle display data" to realize the continuous and smooth following display of the subtitle.

[0158] The above describes the method for generating a virtual exhibition hall based on a curved screen in the embodiments of the present application. Next, a virtual exhibition hall generating device 10 based on a curved screen in the embodiments of the present application will be described. Please refer to Figure 2 In one embodiment, the virtual exhibition hall generating device 10 based on a curved screen in the embodiments of the present application includes:

[0159] A projection module 11, configured to project the subtitle content pre-input in the virtual exhibition hall onto the surface of the display screen through a curved LED display screen, and obtain a subtitle display area mapping matrix and an initial subtitle image;

[0160] The feature extraction module 12 is used to perform feature extraction and adaptive fusion operations on the display area based on the subtitle display area mapping matrix and the initial subtitle image, and obtain the target mapping matrix and the hierarchical feature atlas;

[0161] The interaction analysis module 13 is used to perform LED module interaction analysis on the target mapping matrix and the hierarchical feature atlas, and obtain the display parameter matrix;

[0162] The training module 14 is used to input the display parameter matrix into the deep residual teacher network and the lightweight student network for multi-angle display consistency loss training, and obtain the subtitle display model;

[0163] The output module 15 is used to jointly predict the subtitle position and display brightness by using the subtitle display model, obtain the target display parameter set, and perform gradient domain fusion and temporal smoothing on the subtitle edge according to the target display parameter set, and output the target subtitle display data.

[0164] Through the collaborative cooperation of the above-mentioned various components, by establishing an accurate three-dimensional space coordinate system and a multi-angle scanning mechanism, an accurate modeling of the surface characteristics of the curved LED display screen is realized. A multi-scale feature hierarchical network is used for display area feature extraction and fusion, combined with self-attention mechanism and feature pyramid integration, effectively improving the refinement degree of subtitle display. The graph convolutional network is introduced for LED module interaction analysis, and by constructing the connection relationship matrix between modules, the problem of discontinuous display between adjacent LED modules is solved. A collaborative training framework of the deep residual teacher network and the lightweight student network is designed. Through knowledge distillation and multi-angle consistency loss constraints, the consistency of display effects under different perspectives is ensured. The joint optimization of subtitle position and display brightness is realized. Through the collaborative effect of the dual-path prediction branches, the rationality of subtitle layout and the uniformity of display brightness are ensured. A gradient domain fusion scheme based on the Poisson equation is proposed, combined with ambient light compensation and temporal smoothing processing, effectively eliminating the problems of sudden change in subtitle edge display and temporal jitter.

[0165] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the above-described system, device, and unit can refer to the corresponding processes in the foregoing method embodiments, and will not be described in detail here.

[0166] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable an electronic device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0167] As described above, the above embodiments are only used to illustrate the technical solutions of this application and are not intended to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of various embodiments of this application.

Claims

1. A method for generating a virtual exhibition hall based on an arc-shaped screen, characterized in that, The method includes: Projecting the pre - input subtitle content in the virtual exhibition hall onto the surface of the arc - shaped LED display screen, and obtaining the subtitle display area mapping matrix and the initial subtitle image; Based on the subtitle display area mapping matrix and the initial subtitle image, performing feature extraction and adaptive fusion operation on the display area to obtain the target mapping matrix and the hierarchical feature atlas; Performing LED module interaction analysis on the target mapping matrix and the hierarchical feature atlas to obtain the display parameter matrix; specifically including: dividing the area of the target mapping matrix, dividing the display area into multiple LED display modules to obtain the module division matrix, constructing a display module adjacency graph based on the module division matrix, taking each LED display module as a graph node, calculating the spatial distance between nodes to obtain the spatial adjacency matrix; performing area aggregation operation on each feature map in the hierarchical feature atlas to obtain the module feature matrix, and performing message passing and updating on the node features based on the spatial adjacency matrix and the module feature matrix to obtain the node state matrix; performing edge attention operation on the node state matrix to obtain the module - to - module connection relationship matrix, and co - optimizing the brightness and contrast of the LED display modules according to the module - to - module connection relationship matrix to obtain the module control parameters; using the module control parameters to perform dense sampling on the display area to obtain the display unit parameter distribution map, and performing boundary constraint on the display unit parameter distribution map to obtain the display parameter matrix; Inputting the display parameter matrix into the deep residual teacher network and the lightweight student network for multi - angle display consistency loss training to obtain the subtitle display model; Using the subtitle display model to jointly predict the subtitle position and display brightness to obtain the target display parameter set, and performing gradient - domain fusion and temporal smoothing on the subtitle edge according to the target display parameter set to output the target subtitle display data.

2. The method for generating a virtual exhibition hall based on a curved screen according to claim 1, wherein The step of projecting the pre - input subtitle content in the virtual exhibition hall onto the surface of the arc - shaped LED display screen and obtaining the subtitle display area mapping matrix and the initial subtitle image includes: Dividing the surface of the arc - shaped LED display screen into multiple grid points, and assigning a unique identifier to each grid point to obtain the grid point identification matrix; Using a laser rangefinder to perform multi - angle scanning and distance data analysis on the grid points within a preset angle range to obtain the grid point distance matrix; Based on the grid point identification matrix and the grid point distance matrix, calculating the three - dimensional spatial coordinate values of each grid point, establishing the surface point cloud data set of the arc - shaped LED display screen, and performing surface fitting operation on the surface point cloud data set to obtain the display screen surface equation; Converting the pre - input subtitle content in the virtual exhibition hall into a font contour vector, calculating the projection position of the font contour on the surface according to the display screen surface equation to obtain the subtitle contour projection data, and performing rasterization processing on the subtitle contour projection data to map the continuous projection contour to the discrete LED module positions to obtain the subtitle display area mapping matrix; Perform font anti-aliasing and rasterization processing based on the subtitle contour projection data to obtain a high-resolution font bitmap, and perform a resolution adaptive transformation on the high-resolution font bitmap to scale the font bitmap to the display resolution to obtain an initial subtitle image.

3. The method for generating a virtual exhibition hall based on an arc-shaped screen according to claim 2, wherein Based on the subtitle display area mapping matrix and the initial subtitle image, perform feature extraction and adaptive fusion operations on the display area to obtain a target mapping matrix and a hierarchical feature atlas, including: Input the subtitle display area mapping matrix and the initial subtitle image into a multi-scale feature hierarchical network for multi-scale decomposition to obtain n feature layers of different scales, and perform convolution operations and local feature extraction on the n feature layers of different scales respectively to obtain n sets of original feature maps; Construct a feature correlation graph according to the n sets of original feature maps, calculate the similarity weights between feature nodes to obtain a feature correlation matrix, and use the feature correlation matrix to perform weighted integration on the n sets of original feature maps to obtain a fused feature map; Perform self-attention operation on the fused feature map to obtain an attention weight matrix, and perform spatial re-weighting on the fused feature map based on the attention weight matrix to obtain a target mapping matrix; Input the fused feature map and the attention weight matrix into a transposed convolution network for feature upsampling and reconstruction to obtain feature representations of different resolutions, and perform feature pyramid integration on the feature representations of different resolutions to obtain a hierarchical feature atlas.

4. The method for generating a virtual exhibition hall based on an arc screen according to claim 1, wherein Input the display parameter matrix into a deep residual teacher network and a lightweight student network for multi-angle display consistency loss training to obtain a subtitle display model, including: Input the display parameter matrix into the deep residual teacher network to perform multi-level feature transformation on the display parameters to obtain a teacher network feature map, and perform data augmentation on the display parameter matrix to generate a multi-view training dataset; Input the multi-view training dataset into the lightweight student network for feature extraction and transformation to obtain a student network feature map, and perform feature alignment operation on the teacher network feature map and the student network feature map to obtain a knowledge distillation loss; Calculate the display effect differences under different views based on the multi-view training dataset, construct a multi-angle display consistency loss value, and perform weighted combination of the knowledge distillation loss and the multi-angle display consistency loss value to obtain an overall loss function; Perform backpropagation training on the lightweight student network according to the overall loss function, update the network parameters to obtain an optimized student network, and perform model pruning and quantization processing on the optimized student network to obtain a subtitle display model.

5. The method for generating a virtual exhibition hall based on an arc-shaped screen according to claim 4, wherein Use the subtitle display model to jointly predict the subtitle position and display brightness to obtain a target display parameter set, and perform gradient domain fusion and temporal smoothing on the subtitle edge according to the target display parameter set to output target subtitle display data, including: Set the subtitle display model to the inference mode, input the current subtitle content and display area information, obtain the initial prediction result, and perform decoding analysis on the initial prediction result to separate the output data of the position prediction branch and the brightness prediction branch, obtaining dual-path prediction features; Perform coordinate transformation on the position prediction branch in the dual-path prediction features, map the predicted position to the physical coordinate system of the display screen to obtain subtitle layout coordinates. At the same time, calculate the brightness distribution of the display unit based on the brightness prediction branch in the dual-path prediction features, and perform brightness normalization operation to obtain a brightness control sequence; Calculate the position-brightness coupling coefficient according to the subtitle layout coordinates and the brightness control sequence to obtain a fusion weight, and use the fusion weight to perform weighted combination on the subtitle layout coordinates and the brightness control sequence to obtain a target display parameter set; Perform ambient light compensation operation according to the target display parameter set, perform gradient domain fusion and temporal smoothing processing on the subtitle edge, and output target subtitle display data.

6. The virtual exhibition hall generation method based on an arc-shaped screen according to claim 5, wherein The performing ambient light compensation operation according to the target display parameter set, performing gradient domain fusion and temporal smoothing processing on the subtitle edge, and outputting target subtitle display data includes: Perform local gradient calculation on the subtitle edge area according to the target display parameter set to obtain a gradient value matrix and edge contour data, and construct a Poisson boundary equation according to the gradient value matrix and the edge contour data to obtain an edge fusion function; Use the edge fusion function to perform pixel-level gradual change processing on the subtitle edge area and the background area to obtain edge transition area data, and calculate a region fusion coefficient according to the edge transition area data and the display area distribution characteristics to obtain a region weight matrix; Perform multi-point acquisition and spatial distribution analysis on the ambient light of the display area to obtain an ambient light intensity map, and perform block brightness compensation on the display area according to the ambient light intensity map to obtain a light response matrix; Perform joint optimization operation on the region weight matrix and the light response matrix to construct a display compensation model to obtain region compensation parameters, and perform adaptive brightness adjustment on the subtitle display area according to the region compensation parameters to obtain compensated display data; Reconstruct the frame sequence of the compensated display data according to the display timing to obtain a dynamic display sequence, and perform time-domain filtering operation on the dynamic display sequence to eliminate inter-frame jitter to obtain smooth sequence data; Input the smooth sequence data into the display buffer for frame synchronization control, and perform data reorganization according to the refresh timing and synchronization signal of the display screen to obtain target subtitle display data.

7. A virtual exhibition hall generation device based on a curved screen, characterized in that, For implementing the virtual exhibition hall generation method based on an arc screen as described in claim 1, the virtual exhibition hall generation device based on an arc screen includes: A projection module for projecting the pre-input subtitle content in the virtual exhibition hall onto the surface of the display screen through an arc-shaped LED display screen, and obtaining a subtitle display area mapping matrix and an initial subtitle image; A feature extraction module for performing feature extraction and adaptive fusion operation on the display area based on the subtitle display area mapping matrix and the initial subtitle image to obtain a target mapping matrix and a hierarchical feature atlas; An interaction analysis module, which is used to perform LED module interaction analysis on the target mapping matrix and the hierarchical feature map set to obtain a display parameter matrix; specifically including: performing regional segmentation on the target mapping matrix, dividing the display area into multiple LED display modules to obtain a module division matrix, and constructing a display module adjacency graph based on the module division matrix. Taking each LED display module as a graph node, calculating the spatial distance between nodes to obtain a spatial adjacency matrix; performing regional aggregation operations on each feature map in the hierarchical feature map set to obtain a module feature matrix, and performing message passing and updating on the node features based on the spatial adjacency matrix and the module feature matrix to obtain a node state matrix; performing edge attention operations on the node state matrix to obtain an inter-module connection relationship matrix, and co-optimizing the brightness and contrast of the LED display modules according to the inter-module connection relationship matrix to obtain module control parameters; using the module control parameters to densely sample the display area to obtain a display unit parameter distribution map, and performing boundary constraints on the display unit parameter distribution map to obtain a display parameter matrix; A training module, which is used to input the display parameter matrix into a deep residual teacher network and a lightweight student network for multi-angle display consistency loss training to obtain a subtitle display model; An output module, which is used to jointly predict the subtitle position and display brightness using the subtitle display model to obtain a target display parameter set, and perform gradient domain fusion and temporal smoothing on the subtitle edge according to the target display parameter set, and output target subtitle display data.

Citation Information

Patent Citations

  • Visual enhancement method and system of LED creative spherical display screen

    CN119417718A