Curved screen multi-view 3D image generation method and device
By collecting the viewing angle parameters of the arc screen and combining the graph neural network and the multi-head self-attention mechanism, the viewing angle characteristics are compared and learned to generate viewing angle mapping parameters, which solves the problem of viewing angle switching delay and image distortion when displaying multi-view images of arc screens, and achieves high-quality 3D image reconstruction and visual experience.
Patent Information
- Application Number
- CN202510200250.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-02-24
AI Technical Summary
The curved surface characteristics of the arc-shaped screen lead to problems such as viewing angle switching delay and image distortion when displaying multi-view images, affecting the user's viewing experience and interaction effects.
By collecting the viewing angle parameters of the arc screen, performing residual feature processing and tensor product operations, combining the graph neural network and the multi-head self-attention mechanism, comparative learning of the viewing angle features, generating viewing angle mapping parameters, and 3D image reconstruction through the U-Net network to output a 3D image sequence with continuous viewing angles.
It improves the consistency of feature during viewing angle conversion, reduces image distortion during viewing angle switching, and significantly improves the visual quality of 3D images.
Smart Images

Figure CN119671885B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image generation technology, and in particular to a method and device for generating multi-viewing angle 3D images on a curved screen. Background Art
[0002] Curved screens are widely used in games, virtual reality and other fields due to their unique immersive viewing experience and wide field of view. However, due to the curved surface characteristics of curved screens, problems such as perspective switching delay and image distortion often occur during the multi-perspective image display process, which seriously affects the user's viewing experience and interactive effect.
[0003] Traditional flat display algorithms cannot effectively handle the curvature characteristics of curved screens, and often cause serious delays when switching perspectives. Especially in large perspective switching scenarios, the system response speed drops significantly due to the need to process and reconstruct a large amount of image data, resulting in image freezes and discontinuities, which has a serious impact on applications that require high real-time interactivity. Most current delay compensation methods are based on simple linear interpolation or prediction models, which cannot accurately capture the nonlinear transformation characteristics of curved screens at different perspectives, and lack modeling of temporal correlations during perspective conversion, resulting in obvious unnaturalness and distortion in the generated 3D images during continuous perspective changes. Summary of the invention
[0004] The present invention provides a method and device for generating multi-viewing angle 3D images on a curved screen, which improves the feature consistency during the viewing angle conversion process, reduces the image distortion during the viewing angle switching, and thus improves the visual quality of the 3D image.
[0005] In a first aspect, the present invention provides a method for generating a multi-view 3D image on a curved screen, the method comprising:
[0006] Collect the viewing angle parameter set of the curved screen, and perform residual feature processing on the original image sequence to obtain a multi-dimensional spatial feature map;
[0007] Performing a tensor product operation on the multidimensional spatial feature map and the viewing angle parameter set to obtain viewing angle conversion delay data;
[0008] The multidimensional space feature graph is set as a node attribute of a graph neural network, and the perspective conversion delay data is set as an edge attribute to perform a graph convolution operation to obtain a spatiotemporal dependency feature;
[0009] Performing multi-head self-attention calculation on the spatiotemporal dependent features to obtain delay compensation features, and performing comparative learning on the perspective features based on the delay compensation features to obtain perspective mapping parameters;
[0010] The perspective mapping parameters, the delay compensation features and the original image sequence are input into a U-Net network for 3D image reconstruction, and a 3D image sequence with a continuous perspective is output.
[0011] In a second aspect, the present invention provides a curved screen multi-view 3D image generation device, the curved screen multi-view 3D image generation device comprising:
[0012] An acquisition module is used to acquire a set of viewing angle parameters of the curved screen and perform residual feature processing on the original image sequence to obtain a multi-dimensional spatial feature map;
[0013] A tensor product operation module, used for performing a tensor product operation on the multidimensional spatial feature map and the viewing angle parameter set to obtain viewing angle conversion delay data;
[0014] A graph convolution operation module, used to set the multidimensional space feature map as a node attribute of a graph neural network and the perspective conversion delay data as an edge attribute to perform a graph convolution operation to obtain a spatiotemporal dependency feature;
[0015] A perspective mapping module, configured to perform multi-head self-attention calculation on the spatiotemporal dependent features to obtain delay compensation features, and perform comparative learning on perspective features based on the delay compensation features to obtain perspective mapping parameters;
[0016] The image reconstruction module is used to input the perspective mapping parameters, the delay compensation features and the original image sequence into a U-Net network to perform 3D image reconstruction and output a 3D image sequence with a continuous perspective.
[0017] In the technical solution provided by the present invention, by constructing a surface adaptive convolution layer, the size of the receptive field is dynamically adjusted for different surface positions, which effectively solves the problem of uneven image feature extraction caused by the curvature characteristics of the curved screen and improves the accuracy of feature expression. The multi-perspective conversion delay is modeled, and the dual-path feature extraction structure of the branch network and the trunk network is combined to achieve an accurate description of the nonlinear perspective conversion characteristics, effectively reducing the delay phenomenon during perspective switching. By introducing a graph neural network structure, spatial features and delay data are modeled as node and edge attributes respectively, which enhances the ability to characterize the spatiotemporal dependency between perspectives and ensures the continuity of perspective conversion. Based on the multi-head self-attention mechanism and contrast learning strategy, accurate alignment and mapping of cross-perspective features are achieved, feature consistency during perspective conversion is improved, and image distortion during perspective switching is reduced. Through the multi-scale feature fusion and spatial adaptive normalization processing of the U-Net network, smooth perspective conversion is achieved while maintaining image details, significantly improving the visual quality of 3D images. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.
[0019] Figure 1 Schematic diagram of the steps of the method for generating multi-view 3D images on a curved screen in an embodiment of the present invention;
[0020] Figure 2 Schematic diagram of the structure of a curved screen multi-viewing angle 3D image generating device in an embodiment of the present invention. DETAILED DESCRIPTION
[0021] The embodiments of the present invention provide a method and device for generating multi-view 3D images for a curved screen. The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0022] For ease of understanding, the specific process of the embodiment of the present invention is described below. Figure 1 , an embodiment of the method for generating a multi-view 3D image on a curved screen in an embodiment of the present invention includes:
[0023] Step S1, collecting a viewing angle parameter set of the curved screen, and performing residual feature processing on the original image sequence to obtain a multi-dimensional space feature map;
[0024] It is understandable that the execution subject of the present invention may be a curved screen multi-view 3D image generating device, or a terminal or a server, which is not limited here. The embodiment of the present invention is described by taking a server as the execution subject as an example.
[0025] Specifically, the key geometric parameters of the curved screen are measured to obtain its curvature radius R, screen width W and viewing angle range θ, and a three-dimensional coordinate system suitable for the curved screen is established based on these parameters. In this coordinate system, the horizontal direction is defined as the x-axis, the vertical direction is defined as the y-axis, and the depth of field direction is defined as the z-axis. By modeling the curved screen at different spatial positions, a set of viewing angle parameters is obtained. A five-layer residual network structure is constructed, and each residual block contains two 3×3 convolution layers and is equipped with a ReLU activation function to effectively alleviate the gradient vanishing problem while ensuring the depth of the model and enhance the ability to extract feature information. After the original image sequence passes through the five-layer residual network, a basic feature map with rich local information is formed, which contains multi-dimensional information such as spatial structure, edge information and texture features. Due to the special morphology of the curved screen, ordinary convolution operations cannot accurately capture the pixel distribution characteristics on the curved screen in a plane coordinate system, so a surface adaptive convolution layer is introduced to meet the feature extraction requirements of different curvature positions. Based on the curvature radius R, an adaptive convolution kernel structure is constructed so that the size of the receptive field at different positions can be dynamically adjusted as the surface changes. In this process, the neighborhood relationship of each pixel is transformed according to the curvature characteristics of the curved screen to ensure that the directionality of feature extraction can adapt to the geometric form of the screen and maintain a high feature consistency in the entire image space. After the basic feature map passes through the surface adaptive convolution layer, a surface feature map is generated. This feature map not only contains the information of the original image, but also corrects it based on the surface geometry, so that the feature distribution is more consistent with the actual form of the physical screen. Global average pooling and maximum pooling operations are performed on the surface feature map to capture the average features and the most significant features in the global range, respectively, and construct spatial attention weights. By performing matrix multiplication operations on the surface feature map and the spatial attention weight, a weighted feature map is generated, which enables the network to more effectively focus on the feature areas with high importance and improve the accuracy of perspective conversion. Batch normalization is performed on the weighted feature map to ensure that the data maintains a stable distribution between different batches and prevent gradient vanishing or gradient explosion problems. Batch normalization can speed up network training and improve the generalization ability of the model, so that the features can remain stable when multiple perspectives change. Output multi-dimensional spatial feature map.
[0026] Step S2, performing a tensor product operation on the multidimensional spatial feature map and the viewing angle parameter set to obtain viewing angle conversion delay data;
[0027] Specifically, a computational framework consisting of a branch network and a trunk network is constructed, wherein the branch network is used for dimensionality reduction mapping of multidimensional spatial feature maps, while the trunk network is used for temporal feature extraction of a set of view parameters. In this structure, the branch network consists of four fully connected layers, while the trunk network contains three fully connected layers. This structural design ensures that features are fully learned while reducing computational complexity. The multidimensional spatial feature map is input into the branch network for feature dimensionality reduction mapping. Since the multidimensional spatial feature map has a high dimension, the dimension is reduced through the branch network. The four fully connected layers gradually extract and compress the feature dimensions to generate a more compact feature representation and obtain a feature mapping vector. At the same time, the view parameter set is preprocessed for temporal feature extraction. Since the view change is a dynamic process, the current view and the target view are timestamped to represent the temporal relationship between different viewpoints. The introduction of timestamps ensures that the model understands the temporal dependency of view conversion and provides structured information for subsequent temporal feature extraction. After the timestamp is marked, the marked view parameters are input into the backbone network. The three-layer fully connected structure of the network gradually extracts the temporal characteristics of the view parameters and generates a temporal mapping vector. This vector effectively encodes the temporal correlation between different viewpoints. A tensor product operation is performed on the feature mapping vector and the temporal mapping vector to calculate the view conversion feature vector. The role of the tensor product operation is to interactively fuse the data of two different modalities so that the spatial features and the temporal features work together to capture the multi-dimensional view change law. The mean square error loss function is used to supervise the view conversion feature vector to minimize the error between the conversion feature and the target conversion feature. At the same time, in order to prevent overfitting and improve the generalization ability of the model, the L2 regularization term is introduced to smooth the feature vector, thereby limiting the excessive fluctuation of the model parameters and making the prediction of the view conversion more stable and reliable. After the optimization training is completed, the optimized feature vector is input into the nonlinear mapping layer for view conversion delay calculation. The nonlinear mapping layer maps the optimized feature vector through a multi-layer perceptron or other nonlinear transformation structure to calculate the final view conversion delay data. This data is used to describe the temporal relationship between different perspectives during the perspective conversion process and is applied to multi-perspective 3D image generation to ensure timing consistency and visual smoothness when switching perspectives.
[0028] Step S3, setting the multidimensional space feature map as the node attribute of the graph neural network, setting the perspective conversion delay data as the edge attribute to perform graph convolution operation, and obtain the spatiotemporal dependency feature;
[0029] Specifically, the multidimensional spatial feature map is used as the node feature of the graph neural network, and the perspective conversion delay data is used as the edge weight to construct a complete graph structure. On this basis, an adjacency matrix is constructed, which is used to describe the connection relationship between nodes in the graph. The degree matrix is calculated for the adjacency matrix to obtain the angle matrix. The degree matrix is a diagonal matrix whose diagonal elements represent the number of connections of each node, so that the local structural information of the node can be reflected. In order to avoid the influence of the difference in the number of connections of different nodes on the calculation, the angle matrix is normalized to generate a normalized matrix, so as to ensure that the influence of each node can be consistent during feature propagation and avoid instability caused by too large or too small degree during the calculation process. The normalized matrix is matrix multiplied with the adjacency matrix and the node feature to calculate the first-order feature propagation result. This calculation process is an information aggregation based on the graph structure, that is, each node not only retains its own feature information, but also receives the information of its neighboring nodes, and performs weighted summation to form an update result that integrates the features of surrounding nodes. The feature propagation method based on the adjacency matrix can effectively capture local spatial relationships, making the feature expression more consistent with the spatial dependency pattern of the perspective conversion. The first-order feature propagation result is transformed to enhance the model's expressiveness. The result is linearly transformed to adjust the feature dimension to adapt it to the subsequent calculation requirements. The linearly transformed data is nonlinearly mapped through the ReLU activation function to increase the model's learning ability and enable it to capture more complex nonlinear relationships. After this process, a single-layer graph convolution output is obtained, which contains the feature representation after one information propagation and can reflect the perspective feature fusion in the local area. Multi-layer feature propagation is performed by stacking three layers of graph convolution layers to perform feature propagation in a wider range. By stacking three layers of graph convolution layers, information is ensured to be propagated to more distant nodes, so that each node can integrate neighborhood features from a wider range. At the same time, in order to prevent information from being lost during the propagation process, a jump connection mechanism is introduced to enable the features of different layers to be fused, thereby improving the model's feature expression ability. The multi-layer feature fusion result is obtained by fusion of multi-layer graph convolution and jump connection. A global attention layer is added based on the multi-layer feature fusion result to calculate the attention weights between different nodes. The global attention mechanism adaptively distributes the importance of different nodes according to feature relevance, thereby enhancing the influence of key nodes and suppressing the interference of irrelevant features. Through the calculation of the global attention layer, the spatiotemporal dependency feature is output, which has spatial continuity and can capture the temporal correlation during the perspective conversion process.
[0030] Step S4, performing multi-head self-attention calculation on the spatiotemporal dependent features to obtain delay compensation features, and performing comparative learning on the perspective features based on the delay compensation features to obtain perspective mapping parameters;
[0031] Specifically, the spatiotemporal dependent features are linearly mapped to obtain the query matrix (Query), key matrix (Key), and value matrix (Value). The query matrix is used to indicate how the current feature vector matches other features, while the key matrix is used to store the matching basis of all features, and the value matrix is used to store the actual information content, so that the subsequent attention calculation can extract key information from the value matrix according to the matching of the query and the key. The query matrix and the key matrix are matrix multiplied and scaled to ensure that the calculation result has good numerical stability. The result of the matrix multiplication is normalized by dividing it by the square root of the dimension to avoid the problem of gradient vanishing or gradient exploding due to excessive values, thereby ensuring that the numerical range of the attention calculation is within a reasonable range. After scaling, the obtained attention score is softmax normalized to convert it into a normalized weight distribution, so that the attention can be more focused on important information while suppressing irrelevant interference information. The normalized attention score is matrix multiplied with the value matrix to calculate the single-head attention output, which reflects how each feature interacts with other features, thereby extracting key information in the spatiotemporal features. The spatiotemporal dependent features are projected into 8 different subspaces, and independent attention calculations are performed on each subspace to generate 8 groups of attention outputs, ensuring that different attention heads can focus on different types of spatiotemporal features respectively, so that the final attention results have stronger expressiveness and higher robustness. After completing the 8 groups of attention calculations, all attention outputs are spliced to obtain a complete feature representation. The linear mapping layer is used to perform dimension conversion to ensure that the final multi-head attention results are adapted to the subsequent network structure. The multi-head attention results are feature processed to improve the stability and expressiveness of the model. The multi-head attention results are input into a two-layer feedforward neural network and processed by residual connection and layer normalization to ensure the distribution stability of the features and improve the nonlinear expression ability of the model. The role of the feedforward neural network is to transform the attention features to improve its modeling ability for perspective conversion, and the introduction of residual connection ensures that the gradient is better propagated in the deep network to avoid the gradient vanishing problem. Through the layer normalization operation, the data of different batches are ensured to remain consistent during the calculation process, thereby improving the generalization ability of the model. Through calculation, the delay compensation feature is obtained, which can effectively describe the time dependency in the process of perspective conversion and provide accurate compensation information. Based on the delay compensation feature, the perspective feature is contrastively learned to obtain the final perspective mapping parameters. Contrastive learning constructs positive and negative samples so that the model can learn the features of perspective mapping more accurately. In this process, the similarity of features under different perspectives is calculated, and a stable mapping model is trained by maximizing the similarity between positive samples and minimizing the similarity between negative samples.By contrastive learning of the delay compensation features, the perspective mapping parameters are finally obtained. The parameters are used to control the conversion between different perspectives in the process of multi-perspective 3D image generation, so that the generated image sequence maintains coherence between multiple perspectives and ensures the smoothness and authenticity of the perspective change.
[0032] The delay compensation features are mapped to a low-dimensional space. They are converted into compact low-dimensional feature vectors by dimensionality reduction to retain key features while reducing redundant information. The delay compensation features are mapped using a fully connected layer or dimensionality reduction projection operation to reduce the data dimension to a range suitable for subsequent comparative learning and ensure that the feature vector can still effectively express the most critical information in the perspective conversion process. After completing the low-dimensional space mapping, a feature memory queue is constructed based on the low-dimensional feature vector to store features of multiple perspectives and reasonably organize positive and negative sample pairs. Since perspective conversion is a continuous process, adjacent perspectives have high similarity, while perspectives that are far apart have large differences in feature space. Temporally continuous adjacent perspective feature pairs are set as positive sample pairs to learn smooth conversion patterns between similar perspectives, and non-adjacent perspective feature pairs are set as negative sample pairs to help the model identify global change relationships between perspectives. The role of the feature memory queue is to maintain a continuously updated sample library so that the model can continuously introduce new feature pairs during training, thereby improving generalization ability and ensuring that the learned perspective mapping parameters are applicable to a wider range of scenarios. After constructing the training sample set, the similarity between feature pairs is calculated to establish a preliminary feature mapping function. The purpose of similarity calculation is to measure the degree of feature difference between different perspectives, so as to provide an optimization target for contrastive learning. By calculating the similarity between perspective features, it is determined which perspectives are closer in the feature space and which perspectives are more difficult to convert. In this process, the similarity score is used to guide model optimization so that the features of positive sample pairs are closer and the features of negative sample pairs are more separated, thereby enhancing the predictability of perspective conversion and improving the stability of generated 3D images. However, the initial feature mapping function is calculated only based on feature similarity, without considering the difference in perspective conversion difficulty, which is uneven in practical applications. For example, perspective switching within a smaller angle range is usually smoother, while switching at a larger angle involves more complex geometric transformations. In order to improve the adaptability of the model, a dynamic weight mechanism is introduced to assign different weight coefficients to different sample pairs according to the complexity of perspective conversion. Samples with higher difficulty perspective conversion are given greater influence during training, so that the model pays more attention to the optimization of high-difficulty conversions, while samples with lower difficulty perspective conversion are given lower weights to prevent the model from overfitting to simple conversion patterns. Through the dynamic weight mechanism, the model gradually adjusts its mapping ability during the learning process, so that the perspective conversion can maintain high accuracy and high stability in different scenarios. Based on the weighted feature mapping function, the perspective mapping parameter optimization calculation of gradient descent is performed to calculate the final perspective mapping parameters. In this process, the optimization goal is to minimize the gap between positive sample pairs and maximize the distinction between negative sample pairs, so as to ensure that the learned perspective mapping parameters can accurately describe the relationship between different perspectives.By continuously optimizing the perspective mapping parameters, the model gradually adjusts the mapping rules during the training process to make it more consistent with the actual perspective conversion rules, and finally obtains the optimized perspective mapping parameters.
[0033] Step S5: input the perspective mapping parameters, delay compensation features and the original image sequence into the U-Net network to perform 3D image reconstruction, and output a 3D image sequence with a continuous perspective.
[0034] Specifically, the encoder part of the U-Net network is constructed to extract multi-scale features from the original image sequence. Since the U-Net network has the ability to encode and decode high-resolution images layer by layer, in the encoder part, four downsampling processes are performed in sequence through the downsampling convolution module, so that the network can learn image features at different scales. In this process, each downsampling not only reduces the spatial resolution of the feature map, but also increases the channel dimension, so that the network can capture more abstract feature information at a deep level and extract multi-level feature maps, including different levels of data from local edge information to global semantic information. A spatial adaptive normalization layer is constructed based on the delay compensation feature to adjust the feature distribution of the multi-level feature map at different scales. The spatial adaptive normalization layer normalizes the multi-layer features generated by the U-Net encoder according to the dynamic changes of the delay compensation feature, so that the features under different perspectives maintain a consistent distribution pattern. The normalization operation can reduce the brightness difference between different perspectives and improve the model's adaptability to non-uniform feature changes during perspective conversion, thereby ensuring that the features extracted under different perspectives are more stable and reliable. After completing feature normalization, the perspective mapping parameter is introduced to modulate the features in a targeted manner. Construct a feature modulation module, input the view mapping parameters into the module, and dynamically adjust the channel weights of the normalized feature map. In this way, the network can adaptively adjust the feature expression mode under different view conditions, so that the view change can be captured and reconstructed more accurately. The dynamic modulation mechanism ensures that important features can be fully retained during the view conversion process, while irrelevant features are appropriately weakened, thereby improving the accuracy and naturalness of the view conversion. After this stage of processing, the network obtains a view-sensitive feature map. Construct the decoder part of the U-Net network to achieve the final 3D image reconstruction. The decoder performs four feature amplifications on the view-sensitive feature map through the upsampling transposed convolution module, so that it gradually recovers to the same resolution as the original image. In order to avoid the loss of information during the upsampling process and ensure the integrity of the details of the generated image, a jump connection is introduced in each amplification process to splice and fuse the encoding features and decoding features of the corresponding level. The role of the jump connection is to directly transfer high-resolution local detail information from the encoder part to the decoder part, so that when the image is restored, the edge information and texture details of the original image can be better maintained, thereby improving the quality of the reconstructed image. After feature amplification and fusion, the generated reconstructed feature map is subjected to brightness equalization and color correction to ensure that the final generated 3D image sequence is more natural and realistic visually. The purpose of brightness equalization is to eliminate the local brightness difference caused by the change of viewing angle, so that the brightness of images from different viewing angles remains consistent throughout the sequence, while the role of color correction is to adjust the color distribution between different viewing angles so that the generated image will not have obvious color cast or incoherence problems.This adjustment step ensures that the final 3D image sequence has stable brightness and color performance at different viewing angles, thereby enhancing the visual experience. The initial 3D image sequence is spatially enhanced to improve the clarity and sharpness of the image. Since the image may be slightly blurred or have unclear boundaries during the viewing angle conversion process, the spatial edge enhancement method is used to make the image outline clearer and improve the contrast of local areas, thereby enhancing the final 3D image quality. A set of 3D image sequences with continuous viewing angles is generated, so that when viewing from multiple viewing angles on a curved screen, the image transition is smooth and natural, while maintaining high-quality visual details and stereoscopic effects, thereby enhancing the immersion and realism of the 3D display.
[0035] In the embodiment of the present invention, by constructing a surface adaptive convolution layer, the size of the receptive field is dynamically adjusted for different surface positions, which effectively solves the problem of uneven image feature extraction caused by the curvature characteristics of the curved screen and improves the accuracy of feature expression. The multi-perspective conversion delay is modeled, and the dual-path feature extraction structure of the branch network and the trunk network is combined to achieve an accurate description of the nonlinear perspective conversion characteristics, effectively reducing the delay phenomenon during perspective switching. By introducing a graph neural network structure, spatial features and delay data are modeled as node and edge attributes respectively, which enhances the ability to characterize the spatiotemporal dependency between perspectives and ensures the continuity of perspective conversion. Based on the multi-head self-attention mechanism and contrast learning strategy, accurate alignment and mapping of cross-perspective features are achieved, feature consistency during perspective conversion is improved, and image distortion during perspective switching is reduced. Through the multi-scale feature fusion and spatial adaptive normalization processing of the U-Net network, smooth perspective conversion is achieved while maintaining image details, significantly improving the visual quality of 3D images.
[0036] In a specific embodiment, the process of executing step S1 may specifically include the following steps:
[0037] Measure the parameters of the curved screen to obtain the curvature radius R, screen width W, and viewing angle range θ, establish a three-dimensional coordinate system with the horizontal direction as the x-axis, the vertical direction as the y-axis, and the depth of field direction as the z-axis, and obtain a viewing angle parameter set;
[0038] Construct a five-layer residual network structure, each of which contains two 3×3 convolutional layers and a ReLU activation function. Input the original image sequence into the five-layer residual network structure for basic feature processing to obtain the basic feature map.
[0039] A surface adaptive convolution layer is constructed according to the curvature radius R, and the corresponding receptive field size is set for different surface positions. The basic feature map is input into the surface adaptive convolution layer to obtain the surface feature map.
[0040] Perform global average pooling and maximum pooling operations on the surface feature map to obtain the spatial attention weight, and perform matrix multiplication operation on the surface feature map and the spatial attention weight to obtain the weighted feature map;
[0041] Batch normalization is performed on the weighted feature map to output a multidimensional spatial feature map.
[0042] Specifically, the parameters of the curved screen are measured to obtain the key parameters for constructing the coordinate system. Since the geometric shape of the curved screen determines the way the viewing angle is converted, the radius of curvature is measured and obtained. , Screen Width And the viewing angle range Among them, the radius of curvature Defined as the radius of the curved screen, indicating the curvature of the screen, the screen width Refers to the total length of the curved screen in the horizontal direction, while the viewing angle range It represents the coverage of the observation angle and is used to describe the limit of the change of the viewing angle. Based on these parameters, a horizontal direction is established. Axis, vertical direction is Axis, depth of field direction is The three-dimensional coordinate system of the axis is established to obtain the set of viewing angle parameters. After the coordinate system is established, feature extraction is performed on the input original image sequence to extract key image information and provide high-quality basic features for subsequent processing steps. A five-layer residual network structure is constructed, and each layer of residual blocks contains two The convolutional layer is used to extract local features, while the ReLU activation function is used to introduce nonlinear characteristics to enhance the model's expressiveness. In each residual block, the input features are processed by two After the convolution layer is calculated, an intermediate feature map is generated and output after being activated by ReLU. To avoid information loss, each residual block contains a jump connection so that the input is directly added to the output, thus forming a residual learning mechanism. For each layer of residual blocks, it is expressed by the following formula:
[0043] ;
[0044] in, Indicates The output feature map of the layer, and Represent the weight matrices of the two convolutional layers respectively, and is the corresponding bias term, represents the ReLU activation function, and Represents the convolution operation. The advantage of the residual structure is that even if the number of network layers increases, the gradient can still be effectively propagated, thereby avoiding the gradient vanishing problem and improving the training efficiency of the model. After calculation by the five-layer residual network, the basic feature map is obtained, which contains rich local features and can provide input data for subsequent surface adaptive calculations. However, since the geometric shape of the curved screen is different from that of the ordinary flat screen, directly using the traditional convolution layer will cause the local receptive field of feature extraction to not match, thereby affecting the accuracy of the perspective conversion. Therefore, according to the radius of curvature Construct a surface adaptive convolution layer so that different surface positions have corresponding receptive field sizes. Define the receptive field radius As the size of the surface adaptive convolution kernel, the radius is adjusted as the surface position changes, and its calculation method is as follows:
[0045] ;
[0046] in, Indicates the horizontal coordinate value of the current calculated position. Represents the tangent function, which is used to calculate the receptive field range corresponding to different positions on the screen. In areas where the surface is more curved, the receptive field will increase accordingly to capture more spatial information, while in areas with smaller curvature, the receptive field will decrease to maintain accurate feature extraction. In this way, it is ensured that the network can effectively extract features on the entire surface and generate a surface feature map. Spatial attention calculation is performed on the surface feature map to enhance the feature expression ability of key areas. Global average pooling and maximum pooling operations are performed on the surface feature map to obtain spatial attention weights. Among them, the role of global average pooling is to calculate the global information of the entire feature map, while the role of maximum pooling is to highlight the most significant local features. Assume that the size of the feature map is ,in is the height, is the width, is the number of channels, then the global average pooling calculation formula is as follows:
[0047] ;
[0048] in, Representative The global average pooling result of channels, Represents the surface feature map in The pixel value at each channel. Similarly, the maximum pooling calculation formula is as follows:
[0049] ;
[0050] in, represents the maximum pooling result, which can highlight the most representative eigenvalues. The two pooling results are fused to generate the final spatial attention weight. The weighted feature map is obtained by performing matrix multiplication operation on the surface feature map and the spatial attention weight:
[0051] ;
[0052] in, is the spatial attention weight matrix, is the surface feature map, and This is a weighted feature map that can highlight the features of important areas while suppressing the interference of irrelevant areas. After the weighted feature calculation is completed, batch normalization is performed on it to ensure that the data has a stable distribution between different batches and prevent the gradient disappearance or gradient explosion problem. The calculation formula of batch normalization is as follows:
[0053] ;
[0054] in, represents the mean of the current batch, Represents the standard deviation. The role of batch normalization is to ensure the stability of feature distribution, so that the neural network can converge faster during training and improve generalization ability. Through the above steps, a multidimensional spatial feature map is output.
[0055] This embodiment also includes: performing curvature segmentation on the curved screen surface, dividing the curved screen into N curvature intervals according to the curvature radius R, establishing a local coordinate system for each curvature interval, and obtaining a curvature distribution map; inputting the curvature distribution map into a surface adaptive deformation network, extracting surface deformation features through three convolution layers, and obtaining a surface deformation tensor; performing surface reprojection on the multi-dimensional space feature map, and performing local geometric correction on the feature map based on the surface deformation tensor to obtain a surface correction feature map; constructing a perspective perception attention module, performing feature fusion on the perspective parameter set and the surface correction feature map, and highlighting the feature components related to the perspective through the channel attention mechanism to obtain a perspective enhancement feature map; setting a depth perception convolution layer for the perspective enhancement feature map, Features at different depths are adaptively sampled through deformable convolution to obtain a depth-sensitive feature map; a multi-scale feature pyramid is constructed based on the depth-sensitive feature map, and the spatial dependency from local to global is captured through feature maps of different scales to obtain a multi-scale feature sequence; the multi-scale feature sequence is fused through a cross-scale feature aggregation module, and features of different scales are adaptively combined using dynamic weights to obtain an aggregated feature map; a disparity consistency check is performed on the aggregated feature map, and feature matching between adjacent viewpoints is calculated to correct features in inconsistent areas, and a viewpoint compensation feature map is output; batch normalization and feature map fusion processing are performed on the viewpoint compensation feature map and the weighted feature map to output a multi-dimensional spatial feature map.
[0056] In a specific embodiment, the process of executing step S2 may specifically include the following steps:
[0057] Build a branch network and a backbone network, where the branch network contains four fully connected layers and the backbone network contains three fully connected layers;
[0058] Input the multidimensional space feature map into the branch network for feature dimension reduction mapping to obtain a feature mapping vector;
[0059] Timestamp the current view and the target view in the view parameter set to obtain the marked view parameters, and input the marked view parameters into the backbone network to extract the timing characteristics to obtain the timing mapping vector;
[0060] Perform a tensor product operation on the feature mapping vector and the time series mapping vector to obtain a view conversion feature vector;
[0061] The mean square error loss function is used to optimize the view conversion feature vector, and the L2 regularization term is introduced for smoothness constraint to obtain the optimized feature vector.
[0062] The optimized feature vector is input into the nonlinear mapping layer to calculate the view conversion delay and output the view conversion delay data.
[0063] Specifically, two independent neural network structures are constructed, in which the main function of the branch network is to reduce the dimension of the input multi-dimensional spatial feature map, while the backbone network is used to process the temporal information in the perspective parameter set to calculate the perspective conversion feature vector. The branch network consists of four fully connected layers, each of which gradually reduces the dimension of the input features and enhances the expressive power of the model through nonlinear activation functions, while the backbone network contains three fully connected layers, which are used to learn the temporal dependency between different perspective parameters and provide temporal features for the final perspective conversion calculation. The multi-dimensional spatial feature map is input into the branch network for feature dimensionality reduction mapping. Assume that the input multi-dimensional spatial feature map is represented as a matrix ,in represents the sample size, Represents the dimension of the original feature, then the first layer of the branch network is fully connected as:
[0064] ;
[0065] in, represents the output features of the first layer, is the weight matrix of the first layer, is the bias term, is the output dimension of the layer, and Represents a nonlinear activation function, such as ReLU or Leaky ReLU. Similarly, after four layers of full connection transformation, the reduced-dimensional feature map vector is obtained:
[0066] ;
[0067] in, represents the final feature mapping vector, which has been dimensionally reduced and can more compactly represent the original spatial feature information while retaining key features. At the same time, in order to learn the temporal characteristics of perspective conversion, the perspective parameter set is timestamped to construct the temporal feature input. Assume that the parameters of the current perspective are , the target viewing angle parameter is ,in Represent the three-dimensional coordinates, Represents the viewing angle, so a timestamp variable is introduced To mark:
[0068] ;
[0069] in, and Represent the timestamp information of the current view and the target view respectively. The marked view parameters are input into the backbone network, and after three layers of full connection mapping, the temporal mapping vector is extracted:
[0070] ;
[0071] in, Represents the final temporal feature, which is used to represent the perspective change pattern at different time points during the perspective conversion process. and the timing mapping vector The tensor product operation is performed on these two vectors to calculate the final perspective transformation feature vector. The calculation of the tensor product is defined as follows:
[0072] ;
[0073] in, Represents the tensor product operation, which is used to fuse spatial features with temporal features to learn the transformation relationship between different perspectives. It is calculated using element-by-element multiplication:
[0074] ;
[0075] This operation ensures the close integration of spatial features and temporal features, so that the spatial structure and temporal dependency can be comprehensively considered during the perspective conversion process. The mean square error loss function is used to optimize the perspective conversion feature vector to minimize the conversion error. The mean square error loss function is defined as follows:
[0076] ;
[0077] in, represents the predicted view transformation feature vector, represents the true conversion characteristics, is the number of samples. The purpose of this loss function is to minimize the square error between the predicted value and the true value, thereby improving the accuracy of the perspective conversion. In order to prevent the model from overfitting and ensure the stability of parameter optimization, we introduce Regularization term to impose smoothness constraints on the weights. The regularization is calculated as follows:
[0078] ;
[0079] in, is the regularization coefficient, Represents the weight parameter in the model. The role of this regularization term is to limit excessive fluctuations in the parameters, thereby improving the generalization ability of the model and preventing numerical instability problems during training. After completing the optimization training of the perspective conversion feature, the optimized feature vector is input into the nonlinear mapping layer to calculate the final perspective conversion delay data. The role of the nonlinear mapping layer is to make the final conversion feature better adapt to different perspective change modes through a series of nonlinear transformations. Assume that the calculation formula of the nonlinear mapping layer is as follows:
[0080] ;
[0081] in, Represents the final perspective conversion delay data, and are the weights and biases of the mapping layer, and Functions are used to introduce nonlinear characteristics to ensure dynamic adjustment capabilities of the output data.
[0082] In this embodiment, before performing a tensor product operation on the multidimensional space feature map and the perspective parameter set to obtain the perspective conversion delay data, setting the multidimensional space feature map as the node attribute of the graph neural network and the perspective conversion delay data as the edge attribute for graph convolution operation, it also includes: constructing a nonlinear affine state equation for the perspective conversion delay data, setting the curvature characteristics and the perspective conversion characteristics of the curved screen as state variables, and obtaining a perspective switching state model; constructing a nonlinear disturbance observer based on the perspective switching state model, and online observing the matching interference items and non-matching interference items in the perspective conversion process to obtain the disturbance observation results; dynamically compensating the disturbance observation results, calculating the compensation gain through a homogeneous second-order quasi-continuous algorithm, and obtaining a delay compensation vector ; Construct an integral sliding mode controller based on the delay compensation vector, design a dynamic switching surface including the integral term of the state error, and obtain the switching control parameters; input the switching control parameters into the quasi-continuous controller, perform nonlinear correction on the perspective conversion delay data through the continuous reaching law, and obtain the corrected delay data; construct a Lyapunov function based on the corrected delay data, design stability constraints through quadratic functions, and obtain stability constraint parameters; perform terminal convergence control on the corrected delay data according to the stability constraint parameters, optimize the parameters through the finite time convergence criterion, and obtain the optimized delay data; perform jitter suppression processing on the optimized delay data, eliminate high-frequency disturbances in the perspective switching process through an adaptive smoothing algorithm, and output the perspective conversion delay data.
[0083] In a specific embodiment, the process of executing step S3 may specifically include the following steps:
[0084] The multidimensional space feature graph is set as the node feature, the view conversion delay data is set as the edge weight, and the adjacency matrix is constructed;
[0085] Perform degree matrix calculation on the adjacency matrix to obtain an angle matrix, and perform normalization operation on the angle matrix to obtain a normalized matrix;
[0086] Perform matrix multiplication on the normalized matrix, the adjacency matrix and the node features to obtain the first-order feature propagation result;
[0087] The first-order feature propagation result is linearly transformed and nonlinearly mapped through the ReLU activation function to obtain a single-layer graph convolution output;
[0088] By stacking three layers of graph convolutional layers, multi-layer feature propagation is performed, and the features of each layer are fused through jump connections to obtain a multi-layer feature fusion result;
[0089] Based on the multi-layer feature fusion results, a global attention layer is added to calculate the attention weights between different nodes and output spatiotemporal dependency features.
[0090] Specifically, the multidimensional space feature map is set as the node feature of the graph neural network, and the view conversion delay data is set as the edge weight to construct an adjacency matrix for feature propagation. By node set and edge sets Composition, where each node Represents the spatial features corresponding to a perspective, edge represents the conversion relationship between two perspectives, and the weight of the edge is determined by the perspective conversion delay data. Assume that the node feature matrix of the multidimensional space feature graph is ,in is the number of nodes, is the feature dimension of each node, and the adjacency matrix It represents the connection relationship between nodes and contains the corresponding edge weights. In order to perform feature propagation, the corresponding degree matrix is calculated , which is defined as a diagonal matrix with the diagonal elements Representation Node The degree of a node is the sum of the weights of all the edges connecting it:
[0091] ;
[0092] in Represents a slave node To Node To ensure that information is not unbalanced due to degree differences during feature propagation, the degree matrix is normalized to obtain the normalized matrix ,in:
[0093] ;
[0094] The purpose of this normalization operation is to ensure that information will not be unevenly distributed due to the large or small degree of some nodes during feature propagation. With the adjacency matrix And the node feature matrix Perform matrix multiplication to calculate the first-order feature propagation result. Perform normalized adjacency matrix calculation:
[0095] ;
[0096] Compute the first-order eigenpropagation:
[0097] ;
[0098] in, is the trainable weight matrix of the first layer of graph convolution, is the output of the first-order feature propagation, which represents the information of each node on the graph after one adjacency propagation. In order to enhance the feature expression capability, the first-order feature propagation result is linearly transformed and nonlinearly mapped through the ReLU activation function to introduce more complex feature representation capabilities:
[0099] ;
[0100] in, ReLU stands for the ReLU activation function, which removes negative values so that the neural network can learn nonlinear relationships more effectively. A single-layer graph convolution can only capture the information of first-order neighbors, but in the process of perspective conversion, long-distance perspective changes also need to be modeled. Therefore, three layers of graph convolution are stacked to perform deeper feature propagation and gradually extract multi-layer features. The graph convolution calculation of the layer is expressed as:
[0101] ;
[0102] in, Representative The trainable weight matrix of the layer, and It means after Feature vector after layer graph convolution. In the process of multi-layer feature propagation, since information will be lost in the deep propagation process, the skip connection mechanism is introduced to fuse the features of different layers to retain the low-level local information and extract higher-level global patterns. The calculation method of the skip connection is as follows:
[0103] ;
[0104] in, Represents the final multi-layer feature fusion result, which integrates the first-order, second-order, and third-order features and can maintain a good information flow, thereby improving the representation ability of spatiotemporal features. In order to improve the model's ability to understand different perspectives, a global attention layer is added on the basis of multi-layer feature fusion to calculate the attention weights between different nodes. The calculation method of global attention is as follows:
[0105]
[0106] in, Representative Node and nodes The attention score between is defined as follows:
[0107] ;
[0108] in, is a trainable attention parameter, Represents a vector concatenation operation, and Respectively represent nodes and In the The feature representation of the layer. The calculated attention weights Used to weight node features to generate the final spatiotemporal dependency features:
[0109] ;
[0110] in, Representative Node The final spatiotemporal dependency feature of , which integrates the information of multiple neighbor nodes through the global attention mechanism, enables the model to better capture the global relationship in perspective conversion.
[0111] Among them, before obtaining the spatiotemporal dependent features and performing multi-head self-attention calculations, it also includes: constructing a perspective constraint network, inputting the spatiotemporal dependent features into the perspective constraint network, extracting the spatial position relationship between adjacent perspectives, and obtaining a perspective constraint matrix; performing multi-point optical sampling on the curved screen, obtaining reflectivity and scattering rate data on the screen surface, establishing a light propagation model, and obtaining perspective optical features; performing feature fusion on the perspective constraint matrix and the perspective optical features, and obtaining a fused feature vector through weighted average operation; constructing a constrained optimization model, setting the fused feature vector as a constraint condition, performing nonlinear optimization on the spatiotemporal dependent features, and obtaining initial optimized features; constructing a noise model based on the perspective optical features, performing noise modeling on the initial optimized features, and obtaining a noise feature distribution; using a variational inference method, performing probability optimization on the initial optimized features based on the noise feature distribution, and obtaining optimized spatiotemporal dependent features; using a Kalman filter, performing time series filtering on the optimized spatiotemporal dependent features, eliminating jitter during the perspective conversion process, and obtaining smooth spatiotemporal features; inputting the smooth spatiotemporal features into a feature calibration network, correcting feature deviations through residual learning, and outputting calibrated spatiotemporal dependent features.
[0112] In a specific embodiment, the process of executing step S4 may specifically include the following steps:
[0113] Linearly map the spatiotemporal dependency features to obtain the query matrix, key matrix, and value matrix;
[0114] Perform matrix multiplication of the query matrix and the key matrix and perform scaling, and normalize by dividing by the square root of the dimension to get the attention score;
[0115] Perform softmax operation on the attention score and perform matrix multiplication operation with the value matrix to get the single-head attention output;
[0116] Project the spatiotemporal dependent features into 8 different subspaces, perform attention calculation on each subspace, and obtain 8 sets of attention outputs;
[0117] The 8 groups of attention outputs are concatenated and dimensionally transformed through a linear mapping layer to obtain the multi-head attention result;
[0118] The multi-head attention results are input into a two-layer feedforward neural network, feature processing is performed through residual connection and layer normalization, and delay compensation features are output;
[0119] The perspective features are contrastively learned based on the delay compensation features to obtain the perspective mapping parameters.
[0120] Specifically, the spatiotemporal dependency features are linearly mapped to generate query matrix, key matrix and value matrix. Assume that the spatiotemporal dependency features are represented as matrix ,in is the sequence length of the input feature, If is the feature dimension, the query matrix is generated through three independent linear transformations , key matrix Sum Matrix :
[0121] ;
[0122] in, , , are the linear transformation matrices for query, key, and value, respectively, The projection dimension for attention calculation. Query matrix For use with key matrix To calculate the similarity, the value matrix It is used for the final weighted sum of information. After generating the query matrix and the key matrix, the similarity between them is calculated to obtain the attention score. The similarity calculation is implemented by matrix multiplication:
[0123] ;
[0124] in, represents the unnormalized attention score matrix, As a scaling factor, it is used to prevent the calculation result from being too large, causing gradient vanishing or gradient exploding. This scaling operation ensures that the numerical range of the attention score is not too large, thereby improving training stability. After calculating the attention score, it is normalized by softmax to convert the score into an attention weight:
[0125] ;
[0126] in, Represents the normalized attention weight matrix, each row of which represents the attention distribution of the query position to all key positions. Perform matrix multiplication to compute the output of a single-head attention:
[0127] ;
[0128] in, Represents the calculation result of single-head attention, which extracts the most relevant information through the attention mechanism and assigns different weights to features at different time steps. The spatiotemporal dependent features are projected into multiple different subspaces, and attention calculations are performed on each subspace separately to obtain richer feature expressions. Assuming that 8 different attention heads are used, the calculation method of each attention head is the same as that of single-head attention, except that different linear transformation matrices are used:
[0129] ;
[0130] ;
[0131] ;
[0132] ;
[0133] in, Representative A head of attention, For the dimension of each subspace, it usually satisfies In this way, each attention head learns different information representations and extracts different types of spatiotemporal features respectively. After obtaining 8 sets of attention outputs, they are concatenated to merge the information from different attention heads:
[0134] ;
[0135] This concatenation operation connects the outputs of all attention heads together to form a complete attention representation. The linear mapping layer performs dimension transformation so that the final multi-head attention result is still in the same space as the input feature:
[0136] ;
[0137] in, is a mapping matrix, which transforms the multi-head attention output back to the original feature space to keep the features consistent. After completing the multi-head attention calculation, the multi-head attention result is input into the two-layer feedforward neural network to improve the expression ability of the model. The calculation method of the feedforward network is as follows:
[0138] ;
[0139] in, is the weight matrix of the feedforward neural network, is the bias term, is a nonlinear activation function (such as ReLU). In order to ensure that the gradient can be effectively propagated and maintain the stability of the features, residual connections and layer normalization are introduced:
[0140] ;
[0141] Among them, LayerNorm represents the layer normalization operation, which can ensure that the features of different batches maintain a stable distribution, thereby improving the generalization ability of the model. After the above calculation, the delay compensation feature is obtained. Based on the delay compensation feature, the contrastive learning of the view feature is performed to obtain the view mapping parameter. Contrastive learning trains a stable mapping model by maximizing the similarity of positive sample pairs and minimizing the similarity of negative sample pairs. Assume that the view feature is and Represent the features of the current view and the target view respectively, and calculate the similarity between the two:
[0142] ;
[0143] in, Represents cosine similarity. The closer the value is to 1, the closer the mapping relationship between the two perspectives is. In order to optimize the perspective mapping parameters, a contrast loss function is constructed:
[0144] ;
[0145] The denominator represents all possible pairs of perspectives, and the numerator represents the similarity between the current perspective and the target perspective. By maximizing the similarity of the target perspective while minimizing the influence of other perspectives, a more robust mapping parameter is trained.
[0146] In a specific embodiment, the execution step performs comparative learning on the view feature based on the delay compensation feature to obtain the view mapping parameter, which may specifically include the following steps:
[0147] Performing low-dimensional space mapping on the delay compensation feature to obtain a low-dimensional feature vector;
[0148] Based on the low-dimensional feature vector, a feature memory queue is constructed, and time-continuous adjacent view feature pairs are set as positive sample pairs, and non-adjacent view feature pairs are set as negative sample pairs to obtain a training sample set;
[0149] Calculate the cosine similarity of the feature pairs in the training sample set to obtain the feature similarity score, and construct the initial feature mapping function based on the feature similarity score;
[0150] A dynamic weight mechanism is introduced based on the initial feature mapping function. Weight coefficients are set for different sample pairs according to the difficulty of perspective conversion to obtain a weighted feature mapping function. Based on the weighted feature mapping function, a gradient descent perspective mapping parameter optimization calculation is performed to output the perspective mapping parameters.
[0151] Specifically, the delay compensation feature is mapped into a low-dimensional space to obtain a more compact feature representation and reduce the computational complexity. Assume that the delay compensation feature is represented as ,in is the sample size, is the original feature dimension, then a linear transformation is used to project it into a low-dimensional feature space:
[0152] ;
[0153] in, is the dimension reduction mapping matrix, is the bias term, is the feature dimension after dimensionality reduction, satisfying To reduce computational overhead. Through this mapping, a low-dimensional feature vector is obtained , this feature vector can reduce data redundancy and improve computational efficiency while retaining key information. A feature memory queue is constructed based on the low-dimensional feature vector to store the view features of different time steps and construct positive and negative sample pairs. Feature memory queue Set to a size of A FIFO (first-in-first-out) buffer, which stores the past For the construction of the view feature pair, let the feature of the current time step be ,time The characteristics of the moment are , then the adjacent view feature pairs that are continuous in time are defined as positive sample pairs:
[0154] ;
[0155] For non-adjacent view feature pairs with a large time interval, they are taken as negative sample pairs:
[0156] ;
[0157] in, Indicates the time interval, Take a larger value to ensure that the features of the negative sample pairs are significantly different. By constructing positive and negative sample pairs, ensure that the model can learn the smooth changes between temporally continuous perspectives during training, while distinguishing the conversion mode of long-distance perspectives. After constructing the training sample set, perform similarity calculations on the feature pairs to obtain feature similarity scores, and build the initial feature mapping function based on this. A common similarity measure is cosine similarity, which is calculated as follows:
[0158] ;
[0159] in, Representation characteristics and The similarity between them, represents the dot product operation, Represents a vector The similarity metric range is between [-1,1]. The closer the value is to 1, the more similar the features of the two perspectives are, and the closer the value is to -1, the farther the perspective relationship between the two is. Based on the feature similarity score, an initial feature mapping function is constructed. The function of this function is to learn the mapping relationship between perspectives and predict the target perspective feature corresponding to the given perspective feature. The initial feature mapping function is defined as follows:
[0160] ;
[0161] in, represents the learned mapping function, According to The predicted viewing angle features should ideally be as close to the real features as possible During the training process, the mapping function optimizes the loss function to make the predicted features as consistent as possible with the real features. In practical applications, the difficulty of different perspective conversions is uneven. For example, the conversion from a close perspective is usually simpler, while the process of switching from a long-distance perspective to a target perspective is more complicated. Therefore, a dynamic weight mechanism is introduced to assign different weight coefficients to different sample pairs according to the difficulty of perspective conversion. Let the similarity score of the feature pair be , then define the weight coefficient for:
[0162] ;
[0163] The meaning of this weight coefficient is that the lower the similarity of the sample pair, the greater the difficulty of conversion, and needs to be given a higher weight during the training process, while the sample pair with higher similarity has a lower conversion difficulty, so the weight is lower. Based on this weight, construct a weighted feature mapping function:
[0164] ;
[0165] in, represents the training loss, represents all training sample pairs, Represents the Euclidean distance loss. The optimization goal of this loss function is to make the prediction error of high-difficulty conversion as small as possible, while ensuring that low-difficulty conversion does not affect the overall training process. After completing the construction of the weighted feature mapping function, perform gradient descent optimization to calculate the final perspective mapping parameters. Assume that the perspective mapping parameters are , then the gradient descent method is used to update the parameters:
[0166] ;
[0167] in, is the learning rate, is the gradient of the loss function to the parameter. Through continuous iterative optimization, the perspective mapping function Learn more accurate mapping relationships to improve the accuracy of perspective conversion.
[0168] In a specific embodiment, the process of executing step S5 may specifically include the following steps:
[0169] Construct the encoder part of the U-Net network, perform multi-scale feature extraction on the original image sequence, and perform four downsampling processes in sequence through the downsampling convolution module to obtain a multi-level feature map;
[0170] A spatial adaptive normalization layer is constructed based on the delay compensation feature, and the feature distribution of the multi-level feature map is adjusted at different scales to obtain a normalized feature map;
[0171] The view mapping parameters are input into the feature modulation module, and the channel weights of the normalized feature map are dynamically adjusted to obtain a view-sensitive feature map;
[0172] Construct the decoder part of the U-Net network, perform four feature amplifications on the view-sensitive feature map through the upsampling transposed convolution module, and introduce a jump connection in each amplification process to concatenate and fuse the encoding features and decoding features of the corresponding levels to obtain the reconstructed feature map;
[0173] The reconstructed feature map is subjected to brightness equalization and color correction to obtain an initial 3D image sequence, and the initial 3D image sequence is subjected to spatial edge enhancement processing to output a 3D image sequence with continuous viewing angles.
[0174] Specifically, the encoder part of the U-Net network is constructed to extract multi-scale features of the input original image sequence and extract deep features by gradually downsampling. Assume that the input original image sequence is ,in and Represents the height and width of the image respectively, is the number of channels, the encoder part extracts preliminary features through a series of convolutional layers:
[0175] ;
[0176] in, is the convolution kernel, represents the convolution operation, is the bias term, ReLU represents the ReLU activation function, which is used to introduce nonlinear characteristics. In order to obtain multi-scale features, a downsampling convolution module is constructed and four downsampling processes are performed in sequence, so that the network can gradually extract global features. The downsampling of a layer is calculated as follows:
[0177] ;
[0178] Among them, Down It represents the downsampling operation, using the maximum pooling or convolution operation with a step size of 2 to reduce the spatial resolution of the feature map while increasing the number of channels so that the network can extract a more abstract representation. After four downsamplings, the resulting multi-level feature map is:
[0179] ;
[0180] This feature map has the smallest spatial resolution, but contains global deep features for subsequent image reconstruction. A spatial adaptive normalization layer is constructed based on the delay compensation feature to normalize the multi-level feature map so that it maintains a stable feature distribution when different perspectives are transformed. Assume that the delay compensation feature is , each feature channel is normalized by adaptive normalization operation:
[0181] ;
[0182] in, and Represents the feature maps The mean and standard deviation of and The scaling and offset parameters controlled by the delay compensation feature make the feature map more adaptable at different time steps. Through normalization, the multi-scale feature map can better adapt to the dynamic perspective transformation and improve the stability of the generated image. After normalization, the feature is modulated to enhance the influence of the perspective mapping parameters on the feature. Construct a feature modulation module and use the perspective mapping parameters As input, the channel weights of the normalized feature map are dynamically adjusted. The feature modulation is calculated as follows:
[0183] ;
[0184] in, is a weight matrix controlled by the perspective mapping parameters. The role of this matrix is to adjust the expression of features according to different perspectives, so that the final perspective-sensitive feature map can more accurately match the target image of different perspectives. After feature modulation, the network adaptively adjusts the distribution of the feature map to ensure that the reconstructed image can conform to the physical laws of perspective conversion. Construct the decoder part of the U-Net network to gradually restore the spatial resolution of the image through upsampling operations and ensure that the reconstructed 3D image has high-quality details. During the decoding process, each upsampling operation uses a transposed convolution for feature amplification, and introduces a jump connection after each upsampling to combine the multi-scale features in the encoder:
[0185] ;
[0186] in, is the transposed convolution kernel, represents the upsampling operation, and Represents the reconstructed features in the decoder. After each upsampling, it is concatenated and fused with the features of the corresponding level of the encoder:
[0187] ;
[0188] Through the jump connection method, high-resolution detail information is directly passed to the decoder, thereby improving the clarity of the final generated image. After completing the feature amplification, the reconstructed feature map is brightness balanced and color corrected to ensure that the generated 3D image has a good visual effect. The calculation method of brightness balance is as follows:
[0189]
[0190] in, is the initially generated 3D image, and This is an image after brightness balance. This operation ensures that the overall brightness of the image is evenly distributed to avoid overexposure or darkening. Perform color correction to ensure that the colors at different viewing angles are consistent. The correction calculation is as follows:
[0191] ;
[0192] in, and is a color correction parameter used to adjust color uniformity and contrast. In order to improve the sharpness of the generated 3D image, it is subjected to spatial edge enhancement processing to ensure that the edges of objects are clearer. The calculation formula is as follows:
[0193] ;
[0194] in, represents the image edge information calculated by the Laplacian operator, and To enhance the intensity coefficient, this operation enhances the local contrast of the image, making the details more obvious, and finally outputs a 3D image sequence with continuous viewing angles.
[0195] The above describes the method for generating a multi-view 3D image on a curved screen in an embodiment of the present invention. The following describes the device for generating a multi-view 3D image on a curved screen in an embodiment of the present invention. Figure 2 In one embodiment of the present invention, a device for generating a multi-view 3D image on a curved screen includes:
[0196] An acquisition module is used to acquire a set of viewing angle parameters of the curved screen and perform residual feature processing on the original image sequence to obtain a multi-dimensional spatial feature map;
[0197] A tensor product operation module is used to perform a tensor product operation on a multi-dimensional spatial feature map and a view parameter set to obtain view conversion delay data;
[0198] The graph convolution operation module is used to set the multi-dimensional spatial feature map as the node attribute of the graph neural network and the perspective conversion delay data as the edge attribute to perform graph convolution operation to obtain spatiotemporal dependency features;
[0199] A perspective mapping module is used to perform multi-head self-attention calculation on the spatiotemporal dependent features to obtain delay compensation features, and to perform comparative learning on the perspective features based on the delay compensation features to obtain perspective mapping parameters;
[0200] The image reconstruction module is used to input the perspective mapping parameters, delay compensation features and original image sequence into the U-Net network for 3D image reconstruction and output a 3D image sequence with continuous perspective.
[0201] Through the collaboration of the above components, by constructing a surface adaptive convolution layer, the size of the receptive field is dynamically adjusted for different surface positions, effectively solving the problem of uneven image feature extraction caused by the curvature characteristics of the curved screen and improving the accuracy of feature expression. The delay of multi-perspective conversion is modeled, and the dual-path feature extraction structure of the branch network and the trunk network is combined to achieve an accurate description of the nonlinear perspective conversion characteristics, effectively reducing the delay phenomenon during perspective switching. By introducing the graph neural network structure, spatial features and delay data are modeled as node and edge attributes respectively, enhancing the ability to characterize the spatiotemporal dependency relationship between perspectives and ensuring the continuity of perspective conversion. Based on the multi-head self-attention mechanism and contrast learning strategy, accurate alignment and mapping of cross-perspective features are achieved, which improves the feature consistency during perspective conversion and reduces image distortion during perspective switching. Through the multi-scale feature fusion and spatial adaptive normalization processing of the U-Net network, smooth perspective conversion is achieved while maintaining image details, significantly improving the visual quality of 3D images.
[0202] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, systems and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0203] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or the whole or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store program code.
[0204] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for generating multi-view 3D images on a curved screen, characterized in that: The method comprises: Collect the viewing angle parameter set of the curved screen, and perform residual feature processing on the original image sequence to obtain a multi-dimensional spatial feature map; Performing a tensor product operation on the multidimensional spatial feature map and the viewing angle parameter set to obtain viewing angle conversion delay data; The multidimensional space feature graph is set as a node attribute of a graph neural network, and the perspective conversion delay data is set as an edge attribute to perform a graph convolution operation to obtain a spatiotemporal dependency feature; Perform multi-head self-attention calculation on the spatiotemporal dependent features to obtain delay compensation features, and perform comparative learning on the perspective features based on the delay compensation features to obtain perspective mapping parameters; specifically including: linearly mapping the spatiotemporal dependent features to obtain a query matrix, a key matrix and a value matrix; performing matrix multiplication operation on the query matrix and the key matrix and performing scaling, and normalizing by dividing by the square root of the dimension to obtain an attention score; performing a softmax operation on the attention score, and performing a matrix multiplication operation with the value matrix to obtain a single-head attention output; projecting the spatiotemporal dependent features into 8 different subspaces, performing attention calculation on each subspace respectively, and obtaining 8 groups of attention outputs; splicing the 8 groups of attention outputs, and performing dimensionality conversion through a linear mapping layer to obtain a multi-head attention result; The head attention result is input into a two-layer feedforward neural network, and feature processing is performed through residual connection and layer normalization to output delay compensation features; the delay compensation features are mapped into a low-dimensional space to obtain a low-dimensional feature vector; a feature memory queue is constructed based on the low-dimensional feature vector, and time-continuous adjacent view feature pairs are set as positive sample pairs, and non-adjacent view feature pairs are set as negative sample pairs to obtain a training sample set; cosine similarity is calculated for the feature pairs in the training sample set to obtain feature similarity scores, and an initial feature mapping function is constructed based on the feature similarity scores; a dynamic weight mechanism is introduced based on the initial feature mapping function, and weight coefficients are set for different sample pairs according to the difficulty of view conversion to obtain a weighted feature mapping function, and based on the weighted feature mapping function, a gradient descent view mapping parameter optimization calculation is performed to output the view mapping parameters; The perspective mapping parameters, the delay compensation features and the original image sequence are input into a U-Net network for 3D image reconstruction, and a 3D image sequence with a continuous perspective is output.
2. The method for generating multi-view 3D images on a curved screen according to claim 1, characterized in that: The method collects a viewing angle parameter set of the curved screen and performs residual feature processing on the original image sequence to obtain a multi-dimensional space feature map, including: Measure the parameters of the curved screen to obtain the curvature radius R, screen width W, and viewing angle range θ, establish a three-dimensional coordinate system with the horizontal direction as the x-axis, the vertical direction as the y-axis, and the depth of field direction as the z-axis, and obtain a viewing angle parameter set; Construct a five-layer residual network structure, each of which includes two 3×3 convolutional layers and a ReLU activation function, input the original image sequence into the five-layer residual network structure for basic feature processing, and obtain a basic feature map; Constructing a surface adaptive convolution layer according to the curvature radius R, setting corresponding receptive field sizes for different surface positions, inputting the basic feature map into the surface adaptive convolution layer, and obtaining a surface feature map; Performing global average pooling and maximum pooling operations on the surface feature map to obtain a spatial attention weight, and performing a matrix multiplication operation on the surface feature map and the spatial attention weight to obtain a weighted feature map; Batch normalization is performed on the weighted feature map to output a multidimensional spatial feature map.
3. The method for generating multi-view 3D images on a curved screen according to claim 1, characterized in that: The performing a tensor product operation on the multidimensional spatial feature map and the viewing angle parameter set to obtain viewing angle conversion delay data includes: Constructing a branch network and a trunk network, wherein the branch network includes four fully connected layers, and the trunk network includes three fully connected layers; Inputting the multidimensional space feature map into the branch network to perform feature dimension reduction mapping to obtain a feature mapping vector; Timestamping the current view and the target view in the view parameter set to obtain marked view parameters, and inputting the marked view parameters into the backbone network to extract timing characteristics to obtain a timing mapping vector; Performing a tensor product operation on the feature mapping vector and the time series mapping vector to obtain a perspective conversion feature vector; The perspective conversion feature vector is optimized and trained by using a mean square error loss function, and an L2 regularization term is introduced for smoothing constraints to obtain an optimized feature vector; The optimized feature vector is input into a nonlinear mapping layer to perform view conversion delay calculation, and view conversion delay data is output.
4. The method for generating multi-view 3D images on a curved screen according to claim 1, characterized in that: The step of setting the multidimensional spatial feature graph as a node attribute of a graph neural network and setting the perspective conversion delay data as an edge attribute to perform a graph convolution operation to obtain a spatiotemporal dependency feature includes: The multidimensional space feature graph is set as a node feature, the view conversion delay data is set as an edge weight, and an adjacency matrix is constructed; Performing degree matrix calculation on the adjacency matrix to obtain an angle matrix, and performing normalization operation on the angle matrix to obtain a normalized matrix; Performing a matrix multiplication operation on the normalized matrix, the adjacency matrix and the node features to obtain a first-order feature propagation result; The first-order feature propagation result is linearly transformed, and nonlinearly mapped through a ReLU activation function to obtain a single-layer graph convolution output; By stacking three layers of graph convolutional layers, multi-layer feature propagation is performed, and the features of each layer are fused through jump connections to obtain a multi-layer feature fusion result; Based on the multi-layer feature fusion results, a global attention layer is added to calculate the attention weights between different nodes, and the spatiotemporal dependency features are output.
5. The method for generating multi-view 3D images on a curved screen according to claim 1, characterized in that: The step of inputting the perspective mapping parameter, the delay compensation feature and the original image sequence into a U-Net network for 3D image reconstruction and outputting a 3D image sequence with a continuous perspective includes: Constructing the encoder part of the U-Net network, performing multi-scale feature extraction on the original image sequence, and performing four downsampling processes in sequence through the downsampling convolution module to obtain a multi-level feature map; Constructing a spatial adaptive normalization layer based on the delay compensation feature, adjusting the feature distribution of the multi-level feature map at different scales to obtain a normalized feature map; Inputting the viewing angle mapping parameters into a feature modulation module, dynamically adjusting the channel weights of the normalized feature map, and obtaining a viewing angle sensitive feature map; Constructing the decoder part of the U-Net network, amplifying the view-sensitive feature map four times through the upsampling transposed convolution module, and introducing a skip connection in each amplification process, splicing and fusing the encoding features and decoding features of the corresponding levels to obtain a reconstructed feature map; The reconstructed feature map is subjected to brightness equalization and color correction to obtain an initial 3D image sequence, and the initial 3D image sequence is subjected to spatial edge enhancement processing to output a 3D image sequence with a continuous viewing angle.
6. A curved screen multi-view 3D image generation device, characterized in that: The device is used to execute the method for generating a curved screen multi-viewing angle 3D image according to any one of claims 1 to 5, the device comprising: An acquisition module is used to acquire a set of viewing angle parameters of the curved screen and perform residual feature processing on the original image sequence to obtain a multi-dimensional spatial feature map; A tensor product operation module, used for performing a tensor product operation on the multidimensional spatial feature map and the viewing angle parameter set to obtain viewing angle conversion delay data; A graph convolution operation module, used to set the multidimensional space feature map as a node attribute of a graph neural network and the perspective conversion delay data as an edge attribute to perform a graph convolution operation to obtain a spatiotemporal dependency feature; A perspective mapping module, configured to perform multi-head self-attention calculation on the spatiotemporal dependent features to obtain delay compensation features, and perform comparative learning on perspective features based on the delay compensation features to obtain perspective mapping parameters; The image reconstruction module is used to input the perspective mapping parameters, the delay compensation features and the original image sequence into a U-Net network to perform 3D image reconstruction and output a 3D image sequence with a continuous perspective.
Citation Information
Patent Citations
Indoor dynamic environment map construction method and system based on image restoration and completion
CN118822906A
Visual enhancement method and system of LED creative spherical display screen
CN119417718A