A naked eye 3D interaction method and system based on OSG
By using OSG and deep learning models in naked-eye 3D display technology, adjusting the camera and model poses in real time and optimizing parallax images, the problems of limited viewing angle, low image quality and poor three-dimensionality in the existing technology are solved, and high-quality naked-eye 3D display effect is achieved.
Patent Information
- Application Number
- CN202510109346.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-23
AI Technical Summary
The existing naked-eye 3D display technology has problems such as limited viewing angle, low image quality, and poor three-dimensional sense, and it is impossible to dynamically adjust the scene according to the user's interactive behavior.
By loading the OSG library and deep learning model, user interaction data is collected in real time, the obtained viewing angle parameters and model parameters are analyzed, and the camera position and model pose in the OSG scene are adjusted. Use computer vision technology to generate parallax images, and optimize parallax images through deep learning algorithms to improve the reality and three-dimensionality of 3D effects.
Real-time and intelligent adjustments are achieved based on the actual interactive behavior of users, improving the realism and three-dimensionality of the 3D effect, and avoiding the problems of limited viewing angles and low image quality.
Smart Images

Figure CN119559076B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of naked-eye 3D technology, and in particular to an OSG-based naked-eye 3D interaction method and system. Background Art
[0002] With the rapid development of computer graphics technology and display technology, 3D display technology has been widely used in entertainment, education, medical treatment, industrial design and other fields. Traditional 3D display technology usually requires the use of specific auxiliary equipment (such as 3D glasses) to achieve stereoscopic visual effects, which not only limits the user's viewing experience, but also increases additional costs and usage complexity. Therefore, naked-eye 3D technology has gradually become a hot topic of research. It uses special optical design or image processing technology to enable users to directly observe the 3D effect without wearing any auxiliary equipment.
[0003] In the implementation of naked-eye 3D technology, the construction and rendering of three-dimensional models are key links. As a high-performance open source scene graph management development library, OpenSceneGraph (OSG) is widely used in the field of three-dimensional visualization. It provides rich functions to load, render and interact with three-dimensional models, and is one of the important tools for realizing naked-eye 3D display. However, relying solely on OSG for basic 3D rendering is not enough to achieve high-quality naked-eye 3D effects. Currently, it is impossible to dynamically adjust the scene according to the user's interactive behavior, and the existing naked-eye 3D display technology often has problems such as limited viewing angle, low image quality, and weak stereoscopic effect. Therefore, it is necessary to provide a naked-eye 3D interaction method and system based on OSG to solve the above problems. Summary of the invention
[0004] In view of the deficiencies in the prior art, the object of the present invention is to provide an OSG-based naked-eye 3D interaction method and system to solve the problems existing in the above-mentioned background technology.
[0005] The present invention is implemented as follows: a naked eye 3D interaction method based on OSG, the method comprising the following steps:
[0006] Load the OSG library and deep learning model, configure naked-eye 3D display parameters, use OSG to load 3D models, textures and light sources, and obtain the OSG scene;
[0007] Collect user interaction data in real time, input the interaction data into a deep learning model, and perform real-time analysis;
[0008] Output the view parameters and model parameters obtained by the analysis, and adjust the camera position and model posture in the OSG scene according to the view parameters and model parameters;
[0009] Based on the optimized OSG scene, computer vision technology is used to generate parallax images, and deep learning algorithms are used to optimize the parallax images to improve the realism and stereoscopic effect of the 3D effect.
[0010] The optimized parallax images are synthesized into naked-eye 3D images and output to a display device.
[0011] Another object of the present invention is to provide a naked eye 3D interactive system based on OSG, the system comprising:
[0012] Initialize the configuration module, which is used to load the OSG library and deep learning model, configure the naked eye 3D display parameters, use OSG to load the 3D model, texture and light source, and obtain the OSG scene;
[0013] An interactive data processing module is used to collect user interactive data in real time, input the interactive data into a deep learning model, and perform real-time analysis;
[0014] A position and posture adjustment module, used to output the perspective parameters and model parameters obtained by analysis, and adjust the camera position and model posture in the OSG scene according to the perspective parameters and model parameters;
[0015] The parallax image optimization module is used to generate parallax images based on the optimized OSG scene using computer vision technology and optimize the parallax images using deep learning algorithms to improve the realism and stereoscopic effect of the 3D effect;
[0016] The naked-eye 3D image module is used to synthesize the optimized parallax images into a naked-eye 3D image and output it to a display device.
[0017] Compared with the prior art, the present invention has the following beneficial effects:
[0018] The present invention inputs the interactive data into the deep learning model, analyzes the obtained viewing angle parameters and model parameters, and adjusts the camera position and model posture in the OSG scene according to the viewing angle parameters and model parameters. In this way, the present invention can make real-time and intelligent adjustments according to the actual interactive behavior of the user. It also generates parallax images based on the optimized OSG scene using computer vision technology, and optimizes the parallax images using deep learning algorithms to improve the realism and stereoscopic sense of the 3D effect, so that the viewing angle is not limited and the image quality is high. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 The figure is a flow chart of a naked-eye 3D interaction method based on OSG.
[0020] Figure 2 The present invention is a flowchart for collecting user interaction data in an OSG-based naked-eye 3D interaction method.
[0021] Figure 3 The flowchart is for adjusting the camera position and model posture in an OSG-based naked-eye 3D interaction method.
[0022] Figure 4 The present invention is a flowchart for generating parallax images in a naked-eye 3D interaction method based on OSG.
[0023] Figure 5 The present invention is a flowchart of synthesizing a naked-eye 3D image in a naked-eye 3D interaction method based on OSG.
[0024] Figure 6 This is a structural diagram of an OSG-based naked-eye 3D interactive system. DETAILED DESCRIPTION
[0025] In order to make the purpose, technical solution and advantages of the present invention clearer, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0026] The specific implementation of the present invention is described in detail below in conjunction with specific embodiments.
[0027] like Figure 1 As shown, an embodiment of the present invention provides a naked eye 3D interaction method based on OSG, and the method comprises the following steps:
[0028] S100, load the OSG library and deep learning model, configure naked eye 3D display parameters, use OSG to load 3D models, textures and light sources, and obtain the OSG scene;
[0029] S200, collecting user interaction data in real time, inputting the interaction data into a deep learning model, and performing real-time analysis;
[0030] S300, outputting the perspective parameters and model parameters obtained by analysis, and adjusting the camera position and model posture in the OSG scene according to the perspective parameters and model parameters;
[0031] S400, based on the optimized OSG scene, uses computer vision technology to generate parallax images, and uses deep learning algorithms to optimize the parallax images to improve the realism and stereoscopic effect of the 3D effect;
[0032] S500: synthesize the optimized parallax images into a naked-eye 3D image, and output it to a display device.
[0033] It should be noted that OpenSceneGraph (OSG), as a high-performance open source scene graph management and development library, is widely used in the field of 3D visualization. It provides rich functions for loading, rendering and interacting with 3D models, and is one of the important tools for achieving naked-eye 3D display. However, relying solely on OSG for basic 3D rendering is not enough to achieve high-quality naked-eye 3D effects. Currently, the scene cannot be dynamically adjusted according to the user's interactive behavior, and the existing naked-eye 3D display technology often has problems such as limited viewing angle, low image quality, and weak stereoscopic effect. The embodiments of the present invention are intended to solve the above problems.
[0034] In the embodiment of the present invention, it is first necessary to install and configure OpenSceneGraph (OSG) and its dependent libraries, install a deep learning framework (such as TensorFlow or PyTorch) and a computer vision library (such as OpenCV). And equipped with a display or projection device that supports naked-eye 3D display, configure high-performance computing resources, including GPU acceleration. Then you can load the OSG library and deep learning model, the deep learning model is trained and constructed in advance, and the naked-eye 3D display parameters such as parallax, focal length, etc. are configured. Use OSG to load 3D models, maps and light sources to obtain an initialized OSG scene graph. Then the user's interaction data will be collected in real time, and the interaction data will be input into the deep learning model for real-time analysis. The deep learning model will output the perspective parameters and model parameters obtained by the analysis, and then adjust the camera position and model posture in the OSG scene according to the perspective parameters and model parameters. In this way, the present invention can make real-time and intelligent adjustments according to the actual interaction behavior of the user. Then, based on the optimized OSG scene, computer vision technology will be used to generate parallax images, and deep learning algorithms will be used to optimize the parallax images to improve the realism and stereoscopic effect of the 3D effect. The optimized parallax images will be synthesized into naked-eye 3D images and output to the display device. The naked-eye 3D synthesis parameters can also be automatically adjusted according to the characteristics of the display device and environmental changes to ensure the best 3D display effect.
[0035] like Figure 2 As shown, as a preferred embodiment of the present invention, the step of collecting user interaction data in real time and inputting the interaction data into the deep learning model specifically includes:
[0036] S201, collecting user interaction data in real time through a camera and a sensor, where the interaction data includes head posture and gesture position;
[0037] S202, preprocessing the interaction data, where the preprocessing includes denoising and normalization;
[0038] S203, inputting the preprocessed interaction data into a deep learning model, wherein the deep learning model is a convolutional neural network or a recurrent neural network.
[0039] In an embodiment of the present invention, the user's interaction data will be collected in real time through cameras and sensors. The interaction data includes head posture and gesture position. The head posture includes pitch angle, yaw angle and roll angle. The interaction data will then be preprocessed. The preprocessing includes denoising and normalization. Finally, the preprocessed interaction data will be input into a deep learning model. The deep learning model is a convolutional neural network (CNN) and a recurrent neural network (RNN). CNN and RNN need to be trained and constructed in advance.
[0040] As a preferred embodiment of the present invention, the interaction data is preprocessed, and the preprocessing includes the steps of denoising and normalization, specifically including:
[0041] Apply wavelet transform to the interaction data to obtain low-frequency subband and high-frequency subband;
[0042] Calculate the soft threshold according to the standard deviation of the high frequency sub-band, calculate the absolute value of the high frequency sub-band, and compare the absolute value with the soft threshold, and filter the processed high frequency sub-band according to the comparison result;
[0043] The low-frequency sub-band and the processed high-frequency sub-band are combined to obtain the denoised interaction data;
[0044] Given a reference trajectory, calculate the Euclidean distance matrix between the denoised interaction data and the reference trajectory;
[0045] According to the Euclidean distance matrix, the DTW algorithm is used to calculate and determine the optimal alignment path between the denoised interaction data and the reference trajectory;
[0046] According to the optimal alignment path, the user data is interpolated to the same time step as the reference trajectory to obtain the aligned interaction data;
[0047] The median of the aligned interaction data was calculated, and the interquartile range was calculated based on the median;
[0048] According to the interquartile range, outliers are removed from the aligned interaction data to obtain data after outliers are removed;
[0049] The data after removing outliers is mapped to a range through dynamic mapping to perform dynamic normalization to obtain normalized data.
[0050] In an embodiment of the present invention, a wavelet transform is used to simultaneously analyze the local and overall characteristics of the data. Compared with traditional low-pass filters or mean smoothing, wavelet denoising does not simply "over-smooth" the data, but retains important low-frequency information while removing high-frequency noise. It can effectively process complex interactive data, especially time series containing noise. And robust normalization is achieved based on the median and interquartile range, which is more advantageous when processing outliers and is more stable than traditional mean normalization and maximum and minimum normalization. DTW can perform nonlinear alignment of the user's motion data in time series, so that the differences in the motion trajectories of different users in the time dimension are eliminated. It is particularly suitable for scenarios that require comparative analysis with reference trajectories, such as standard motion evaluation, posture calibration, etc.
[0051] like Figure 3 As shown, as a preferred embodiment of the present invention, the step of adjusting the camera position and model posture in the OSG scene according to the viewing angle parameters and model parameters specifically includes:
[0052] S301, adjusting the position of the camera according to the translation vector in the viewing angle parameter, and updating the posture of the camera using the rotation matrix in the viewing angle parameter;
[0053] S302, traversing all model nodes in the OSG scene to determine the model that needs to be adjusted;
[0054] S303, adjusting the position of the model according to the translation vector in the model parameters, and updating the posture of the model according to the scaling ratio and rotation angle in the model parameters.
[0055] In an embodiment of the present invention, the viewing angle parameters include a rotation matrix (describing the yaw angle, pitch angle, and roll angle of the camera) and a translation vector (describing the position movement of the camera in three-dimensional space), and the model parameters include a scaling factor, a rotation angle, and a translation vector. When adjusting the camera position and model posture, the camera position is first adjusted according to the translation vector in the viewing angle parameters, which usually involves modifying the camera's osg::Vec3 position attribute. If the camera has a parent node (such as a view matrix transformation node), the parent node's transformation may need to be updated to reflect the camera's new position. The camera's posture is updated using a rotation matrix by modifying the camera's view matrix. In OSG, the osg::Matrix class can be used to represent and manipulate transformation matrices. The rotation matrix is applied to the current camera's view matrix to update the camera's direction. In addition, if the camera uses a custom projection matrix or view matrix stack, these matrices may need to be updated accordingly. Then traverse all model nodes in the OSG scene, determine the model that needs to be adjusted, and adjust the model's position according to the translation vector in the model parameters. This is achieved by modifying the translation property of the model's osg::Transform node. If the model does not have a direct transformation node, you need to create a new osg::Transform node to wrap the model and apply the translation transformation. Finally, update the model's posture according to the scale and rotation angle in the model parameters. This is achieved by modifying the model's transformation matrix. In OSG, use osg::MatrixTransform to apply the transformation matrix, which includes rotation and scaling transformations to reflect the model's new posture.
[0056] As a preferred embodiment of the present invention, the step of adjusting the position of the model according to the translation vector in the model parameters and updating the posture of the model according to the scaling ratio and rotation angle in the model parameters specifically includes the following sub-steps:
[0057] Adjust the position, scaling and rotation angle of the model according to the model parameters to obtain an adjusted model;
[0058] The error between each adjusted model and model parameters in terms of model position, scaling, and rotation angle is used as the optimization target;
[0059] Define the collision threshold, traverse all models in the scene, calculate the distance between every two models, and compare the obtained distance with the collision threshold to construct the collision penalty term;
[0060] Calculate the angle between the surface normal vector of the model and the direction of the light source, evaluate the degree of light occlusion by the model, and construct a light occlusion penalty term;
[0061] The optimization objective, collision penalty term and light occlusion penalty term are weighted respectively to form a multi-objective optimization function;
[0062] The gradient descent method is used to iteratively optimize the model parameters. In each iteration, the improvement direction of the model parameters is calculated according to the gradient direction of the multi-objective optimization function. After the optimization is completed, the optimized model parameters are obtained.
[0063] The final model transformation matrix is generated using the optimized model parameters, and the posture and position of the model are updated according to the final model transformation matrix.
[0064] In the embodiment of the present invention, the objective function is used to integrate multiple adjustment targets (such as model position, scale, direction) with constraint conditions (such as collision, light occlusion), and the optimal adjustment parameters are automatically calculated through the optimization algorithm, thereby reducing the complexity of manual adjustment. The collision penalty term effectively guides the optimization algorithm to adjust the position and size of the model to ensure that the adjusted model does not overlap with other models in the scene. The light occlusion penalty term dynamically adjusts the final position and posture of the model by calculating the angle between the light source direction and the model normal vector to avoid the important light source from being blocked, thereby optimizing the overall lighting effect of the scene.
[0065] like Figure 4 As shown, as a preferred embodiment of the present invention, the steps of generating a disparity image using computer vision technology according to the optimized OSG scene and optimizing the disparity image using a deep learning algorithm specifically include:
[0066] S401, extracting depth information from the optimized OSG scene by rendering a depth map of the scene, wherein the depth map records distance information of each point in the scene;
[0067] S402, using a stereo matching algorithm to generate left-eye and right-eye disparity images according to the depth information, where the disparity images are used to simulate the difference between left-eye and right-eye images seen by human eyes;
[0068] S403, inputting the generated disparity image into a generative adversarial network to optimize the quality of the disparity image and reduce noise and artifacts.
[0069] In an embodiment of the present invention, in order to generate a disparity image, depth information is first extracted from an optimized OSG scene by rendering a depth map of the scene, wherein the depth map records the distance information of each point in the scene; then, a stereo matching algorithm is used to generate left-eye and right-eye disparity images based on the depth information, wherein the disparity images are used to simulate the difference between the left-eye and right-eye images seen by the human eye; finally, the generated disparity images are input into a generative adversarial network to optimize the quality of the disparity images, reduce noise and artifacts, and improve the realism and stereoscopic sense of the 3D effect. The generative adversarial network (GAN) needs to be trained and constructed in advance.
[0070] As a preferred embodiment of the present invention, the step of inputting the generated disparity image into a generative adversarial network to optimize the quality of the disparity image and reduce noise and artifacts specifically includes:
[0071] Construct a three-branch generator with global branch, local branch and depth branch;
[0072] The disparity image is input into the global branch, and the global representation features of the disparity image are extracted through convolution operation;
[0073] The global representation features are fed into the Transformer module, and the global features are obtained by modeling long-distance deep dependencies.
[0074] Decode the global features to obtain a globally optimized image;
[0075] Input the disparity image into the local branch, and crop the disparity image into several local areas;
[0076] The local area is passed through a dense convolutional network to extract local detail features and obtain a local feature map;
[0077] Deconvolve the local feature maps and then splice them to obtain the local optimized image;
[0078] The disparity image is input into the depth branch and multi-scale convolution operation is performed to extract the deep feature tensor containing global and local geometric information;
[0079] Perform geometric consistency alignment on the deep feature tensor to obtain an aligned deep feature tensor;
[0080] The aligned deep feature tensor is subjected to several deconvolution operations in an iterative manner, and in each deconvolution operation, a skip connection is used to fuse it with the corresponding deep feature tensor as the input of the next deconvolution operation. After the iteration is completed, the output result is smoothed to obtain an optimized depth perception image.
[0081] Perform weighted fusion on the global optimized image, the local optimized image and the depth perception optimized image to obtain an optimized disparity image;
[0082] Inputting the quality of the optimized disparity image into the discriminator to obtain a score of the optimized disparity image;
[0083] The discriminator loss function is constructed by evaluating the discriminator's ability to distinguish between real images and generated images, and the generator loss function is constructed by evaluating the generator's ability to deceive the discriminator;
[0084] The discriminator loss function and the generator loss function are used to construct the adversarial loss function. The perceptual loss function is constructed based on the difference between the optimized disparity image and the high-level features of the real image and the image. The geometric consistency loss function is constructed based on the consistency difference of the optimized disparity image in the geometric space. The stereo smoothness loss function is constructed based on the depth change difference of the optimized disparity image.
[0085] The adversarial loss function, the perceptual loss function, the geometric consistency loss function and the stereo smoothness loss function are weightedly summed to form a multimodal loss function;
[0086] The discriminator and the generator are trained adversarially, and the parameters of the discriminator and the generator are iteratively optimized by minimizing the multimodal loss function. After the optimization is completed, the final disparity image is obtained.
[0087] In an embodiment of the present invention, a three-branch generator is used for multi-dimensional feature optimization, wherein the global branch can well capture the relationship between long-distance pixels in the image and the overall depth trend by introducing the Transformer module. It makes up for the shortcomings of traditional convolutional neural networks (CNNs) in modeling long-distance dependencies, ensuring that the optimized disparity image is more natural and coherent. The local branch focuses on detail repair tasks within a small range, and handles edge blur, noise or artifact problems through dense convolution operations. This refinement processing method can effectively improve the clarity and realism of the image. The depth branch specifically processes the depth maps of the left and right eyes to ensure the consistency of the disparity image in the geometric space. This geometric optimization can significantly improve the stereoscopic effect and make the naked eye 3D perception more realistic. The results of the last three branches are weighted fused, and weights can be dynamically assigned according to the importance of global, local and depth features. This not only retains the authenticity of the global structure, but also strengthens the local details, and the quality of the disparity image output in the end is more comprehensive and balanced.
[0088] The generator and the discriminator are trained adversarially. The generator continuously tries to generate parallax images that are closer to reality, while the discriminator improves its ability to judge the authenticity of the optimized images. This dynamic game mechanism enables the model's optimization ability to surpass the traditional single-target training method. In addition, geometric consistency and stereo smoothness constraints are introduced during the adversarial training process. The geometric consistency loss ensures that the parallax images are aligned in the geometric space by constraining the relationship between the depth maps of the left and right eyes. This constraint significantly improves the coordination of the left and right views and enhances the stereoscopic perception of naked-eye 3D. The stereo smoothness loss constrains the smoothness of depth changes, eliminates abrupt depth jumps or artifacts, and makes the optimized image more natural.
[0089] In addition, for better results, perceptual loss is introduced. Perceptual loss captures high-level semantic features of images through pre-trained feature extraction networks (such as VGG networks). This feature-level optimization method not only focuses on pixel-level similarity, but also ensures that the optimized image is consistent with the real image in content, achieving higher visual quality. Therefore, during the training process, the quality of disparity images can be improved from multiple levels and dimensions through the joint optimization of adversarial loss, perceptual loss, geometric consistency loss, and stereo smoothness loss. Compared with a single loss function, the comprehensive loss strategy can more comprehensively solve various problems that may exist in image optimization.
[0090] like Figure 5 As shown, as a preferred embodiment of the present invention, the step of synthesizing the optimized parallax images into a naked-eye 3D image specifically includes:
[0091] S501, performing parallax image alignment so that the left and right eye parallax images are aligned at the pixel level;
[0092] S502, synthesizing the left and right eye parallax images into a naked eye 3D image using a naked eye 3D synthesis algorithm;
[0093] S503, post-processing the synthesized naked-eye 3D image, where the post-processing includes color correction and brightness adjustment.
[0094] In the embodiment of the present invention, for subsequent synthesis, it is necessary to align the parallax images so that the left and right eye parallax images are aligned at the pixel level, and then the left and right eye parallax images are synthesized into a naked eye 3D image by a naked eye 3D synthesis algorithm. The naked eye 3D synthesis algorithm can use a cylindrical grating method, a slit grating method or a lens array method. Finally, the synthesized naked eye 3D image is post-processed, and the post-processing includes color correction and brightness adjustment. The processed naked eye 3D image is output to a device that supports naked eye 3D display, and the display parameters such as parallax size, focal length, etc. are adjusted according to the characteristics of the display device and user feedback to obtain the best 3D display effect.
[0095] As a preferred embodiment of the present invention, the step of synthesizing the left and right eye parallax images into a naked eye 3D image by a naked eye 3D synthesis algorithm specifically includes:
[0096] Calculate the difference between the pixel value of the left eye image and the corresponding pixel value of the right eye image for each pixel point;
[0097] Obtaining a depth value of the depth map, and using the depth value of the depth map to normalize the difference of corresponding pixel values to obtain a disparity map;
[0098] Input the depth value into a smooth nonlinear function to obtain the depth weight;
[0099] Compare the difference in the edge between the left and right eye disparity images, and obtain the edge weight according to the difference;
[0100] The sliding window method is used to calculate the average difference of the corresponding pixel values of the left and right eye disparity images, and the context weight is obtained according to the difference size;
[0101] The depth weight, edge weight and context weight are weighted and combined to obtain the fusion weight;
[0102] Each pixel value of the left and right eye disparity images is fused according to a fusion ratio dynamically adjusted by a fusion weight to obtain a naked eye 3D image;
[0103] The naked-eye 3D image is corrected by using parallax mapping to obtain the final naked-eye 3D image.
[0104] In an embodiment of the present invention, a dynamic weight matrix is constructed by jointly calculating the depth weight, edge weight and context weight. The weight matrix fully considers the differences between the left and right eye images in different scenes, and combines the depth information and context features to improve the accuracy and consistency of the synthesized image, ensuring that the synthesized image conforms to the visual laws of the human eye, and finally dynamically adjusts the fusion ratio of the left and right eye images according to the fusion weight to synthesize the naked eye 3D image. The present invention ensures that the generated naked eye 3D image has clear edges, realistic stereoscopic sense and natural colors through a dynamic fusion mechanism and a multi-stage optimization process.
[0105] like Figure 6 As shown, an embodiment of the present invention further provides a naked eye 3D interactive system based on OSG, the system comprising:
[0106] Initialization configuration module 100, used to load OSG library and deep learning model, configure naked eye 3D display parameters, use OSG to load 3D model, texture and light source, and obtain OSG scene;
[0107] The interactive data processing module 200 is used to collect the user's interactive data in real time, input the interactive data into the deep learning model, and perform real-time analysis;
[0108] A position and posture adjustment module 300 is used to output the perspective parameters and model parameters obtained by analysis, and adjust the camera position and model posture in the OSG scene according to the perspective parameters and model parameters;
[0109] The parallax image optimization module 400 is used to generate a parallax image using computer vision technology according to the optimized OSG scene, and optimize the parallax image using a deep learning algorithm to improve the realism and stereoscopic effect of the 3D effect;
[0110] The naked-eye 3D image module 500 is used to synthesize the optimized parallax images into a naked-eye 3D image and output it to a display device.
[0111] As a preferred embodiment of the present invention, the interactive data processing module 200 includes:
[0112] An interaction data collection unit, used to collect user interaction data in real time through a camera and a sensor, wherein the interaction data includes head posture and gesture position;
[0113] A data preprocessing unit, used for preprocessing the interaction data, wherein the preprocessing includes denoising and normalization;
[0114] The learning model analysis unit is used to input the preprocessed interaction data into a deep learning model, wherein the deep learning model is a convolutional neural network or a recurrent neural network.
[0115] As a preferred embodiment of the present invention, the position and posture adjustment module 300 includes:
[0116] A camera position adjustment unit, used to adjust the position of the camera according to the translation vector in the viewing angle parameter, and to update the camera's posture using the rotation matrix in the viewing angle parameter;
[0117] The model node traversal unit is used to traverse all model nodes in the OSG scene and determine the model that needs to be adjusted;
[0118] The model posture adjustment unit is used to adjust the position of the model according to the translation vector in the model parameters, and update the posture of the model according to the scaling ratio and rotation angle in the model parameters.
[0119] As a preferred embodiment of the present invention, the parallax image optimization module 400 includes:
[0120] A depth information extraction unit, used to extract depth information from the optimized OSG scene by rendering a depth map of the scene, wherein the depth map records the distance information of each point in the scene;
[0121] A disparity image generating unit, used to generate left-eye and right-eye disparity images according to depth information using a stereo matching algorithm, wherein the disparity images are used to simulate the difference between left-eye and right-eye images seen by human eyes;
[0122] The disparity image optimization unit is used to input the generated disparity image into the generative adversarial network to optimize the quality of the disparity image and reduce noise and artifacts.
[0123] As a preferred embodiment of the present invention, the naked eye 3D image module 500 includes:
[0124] A disparity image alignment unit, used for performing disparity image alignment so that the left and right eye disparity images are aligned at the pixel level;
[0125] A 3D image synthesis unit, used for synthesizing the left and right eye parallax images into a naked eye 3D image through a naked eye 3D synthesis algorithm;
[0126] The image post-processing unit is used to perform post-processing on the synthesized naked-eye 3D image, and the post-processing includes color correction and brightness adjustment.
[0127] The above only describes in detail the preferred embodiments of the present invention, which is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
[0128] It should be understood that, although each step in the flow chart of each embodiment of the present invention is shown in sequence according to the indication of the arrow, these steps are not necessarily performed in sequence according to the order indicated by the arrow. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be performed in other orders. Moreover, at least a portion of the steps in each embodiment may include a plurality of sub-steps or a plurality of stages, and these sub-steps or stages are not necessarily performed at the same time, but can be performed at different times, and the execution order of these sub-steps or stages is not necessarily performed in sequence, but can be performed in turn or alternately with at least a portion of other steps or sub-steps or stages of other steps.
[0129] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0130] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the disclosure in the specification and examples. This application is intended to cover any variations, uses or adaptations of the present disclosure, which follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present disclosure are indicated by the claims.
Claims
1. A naked eye 3D interaction method based on OSG, characterized in that: The method comprises the following steps: Load the OSG library and deep learning model, configure naked-eye 3D display parameters, use OSG to load 3D models, textures and light sources, and obtain the OSG scene; Collect user interaction data in real time, input the interaction data into a deep learning model, and perform real-time analysis; Output the view parameters and model parameters obtained by the analysis, and adjust the camera position and model posture in the OSG scene according to the view parameters and model parameters; Based on the optimized OSG scene, computer vision technology is used to generate parallax images, and deep learning algorithms are used to optimize the parallax images to improve the realism and stereoscopic effect of the 3D effect. synthesizing the optimized parallax images into a naked-eye 3D image, and outputting the image to a display device; The step of collecting user interaction data in real time and inputting the interaction data into a deep learning model specifically includes: Collect user interaction data in real time through cameras and sensors, the interaction data including head posture and gesture position; Preprocessing the interaction data, wherein the preprocessing includes denoising and normalization; The preprocessed interaction data is input into a deep learning model, which is a convolutional neural network or a recurrent neural network.
2. The naked eye 3D interaction method based on OSG according to claim 1, characterized in that: The preprocessing of the interaction data includes the steps of denoising and normalization, specifically including: Apply wavelet transform to the interaction data to obtain low-frequency subband and high-frequency subband; Calculate the soft threshold according to the standard deviation of the high frequency sub-band, calculate the absolute value of the high frequency sub-band, and compare the absolute value with the soft threshold, and filter the processed high frequency sub-band according to the comparison result; The low-frequency sub-band and the processed high-frequency sub-band are combined to obtain the denoised interaction data; Given a reference trajectory, calculate the Euclidean distance matrix between the denoised interaction data and the reference trajectory; According to the Euclidean distance matrix, the DTW algorithm is used to calculate and determine the optimal alignment path between the denoised interaction data and the reference trajectory; According to the optimal alignment path, the user data is interpolated to the same time step as the reference trajectory to obtain the aligned interaction data; The median of the aligned interaction data was calculated, and the interquartile range was calculated based on the median; According to the interquartile range, outliers are removed from the aligned interaction data to obtain data after outliers are removed; The data after removing outliers is mapped to a range through dynamic mapping to perform dynamic normalization to obtain normalized data.
3. The naked eye 3D interaction method based on OSG according to claim 2, characterized in that: The step of adjusting the camera position and the model posture in the OSG scene according to the viewing angle parameters and the model parameters specifically includes: Adjust the camera's position according to the translation vector in the viewing angle parameter, and update the camera's posture using the rotation matrix in the viewing angle parameter; Traverse all model nodes in the OSG scene and determine the model that needs to be adjusted; Adjust the model's position according to the translation vector in the model parameters, and update the model's pose according to the scale and rotation angle in the model parameters.
4. The naked eye 3D interaction method based on OSG according to claim 3, characterized in that: The step of adjusting the position of the model according to the translation vector in the model parameters and updating the posture of the model according to the scaling ratio and rotation angle in the model parameters specifically includes the following sub-steps: Adjust the position, scaling and rotation angle of the model according to the model parameters to obtain an adjusted model; The error between each adjusted model and model parameters in terms of model position, scaling, and rotation angle is used as the optimization target; Define the collision threshold, traverse all models in the scene, calculate the distance between every two models, and compare the obtained distance with the collision threshold to construct the collision penalty term; Calculate the angle between the surface normal vector of the model and the direction of the light source, evaluate the degree of light occlusion by the model, and construct a light occlusion penalty term; The optimization objective, collision penalty term and light occlusion penalty term are weighted respectively to form a multi-objective optimization function; The gradient descent method is used to iteratively optimize the model parameters. In each iteration, the improvement direction of the model parameters is calculated according to the gradient direction of the multi-objective optimization function. After the optimization is completed, the optimized model parameters are obtained. The final model transformation matrix is generated using the optimized model parameters, and the posture and position of the model are updated according to the final model transformation matrix.
5. The naked eye 3D interaction method based on OSG according to claim 4, characterized in that: The step of generating a disparity image using computer vision technology according to the optimized OSG scene and optimizing the disparity image using a deep learning algorithm specifically includes: Extracting depth information from the optimized OSG scene by rendering a depth map of the scene, wherein the depth map records the distance information of each point in the scene; Using a stereo matching algorithm, a left-eye disparity image is generated based on depth information. The disparity image is used to simulate the difference between the left-eye and right-eye images seen by the human eye. The generated disparity image is input into the generative adversarial network to optimize the quality of the disparity image and reduce noise and artifacts.
6. The naked eye 3D interaction method based on OSG according to claim 5, characterized in that: The step of inputting the generated disparity image into the generative adversarial network to optimize the quality of the disparity image and reduce noise and artifacts specifically includes: Construct a three-branch generator with global branch, local branch and depth branch; The disparity image is input into the global branch, and the global representation features of the disparity image are extracted through convolution operation; The global representation features are fed into the Transformer module, and the global features are obtained by modeling long-distance deep dependencies. Decode the global features to obtain a globally optimized image; Input the disparity image into the local branch, and crop the disparity image into several local areas; The local area is passed through a dense convolutional network to extract local detail features and obtain a local feature map; Deconvolve the local feature maps and then splice them to obtain the local optimized image; The disparity image is input into the depth branch and multi-scale convolution operation is performed to extract the deep feature tensor containing global and local geometric information; Perform geometric consistency alignment on the deep feature tensor to obtain an aligned deep feature tensor; The aligned deep feature tensor is subjected to several deconvolution operations in an iterative manner, and in each deconvolution operation, a skip connection is used to fuse it with the corresponding deep feature tensor as the input of the next deconvolution operation. After the iteration is completed, the output result is smoothed to obtain an optimized depth perception image. Perform weighted fusion on the global optimized image, the local optimized image and the depth perception optimized image to obtain an optimized disparity image; Inputting the quality of the optimized disparity image into the discriminator to obtain a score of the optimized disparity image; The discriminator loss function is constructed by evaluating the discriminator's ability to distinguish between real images and generated images, and the generator loss function is constructed by evaluating the generator's ability to deceive the discriminator; The discriminator loss function and the generator loss function are used to construct the adversarial loss function. The perceptual loss function is constructed based on the difference between the optimized disparity image and the high-level features of the real image and the image. The geometric consistency loss function is constructed based on the consistency difference of the optimized disparity image in the geometric space. The stereo smoothness loss function is constructed based on the depth change difference of the optimized disparity image. The adversarial loss function, the perceptual loss function, the geometric consistency loss function and the stereo smoothness loss function are weightedly summed to form a multimodal loss function; The discriminator and the generator are trained adversarially, and the parameters of the discriminator and the generator are iteratively optimized by minimizing the multimodal loss function. After the optimization is completed, the final disparity image is obtained.
7. The naked eye 3D interaction method based on OSG according to claim 6, characterized in that: The step of synthesizing the optimized parallax images into naked-eye 3D images specifically includes: Perform parallax image alignment so that the left and right eye parallax images are aligned at the pixel level; The left and right eye parallax images are synthesized into a naked eye 3D image through a naked eye 3D synthesis algorithm; The synthesized naked-eye 3D image is post-processed, and the post-processing includes color correction and brightness adjustment.
8. The naked eye 3D interaction method based on OSG according to claim 7, characterized in that: The step of synthesizing the left and right eye parallax images into a naked eye 3D image by using a naked eye 3D synthesis algorithm specifically includes: Calculate the difference between the pixel value of the left eye image and the corresponding pixel value of the right eye image for each pixel point; Obtaining a depth value of the depth map, and using the depth value of the depth map to normalize the difference of corresponding pixel values to obtain a disparity map; Input the depth value into a smooth nonlinear function to obtain the depth weight; Compare the difference in the edge between the left and right eye disparity images, and obtain the edge weight according to the difference; The sliding window method is used to calculate the average difference of the corresponding pixel values of the left and right eye disparity images, and the context weight is obtained according to the difference size; The depth weight, edge weight and context weight are weighted and combined to obtain the fusion weight; Each pixel value of the left and right eye disparity images is fused according to a fusion ratio dynamically adjusted by a fusion weight to obtain a naked eye 3D image; The naked-eye 3D image is corrected by using parallax mapping to obtain the final naked-eye 3D image.
9. A naked-eye 3D interaction system based on OSG, the system being applied to the naked-eye 3D interaction method based on OSG according to any one of claims 1 to 8, characterized in that: The system comprises: Initialize the configuration module, which is used to load the OSG library and deep learning model, configure the naked eye 3D display parameters, use OSG to load the 3D model, texture and light source, and obtain the OSG scene; An interactive data processing module is used to collect user interactive data in real time, input the interactive data into a deep learning model, and perform real-time analysis; A position and posture adjustment module, used to output the perspective parameters and model parameters obtained by analysis, and adjust the camera position and model posture in the OSG scene according to the perspective parameters and model parameters; The parallax image optimization module is used to generate parallax images based on the optimized OSG scene using computer vision technology and optimize the parallax images using deep learning algorithms to improve the realism and stereoscopic effect of the 3D effect; The naked-eye 3D image module is used to synthesize the optimized parallax images into a naked-eye 3D image and output it to a display device.
Citation Information
Patent Citations
Three-dimensional dynamic simulation method for water outlet process of submarine-launched missiles
CN103577656A
Implementation method of attitude interaction naked-eye three-dimensional hybrid virtual reality system
CN111679743A