A method and device for generating spatial visual interactive medium based on spatial calculation
Through technical means such as preprocessing, focal plane simulation and depth estimation, three-dimensional interactive space video is generated, which solves the problem of lack of immersive three-dimensional interaction in the existing technology, and realizes the free exploration and interaction of users in the three-dimensional space.
Patent Information
- Application Number
- CN202411522084.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-29
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2044-10-29
AI Technical Summary
The existing technology cannot provide a rich and immersive three-dimensional interactive experience, and users lack effective reconstruction of depth information and spatial structure when converting two-dimensional information into three-dimensional space.
By preprocessing the original image, focal plane simulation, complex data wavelet transformation, depth estimation and view synthesis, a new viewpoint depth map is generated, a three-dimensional scene is constructed and spatial visual interaction is performed, and three-dimensional interactive spatial video is automatically generated.
The transformation from two-dimensional information to three-dimensional interactive space is realized, and users can freely explore and interact, improving the interactive experience.
Smart Images

Figure CN119363956B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of spatial intelligence and spatial audio-visual, and in particular to a method and device for generating spatial visual interactive media based on spatial calculation. Background Art
[0002] With the rapid development of science and technology, the field of AI is experiencing unprecedented prosperity, with many remarkable innovative achievements emerging, such as the powerful natural language processing capabilities demonstrated by ChatGPT, the widespread application of AIGC (AI Generated Content) technology, and the rise of creative industries such as AI comics.
[0003] Existing 3D reconstruction relies primarily on advanced image processing and computer vision techniques to extract rich depth information and spatial structure from a single 2D image. Combined with 3D modeling, rendering, and animation technologies, AI transforms this 2D information into a realistic 3D interactive space, allowing users to freely explore and interact within it. However, due to complex technical requirements such as video synthesis, motion planning, and light and shadow simulation, existing technologies cannot provide a rich and immersive interactive experience. Summary of the Invention
[0004] The purpose of the embodiments of the present invention is to provide a method and device for generating a spatial visual interactive medium based on spatial calculation, which converts two-dimensional information into three-dimensional interactive spatial video, thereby improving the user's interactive experience.
[0005] To achieve the above objectives, an embodiment of the present invention provides a method for generating a spatial visual interactive medium based on spatial calculation, the method comprising:
[0006] Preprocessing the original image to obtain first information;
[0007] performing focal plane simulation processing on the first information to obtain second information;
[0008] extracting complex data from the first information, and performing wavelet transform on the complex data to obtain a feature set;
[0009] performing depth estimation and view synthesis on the second information to obtain third information;
[0010] Performing spatial visual interaction on the complex data and the third information to generate a new viewpoint depth map;
[0011] A spatial video is determined according to the first information, the feature set, and the new viewpoint depth map.
[0012] Optionally, performing focal plane simulation processing on the first information to obtain second information includes:
[0013] The first information is captured by simulating a camera at different focal lengths and / or focus positions to decompose depth information of the scene.
[0014] Optionally, extracting complex data from the first information includes:
[0015] Data in the first information that exceeds a threshold in at least one of structure, content, and dimension is extracted as complex data.
[0016] Optionally, performing depth estimation and view synthesis on the second information to obtain third information includes:
[0017] Analyzing the texture gradient and perspective distortion of the second information to obtain a depth detail map;
[0018] The spatial data of the depth detail map is subjected to classification analysis processing to obtain third information, wherein the classification analysis processing includes time series sorting, video encoding and data synchronization.
[0019] Optionally, the performing texture gradient and perspective distortion analysis on the second information to obtain a depth detail map includes:
[0020] Analyzing the gradient direction and speed of the texture in the second information to obtain image depth, and determining the depth level according to the occlusion between objects in the image;
[0021] determining, based on the principle of perspective, the convergence point and / or the vanishing point of parallel lines in the second information;
[0022] Statistical learning and / or deep learning are performed on the convergence points and / or vanishing points according to the image depth and depth level to obtain a depth detail map.
[0023] Optionally, performing spatial visual interaction on the complex data and the third information to generate a new viewpoint depth map includes:
[0024] performing a data set analysis on the complex data and the third information to obtain disparity and depth information, and determining a new viewpoint position according to the disparity and depth information;
[0025] Generating pixel values at the new viewpoint position according to at least one method selected from image interpolation technology, bilinear interpolation, cubic interpolation, and deep learning interpolation;
[0026] A new viewpoint depth map is determined according to the new viewpoint position and the pixel values of the new viewpoint position.
[0027] Optionally, determining the spatial video according to the first information, the feature set, and the new viewpoint depth map includes:
[0028] constructing a three-dimensional scene based on the first information, the feature set, and the new viewpoint depth map, and performing spatial region coloring on the three-dimensional scene;
[0029] Determining the spatial data according to the spatial visual interactive medium of the three-dimensional scene, wherein the spatial data includes spatial position, spatial relationship, spatial query, spatial analysis, spatial editing and spatial area coloring;
[0030] determining a three-dimensional interactive space according to the spatial data;
[0031] Perform visual interactive conversion on the data attribute medium in the three-dimensional interactive space to obtain a spatial video.
[0032] Optionally, the original image is at least one monocular image captured by an external source;
[0033] The preprocessing includes at least denoising, distortion correction, color adjustment, contrast enhancement and depth processing, and the depth processing is used to obtain the spatial position, spatial relationship, spatial query, spatial analysis, spatial editing and spatial area coloring of the original image.
[0034] On the other hand, the present invention also provides a device for generating a spatial visual interactive medium based on spatial calculation, the device comprising:
[0035] A first processing module, configured to pre-process the original image to obtain first information;
[0036] a second processing module, configured to perform focal plane simulation processing on the first information to obtain second information;
[0037] a third processing module, configured to extract complex data from the first information and perform wavelet transform on the complex data to obtain a feature set;
[0038] a fourth processing module, configured to perform depth estimation and view synthesis on the second information to obtain third information;
[0039] a fifth processing module, configured to perform spatial visual interaction on the complex data and the third information to generate a new viewpoint depth map;
[0040] The sixth processing module is configured to determine a spatial video based on the first information, the feature set, and the new viewpoint depth map.
[0041] Optionally, performing focal plane simulation processing on the first information to obtain second information includes:
[0042] The first information is captured by simulating a camera at different focal lengths and / or focus positions to decompose depth information of the scene.
[0043] Optionally, performing depth estimation and view synthesis on the second information to obtain third information includes:
[0044] Analyzing the texture gradient and perspective distortion of the second information to obtain a depth detail map;
[0045] The spatial data of the depth detail map is subjected to classification analysis processing to obtain third information, wherein the classification analysis processing includes time series sorting, video encoding and data synchronization.
[0046] Optionally, determining the spatial video according to the first information, the feature set, and the new viewpoint depth map includes:
[0047] constructing a three-dimensional scene based on the first information, the feature set, and the new viewpoint depth map, and performing spatial region coloring on the three-dimensional scene;
[0048] Determining the spatial data according to the spatial visual interactive medium of the three-dimensional scene, wherein the spatial data includes spatial position, spatial relationship, spatial query, spatial analysis, spatial editing and spatial area coloring;
[0049] determining a three-dimensional interactive space according to the spatial data;
[0050] Perform visual interactive conversion on the data attribute medium in the three-dimensional interactive space to obtain a spatial video.
[0051] The present invention provides a method for generating a spatial visual interactive medium based on spatial computing, the method comprising: preprocessing an original image to obtain first information; performing focal plane simulation processing on the first information to obtain second information; extracting complex data from the first information and performing wavelet transform on the complex data to obtain a feature set; performing depth estimation and view synthesis on the second information to obtain third information; performing spatial visual interaction on the complex data and the third information to generate a new viewpoint depth map; and determining a spatial video based on the first information, the feature set, and the new viewpoint depth map. The method restores a three-dimensional interactive space from a 2D image of a single viewpoint and automatically generates video content related thereto, transforming this two-dimensional information into a realistic three-dimensional interactive space, allowing users to freely explore and interact therein, thereby improving the user's interactive experience.
[0052] Other features and advantages of the embodiments of the present invention will be described in detail in the subsequent detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] The accompanying drawings are used to provide a further understanding of the embodiments of the present invention and constitute a part of the specification. Together with the following detailed description, they are used to explain the embodiments of the present invention, but do not constitute a limitation of the embodiments of the present invention. In the accompanying drawings:
[0054] Figure 1 It is a flowchart of a method for generating a spatial visual interactive medium based on spatial calculation according to the present invention;
[0055] Figure 2 It is a flow chart of a specific embodiment of the present invention;
[0056] Figure 3 Schematic diagram of new viewpoint images at different focal planes in the present invention;
[0057] Figure 4 is a schematic diagram of focal plane simulation of the present invention;
[0058] Figure 5 It is a schematic diagram of spatial calculation of focal plane transformation of the present invention.
[0059] Description of Reference Numerals
[0060] 1-Original viewpoint;
[0061] 2-first simulation viewpoint;
[0062] 3- Second simulation viewpoint;
[0063] 4- Third simulation viewpoint;
[0064] 5-first simulated focal plane;
[0065] 6-second simulated focal plane;
[0066] 7-third simulated focal plane;
[0067] A-first simulated viewpoint depth space;
[0068] B-second simulated viewpoint depth space;
[0069] C-third simulated viewpoint depth space;
[0070] D- fourth simulated viewpoint depth space;
[0071] E-Fifth simulated viewpoint depth space. DETAILED DESCRIPTION
[0072] The following describes the specific implementation of the embodiment of the present invention in detail with reference to the accompanying drawings. It should be understood that the specific implementation described herein is only used to illustrate and explain the embodiment of the present invention and is not used to limit the embodiment of the present invention.
[0073] like Figure 1 As shown, a method for generating a spatial visual interactive medium based on spatial calculation of the present invention includes: step S101 is to pre-process the original image to obtain first information.
[0074] The original image is at least one monocular image captured by an external source. Specifically, the original image (also called the original viewpoint image) is an image captured by the acquisition module from an external viewpoint. This original image is the basis for generating spatial domain color and spatial video in the present invention. The acquisition module can be considered a component or functional module responsible for collecting, receiving, or capturing raw data (such as images, video frames, etc.). It is the first step in the entire processing flow and provides the necessary input for subsequent data processing, analysis, modeling, and other steps.
[0075] The preprocessing includes at least denoising, distortion correction, color adjustment, contrast enhancement, and depth processing. The depth processing is used to obtain the spatial position, spatial relationships, spatial query, spatial analysis, spatial editing, and spatial region coloring of the original image. Specifically, the acquisition module performs preliminary processing on the captured image, such as image depth denoising, resizing, and color correction. A depth model of the original viewpoint image is constructed to obtain the original viewpoint image spatial model. The model parameters are optimized using a backpropagation algorithm to recover clean data from noisy data and restore distorted data to undistorted data. The preprocessing includes depth processing to obtain spatial attributes such as the spatial position, spatial relationships, spatial query, spatial analysis, spatial editing, and spatial region coloring of the original image. Spatial position refers to the relative position of pixel atoms in the original viewpoint image in three or four dimensions. Coordinates are used to describe spatial position and locate pixel atoms. The relative position, direction, distance, and topological relationship between pixel atoms constitute the spatial relationship of pixel atoms. Spatial query is the process of querying a pixel atom based on certain conditions, such as spatial position and shape, to find and locate other pixel atoms related to specific spatial elements (spatial elements include points, lines, and surfaces). Spatial analysis refers to the ability to divide and classify the pixel atomic space formed by the original viewpoint image into many different spatial types such as points, lines, and surfaces, which are then combined with the new viewpoint image space to generate a new pixel atomic space. Spatial editing and spatial domain coloring are editing properties that can improve the quality of the original viewpoint image space and convert erroneous colors, so that the pixel atoms in the original image space can be accurately restored. The acquisition module also performs preliminary screening of the captured images to exclude blurred, damaged, or non-compliant images to ensure the quality and consistency of the input data. If the format of the original data is incompatible with the subsequent processing module, the acquisition module may also be responsible for converting it to an appropriate format.
[0076] Step S102 is performing focal plane simulation processing on the first information to obtain second information.
[0077] According to a specific embodiment, the performing focal plane simulation processing on the first information to obtain the second information includes: capturing images of the first information at different focal lengths and / or focus positions by a simulated camera to decompose depth information of the scene.
[0078] Specifically, the content and space of the original viewpoint image are analyzed, including but not limited to objects in the space, lighting conditions, spatial complexity, etc., to clarify the goal of the focal plane simulation processing, decompose the depth information of the spatial scene, and automatically calculate the relative distances of different objects or regional spaces composed of different pixel atoms. The parameter settings of the simulated camera are set. The camera focal length can be the default focal length of the simulated camera or the focal length calculated based on the refraction line and the incident line. When the original viewpoint image is cropped, scaled, blurred, etc., the imaging effects at different focal lengths can be simulated. After continuously determining whether the focal length is consistent with the original viewpoint image, the image focal length is calculated. The images simulated at different focal lengths are compared.
[0079] Focal plane simulation, in computer vision or image processing, decomposes a scene's depth information by simulating images captured by a camera at different focal lengths or focal positions. This method doesn't directly perform mathematical transformations on the image, but instead changes the parameters of the simulated camera to observe the scene's appearance at different depth levels. This is a critical step for depth estimation and 3D / 4D reconstruction.
[0080] Specifically, multiple images with different depths of field are simulated using the original viewpoint image. The foreground and background depths of field of the original viewpoint image are transformed to obtain multiple images with different depths of field. An acquisition module model is established within the spatial visual interaction system. This model simulates various parameters of a real camera, such as focal length and focus. Based on the simulated focal length and focus position, the corresponding parameters of the acquisition module model are adjusted. The scene is rendered using the adjusted acquisition module model parameters to generate images captured by the simulated camera at different focal lengths, focal points, and focal plane positions. The above simulation steps are repeated to generate a series of images at different focal lengths, focal points, and focal plane positions. These images will demonstrate changes in detail at different depth levels, thus generating an image sequence. Different three-dimensional spaces are generated based on different viewpoints, constructing a four-dimensional spatial architecture. The generated three-dimensional spatial image sequences are then compared and analyzed to observe changes in the clarity (in-focus and out-of-focus) of objects in the scene as the focal length, focus, and focal plane position change. Depth map generation uses image processing (such as deep learning, stereo matching, and optical flow) to analyze these changes in the image sequence, inferring the depth information corresponding to each pixel and generating a depth map. Depth information optimization performs denoising and smoothing on the generated depth map to reduce image noise or artifacts introduced by algorithm errors. Through methods such as spatial consistency checks, the depth values of adjacent pixels in the depth map are ensured to be logically reasonable. The extracted depth information is combined with image feature data to reconstruct 3D / 4D scenes, such as Figure 3As shown, one enters a space from the original viewpoint image (monocular), and enters another space from another viewpoint. Among them, viewpoint A1 is the viewpoint simulation camera that performs three-dimensional restoration of the original viewpoint image through multi-level decomposition, focal plane simulation and depth processing. The camera calculates the simulated viewpoint S1 and the simulated viewpoint S2 by continuously simulating the focal plane to obtain the focal length. The viewpoint space K1, viewpoint space K2 and viewpoint space K3 are analyzed at the same level to obtain a visually interactive mobile space. The viewpoint space K1, viewpoint space K2 and viewpoint space K3 are all three-dimensional interactive spaces. You can walk into the viewpoint space K2 from the viewpoint space K1. The viewpoint space is a three-dimensional interactive space that can be freely entered and exited by the person in need. In the present invention, n simulated viewpoints can be simulated and then the right viewpoints of the n simulated viewpoints can be restored, and n three-dimensional interactive spaces can be restored to form a four-dimensional field. You can freely shuttle through any three-dimensional interactive space. Thus, a four-dimensional scene effect is achieved. As shown Figure 4 As shown, the first simulated viewpoint depth space A, the second simulated viewpoint depth space B, the third simulated viewpoint depth space C, the fourth simulated viewpoint depth space D, and the fifth simulated viewpoint depth space E are five three-dimensional spaces, that is, the three-dimensional space of the simulated binocular is restored. In the present invention, the original viewpoint image is an image captured by a monocular camera and a series of subsequent operations are performed. Specifically, another viewpoint is restored by a monocular (single viewpoint). For example, a left viewpoint is selected to restore a right viewpoint. The restoration of this right viewpoint is the right viewpoint calculated by the focal length obtained by the simulated camera through the continuous simulation of the focal plane. After obtaining the left and right viewpoints, the three-dimensional space is restored, allowing people to walk into and out of this space. The first simulated viewpoint depth space A, the second simulated viewpoint depth space B, the third simulated viewpoint depth space C, the fourth simulated viewpoint depth space D, and the fifth simulated viewpoint depth space E are five different three-dimensional spaces. By continuously adding different original viewpoints, different three-dimensional spaces are obtained to construct a four-dimensional world. The video generated by walking from one space to another is a spatial video, and the spatial production video is connected through a spatial visual interactive medium. By comparing the depth data of the original viewpoint image with the 3D depth data, the accuracy and reliability of the depth information extracted by the focal plane simulation method were verified. Based on the verification results, the camera model, rendering parameters, and depth extraction algorithm were adjusted and optimized to improve the performance and accuracy of the overall spatial visual interaction system.
[0081] Feature data is data information obtained through complex data processing and wavelet transforms. Feature data is used to describe data characteristics. In image processing, feature data can include image corners, edges, textures, color histograms, and other features, and is used for image recognition, classification, and retrieval. By extracting the dimensionality and complexity of feature data, while classifying and retaining key information, feature extraction requires corner detection, edge detection, texture analysis, and color histogram analysis. First, high-frequency information obtained after wavelet transform is used to detect corners in the image. Corners are typically the points where local features in an image are most prominent and are important for image matching and recognition. By analyzing the deep detail layer of the wavelet transform, edges in the image can be identified. Edges are areas in an image where pixel values vary dramatically. Texture analysis is performed by using the statistical properties of the wavelet coefficients (such as energy and entropy) of the wavelet transform to describe the texture characteristics of the image. If the extracted features are too high in dimensionality, dimensionality reduction, such as principal component analysis or linear discriminant analysis, is required to reduce computational effort and improve classification efficiency. The most representative features are selected based on the subject's needs. The spatial visual interaction system uses a classifier trained by extracting feature data to classify and identify multiple images processed in depth. Based on feature extraction, it simulates multiple images in a new viewpoint range to obtain a multi-view set image.
[0082] Step S103 is to extract complex data from the first information and perform wavelet transform on the complex data to obtain a feature set.
[0083] According to a specific implementation, extracting complex data from the first information includes extracting data whose structure, content, and dimension in the first information exceed a threshold as complex data, where the threshold of the dimension is four dimensions.
[0084] Specifically, complex data generally refers to data sets that exceed conventional and / or simple data types in terms of structure, content, dimensionality, or processing. Complex data structures include unstructured data, semi-structured data, and other complex data. Compared to structured data (such as tables in a multi-level decomposition database), unstructured data (such as text, images, audio, and video) generally lacks a fixed format or schema. Semi-structured data refers to data with a specific format, but this structure is not as strict and standardized as that of a relational database, and therefore also has a certain degree of complexity in processing. The complexity of complex data content includes diversity, high dimensionality, and dynamic nature. Complex data contains multiple types of information, such as text, numbers, images, and videos, which may come from different sources and have different representations and semantics. The high dimensionality of complex data refers to the fact that when converting original viewpoint images into high-dimensional three-dimensional videos, the data has multiple dimensions or attributes. These dimensions may have complex correlations and dependencies, making data analysis and understanding difficult. In the present invention, data analysis is structured. The dynamic nature of complex data means that in the process of dynamic changes of complex data, new data points may be generated at any time and old data points may become obsolete or invalid. This requires the data processing system in the spatial interaction system to be able to process this data in real time or near real time.
[0085] The wavelet transform is used to decompose complex data into components of different frequencies. Unlike the Fourier transform, the wavelet transform has good localization characteristics in both the time domain and the frequency domain, and can simultaneously reveal the characteristics of the signal in time and frequency. In image processing, wavelet transform is often used for tasks such as image compression, denoising, and feature extraction. Specifically, the image file to be processed is first read, which usually includes a two-dimensional array and / or a three-dimensional array (color image, including three RGB channels and constructing a four-dimensional scene), and the image file is converted into grayscale (if the image being processed is a color image and the task does not require color information, the image can be converted into a grayscale image first for subsequent processing).
[0086] The effect of wavelet transform depends largely on the selected wavelet basis (such as Haar wavelet, Daubechies wavelet, Symlets wavelet, etc.). Different wavelet bases have different localization characteristics in the time domain and frequency domain, and are suitable for different application scenarios. Determine the decomposition level, and determine the number of wavelet decomposition layers according to the multi-level decomposition processing requirements in the present invention. The more layers, the finer the frequency domain division, but the higher the computational complexity. When performing wavelet transform, a two-dimensional wavelet transform is first performed on the image after preprocessing the original viewpoint image. The two-dimensional wavelet transform refers to performing a one-dimensional wavelet transform on each row of extracted data and a one-dimensional wavelet transform on each column to obtain wavelet coefficients. After the transformation, the image is decomposed into multiple coefficient matrices, including a low-frequency coefficient matrix and multiple high-frequency coefficient matrices. The processing of wavelet coefficients can reduce the amount of data by quantization (reducing the precision of the coefficients) and threshold processing (setting coefficients less than a certain threshold to zero). In denoising applications, high-frequency noise can be removed by setting a threshold, retaining only important edge and texture information. In feature extraction, specific wavelet coefficients can be selected as features based on the task requirements. When reconstructing a new viewpoint image, the inverse wavelet transform requires reconstructing the wavelet coefficients to restore the image after compression, denoising, and multi-level decomposition. This typically involves performing an inverse two-dimensional wavelet transform on the processed low-frequency and high-frequency coefficient matrices. After obtaining the processed image, feature extraction is performed to obtain a feature set for further analysis. This method optimizes the process by adjusting the wavelet basis, decomposition level, quantization parameters, thresholds, and other factors.
[0087] Multi-level decomposition refers to the process of gradually breaking down a complex dataset or image into simpler, more easily processed or analyzed components through a series of steps or levels. This multi-level decomposition is divided into focal plane simulation, depth information acquisition, and information integration and analysis. These decomposition levels are typically organized according to frequency, scale, space, resolution, and other factors to extract information at different levels of abstraction. Multi-level decomposition includes three steps. First, focal plane simulation is required in the present invention. Multi-level decomposition is no longer a direct mathematical transformation of the image, but rather decomposes the scene's depth information by simulating different focal planes. That is, the spatial visual interaction system does not directly analyze the complexity of the image content, but instead captures information at different depth levels of the scene by changing the focal length or focus position of the simulated camera. Secondly, by simulating multiple focal planes, the spatial visual interaction system can obtain clear images (or at least relatively clearer representations) of objects at different depths in the scene. These images at different focal planes can be viewed as a multi-level representation of the scene's depth information, with each level corresponding to the clarity of objects within a specific depth range.
[0088] Step S104 is to perform depth estimation and view synthesis on the second information to obtain third information.
[0089] According to a specific implementation, the depth estimation and view synthesis of the second information to obtain the third information includes: analyzing the texture gradient and perspective distortion analysis of the second information to obtain a depth detail map; and performing time series sorting, video encoding, and data synchronization on the depth map corresponding to the depth detail map to obtain the third information.
[0090] The texture gradient and perspective distortion analysis of the second information is performed to obtain a depth detail map, including: analyzing the gradient direction and speed of the texture in the second information to obtain image depth, and determining the depth level according to the occlusion between objects in the image; determining the convergence point and / or vanishing point of parallel lines in the second information according to the principle of perspective; and performing statistical learning and / or deep learning on the convergence point and / or vanishing point according to the image depth and depth level to obtain a depth detail map.
[0091] Specifically, the depth estimation refers to inferring the three-dimensional depth information of objects in the scene from a two-dimensional image or image sequence. This can be achieved through a variety of methods, such as stereo vision, structured light, monocular vision, etc. In monocular vision, depth estimation usually relies on clues in the image, such as texture gradients, occlusion relationships, perspective distortion, etc. Depth cues analyze texture gradients and infer surface depth changes by analyzing the gradient direction and speed of the texture in the image. Areas with dense and rapidly changing textures usually represent nearby objects, while areas with sparse or slowly changing textures represent distant objects. Based on the occlusion relationship obtained by depth estimation, the spatial visual interaction system observes the mutual occlusion between objects. The occluded objects are usually located behind the occluders (i.e., deeper layers). Depth estimation performs perspective distortion and uses the principle of perspective to analyze the convergence or vanishing points of parallel lines in the image (such as roads, building edges, etc.), which can help depth estimate the depth hierarchy in the scene. Depth detail map generation: A training model is trained based on the combination of statistical learning and deep learning of the spatial visual interaction system. The three-dimensional structure is restored from the complex mapping relationship between spatial structure and depth layer information in the training model through multi-view geometry obtained by simulating the focal plane. By repeating the above steps, a four-dimensional structure is obtained to achieve interactive connection between spaces and can be entered and exited according to user needs, thereby obtaining accurate depth layer information.
[0092] View synthesis refers to generating videos from new perspectives and / or viewpoints based on known original viewpoint images. The present invention preferably utilizes techniques such as image interpolation, view warping, and illumination models in this process. View synthesis enables virtual reality, augmented reality, and video simulation, allowing users to observe the same scene from different perspectives or viewpoints, resulting in a more immersive and rich experience. A dataset of feature data is analyzed for parallax and depth information, leveraging this information to determine the precise location of the new viewpoint. Image interpolation techniques, including bilinear interpolation, cubic interpolation, and more advanced deep learning interpolation, are then used to generate pixel values at the new viewpoint. Based on the new viewpoint's location, a graphics transformation and fusion system transforms and warps the interpolated image to simulate the perspective effect observed from the new viewpoint. The interpolated and warped image components are then fused to ensure a smooth transition. Lighting and shadow adjustments to simulate the perspective effect of the new viewpoint are implemented using illumination models based on the scene's lighting conditions. Lighting factors such as ambient light, diffuse reflection, and specular reflection are considered to achieve a more accurate perspective space for the new viewpoint. After a series of processing, the synthesized multiple depth level images are sharpened, contrasted, etc. to improve the quality of the video space.
[0093] Step S105 is to generate a new viewpoint depth map by performing spatial visual interaction on the complex data and the third information.
[0094] According to a specific implementation, the spatial visual interaction of the complex data and the third information to generate a new viewpoint depth map includes: performing data set analysis on the complex data and the third information to obtain disparity and depth information, and determining a new viewpoint position based on the disparity and depth information; generating pixel values of the new viewpoint position based on at least one method selected from image interpolation technology, bilinear interpolation, cubic interpolation, and deep learning interpolation; and determining a new viewpoint depth map based on the new viewpoint position and the pixel values of the new viewpoint position.
[0095] Step S106 is to determine a spatial video according to the first information, the feature set and the new viewpoint depth map.
[0096] According to a specific embodiment, determining a spatial video based on the first information, feature set, and new viewpoint depth map includes: constructing a three-dimensional scene based on the first information, feature set, and new viewpoint depth map, and performing spatial region coloring on the three-dimensional scene; determining the spatial data based on the spatial visual interactive medium of the three-dimensional scene, the spatial data including spatial position, spatial relationship, spatial query, spatial analysis, spatial editing, and spatial region coloring; determining a three-dimensional interactive space based on the spatial data; and performing visual interactive conversion on the data attribute medium in the three-dimensional interactive space to obtain a spatial video. The three-dimensional interactive space is composed of multiple three-dimensional fields, also known as a four-dimensional field.
[0097] Specifically, the spatial video generally refers to video content that can express spatial position, direction, and depth information. Unlike traditional two-dimensional videos, spatial videos may contain image data from multiple perspectives or viewpoints, as well as depth maps, light field information, etc. associated with these images. It is a process of combining new viewpoint image spaces synthesized from multiple frames into spatial videos in a time sequence. Spatial video is used to provide users with a more immersive and interactive video experience, allowing users to observe video content from different perspectives, and even enter the three-dimensional space presented by the video to explore. With the development of virtual reality and augmented reality technologies, the application prospects of spatial videos are becoming increasingly broad.
[0098] In this invention, AI technology can also automatically generate video content based on this technology. This involves multiple complex processes, including video synthesis, motion planning, and light and shadow simulation. The spatial visual interaction system possesses a high degree of creativity and a deep understanding of the physical world. Through the integrated application of these technologies, AI can generate coherent and expressive videos, scalably showcasing dynamic changes and interactive processes in three-dimensional space.
[0099] In short, with the continuous advancement of AI technology, it has become possible to restore three-dimensional interactive space and automatically generate video content through 2D images from a single viewpoint. This not only brings new creative methods and unlimited possibilities to film and television production, game development and other fields, but also provides people with richer and more immersive entertainment and interactive experiences.
[0100] Figure 2 It is a flow chart of a specific embodiment of the present invention, such as Figure 2 As shown, this embodiment includes: acquiring an original viewpoint image (ie, original image) through an acquisition module, connecting the acquisition module and the spatial visual interaction system to transmit the original viewpoint image to the spatial visual interaction system.
[0101] To define the acquisition module's viewpoint identification and selection, the first step is to determine the viewpoint (or angle) from which to capture the image. This viewpoint can be fixed or movable (the spatial visual interaction system controls the acquisition module's acquisition angle, viewpoint, and distance). Configuring the acquisition device: Based on the selected viewpoint, configure the corresponding image capture device. This includes adjusting the acquisition module's parameters, such as resolution, frame rate, and exposure, to ensure that the captured image quality meets the requirements of subsequent processing. Next, capture the original viewpoint image. Initiating capture: Through software or hardware control, activate the image capture device and begin capturing images from the selected viewpoint. Real-time or timed capture: Depending on actual needs, you can choose to capture images in real time (e.g., a video stream) or at a specific interval (e.g., one frame per second). Data storage is then performed, storing the captured images somewhere in the spatial visual interaction system for subsequent processing (e.g., local storage, remote server, or cloud storage). Next, connect the acquisition module to the spatial visual interaction system. Based on the specific needs of the acquisition module and the spatial visual interaction system, select an appropriate data transmission method. Establish a connection between the acquisition module and the spatial visual interaction system using the selected transmission method. Once the connection is established, the captured raw viewpoint images are transmitted to the spatial visual interaction system in real time or on demand. The spatial visual interaction system then receives the raw viewpoint images transmitted from the acquisition module. The received images are preprocessed, such as by denoising, correcting distortion, and adjusting color, to improve image quality. The preprocessed images are then used to perform 3D reconstruction (if the goal is to build a 3D model) or generate a spatial video (if the goal is to create a video with depth perception). Finally, the 3D model or spatial video is displayed on a visualization interface, and interactive functions are provided so that users can observe, explore, or manipulate the visual content from different angles.
[0102] Before the acquisition module passes the original viewpoint image data, the acquisition module is responsible for performing some processing operations, such as image denoising, resizing, color correction, etc., to perform preliminary screening of the captured images and exclude blurred, damaged or non-compliant images. This ensures the quality and consistency of the input data. If the format of the original data is incompatible with the subsequent processing module, the acquisition module may also be responsible for converting it to an appropriate format. In the process of preprocessing the original image, the original viewpoint image will first be denoised. The spatial visual interaction system will perform mean filtering, median filtering, and Gaussian filtering on the identified original viewpoint image to obtain the first information. The spatial visual interaction system will perform contrast enhancement processing on the original viewpoint image, including histogram equalization, adaptive histogram equalization, contrast stretching, nonlinear enhancement, etc., to obtain the second information.
[0103] After preprocessing, the raw viewpoint image data is transmitted to the multi-level decomposition module for data analysis. Raw viewpoint image preprocessing, image acquisition, and preliminary screening are performed. Raw viewpoint image data is captured and initially screened. Image recognition algorithms or manual inspection are used to eliminate blurry, damaged, or substandard images. Color correction is performed to ensure image color consistency, and the image size is adjusted to meet the requirements of subsequent processing modules. If the raw data format is incompatible, the spatial visual interaction system performs format conversion. This reduces image noise and improves image quality.
[0104] According to the filtering process, contrast enhancement processing is performed after the filtering process to improve the visual effect of the image and make the details in the image clearer. The spatial visual interaction system needs to perform histogram equalization, adaptive histogram equalization, and nonlinear enhancement processing. Histogram equalization: By stretching the histogram of the image to make its distribution more uniform, the contrast of the image is enhanced. Adaptive histogram equalization (such as CLAHE): Based on the histogram equalization, the local area is restricted to avoid noise amplification caused by excessive enhancement. Contrast stretching: Directly adjust the brightness range of the image to make the dark part darker and the bright part brighter. Nonlinear enhancement: Use a nonlinear function (such as a logarithmic function, an exponential function) to map the image pixel values to change the contrast of the image. The first information includes the denoised image data and the contrast-enhanced image data.
[0105] The data is transmitted to the multi-level decomposition module. The pre-processed first information is transmitted to the multi-level decomposition module. The pre-processing module packages or separately transmits the denoised image data and the contrast-enhanced image data in the first information. The data is transmitted to the multi-level decomposition module via wireless or wired transmission. In the multi-level decomposition module, this information is used for further data analysis, such as feature extraction, target detection, 3D reconstruction, etc. for classification and decomposition. Figure 2 As shown in the figure, the original viewpoint image first passes through the preprocessing module (including denoising and contrast enhancement). The processed data (first information and second information) is then transmitted to the multi-level decomposition module for subsequent data analysis. This process ensures the quality and consistency of the input data, providing a solid foundation for subsequent advanced visual processing tasks.
[0106] The first information is extracted from complex data and subjected to wavelet transform to obtain a feature set, and the first information is subjected to focal plane simulation processing to obtain the second information; the first information is processed by a multi-level grading module, and the multi-level decomposition module will identify, classify, pick up, and process the data information using an attention mechanism based on the received first information. In the present invention, multi-level decomposition refers to the creative application of scene depth information simulating different focal planes. Multi-level decomposition is divided into focal plane simulation, depth information acquisition, and information integration and analysis. Multi-level decomposition refers to the gradual decomposition of a complex data set or image into simpler, easier to process or analyze parts through a series of steps or levels. These decomposition levels are usually organized according to frequency, scale, space, resolution, etc., so as to extract information at different levels of abstraction.
[0107] According to one specific embodiment, multi-level decomposition has three steps. Multi-level decomposition is no longer a direct mathematical transformation of the image, but rather decomposes the scene's depth information by simulating different focal planes. This means that the spatial visual interaction system does not directly analyze the complexity of the image content, but instead captures information about the scene at different depth levels by changing the focal length or focus position of the simulated camera. Secondly, to obtain depth information, by simulating multiple focal planes, the spatial visual interaction system can obtain clear images (or at least relatively clearer representations) of objects at different depths in the scene. These images at different focal planes can be viewed as a multi-level representation of the scene's depth information, with each level corresponding to the clarity of objects within a specific depth range. Based on the first information, focal plane simulation is performed to obtain the second information. In the present invention, focal plane simulation refers to the decomposition of a scene's depth information by simulating images captured by a camera at different focal lengths or focus positions in computer vision or image processing. This method does not directly perform mathematical transformations on the image, but instead observes the scene's performance at different depth levels by changing the parameters of the simulated camera. This is the most critical step for the four-dimensional construction of the present invention.
[0108] like Figure 5 As shown, by distributing the intersection planes at different depth planes, the simulated camera in the scene has an adjustable focal length. The simulated camera calculates the focal point by using the incident and refracted light rays, and the intersection of the reflection and refraction of the parallel line from the left viewpoint (the original viewpoint) and the restored right viewpoint (the new viewpoint). The focal length is calculated by changing the focal plane of the simulated camera.
[0109] Multiple images with different depths of field are simulated from the original viewpoint image. The foreground and background depths of field of the original viewpoint image are transformed to obtain multiple images with different depths of field. An acquisition module model is established within the spatial visual interaction system. This model simulates various parameters of a real camera, such as focal length and focus. Based on the simulated focal length and focus position, the corresponding parameters of the acquisition module model are adjusted. The scene is rendered using the adjusted acquisition module model parameters to generate images captured by the simulated camera at different focal lengths, focal points, and focal plane positions. The above simulation steps are repeated to generate a series of images at different focal lengths, focal points, and focal plane positions. These images will show changes in detail at different depth levels, generating an image sequence. The generated image sequences are then compared and analyzed to observe changes in the clarity of objects in the scene (in-focus and out-of-focus) as the focal length, focus, and focal plane position change. Depth map generation uses image processing techniques (such as deep learning, stereo matching, and optical flow) to analyze these changes in the image sequence, inferring the depth information corresponding to each pixel and generating a depth map. Depth information optimization denoises and smoothes the generated depth map to reduce image noise and artifacts introduced by algorithmic errors. Spatial consistency checks and other methods are used to ensure the logical rationality of the depth values of adjacent pixels in the depth map. The extracted depth information, combined with image feature data, is used to reconstruct the 3D scene. The accuracy and reliability of the depth information extracted using the focal plane simulation method are verified by comparing it with the 3D depth data of the original viewpoint image. Based on the verification results, the camera model, rendering parameters, and depth extraction algorithm are adjusted and optimized to improve the performance and accuracy of the overall spatial visual interaction system.
[0110] Finally, these information from different focal planes can be integrated to reconstruct the three-dimensional structure of the scene, perform depth estimation, or implement image processing tasks that rely on depth information. Although this decomposition and reorganization process is not a mathematical transformation in the traditional sense, it still logically embodies the idea of multi-level decomposition. The advantages of using multi-level decomposition in the present invention are enhanced depth perception, improved processing efficiency, and flexibility. By simulating multiple focal planes, the spatial visual interaction system can more accurately capture and represent the depth information of the scene, which is also of great significance to computer vision, augmented reality, robot navigation and other fields. Compared with directly processing the entire complex scene, the multi-level decomposition method can gradually focus on key information, thereby improving computational efficiency while maintaining processing accuracy. This method allows the number and position of focal planes to be adjusted according to different application requirements, thereby flexibly adapting to different scenes and tasks. Multi-level decomposition in the present invention specifically refers to an innovative method of "decomposing" and obtaining the depth information of the scene by simulating different focal planes. It embodies a framework for hierarchical and structured processing of complex problems. The third information and feature set are obtained through the multi-level decomposition module;
[0111] The feature set refers to the data set obtained by wavelet transform through the multi-level decomposition module. In image processing, wavelet transform is often used for tasks such as image compression, denoising, and feature extraction. In the present invention, it is first necessary to read the image file to be processed, which is usually a two-dimensional array or a three-dimensional array (color image, including three RGB channels) and convert it into grayscale. If the image being processed is a color image and the task does not require color information, the image can be converted into a grayscale image first for subsequent processing. Determine the decomposition level and determine the number of wavelet decomposition layers based on the multi-level decomposition processing requirements in the present invention. The more layers there are, the finer the frequency domain division, but the higher the computational complexity.
[0112] After obtaining the pixel values at the new viewpoint, the image transformation and fusion system transforms and warps the interpolated image based on the new viewpoint to simulate the perspective effect from the new viewpoint. The interpolated and warped image components are then fused to ensure a smooth transition. To simulate the perspective effect from the new viewpoint, lighting and shadows are adjusted by applying a lighting model to adjust the lighting and shadow effects of the new viewpoint image based on the scene's lighting conditions. Shadow effects: In view synthesis, shadow effects are typically generated by calculating the relative position of the light source and the object, as well as the interaction of light on the object's surface and other objects in the scene. This involves shadow map generation and rendering techniques such as shadow mapping and screen space shadows. In view synthesis, the effects of ambient light on object surfaces can be simulated by adding an ambient light map or using an ambient light model. This helps increase scene brightness and detail, making the image more realistic. By accounting for lighting factors such as ambient light, diffuse reflection, and specular reflection, the perspective space of the new viewpoint is more accurate. After a series of processing, the synthesized multiple depth level images are sharpened, contrasted, etc. to improve the quality of the video space.
[0113] During training, neural networks learn how to predict missing pixel values based on input image features (such as edges and textures). Due to the powerful representational capabilities of neural networks, these methods are able to capture more complex image statistics than traditional interpolation methods and produce higher-quality interpolation results. The quality of interpolation is quantified by evaluating the network's performance on specific datasets (using metrics such as mean squared error and peak signal-to-noise ratio).
[0114] Based on the new viewpoint, the interpolated image is transformed to simulate the perspective effect observed from the new viewpoint. The perspective transformation matrix M can be calculated using the camera's internal and external parameters.
[0115] Applying a lighting model adjusts the lighting and shading of the new viewpoint image, taking into account factors such as ambient light, diffuse reflection, and specular reflection. Ambient light is often modeled as a global lighting factor that uniformly illuminates all objects in the scene. Mathematically, this can be represented as a constant term added to the brightness value of each pixel.
[0116] The mathematical description of the depth estimation process includes: depth cue analysis, texture gradients, occlusion relationships, perspective distortion, depth detail map generation, and perspective transformation matrices. Specifically, the original viewpoint image and possible auxiliary information (such as depth sensor data) are input; depth is inferred by analyzing the gradient direction and speed of textures in the image; mutual occlusion between objects is observed to infer depth levels; perspective principles are used to analyze the convergence or vanishing points of parallel lines; statistical learning and deep learning are combined to obtain accurate depth level information from the trained model; this is used to transform the image from one viewpoint to another. According to the steps described above, a new viewpoint depth map is generated.
[0117] In the present invention, the construction of a three-dimensional interactive space refers to the use of computer graphics and virtual reality technology to create a three-dimensional environment with which a user number can interact. This environment can be virtual or a three-dimensional reconstruction based on the real world. The three-dimensional interactive space is used in multiple aspects of spatial design, such as three-dimensional modeling, rendering, physical simulation, and interactive design. Users can immerse themselves in the virtual world and interact with virtual objects to obtain a more realistic and rich experience. The spatial visual interaction system constructs details such as scenes, objects, and materials, reconstructs lighting, shadows, texture mapping, etc., and determines the balance between rendering quality and performance. For the interactive space, physical properties such as mass, matcha, and elasticity are determined.
[0118] The mathematical and physical expressions of the space construction process are as follows: new viewpoint depth map, first information of the original viewpoint image (such as color, texture, material, etc.), feature set (such as object recognition, scene segmentation results); construct a three-dimensional scene model based on the depth map and image information; determine lighting, shadows, texture mapping, etc., and render a three-dimensional scene; add physical properties such as mass, friction, elasticity, etc. to the interactive space; design the interaction method between users and virtual objects. The spatial visual interaction system is based on the creation and transformation of the geometric body, such as translation, rotation, scaling, etc. These transformations can achieve three-dimensional reconstruction through matrix operations. According to the mechanical equations, such as Newton's second law F=ma, it is used to calculate the movement of objects. According to the above steps, a constructed three-dimensional interactive space is obtained. According to the above steps, a constructed three-dimensional interactive space is obtained.
[0119] The constructed 3D interactive space is obtained from the spatial visual interaction system, including rendering information such as the scene model, materials, lighting, and shadows. The view synthesis module obtains multiple frames of new viewpoint images and their corresponding depth maps. These images have been rendered and interpolated based on different viewpoints specified by the user or the system. Temporal organization: The multiple frames of new viewpoint images are sorted and organized in a temporal sequence. This typically means the images are captured in chronological order or generated in simulation based on time steps. Video encoding: Each frame is encoded, including compression and format conversion, to reduce file size and improve transmission efficiency. The encoding process may include steps such as color space conversion, motion estimation, transform coding, quantization, and entropy coding. Depth maps and light field information also require corresponding encoding to ensure they can be played synchronously with the image data. Data synchronization: Ensure that the image data, depth map, and light field information are temporally synchronized. This is crucial for providing accurate depth perception and an interactive experience. Video output: The encoded video data, depth map, and light field information are packaged into a video file format such as MP4 or AVI. Output video files for playback on virtual reality devices, augmented reality platforms, or traditional video players. Software implementation of a spatial visual interaction system: Use a 3D rendering engine (such as Unity or Unreal Engine) to build and render 3D scenes. Use a video processing library (such as FFmpeg) for image and video encoding, decoding, and format conversion. Develop specialized scripts or applications to organize time series, synchronize data, and package video files. Use a high-performance graphics processing unit (GPU) to accelerate the 3D rendering and video encoding process. Sufficient storage space is required to store the generated video files and data such as depth maps. View transformation: Use the perspective transformation matrix to project the 3D scene onto the 2D image plane. This involves mathematical operations such as matrix multiplication and coordinate transformation. Image coding: Transform coding (such as the DCT transform) converts the image from the spatial domain to the frequency domain, facilitating subsequent quantization and entropy coding. During quantization, continuous pixel values or transform coefficients are mapped to a limited set of discrete values to reduce the amount of data. Entropy coding (such as Huffman coding and arithmetic coding) exploits the statistical properties of data for lossless or lossy compression. Optical simulation of physical processes: During the three-dimensional rendering process, the propagation and interaction of light in the scene are simulated, including physical phenomena such as reflection, refraction, and shadow. Lighting models (such as the Phong model) are used to calculate the lighting effects on the surface of objects to produce realistic visual effects. Motion simulation: If the spatial video contains dynamic scenes or animation effects, the motion trajectory and dynamic characteristics of the object are simulated. Physics equations (such as Newton's second law) are used to calculate parameters such as the acceleration, velocity, and position of the object. Perception simulation: Depth maps and light field information are used to simulate the human eye's perception of three-dimensional space, including depth perception, motion parallax, etc.
[0120] This invention discloses a method for restoring a three-dimensional interactive space from a single-viewpoint 2D image. Specifically, this method involves restoring the three-dimensional interactive space from a single-viewpoint 2D image and automatically generating related video content, transforming this two-dimensional information into a realistic three-dimensional interactive space within which users can freely explore and interact. By continuously adjusting the distance between the viewpoints to simulate and restore the view images from both viewpoints, the method extracts rich depth information and spatial structure from the single 2D image.
[0121] Through focal plane inference and correction, focal planes at different depths are simulated, and image features at these different focal planes are captured or calculated. Depth estimation is performed on the original image. Using the original image and focal plane information from a single viewpoint, a depth map of the scene is estimated through methods such as stereo matching and deep learning. Based on the estimated depth map and known camera parameters (such as focal length and baseline distance), view transformation and projection techniques from computer graphics, also known as depth view synthesis, are used to generate the right viewpoint image.
[0122] This invention does not directly perform multi-level decomposition of the image (such as wavelet transform), but instead decomposes the depth information of the scene by simulating different focal planes. Based on the depth map of the 2D image, the invention utilizes the psychological stereoscopic vision characteristics of the human eye, extracts the depth information from the flat image of a viewpoint, and combines it with the original viewpoint image to reconstruct the 3D interactive space.
[0123] Through the present invention, users can achieve spatial roaming, that is, free movement and rotation in a three-dimensional interactive space, controlled by input devices such as a mouse, keyboard or other devices, and can move freely in the three-dimensional interactive space and generate interactive feedback, providing visual (such as visual effects during collision) and auditory (such as collision sounds) feedback.
[0124] On the other hand, the present invention also proposes a device for generating a spatial visual interactive medium based on spatial computing, the device including: a first processing module for preprocessing the original image to obtain first information; a second processing module for performing focal plane simulation processing on the first information to obtain second information; a third processing module for extracting complex data of the first information, and performing wavelet transform on the complex data to obtain a feature set; a fourth processing module for performing depth estimation and view synthesis on the second information to obtain third information; a fifth processing module for performing spatial visual interaction on the complex data and the third information to generate a new viewpoint depth map; and a sixth processing module for determining a spatial video based on the first information, the feature set and the new viewpoint depth map.
[0125] The focal plane simulation processing of the first information to obtain the second information includes: capturing images of the first information at different focal lengths and / or focus positions by simulating a camera to decompose the depth information of the scene. The depth estimation and view synthesis processing of the second information to obtain the third information includes: analyzing the texture gradient and perspective distortion analysis of the second information to obtain a depth detail map; and classifying and analyzing the spatial data of the depth detail map to obtain the third information, wherein the classification and analysis processing includes time series sorting, video encoding, and data synchronization. The determining of the spatial video based on the first information, feature set, and new viewpoint depth map includes: constructing a three-dimensional scene based on the first information, feature set, and new viewpoint depth map, and performing spatial region coloring on the three-dimensional scene; determining the spatial data based on the spatial visual interactive medium of the three-dimensional scene, wherein the spatial data includes spatial position, spatial relationship, spatial query, spatial analysis, spatial editing, and spatial region coloring; determining a three-dimensional interactive space based on the spatial data; and performing visual interactive conversion on the data attribute medium in the three-dimensional interactive space to obtain the spatial video.
[0126] The present invention provides a method for generating a spatial visual interactive medium based on spatial computing, the method comprising: preprocessing an original image to obtain first information; performing focal plane simulation processing on the first information to obtain second information; extracting complex data from the first information and performing wavelet transform on the complex data to obtain a feature set; performing depth estimation and view synthesis on the second information to obtain third information; performing spatial visual interaction on the complex data and the third information to generate a new viewpoint depth map; and determining a spatial video based on the first information, the feature set, and the new viewpoint depth map. The method restores a three-dimensional interactive space from a 2D image of a single viewpoint and automatically generates video content related thereto, transforming this two-dimensional information into a realistic three-dimensional interactive space, allowing users to freely explore and interact therein, thereby improving the user's interactive experience.
[0127] The above describes in detail the optional implementation methods of the embodiments of the present invention in conjunction with the accompanying drawings. However, the embodiments of the present invention are not limited to the specific details in the above implementation methods. Within the technical concept of the embodiments of the present invention, various simple modifications can be made to the technical solutions of the embodiments of the present invention, and these simple modifications all fall within the scope of protection of the embodiments of the present invention.
[0128] It should also be noted that the various specific technical features described in the above specific embodiments can be combined in any appropriate manner without contradiction. To avoid unnecessary repetition, the embodiments of the present invention will not further describe various possible combinations.
[0129] Those skilled in the art will understand that all or part of the steps in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a program, which is stored in a storage medium and includes a number of instructions for causing a single-chip microcomputer, chip or processor to execute all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program code.
[0130] In addition, various implementations of the embodiments of the present invention may be arbitrarily combined, and as long as they do not violate the concept of the embodiments of the present invention, they should also be regarded as the contents disclosed in the embodiments of the present invention.
Claims
1. A method for generating spatial visual interactive media based on spatial calculation, characterized in that: The method includes: Preprocessing the original image to obtain first information; performing focal plane simulation processing on the first information to obtain second information; extracting complex data from the first information, and performing wavelet transform on the complex data to obtain a feature set; performing depth estimation and view synthesis on the second information to obtain third information; Performing spatial visual interaction on the complex data and the third information to generate a new viewpoint depth map; Determine a spatial video based on the first information, the feature set, and the new viewpoint depth map; The performing depth estimation and view synthesis on the second information to obtain third information includes: Analyzing the texture gradient and perspective distortion of the second information to obtain a depth detail map; The spatial data of the depth detail map is subjected to classification analysis processing to obtain third information, wherein the classification analysis processing includes time series sorting, video encoding and data synchronization.
2. The method according to claim 1, characterized in that The performing focal plane simulation processing on the first information to obtain second information includes: The first information is captured by simulating a camera at different focal lengths and / or focus positions to decompose depth information of the scene.
3. The method according to claim 1, characterized in that The extracting of complex data of the first information includes: Data in the first information that exceeds a threshold in at least one of structure, content, and dimension is extracted as complex data.
4. The method according to claim 1, wherein The analyzing the texture gradient and perspective distortion of the second information to obtain a depth detail map includes: Analyzing the gradient direction and speed of the texture in the second information to obtain image depth, and determining the depth level according to the occlusion between objects in the image; determining, based on the principle of perspective, the convergence point and / or the vanishing point of parallel lines in the second information; Statistical learning and / or deep learning are performed on the convergence points and / or vanishing points according to the image depth and depth level to obtain a depth detail map.
5. The method according to claim 1, wherein The performing spatial visual interaction on the complex data and the third information to generate a new viewpoint depth map includes: performing a data set analysis on the complex data and the third information to obtain disparity and depth information, and determining a new viewpoint position according to the disparity and depth information; Generating pixel values at the new viewpoint position according to at least one method selected from image interpolation technology, bilinear interpolation, cubic interpolation, and deep learning interpolation; A new viewpoint depth map is determined according to the new viewpoint position and the pixel values of the new viewpoint position.
6. The method according to claim 1, characterized in that The determining of the spatial video according to the first information, the feature set, and the new viewpoint depth map includes: constructing a three-dimensional scene based on the first information, the feature set, and the new viewpoint depth map, and performing spatial region coloring on the three-dimensional scene; Determining spatial data based on the spatial visual interactive medium of the three-dimensional scene, wherein the spatial data includes spatial position, spatial relationship, spatial query, spatial analysis, spatial editing, and spatial area coloring; determining a three-dimensional interactive space according to the spatial data; Perform visual interactive conversion on the data attribute medium in the three-dimensional interactive space to obtain a spatial video.
7. The method according to claim 1, characterized in that The original image is at least one monocular image captured by an external source; The preprocessing includes at least denoising, distortion correction, color adjustment, contrast enhancement and depth processing, and the depth processing is used to obtain the spatial position, spatial relationship, spatial query, spatial analysis, spatial editing and spatial area coloring of the original image.
8. A device for generating spatial visual interactive media based on spatial calculation, characterized in that: The device includes: A first processing module, configured to pre-process the original image to obtain first information; a second processing module, configured to perform focal plane simulation processing on the first information to obtain second information; a third processing module, configured to extract complex data from the first information and perform wavelet transform on the complex data to obtain a feature set; a fourth processing module, configured to perform depth estimation and view synthesis on the second information to obtain third information, including: analyzing the second information for texture gradient and perspective distortion to obtain a depth detail map; and performing classification analysis on spatial data of the depth detail map to obtain third information, wherein the classification analysis includes time series sorting, video encoding, and data synchronization; a fifth processing module, configured to perform spatial visual interaction on the complex data and the third information to generate a new viewpoint depth map; The sixth processing module is configured to determine a spatial video based on the first information, the feature set, and the new viewpoint depth map.
9. The device according to claim 8, characterized in that The performing focal plane simulation processing on the first information to obtain second information includes: The first information is captured by simulating a camera at different focal lengths and / or focus positions to decompose depth information of the scene.
10. The device according to claim 8, characterized in that The determining of the spatial video according to the first information, the feature set, and the new viewpoint depth map includes: constructing a three-dimensional scene based on the first information, the feature set, and the new viewpoint depth map, and performing spatial region coloring on the three-dimensional scene; Determining spatial data based on the spatial visual interactive medium of the three-dimensional scene, wherein the spatial data includes spatial position, spatial relationship, spatial query, spatial analysis, spatial editing, and spatial area coloring; determining a three-dimensional interactive space according to the spatial data; Perform visual interactive conversion on the data attribute medium in the three-dimensional interactive space to obtain a spatial video.
Citation Information
Patent Citations
Multifocal plane based method to produce stereoscopic viewpoints in a dibr system (mfp-dibr)
CN112136324A
Neural radiation field NeRF-based three-dimensional complex scene refined reconstruction method and device
CN118429526A