Scene Adaptive Adjustment Method and System for the Virtual-Reality Integration of Intelligent Internet of Things and Metaverse
By obtaining scene image and environment parameter data, using resolution decomposition and block identification of scene types, building multi-dimensional environment parameter tensors and digital twin models, combining adaptive coordinate mapping networks, the precise fusion of virtual objects and entity space is achieved, solving the problem of poor fusion effect between virtual objects and real scenes, and improving user experience and rendering effect.
Patent Information
- Application Number
- CN202510616040.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-05-14
AI Technical Summary
The existing technology is difficult to achieve the precise fusion of virtual objects and real scenes under different lighting conditions and complex spatial structures, resulting in poor fusion effect of virtual objects and real scenes, affecting user experience, and it is impossible to build an accurate digital twin model based on the multi-source environment data collected by IoT devices. It lacks an adaptive feature optimization mechanism and cannot flexibly adjust the rendering effect.
By obtaining scene image and environment parameter data, using resolution decomposition and block recognition of scene types, building multi-dimensional environment parameter tensors and digital twin models, combining adaptive coordinate mapping networks, virtual objects are accurately projected and rendered in the entity space, and multi-scale feature extraction and iterative optimization techniques are used to generate rendering parameters to achieve the precise fusion of virtual objects and entity space.
It realizes the precise integration of virtual objects and physical scenes, improves the user's immersion and interactive experience, improves the adaptability and realism of virtual objects under different environmental conditions, and enhances the system's real-time response and rendering effect.
Smart Images

Figure CN120125761B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of artificial intelligence and Internet of Things, and particularly to a method and system for adaptively adjusting the virtual-real integration scenario of intelligent Internet of Things and the metaverse. Background Art
[0002] With the rapid development of Internet of Things technology and virtual reality technology, the virtual-real integration scenario of intelligent Internet of Things and the metaverse has become an important direction of the new generation of digital interaction. The seamless integration of virtual objects and the real environment requires the system to accurately identify the real scene, adapt to environmental parameters, and achieve reasonable rendering to provide an immersive user experience.
[0003] However, the existing technologies still have problems such as being difficult to effectively process scenes with different lighting conditions and a relatively high degree of complexity in spatial structure, resulting in a poor fusion effect between virtual objects and the real scene, affecting the user experience, being unable to build an accurate digital twin model based on multi-source environmental data collected by Internet of Things devices, being difficult to achieve the precise interaction and mapping relationship between virtual objects and the physical environment, and lacking an adaptive feature optimization mechanism, and being unable to flexibly adjust the rendering effect according to the scene type and environmental changes, making the virtual objects perform inconsistently in different scenes.
[0004] Therefore, there is an urgent need for a solution to solve the problems existing in the existing technologies. Summary of the Invention
[0005] Embodiments of the present invention provide a method and system for adaptively adjusting the virtual-real integration scenario of intelligent Internet of Things and the metaverse, which can at least solve some of the problems existing in the existing technologies.
[0006] In a first aspect of embodiments of the present invention, a method for adaptively adjusting the virtual-real integration scenario of intelligent Internet of Things and the metaverse is provided, including:
[0007] Obtain scene image data and environmental parameter data; perform resolution decomposition and block division on the scene image data, and identify the scene type through a preset scene recognition model to obtain scene type information;
[0008] Construct a multi-dimensional environmental parameter tensor based on the environmental parameter data, establish an environmental digital twin model, and map the environmental digital twin model to the virtual space by using an adaptive coordinate mapping network;
[0009] Construct multiple scene feature vector subspaces based on the scene image data, each scene feature vector subspace corresponding to a kind of scene type information, select the corresponding scene feature vector subspace according to the scene type information, and map the scene image data to the selected scene feature vector subspace to generate a target scene feature vector;
[0010] Obtain the projection position information of virtual objects in the virtual space in the physical space based on the environmental digital twin model; input the target scene feature vector into the preset feature optimization sub-network for iterative optimization to obtain the optimized feature vector;
[0011] Based on the hyperbolic space manifold features and projection position information constructed from the optimized feature vector, obtain the enhanced feature data and the spatial mapping matrix, and use the preset parameter mapping sub-network to generate the scene rendering parameters; perform rendering processing on the virtual object based on the scene rendering parameters, and superimpose and display the rendered virtual object on the physical space display interface. In an alternative embodiment,
[0012] Perform resolution decomposition and block division on the scene image data, and perform scene type recognition through the preset scene recognition model to obtain the scene type information including:
[0013] Perform hierarchical perception on the scene image data to generate multiple scene image layers with different resolutions;
[0014] Adaptively divide each scene image layer into blocks, and determine the optimal block size of each scene image layer based on the information entropy of the image blocks to obtain the corresponding image block sequence;
[0015] Construct a discrete wavelet transform pyramid for each image block sequence and extract multi-scale frequency domain features;
[0016] Input the multi-scale frequency domain features into the preset heterogeneous convolutional neural network, and the heterogeneous convolutional neural network uses different convolutional kernel sizes for feature extraction of frequency domain features at different scales to obtain the image block feature vectors;
[0017] Input the image block feature vectors into the preset self-calibrating feature enhancement network, and iteratively optimize the feature vectors based on the feedback mechanism to generate enhanced feature vectors;
[0018] Recombine the enhanced feature vectors according to the adaptive attention weights to generate a multi-scale feature sequence; input the multi-scale feature sequence into the preset two-stream neural network, the spatial stream branch of the two-stream neural network extracts the spatial features of the multi-scale feature sequence, and the temporal stream branch of the two-stream neural network extracts the temporal features of the multi-scale feature sequence; fuse the spatial features and the temporal features to generate a fused feature vector;
[0019] Input the fused feature vector into the preset hierarchical scene classifier, and the hierarchical scene classifier performs progressive classification on the fused feature vector based on the scene semantic hierarchical relationship to obtain the scene type information.
[0020] In an alternative embodiment,
[0021] Construct a multi-dimensional environmental parameter tensor based on environmental parameter data, establish an environmental digital twin model, and map the environmental digital twin model to a virtual space using an adaptive coordinate mapping network, including:
[0022] Align the environmental parameter data based on a spatio-temporal synchronization mechanism to construct a multi-dimensional environmental parameter tensor;
[0023] Perform adaptive wavelet packet decomposition on the multi-dimensional environmental parameter tensor to generate a multi-scale environmental feature spectrum;
[0024] Input the multi-scale environmental feature spectrum into a preset spectral embedding network, extract the correlation features between parameters through a multi-head attention mechanism, and generate a high-dimensional environmental feature vector;
[0025] Construct a hybrid memory neural network, input the high-dimensional environmental feature vector into the hybrid memory neural network, store long-term dependence information based on an external memory matrix, and capture transient change features in combination with a short-term dynamic memory unit to generate environmental dynamic prediction data;
[0026] Construct a recurrent neural tensor network, generate a fusion feature tensor based on the high-dimensional environmental feature vector and the environmental dynamic prediction data and input it into the recurrent neural tensor network. The recurrent neural tensor network uses a tensor decomposition method to perform multi-dimensional correlation analysis on the input data, establish a dynamic coupling relationship between environmental elements, and generate an environmental digital twin model;
[0027] Construct an adaptive coordinate mapping network. The adaptive coordinate mapping network includes a spatial transformation sub-network and a deformation compensation sub-network. Map the topological structure of the environmental digital twin model to a virtual space coordinate system through the spatial transformation sub-network, and use the deformation compensation sub-network to dynamically correct the geometric deformation during the mapping process, and map the corrected environmental digital twin model to the rendering coordinate system of the virtual space.
[0028] In an alternative embodiment,
[0029] Construct a recurrent neural tensor network, generate a fusion feature tensor based on the high-dimensional environmental feature vector and the environmental dynamic prediction data and input it into the recurrent neural tensor network. The recurrent neural tensor network uses a tensor decomposition method to perform multi-dimensional correlation analysis on the input data, establish a dynamic coupling relationship between environmental elements, and generate an environmental digital twin model, including:
[0030] Construct a Riemannian manifold mapping for the high-dimensional environmental feature vector to generate manifold features, and project the environmental dynamic prediction data into a curvature tensor space to generate tensor features;
[0031] Input the manifold features and the tensor features into a preset Riemannian optimization model, construct an optimization objective function based on geodesic distance metric, and use the conjugate gradient method with an adaptive step size to iteratively solve the optimization objective function to obtain optimized features;
[0032] Project the optimized features onto the Euclidean space through Riemannian exponential mapping to generate a fused feature tensor;
[0033] Construct a recursive neural tensor network, which includes an orthogonal tensor decomposition unit and a non - negative tensor decomposition unit. The orthogonal tensor decomposition unit performs local feature decomposition on the fused feature tensor, and the non - negative tensor decomposition unit performs global feature decomposition on the fused feature tensor; input the local feature decomposition result and the global feature decomposition result into a preset tensor recursive unit, and establish the correlation relationship between features through recursive feature transfer;
[0034] Construct a tensor reconstruction module based on spectral decomposition, use nuclear norm regularization constraint to reconstruct features, generate an environmental factor correlation tensor, and generate an environmental digital twin model based on the environmental factor correlation tensor.
[0035] In an alternative embodiment,
[0036] Construct multiple scene feature vector subspaces based on scene image data. Each scene feature vector subspace corresponds to a type of scene type information. Select the corresponding scene feature vector subspace according to the scene type information, and map the scene image data to the selected scene feature vector subspace to generate target scene feature vectors, including:
[0037] Input the scene image data into a multi - layer convolutional neural network to extract multi - scale feature data, and perform an attention mechanism operation on the multi - scale feature data to obtain attention feature data;
[0038] Input the attention feature data into a variational auto - encoder network to obtain high - dimensional feature data, and construct the multiple scene feature vector subspaces according to the high - dimensional feature data;
[0039] Perform manifold embedding operation on the multiple scene feature vector subspaces to obtain manifold feature data, and input the manifold feature data into a Riemannian optimization network to obtain multiple feature distribution data;
[0040] Input the multiple feature distribution data into a Dirichlet mixture model to obtain multiple probability parameters, construct distance feature data based on the multiple probability parameters, perform spectral decomposition operation on the distance feature data to obtain feature complexity data, and establish the correspondence between the scene type information and the multiple scene feature vector subspaces according to the feature complexity data;
[0041] Select the corresponding target scene feature vector subspace from the multiple scene feature vector subspaces according to the scene type information, perform Markov chain sampling on the target scene feature vector subspace to obtain the optimal feature subspace, and generate the target scene feature vector.
[0042] In an alternative implementation,
[0043] Input the multiple feature distribution data into the Dirichlet mixture model to obtain multiple probability parameters, construct distance feature data based on the multiple probability parameters, perform spectral decomposition operation on the distance feature data to obtain feature complexity data, and establish the correspondence between the scene type information and the multiple scene feature vector subspaces according to the feature complexity data, including:
[0044] Perform annealing Monte Carlo sampling on the multiple feature distribution data to obtain initial sampling data, input the initial sampling data into the Hamiltonian dynamics model to construct a potential energy function, solve the phase space equations for the potential energy function to obtain phase distribution data, and perform importance sampling on the initial sampling data based on the phase distribution data to obtain sampling feature data;
[0045] Input the sampling feature data into the Dirichlet mixture model to obtain multiple probability parameters, construct a direction distribution matrix according to the multiple probability parameters, and calculate the geodesic distance and the Riemannian distance for the direction distribution matrix respectively to obtain distance feature data;
[0046] Construct a Riemannian metric tensor based on the distance feature data, map the Riemannian metric tensor to the dual space to obtain a dual feature tensor, perform high-order tensor decomposition on the dual feature tensor to obtain eigenvalue data and eigenvector data, and input the eigenvalue data and the eigenvector data into a quantum neural network to obtain feature complexity data;
[0047] Construct a potential function of the conditional random field model based on the feature complexity data, input the scene type information into the conditional random field model for variational inference to obtain a marginal probability distribution, perform variational approximation based on the marginal probability distribution to obtain a probability mapping matrix, and establish the correspondence between the scene type information and the multiple scene feature vector subspaces based on the probability mapping matrix.
[0048] In an alternative implementation,
[0049] Obtain enhanced feature data and a space mapping matrix based on the hyperbolic space manifold features and projection position information constructed by the optimized feature vectors, and generate scene rendering parameters using a preset parameter mapping subnetwork, including:
[0050] Input the optimized feature vector into the feature decomposition network to obtain a multi-dimensional feature tensor, and perform tensor decomposition operations on the multi-dimensional feature tensor to obtain feature component data; input the projection position information into the spatial transformation network to obtain a position feature vector, and perform manifold mapping operations on the position feature vector to obtain a spatial mapping matrix;
[0051] Map the feature component data to the hyperbolic space to obtain initial feature data, perform Poisson manifold embedding operations on the initial feature data to construct a curvature weight function, and perform non-linear transformations on the initial feature data based on the curvature weight function to obtain enhanced feature data;
[0052] Input the enhanced feature data and the spatial mapping matrix into the parameter mapping sub-network, construct a fine-grained decomposer in the parameter mapping sub-network to obtain a multi-scale feature tensor, and establish a feature pyramid network for the multi-scale feature tensor to obtain hierarchical feature data;
[0053] Perform self-attention mechanism operations on the hierarchical feature data to obtain an attention weight matrix, construct a recurrent neural network based on the attention weight matrix to obtain time series feature data, and input the time series feature data into the generative adversarial network to obtain an adversarial feature vector;
[0054] Input the adversarial feature vector into the variational autoencoder for feature reconstruction to obtain reconstructed feature data, perform residual optimization on the reconstructed feature data to obtain a residual feature vector, and generate scene rendering parameters based on the residual feature vector.
[0055] In the second aspect of the embodiments of the present invention, a scene adaptive adjustment system for intelligent IoT and metaverse virtual-real fusion is provided, including:
[0056] The first unit is used to obtain scene image data and environmental parameter data; perform resolution decomposition and block division on the scene image data, and perform scene type recognition through a preset scene recognition model to obtain scene type information;
[0057] The second unit is used to construct a multi-dimensional environmental parameter tensor based on the environmental parameter data, establish an environmental digital twin model, and map the environmental digital twin model to the virtual space using an adaptive coordinate mapping network;
[0058] The third unit is used to construct multiple scene feature vector subspaces based on the scene image data, each scene feature vector subspace corresponding to a scene type information, select the corresponding scene feature vector subspace according to the scene type information, map the scene image data to the selected scene feature vector subspace, and generate a target scene feature vector;
[0059] A fourth unit, configured to obtain the projection position information of a virtual object in a virtual space in a physical space based on an environmental digital twin model; input the target scene feature vector into a preset feature optimization sub-network for iterative optimization to obtain an optimized feature vector;
[0060] A fifth unit, configured to obtain enhanced feature data and a spatial mapping matrix based on the hyperbolic space manifold features constructed from the optimized feature vector and the projection position information, and generate scene rendering parameters by using a preset parameter mapping sub-network; perform rendering processing on the virtual object based on the scene rendering parameters, and superimpose and display the rendered virtual object on the physical space display interface.
[0061] In a third aspect of the embodiments of the present invention, there is provided an electronic device, including:
[0062] A processor; a memory for storing instructions executable by the processor, wherein the processor is configured to call the instructions stored in the memory to execute the foregoing method.
[0063] In a fourth aspect of the embodiments of the present invention, there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the foregoing method is implemented.
[0064] In the present invention, by acquiring scene image data and environmental parameter data, the precise fusion of virtual objects and physical scenes is realized, effectively solving the problem of disharmony between virtual objects and physical environments in traditional virtual-real fusion technologies, significantly enhancing the user's immersion and interaction experience. By adopting multi-dimensional environmental parameter tensors and environmental digital twin models, combined with an adaptive coordinate mapping network, environmental changes are accurately captured and the display effects of virtual objects are adjusted in real time, improving the adaptability and realism of virtual objects under different environmental conditions, enabling virtual elements to naturally integrate into various complex scenes. By constructing a scene feature vector subspace and hyperbolic space manifold features, and using a feature optimization sub-network for iterative optimization, the precise generation and dynamic adjustment of scene rendering parameters are realized, greatly enhancing the real-time response ability and rendering effect of the system, providing high-quality virtual-real fusion technology support for intelligent IoT and metaverse applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] Figure 1 It is a schematic flowchart of the method for scene adaptive adjustment of virtual-real fusion of intelligent IoT and metaverse in the embodiments of the present invention;
[0066] Figure 2 It is a relationship diagram between the number of tensor decomposition factors and the model performance of the method for scene adaptive adjustment of virtual-real fusion of intelligent IoT and metaverse in the embodiments of the present invention;
[0067] Figure 3It is a performance and quality comparison chart of distance feature data for the scene adaptive adjustment method of intelligent IoT and the virtual-real fusion of the metaverse in the embodiments of the present invention;
[0068] Figure 4 It is a schematic structural diagram of the scene adaptive adjustment system of intelligent IoT and the virtual-real fusion of the metaverse in the embodiments of the present invention. Detailed implementation manners
[0069] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part rather than all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0070] The technical solutions of the present invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0071] Figure 1 It is a schematic flowchart of the scene adaptive adjustment method of intelligent IoT and the virtual-real fusion of the metaverse in the embodiments of the present invention. As Figure 1 shown, the method includes:
[0072] Obtain scene image data and environmental parameter data; perform resolution decomposition and block division on the scene image data, and perform scene type recognition through a preset scene recognition model to obtain scene type information;
[0073] Construct a multi-dimensional environmental parameter tensor based on the environmental parameter data, establish an environmental digital twin model, and map the environmental digital twin model to the virtual space by using an adaptive coordinate mapping network;
[0074] Construct multiple scene feature vector subspaces based on the scene image data, each scene feature vector subspace corresponding to a kind of scene type information, select the corresponding scene feature vector subspace according to the scene type information, and map the scene image data to the selected scene feature vector subspace to generate a target scene feature vector;
[0075] Obtain the projection position information of virtual objects in the virtual space in the physical space based on the environmental digital twin model; input the target scene feature vector into a preset feature optimization sub-network for iterative optimization to obtain an optimized feature vector;
[0076] Based on the hyperbolic space manifold features constructed from the optimized feature vectors and the projection position information, enhanced feature data and a spatial mapping matrix are obtained, and a scene rendering parameter is generated using a preset parameter mapping sub-network; based on the scene rendering parameter, a virtual object is rendered, and the rendered virtual object is superimposed and displayed on the entity space display interface. In an alternative embodiment,
[0077] The resolution of the scene image data is decomposed and segmented, and the scene type is recognized through a preset scene recognition model to obtain scene type information including:
[0078] Perform hierarchical perception on the scene image data to generate multiple scene image layers with different resolutions;
[0079] Each of the scene image layers is adaptively segmented, and the optimal segmentation size of each scene image layer is determined based on the information entropy of the image blocks to obtain corresponding image block sequences;
[0080] Construct a discrete wavelet transform pyramid for each of the image block sequences to extract multi-scale frequency domain features;
[0081] Input the multi-scale frequency domain features into a preset heterogeneous convolutional neural network, and the heterogeneous convolutional neural network uses different convolutional kernel sizes for feature extraction of frequency domain features at different scales to obtain image block feature vectors;
[0082] Input the image block feature vectors into a preset self-calibrating feature enhancement network, and iteratively optimize the feature vectors based on a feedback mechanism to generate enhanced feature vectors;
[0083] Recombine the enhanced feature vectors according to adaptive attention weights to generate a multi-scale feature sequence; input the multi-scale feature sequence into a preset two-stream neural network, the spatial stream branch of the two-stream neural network extracts the spatial features of the multi-scale feature sequence, and the temporal stream branch of the two-stream neural network extracts the temporal features of the multi-scale feature sequence; fuse the spatial features and the temporal features to generate a fused feature vector;
[0084] Input the fused feature vector into a preset hierarchical scene classifier, and the hierarchical scene classifier performs progressive classification on the fused feature vector based on the scene semantic hierarchical relationship to obtain scene type information.
[0085] Perform hierarchical perception on the scene image data. Using the Gaussian pyramid downsampling method, the scene image with an original resolution of 1920×1080 pixels is successively downsampled to generate multiple scene image layers with different resolutions including the original resolution. For example, four scene image layers with different resolutions of 1920×1080, 960×540, 480×270, and 240×135 are generated. This can capture the visual information of the scene from different scales and contribute to the accurate recognition of subsequent scene types.
[0086] Perform adaptive block division on each scene image layer. For each scene image layer, calculate the information entropy distribution of the image and determine the optimal block size based on the information entropy threshold. For example, for the 1920×1080 image layer, a block size of 64×64 pixels is used for regions with higher information entropy (such as values greater than 5.5), a block size of 128×128 pixels is used for regions with medium information entropy (such as values between 3.5 - 5.5), and a block size of 256×256 pixels is used for regions with lower information entropy (such as values less than 3.5). This adaptive block division strategy can perform fine-grained block division on information-rich regions and rough block division on regions with less information, thereby improving the calculation efficiency and the accuracy of feature extraction. After completing the block division, the corresponding sequence of image blocks is obtained.
[0087] Construct a discrete wavelet transform pyramid for each sequence of image blocks. Use the Haar wavelet basis function to perform discrete wavelet transform on each image block, decompose it into 3 levels, obtain the low-frequency approximation coefficients (LL) and high-frequency detail coefficients in three directions (LH, HL, HH), and extract multi-scale frequency domain features. For example, for a 64×64 image block, after 3-level wavelet decomposition, 8×8 LL3 coefficients, 8×8 LH3 / HL3 / HH3 coefficients, 16×16 LH2 / HL2 / HH2 coefficients, and 32×32 LH1 / HL1 / HH1 coefficients can be obtained. These coefficients constitute the multi-scale frequency domain features of the image block and contain key information such as the texture, edges, and structure of the image.
[0088] Input the multi-scale frequency domain features into a preset heterogeneous convolutional neural network. The heterogeneous convolutional neural network designs convolutional layers with different convolutional kernel sizes for frequency domain features of different scales. For the LL3 low-frequency coefficients, a 7×7 convolutional kernel is used for feature extraction. For the LH3 / HL3 / HH3 high-frequency coefficients, a 5×5 convolutional kernel is used for feature extraction. For the LH2 / HL2 / HH2 and LH1 / HL1 / HH1 coefficients, 3×3 and 1×1 convolutional kernels are used for feature extraction respectively. Each scale of convolutional layer contains 32 convolutional kernels, with a stride of 1 and a padding method of SAME. After convolutional operations, the feature maps are processed using batch normalization and the ReLU activation function, and the features of different scales are connected through a fully connected layer to finally obtain a 256-dimensional image block feature vector.
[0089] Input the image patch feature vectors into a preset self-calibrating feature enhancement network. Adopt a residual learning structure and an attention mechanism to iteratively optimize the feature vectors. Reduce the dimension of the feature vectors through a fully connected layer of 256×128. Calculate the importance weights of each channel through a channel attention module, with the weight range being 0 - 1, and weight the features. Optimize the features through three residual blocks (each layer contains a fully connected layer of 128×128, batch normalization, and ReLU activation), and restore the feature dimension through a fully connected layer of 128×256. Repeat the entire optimization process 2 times to generate enhanced feature vectors. In the example, the mean squared error of the original features is 0.085, which drops to 0.062 after one optimization and further drops to 0.043 after two optimizations, indicating a significant improvement in the feature quality.
[0090] Recombine the enhanced feature vectors according to the adaptive attention weights to generate a multi-scale feature sequence. Specifically, calculate the attention weights of the image patches at different resolution levels. The weight calculation is based on the L2 norm of the feature vectors and is processed through softmax normalization. For example, in a scenario with four resolution levels, the average attention weights of each layer are 0.35, 0.28, 0.22, and 0.15 respectively. Weight and fuse the features of each layer according to these weights to form a multi-scale feature sequence representing temporal and spatial relationships.
[0091] Input the multi-scale feature sequence into a preset two-stream neural network. This network includes two branches: a spatial stream and a temporal stream. The spatial stream branch uses 5 layers of 3×3 convolutional layers (64 convolutional kernels in each layer) to extract spatial features, and the temporal stream branch uses 3 layers of LSTM units (128 hidden units in each layer) to extract temporal features. The two branches respectively output feature vectors of 512 dimensions, and perform feature fusion through feature concatenation and 1×1 convolution to generate a fused feature vector of 1024 dimensions.
[0092] Input the fused feature vector into a preset hierarchical scene classifier. This classifier adopts a three-level classification structure. The first level distinguishes indoor / outdoor scenes, the second level distinguishes specific scene categories (such as indoor being divided into home, office, commercial, etc.; outdoor being divided into natural, urban, etc.), and the third level identifies specific scene types (such as living room, bedroom, office, etc.). Each level of the classifier is implemented using a fully connected layer, with specific configurations being 1024×512, 512×256, and 256×N (N is the number of categories).
[0093] In this embodiment, Gaussian pyramids are used to generate scene images at multiple resolution levels, which can not only capture the global structural features of the scene, but also retain important local detail information, laying a good foundation for subsequent feature extraction. The pyramid structure is constructed through discrete wavelet transform to extract multi-scale frequency domain features of image patches, effectively expressing the texture, edge and structure information of the image. Adaptive attention weights are used to reorganize the multi-scale features, and spatial features and temporal features are respectively extracted through a two-stream neural network to achieve multi-dimensional fusion of features. The hierarchical scene classifier performs progressive classification based on the semantic hierarchical relationship of the scene, effectively reducing the classification difficulty and improving the recognition accuracy.
[0094] In an alternative embodiment,
[0095] Based on the environmental parameter data, a multi-dimensional environmental parameter tensor is constructed, an environmental digital twin model is established, and the environmental digital twin model is mapped to the virtual space by using an adaptive coordinate mapping network, including:
[0096] Align the environmental parameter data based on the spatio-temporal synchronization mechanism, and construct a multi-dimensional environmental parameter tensor;
[0097] Perform adaptive wavelet packet decomposition on the multi-dimensional environmental parameter tensor to generate a multi-scale environmental feature spectrum;
[0098] Input the multi-scale environmental feature spectrum into a preset spectral embedding network, extract the correlation features between parameters through the multi-head attention mechanism, and generate a high-dimensional environmental feature vector;
[0099] Construct a hybrid memory neural network, input the high-dimensional environmental feature vector into the hybrid memory neural network, store long-term dependence information based on the external memory matrix, and combine the short-term dynamic memory unit to capture transient change features to generate environmental dynamic prediction data;
[0100] Construct a recursive neural tensor network, generate a fusion feature tensor based on the high-dimensional environmental feature vector and the environmental dynamic prediction data, and input it into the recursive neural tensor network. The recursive neural tensor network uses tensor decomposition to perform multi-dimensional correlation analysis on the input data, establish a dynamic coupling relationship between environmental elements, and generate an environmental digital twin model;
[0101] Construct an adaptive coordinate mapping network, which includes a spatial transformation sub-network and a deformation compensation sub-network. Map the topological structure of the environmental digital twin model to the virtual space coordinate system through the spatial transformation sub-network, and use the deformation compensation sub-network to dynamically correct the geometric deformation during the mapping process, and map the corrected environmental digital twin model to the rendering coordinate system of the virtual space.
[0102] Align the environmental parameter data based on the spatio-temporal synchronization mechanism, construct a multi-dimensional environmental parameter tensor. By configuring a clock synchronization module on each acquisition node and adopting the Network Time Protocol (NTP) to achieve millisecond-level time synchronization, sort the multi-dimensional environmental parameters such as temperature (15°C - 30°C), humidity (40% - 80%), light intensity (500 - 2000 lux), air pressure (990 - 1030 hPa), PM2.5 (0 - 75 μg / m³), etc. collected at different time points according to the timestamp (accurate to milliseconds), aggregate the time series at 1-minute intervals, and complete the missing data by the nearest neighbor value averaging interpolation method to form a three-dimensional tensor of size (144×5×N), where 144 represents the number of time points sampled per minute in a day, 5 represents the environmental parameter dimension, and N represents the number of monitoring spatial points.
[0103] Perform adaptive wavelet packet decomposition on the multi-dimensional environmental parameter tensor to generate a multi-scale environmental feature spectrum. Divide the environmental parameter tensor into 4 sub-tensors in the time dimension at 6-hour intervals, and perform 6-layer wavelet packet decomposition on each sub-tensor using the Daubechies wavelet basis function. During the decomposition process, dynamically select the optimal decomposition path through the Shannon entropy criterion, retain the nodes with information content greater than the threshold of 0.85, and discard the noise nodes. For the temperature parameter, focus on retaining the frequency band information of 0.1 - 0.5 Hz; for the humidity parameter, focus on retaining the frequency band information of 0.05 - 0.2 Hz; for the light intensity parameter, focus on retaining the frequency band information of 0.2 - 1 Hz. Form a spectral feature matrix with a size of (64×5×N), where 64 represents the number of retained wavelet packet coefficients.
[0104] Input the multi-scale environmental feature spectrum into a preset graph embedding network, extract the correlation features between parameters through the multi-head attention mechanism, and generate a high-dimensional environmental feature vector. The graph embedding network contains 3 attention heads, the hidden layer dimension of each attention head is 128, and the input layer is a 64×5-dimensional feature matrix. Each attention head calculates the correlation weights between parameters from different angles. For example, the first head focuses on the temperature-humidity relationship, the second head focuses on the light intensity-temperature relationship, and the third head focuses on the air pressure-PM2.5 relationship. Through the kernel attention mechanism, use the radial basis kernel function to calculate the feature similarity, set the kernel width parameter to 0.75, and connect the outputs of the 3 attention heads through a linear transformation layer to form a 384-dimensional high-dimensional environmental feature vector.
[0105] Construct a hybrid memory neural network. Input the high-dimensional environmental feature vector into the hybrid memory neural network. Based on the external memory matrix, store the long-term dependence information, and combine the short-term dynamic memory unit to capture the transient change features, and generate the environmental dynamic prediction data. The network includes an external memory matrix of size (256×256) for storing historical environmental feature information. Set the forgetting rate to 0.05 to ensure that the long-term environmental patterns are retained. The short-term dynamic memory unit is implemented using a gated recurrent unit (GRU) with 128 units. Among them, the reset gate threshold is set to 0.3, and the update gate threshold is set to 0.7. When inputting a 384-dimensional feature vector, calculate the similarity with the external memory through the attention controller, extract the relevant historical information, fuse it with the state vector of the short-term memory unit, predict the change trend of the environmental parameters in the next 30 minutes, and output the environmental dynamic prediction data of size (30×5×N).
[0106] Construct a recurrent neural tensor network. Generate a fusion feature tensor based on the high-dimensional environmental feature vector and the environmental dynamic prediction data and input it into the recurrent neural tensor network to establish the dynamic coupling relationship between environmental elements and generate the environmental digital twin model. Adopt a third-order tensor decomposition structure, splice the 384-dimensional feature vector and the (30×5×N)-dimensional prediction data to form an input tensor. Set the core tensor dimension to (64×64×64), and through 3 projection matrices corresponding to the time, parameter type, and space dimensions respectively, with the matrix sizes being (64×30), (64×5), and (64×N). The recurrent layer adopts a tensor decomposition form, with the number of iterations set to 20, and the convergence threshold for each iteration is 0.001. The output is a four-dimensional tensor of size (30×5×N×M), where M = 8 represents the number of state variables of each environmental element, including statistical features such as mean, variance, and change rate, constituting a complete environmental digital twin model.
[0107] Construct an adaptive coordinate mapping network to map the environmental digital twin model into the virtual space, including a spatial transformation sub-network and a deformation compensation sub-network. The spatial transformation sub-network adopts a 3-layer convolutional structure, with the convolutional kernel sizes being 5×5, 3×3, and 3×3 respectively, and the number of channels being 32, 64, and 128. After extracting the spatial features, output an affine transformation matrix of 9 parameters to map the physical space coordinates (x, y, z) to the virtual space coordinates (x', y', z'). The deformation compensation sub-network adopts a thin plate spline model with 20 control points to correct the non-linear deformation generated during the mapping process, with the correction accuracy better than 0.05 meters. Map the corrected environmental digital twin model to the rendering coordinate system of the virtual space through the OpenGL shader interface to achieve the accurate mapping from the physical environment to the virtual environment.
[0108] In this embodiment, the alignment of environmental parameters and tensor construction are achieved through a spatio-temporal synchronization mechanism, effectively solving the problem of unified expression of multi-source heterogeneous data. The graph embedding network uses a multi-head attention mechanism to deeply explore the complex correlation relationships between environmental parameters. It can not only capture the direct correlations between parameters but also discover potential indirect action mechanisms, enhancing the integrity and accuracy of environmental feature expression. The design of the hybrid memory neural network cleverly combines long-term and short-term memory mechanisms, which can not only retain historical information on environmental changes but also quickly respond to transient changes in the environment, significantly improving the temporal continuity and prediction accuracy of environmental dynamic prediction. The tensor-based modeling method breaks through the limitations of traditional linear analysis and can more realistically reflect the complexity and non-linear characteristics of the environmental system.
[0109] In an alternative embodiment,
[0110] Construct a recurrent neural tensor network, generate a fused feature tensor based on the high-dimensional environmental feature vector and the environmental dynamic prediction data, and input the fused feature tensor into the recurrent neural tensor network. The recurrent neural tensor network uses tensor decomposition to perform multi-dimensional correlation analysis on the input data, establish dynamic coupling relationships between environmental elements, and generate an environmental digital twin model, including:
[0111] Construct a Riemannian manifold mapping for the high-dimensional environmental feature vector to generate manifold features, and project the environmental dynamic prediction data into the curvature tensor space to generate tensor features;
[0112] Input the manifold features and the tensor features into a preset Riemannian optimization model, construct an optimization objective function based on geodesic distance metrics, and use the conjugate gradient method with an adaptive step size to iteratively solve the optimization objective function to obtain optimized features;
[0113] Project the optimized features into the Euclidean space through a Riemannian exponential mapping to generate a fused feature tensor;
[0114] Construct a recurrent neural tensor network, which includes an orthogonal tensor decomposition unit and a non-negative tensor decomposition unit. The orthogonal tensor decomposition unit performs local feature decomposition on the fused feature tensor, and the non-negative tensor decomposition unit performs global feature decomposition on the fused feature tensor; input the local feature decomposition result and the global feature decomposition result into a preset tensor recurrence unit, and establish the correlation relationship between features through recursive feature transfer;
[0115] Construct a tensor reconstruction module based on spectral decomposition, reconstruct the features using nuclear norm regularization constraints, generate an environmental element correlation tensor, and generate an environmental digital twin model based on the environmental element correlation tensor.
[0116] Obtain high-dimensional environmental feature vectors and environmental dynamic prediction data. The high-dimensional environmental feature vectors include multi-dimensional environmental parameters such as temperature, humidity, air pressure, wind speed, PM2.5 concentration, etc. For example, the environmental feature vector collected at a monitoring point in a certain city is [25.3°C, 65%, 101.3 kPa, 3.2 m / s, 35 μg / m³]; the environmental dynamic prediction data contains the change trend data of environmental elements in the next 24 hours, such as the temperature change prediction sequence [24.8°C, 24.5°C,..., 26.2°C].
[0117] When constructing a Riemannian manifold mapping for the high-dimensional environmental feature vectors to generate manifold features, the locally linear embedding algorithm is used to process the high-dimensional environmental feature vectors. For each feature point, its K nearest neighbors (K = 10) are selected to construct a local geometric structure, and the high-dimensional features are mapped to a low-dimensional manifold space by solving the weight matrix, forming a manifold feature vector with a dimension of 64. For example, the manifold features obtained after mapping the original 10-dimensional environmental feature vector can be represented as a 64-dimensional vector [0.31, 0.25, -0.15,..., 0.42].
[0118] When projecting the environmental dynamic prediction data into the curvature tensor space to generate tensor features, a 3rd-order prediction data tensor (time × element type × predicted value) is constructed, and the curvature characteristics in the tensor space are calculated through Riemannian metrics to obtain tensor features reflecting the internal geometric structure of the data. For example, the prediction data of 5 environmental elements in 24 hours can construct a 24×5×3 tensor, and after projection, a tensor feature vector with a dimension of 96 is obtained [0.18, -0.23, 0.37,..., -0.12].
[0119] When fusing features in the Riemannian optimization model, the manifold features and tensor features are used as inputs, and an optimization objective function is constructed based on the geodesic distance. The optimization objective function calculates the distance between the two features on the Riemannian manifold and minimizes it. The conjugate gradient method with an adaptive step size is used for iterative solution. The number of iterations is set to 200, and the convergence threshold is 0.001. The step size is adaptively adjusted according to the gradient direction in each iteration (initial step size 0.01). After the iteration is completed, an optimized feature vector with a dimension of 128 is obtained [0.28, 0.15, -0.22,..., 0.36].
[0120] When projecting the optimized features into the Euclidean space through the Riemannian exponential mapping, a mapping function is defined in the tangent space, and the feature points on the Riemannian manifold are mapped to the Euclidean space along the geodesic direction. The optimized features are transformed and reorganized to form a 4×4×8-dimensional fusion feature tensor, and the numerical range is normalized to the interval [-1, 1].
[0121] When constructing a recursive neural tensor network, a network architecture containing an orthogonal tensor decomposition unit and a non-negative tensor decomposition unit was designed. The orthogonal tensor decomposition unit performs local feature decomposition on the fused feature tensor to maintain the orthogonality between features, and the number of decomposition factors is set to 16; the non-negative tensor decomposition unit is applied to global feature extraction, and non-negative constraints are adopted to ensure the interpretability of the decomposition results, and the number of decomposition factors is set to 32. Taking urban environmental monitoring as an example, local feature decomposition can identify the correlation patterns of temperature and humidity in a specific area, and the numerical representation is [0.42, 0.35,..., -0.18]; global feature decomposition extracts the overall pollutant diffusion law of the city, and the numerical representation is [0.52, 0.38,..., 0.27].
[0122] When inputting the local and global feature decomposition results into a preset tensor recursive unit, a tensor contraction operation is used to establish a recursive transfer relationship between multi-dimensional features. The internal state dimension of the recursive unit is set to 256, a bidirectional information flow mechanism is used to fuse temporal information, and the recursive depth is 3 layers. For example, for air temperature and air quality data, the recursive unit captures the temporal correlation that an increase in temperature leads to an intensification of photochemical reactions and thus affects ozone concentration.
[0123] When constructing a tensor reconstruction module based on spectral decomposition, the singular value thresholding method is used for tensor reconstruction, and the threshold parameter is set to 0.15. At the same time, nuclear norm regularization constraints are introduced (regularization coefficient λ = 0.01) to suppress the influence of noise and improve the stability of the model. The dimension of the environmental factor correlation tensor generated after reconstruction is 10×10×10, where each element represents the correlation strength between different environmental factors, and the value range is [0, 1]. For example, the correlation strength between temperature and ozone concentration is 0.87, indicating a strong correlation; the correlation strength between temperature and PM2.5 is 0.35, indicating a medium correlation.
[0124] An environmental digital twin model is generated based on the environmental factor correlation tensor, which includes an environmental factor attribute layer, a correlation dynamic layer, and an evolution prediction layer. The dynamic coupling relationship between environmental factors is displayed through a visualization interface to achieve digital mapping of complex environmental systems. Taking urban air quality management as an example, the model can predict that when the temperature rises by 2°C and the relative humidity drops by 15%, the ozone concentration will increase by about 28%, providing a scientific basis for environmental governance decisions.
[0125] In this embodiment, a Riemannian manifold mapping mechanism is introduced. By mapping high-dimensional environmental feature vectors to the manifold space, the intrinsic geometric structure of the data is maintained, effectively overcoming the problem of information loss caused by traditional dimensionality reduction methods. An optimization objective function is constructed through geodesic distance metrics and solved using the conjugate gradient method with an adaptive step size, realizing the adaptive fusion of manifold features and tensor features. Precise extraction of local features is achieved through orthogonal tensor decomposition, and the interpretability of global features is ensured through non-negative tensor decomposition. The combination of these two decomposition methods significantly improves the model's expressive ability for environmental systems;
[0126] In the prior art, the construction of environmental digital twin models usually adopts simple feature splicing and linear mapping methods, which are difficult to effectively express the non-linear correlation and dynamic coupling characteristics between environmental elements. When dealing with high-dimensional environmental data, dimensionality reduction is often used, which easily loses important feature information and affects the accuracy and reliability of the model. The prior art generally adopts a fixed-weight fusion strategy in feature fusion and spatial mapping, lacking adaptability and being difficult to cope with complex and changing environmental systems;
[0127] Based on the feature mapping and fusion mechanism of the Riemannian manifold, this embodiment significantly improves the expression accuracy of environmental features, maintains the geometric structure information of the data, and the dual decomposition structure of the recurrent neural tensor network realizes multi-scale modeling of the correlation relationship between environmental elements, enhancing the model's expressive ability. The introduction of the adaptive optimization strategy improves the model's adaptability and robustness to environmental changes. While maintaining the accuracy of the model, it also has good interpretability and practicality, providing a more reliable basis for environmental management decisions.
[0128] Figure 2 This is a graph showing the relationship between the number of tensor decomposition factors and the model performance of the scenario adaptive adjustment method for intelligent IoT and metaverse virtual-real fusion in the embodiment of the present invention, demonstrating the impact of different numbers of tensor decomposition factors on the model accuracy and training time. The graph presents the performance comparison of three different methods: orthogonal tensor decomposition (represented by diamond symbols), non-negative tensor decomposition (represented by circular symbols), and the technical solution of the present invention (represented by solid triangle symbols). The dotted line and square symbols represent the change trend of training time.
[0129] When the number of decomposition factors is 4, the prediction accuracy of the technical solution of the present invention reaches 73.8%, while the accuracies of orthogonal tensor decomposition and non-negative tensor decomposition are 71.2% and 72.3% respectively. As the number of factors increases, the performance gap between the three methods gradually expands. When the number of factors is 32 (marked as "optimal balance point" in the graph), the prediction accuracy of the technical solution of the present invention reaches 89.7%, significantly higher than 81.3% of orthogonal tensor decomposition and 83.4% of non-negative tensor decomposition; when the number of factors increases to 256, the accuracy of the technical solution of the present invention further increases to 96.8%, while the other methods are 87.2% and 89.9% respectively.
[0130] Meanwhile, the training time increases significantly with the increase in the number of factors. As can be seen from the dashed line in the figure, when the number of factors is 4, the training time is only about 10 seconds. When the number of factors is 32, it is about 350 seconds. And when the number of factors increases to 256, the training time surges to about 2000 seconds. Especially after the number of factors exceeds 64, the computational cost increases sharply, but the improvement in accuracy slows down.
[0131] This technical solution realizes the efficient extraction of environmental data features by combining orthogonal and non - negative tensor decomposition techniques, and can obtain high prediction accuracy with a relatively small number of factors. Especially between the number of factors of 16 and 64, the accuracy improvement of this technical solution is the most significant, rapidly increasing from 85.1% to 93.2%. This interval is also the best balance area between performance and computational cost.
[0132] In contrast, the accuracy improvement of orthogonal tensor decomposition and non - negative tensor decomposition is relatively slow when increasing the number of factors, and almost saturates after the number of factors exceeds 128. The marginal benefit brought by further increasing the number of factors is extremely small. It can be observed from the figure that non - negative tensor decomposition is slightly better than orthogonal tensor decomposition under high - factor conditions, but the gap between both of them and this technical solution is still obvious.
[0133] In the figure, a best balance point is clearly marked when the number of decomposition factors is 32. At this time, this technical solution can maintain a relatively high prediction accuracy (89.7%) while maintaining a relatively low training cost (about 350 seconds). This discovery has important guiding significance for the implementation of actual systems and can help decision - makers select the best model configuration parameters.
[0134] Generally speaking, this technical solution utilizes both the local feature extraction ability of orthogonal tensor decomposition and the global feature representation ability of non - negative tensor decomposition, and achieves optimal performance under various factor number configurations. Especially in the case of limited computing resources, its efficiency is more obvious, greatly improving the practical value of the algorithm in actual environmental monitoring and prediction.
[0135] In an alternative implementation,
[0136] Construct multiple scene feature vector sub - spaces based on the scene image data. Each scene feature vector sub - space corresponds to a type of scene type information. Select the corresponding scene feature vector sub - space according to the scene type information, and map the scene image data to the selected scene feature vector sub - space. The generated target scene feature vector includes:
[0137] Input the scene image data into a multi - layer convolutional neural network to extract multi - scale feature data, and perform an attention mechanism operation on the multi - scale feature data to obtain attention feature data;
[0138] Input the attention feature data into a variational autoencoder network to obtain high-dimensional feature data, and construct the multiple scene feature vector subspaces according to the high-dimensional feature data;
[0139] Perform a manifold embedding operation on the multiple scene feature vector subspaces to obtain manifold feature data, and input the manifold feature data into a Riemannian optimization network to obtain multiple feature distribution data;
[0140] Input the multiple feature distribution data into a Dirichlet mixture model to obtain multiple probability parameters, construct distance feature data based on the multiple probability parameters, perform a spectral decomposition operation based on the distance feature data to obtain feature complexity data, and establish the correspondence between the scene type information and the multiple scene feature vector subspaces according to the feature complexity data;
[0141] Select the corresponding target scene feature vector subspace from the multiple scene feature vector subspaces according to the scene type information, perform a Markov chain sampling on the target scene feature vector subspace to obtain an optimal feature subspace, and generate a target scene feature vector.
[0142] Input the scene image data into a multi-layer convolutional neural network to extract multi-scale feature data. A convolutional neural network with five convolutional layers is used, and the number of convolutional kernels in each convolutional layer is 64, 128, 256, 512, and 512 respectively. The input scene image data is a 224×224×3 RGB image. After the first layer of convolution, a feature map of 112×112×64 is obtained. After the second layer of convolution, a feature map of 56×56×128 is obtained. After the third layer of convolution, a feature map of 28×28×256 is obtained. After the fourth layer of convolution, a feature map of 14×14×512 is obtained. After the fifth layer of convolution, a feature map of 7×7×512 is obtained. These five feature maps with different scales constitute the multi-scale feature data.
[0143] Perform an attention mechanism operation on the obtained multi-scale feature data, and apply a spatial attention mechanism and a channel attention mechanism to each feature map. Taking the third-layer feature map (28×28×256) as an example, the spatial attention mechanism calculates the importance weights of each spatial position to generate a 28×28 weight matrix, and multiplies this weight matrix with the original feature map; the channel attention mechanism calculates the importance weights of 256 channels to generate a 1×1×256 weight vector, and multiplies this vector with the original feature map. The results of the two attention mechanisms are obtained through weighted summation to obtain attention feature data, and the weight ratio is 1:1.
[0144] Input the attention feature data into a variational autoencoder network to obtain high-dimensional feature data. The variational autoencoder network consists of an encoder and a decoder. The encoder is composed of three fully connected layers with 4096, 2048, and 1024 neurons respectively, and the activation function is ReLU. The encoder outputs a mean vector and a variance vector, both of which are 512-dimensional. Through the reparameterization trick, a latent vector is sampled from the distribution defined by the mean and variance, and this latent vector is the high-dimensional feature data. The decoder is also composed of three fully connected layers with 1024, 2048, and the original feature dimension respectively, and is used to reconstruct the original features. During training, the weighted sum of the reconstruction loss and the KL divergence loss is used as the total loss function, and the weight ratio is 1:0.5.
[0145] Construct multiple scene feature vector subspaces based on the high-dimensional feature data, and use the K-means clustering algorithm to divide the high-dimensional feature data into K clusters, where each cluster represents a scene feature vector subspace. In this embodiment, the value of K is set to 10, indicating that 10 scene feature vector subspaces are constructed. For each cluster, calculate the mean vector and covariance matrix of all high-dimensional feature data within the cluster. The mean vector represents the center of the subspace, and the covariance matrix describes the distribution shape of the subspace.
[0146] Perform a manifold embedding operation on multiple scene feature vector subspaces to obtain manifold feature data. The t-SNE algorithm is used to reduce the 512-dimensional subspace representation to a 50-dimensional space, preserving the local structural relationship between data points. t-SNE first calculates the similarity between data points in the high-dimensional space and optimizes the positions of data points in the low-dimensional space so that the similarity between points in the low-dimensional space is as close as possible to that in the high-dimensional space. The resulting 50-dimensional representation after dimensionality reduction is the manifold feature data.
[0147] Input the manifold feature data into a Riemannian optimization network to obtain multiple feature distribution data. The Riemannian optimization network performs optimization on a Riemannian manifold, taking into account the geometric structure of the data. The network contains two layers of Riemannian convolutional layers, and each layer contains 32 Riemannian convolutional kernels. For each subspace, the network learns the optimal representation of the subspace on the Riemannian manifold and outputs 32-dimensional feature distribution data.
[0148] Input multiple feature distribution data into a Dirichlet mixture model to obtain multiple probability parameters. The Dirichlet mixture model represents each feature distribution as a mixture of multiple Dirichlet distributions. In this embodiment, a mixture model with 5 components is used, and each component corresponds to a Dirichlet distribution. For each feature distribution data, the model outputs 5 concentration parameter vectors, and each concentration parameter vector contains 32 elements. These parameters describe the probability characteristics of the feature distribution.
[0149] Construct distance feature data based on multiple probability parameters. Calculate the Jensen-Shannon divergence between different subspaces as the distance metric. For any two subspaces i and j, calculate the Jensen-Shannon divergence between their corresponding Dirichlet distributions to obtain a K×K distance matrix, which is the distance feature data.
[0150] Perform spectral decomposition operation on the distance feature data to obtain feature complexity data. Apply eigenvalue decomposition to the distance matrix to obtain its eigenvalues and eigenvectors. The eigenvalues represent the complexity of the subspaces, and the eigenvectors represent the main change directions of the subspaces. Select the first 5 largest eigenvalues and their corresponding eigenvectors to form the feature complexity data.
[0151] Establish the correspondence between the scene type information and multiple scene feature vector subspaces according to the feature complexity data. Analyze the component sizes of the eigenvectors to determine the most matching scene type for each subspace. For example, if subspace 1 has the largest component on the eigenvector related to the "indoor scene", then associate subspace 1 with the "indoor scene" type.
[0152] Select the corresponding target scene feature vector subspace from multiple scene feature vector subspaces according to the scene type information. Taking the input image as an "office scene" as an example, the system identifies its scene type as an "indoor scene" and selects the subspace associated with the "indoor scene" as the target scene feature vector subspace. Perform Markov chain sampling on the target scene feature vector subspace. Randomly select a point in the subspace as the initial point and perform 10,000 iterative samplings according to the transition probability matrix. The mean of the sampled point set is used as the center of the optimal feature subspace to generate a 512-dimensional target scene feature vector.
[0153] In this embodiment, the multi-layer convolutional structure captures scene features from low-level textures to high-level semantics layer by layer. The introduction of the attention mechanism highlights important feature regions, suppresses the interference of irrelevant information, and improves the pertinence and effectiveness of feature extraction. The feature extraction method based on the probabilistic generative model enhances the expression ability and generalization performance of the features, enabling the generated features to better reflect the essential attributes of the scene. Manifold embedding preserves the local structural relationship of the data, and Riemannian optimization takes into account the non-Euclidean nature of the feature space. This combination improves the expression accuracy and discrimination ability of the feature space.
[0154] In an alternative embodiment,
[0155] Input the multiple feature distribution data into a Dirichlet mixture model to obtain multiple probability parameters, construct distance feature data based on the multiple probability parameters, perform spectral decomposition operation on the distance feature data to obtain feature complexity data, and establish the corresponding relationship between the scenario type information and the multiple scenario feature vector subspaces according to the feature complexity data, including:
[0156] Perform annealing Monte Carlo sampling on the multiple feature distribution data to obtain initial sampling data, input the initial sampling data into a Hamiltonian dynamics model to construct a potential energy function, solve the phase space equations for the potential energy function to obtain phase distribution data, and perform importance sampling on the initial sampling data based on the phase distribution data to obtain sampling feature data;
[0157] Input the sampling feature data into a Dirichlet mixture model to obtain multiple probability parameters, construct a direction distribution matrix according to the multiple probability parameters, and calculate the geodesic distance and the Riemannian distance for the direction distribution matrix respectively to obtain distance feature data;
[0158] Construct a Riemannian metric tensor based on the distance feature data, map the Riemannian metric tensor to the dual space to obtain a dual feature tensor, perform high-order tensor decomposition on the dual feature tensor to obtain eigenvalue data and eigenvector data, and input the eigenvalue data and the eigenvector data into a quantum neural network to obtain feature complexity data;
[0159] Construct a potential energy function of a conditional random field model based on the feature complexity data, input the scenario type information into the conditional random field model for variational inference to obtain a marginal probability distribution, perform variational approximation based on the marginal probability distribution to obtain a probability mapping matrix, and establish the corresponding relationship between the scenario type information and the multiple scenario feature vector subspaces based on the probability mapping matrix.
[0160] Perform annealing Monte Carlo sampling on multiple feature distribution data to obtain initial sampling data. Set the initial temperature parameter to 1000, adopt a linear cooling strategy, and set the cooling rate to 0.95 for each round of iteration. Stop the iteration when the temperature drops to 0.01. At each temperature, the system randomly perturbs the current state and calculates the acceptance probability based on the energy difference and the current temperature. For example, for the feature distribution [0.2, 0.3, 0.5], after 1000 iterations of sampling, the obtained initial sampling data may be [[0.22, 0.28, 0.5], [0.19, 0.31, 0.5],...].
[0161] Input the initial sampling data into the Hamiltonian dynamics model to construct the potential energy function. The form of the potential energy function is expressed as a polynomial. For each sampling point, calculate the potential energy based on the position and momentum. For example, for the sampling point [0.22, 0.28, 0.5], the corresponding potential energy may be 13.5. Based on the constructed potential energy function, solve the phase space equations to obtain the phase distribution data. The solution process uses the fourth-order Runge-Kutta method, with a step size set to 0.01 and the number of iterations set to 100. The obtained phase distribution data represents the evolution trajectory of the system in the phase space.
[0162] Perform importance sampling on the initial sampling data based on the phase distribution data. Assign weights to each sample, and the weights are proportional to the density of the sample in the phase space. For example, for the initial sample [0.22, 0.28, 0.5], if its density in the phase space is high, a weight of 0.8 may be assigned; while for the sample [0.19, 0.31, 0.5] with a lower density, a weight of 0.3 may be assigned. Resample according to the weights to obtain the sampling feature data, such as [[0.22, 0.28, 0.5], [0.22, 0.28, 0.5], [0.21, 0.29, 0.5],...].
[0163] Input the sampling feature data into the Dirichlet mixture model to obtain multiple probability parameters. The Dirichlet mixture model is trained using the variational Bayesian method, with the number of components set to 5 and the convergence threshold set to 0.001. After training, the model outputs the weights of each component and the Dirichlet parameters. For example, the weights of 5 components [0.3, 0.25, 0.2, 0.15, 0.1] and the corresponding Dirichlet parameter matrix may be obtained.
[0164] Construct the direction distribution matrix based on multiple probability parameters. Each element in the direction distribution matrix represents the probability of transitioning from one state to another. For example, for a model with 5 components, the size of the constructed direction distribution matrix is 5×5, and the value at position (i, j) in the matrix represents the probability of transitioning from component i to component j.
[0165] Calculate the geodesic distance and the Riemannian distance for the direction distribution matrix respectively to obtain the distance feature data. The geodesic distance is calculated using the Dijkstra algorithm, with the path search depth set to 10. The Riemannian distance is calculated based on the eigenvalue decomposition of the direction distribution matrix. Finally, two 5×5 distance matrices are obtained as the distance feature data.
[0166] Construct the Riemannian metric tensor based on the distance feature data. The Riemannian metric tensor describes the local geometric structure in the feature space and is a fourth-order tensor, whose elements are jointly determined by the geodesic distance and the Riemannian distance. Map the Riemannian metric tensor to the dual space to obtain the dual feature tensor, and the mapping is achieved through the contraction operation.
[0167] Perform high - order tensor decomposition on the dual - feature tensor to obtain eigenvalue data and eigenvector data. The tensor decomposition uses the high - order singular value decomposition method, with the decomposition rank set to 10, the number of iterations to 100, and the convergence threshold to 0.0001. The decomposition results include 10 eigenvalues and the corresponding eigenvectors.
[0168] Input the eigenvalue data and eigenvector data into a quantum neural network to obtain feature complexity data. The quantum neural network has a 3 - layer structure, with each layer containing 5 qubits. The training uses the Adam optimizer, with a learning rate of 0.01 and 500 training iterations. The network outputs 10 feature complexity values, such as [0.78, 0.65, 0.54, 0.47, 0.39, 0.33, 0.28, 0.24, 0.21, 0.19].
[0169] Construct the potential function of the conditional random field model based on the feature complexity data. The potential function uses an exponential form, and the parameters are optimized by the gradient descent method, with a learning rate of 0.05 and 200 iterations. Input the scene type information into the conditional random field model for variational inference to obtain the marginal probability distribution. The variational inference uses the mean - field approximation, with 50 iterations and a convergence threshold of 0.01.
[0170] Perform variational approximation based on the marginal probability distribution to obtain the probability mapping matrix. The variational approximation uses the importance sampling method, with the number of samples being 1000. The final probability mapping matrix size is the number of scene types×the dimension of the eigenvector. For example, for 3 scene types and a 10 - dimensional eigenvector, the obtained matrix size is 3×10.
[0171] Establish the correspondence between the scene type information and multiple scene feature vector sub - spaces based on the probability mapping matrix. For each scene type, select the column indices where the values in the corresponding row of the probability mapping matrix are greater than the threshold of 0.5 as the dimension of the feature vector sub - space corresponding to this scene type. For example, for scene type 1, if the first row of the probability mapping matrix is [0.9, 0.2, 0.8, 0.3, 0.1, 0.7, 0.4, 0.6, 0.3, 0.2], then select the dimensions [0, 2, 5, 7] as its corresponding feature vector sub - space.
[0172] In this embodiment, through the combination of annealed Monte Carlo sampling and the Hamiltonian dynamics model, adaptive sampling of the feature space is realized. By using the Dirichlet mixture model to combine the dual - metric strategy of geodesic distance and Riemannian distance, a more accurate expression of the feature space is constructed. The quantum neural network is introduced to process the results of high - order tensor decomposition, and the advantages of quantum computing are used to handle complex feature relationships;
[0173] In the prior art, the correlation modeling between the scenario feature space and the scenario type usually adopts a simple probability model or distance metric method, ignoring the complexity and dynamic characteristics of the feature distribution. When sampling features and mapping the space, a sampling strategy with a fixed step size and a linear mapping method are mostly used, making it difficult to effectively process the non-linear structure in the high-dimensional feature space, resulting in inaccurate feature expression and limited model generalization ability;
[0174] The annealing mechanism in this embodiment ensures that the sampling process can jump out of the local optimum, while the Hamiltonian dynamics model captures the dynamic evolution characteristics of the feature distribution by constructing a potential energy function and solving the phase space equation, significantly improving the efficiency and accuracy of sampling. The modeling method based on multiple geometric metrics not only considers the probability characteristics of the feature distribution but also maintains the intrinsic geometric structure of the feature space, making the feature expression more complete and accurate. Through dual space mapping and high-order tensor decomposition, the potential associations between features can be explored more deeply. The application of the quantum neural network greatly improves the efficiency and accuracy of complex feature processing.
[0175] Figure 3 This is a performance and quality comparison chart of distance feature data for the scenario adaptive adjustment method of intelligent IoT and the virtual-real fusion of the metaverse in the embodiment of the present invention, showing a comprehensive comparison of the proposed technical solution with the Euclidean distance and the Mahalanobis distance in three key dimensions of distance feature data construction. By integrating the calculation time, feature separation degree, and model stability into the same coordinate system and connecting the performance of each dimension with a line, the comprehensive performance of each method can be intuitively seen.
[0176] In terms of calculation time, the proposed technical solution only needs 30 milliseconds to complete the construction of distance feature data, reducing the calculation overhead by 62.5% and 76.9% compared with the 80 milliseconds of the Euclidean distance method and the 130 milliseconds of the Mahalanobis distance method, respectively. The efficient sampling strategy and the optimized Hamiltonian dynamics model adopted in the proposed technical solution, especially the fourth-order Runge-Kutta method (step size 0.01) used in the potential energy function solution stage, greatly reduce the calculation complexity.
[0177] In terms of feature separation degree, the proposed technical solution achieves a high performance of 87.5%, far exceeding the 67.3% of the Euclidean distance method and the 75.8% of the Mahalanobis distance method. The feature separation degree has increased by 20.2 and 11.7 percentage points, meaning that in the feature space, data points of different scenario types can be more clearly divided with a more distinct boundary. Thanks to the proposed technical solution for constructing the direction distribution matrix by combining the geodesic distance and the Riemann distance, it can more accurately capture the geometric structure and topological relationship in the feature space, especially performing well in dealing with high-dimensional non-linear features.
[0178] In terms of model stability, the variance value of this technical solution is only 1.7, significantly lower than 2.8 of the Euclidean distance method and 3.5 of the Mahalanobis distance method, with a reduction of 39.3% and 51.4% respectively. The lower variance indicates that this technical solution has stronger robustness to different datasets and noise conditions, and the model output is more stable and reliable. The Dirichlet mixture model introduced in this technical solution (with the number of components set to 5 and the convergence threshold to 0.001) accurately models the data distribution, and the strategy of mapping to the dual space through the Riemann metric tensor effectively reduces the uncertainty in the feature space.
[0179] From the overall trend of the connection lines, it can be clearly seen that this technical solution is in the optimal position in all three evaluation dimensions (the lowest calculation time, the highest feature separation degree, and the lowest variance), forming an ideal "high-low-low" performance combination. The traditional Euclidean distance and Mahalanobis distance methods have deficiencies in different dimensions and cannot balance efficiency, accuracy, and stability simultaneously. Especially the Mahalanobis distance method, although it is superior to the Euclidean distance in terms of feature separation degree, its calculation time is the longest and its variance is the highest, presenting an obvious bottleneck in practical applications.
[0180] Integrating the data of the three dimensions, this technical solution has achieved an approximately 67.3% improvement in calculation efficiency, a 20.2% improvement in feature separation degree, and a 45.4% improvement in stability compared with the traditional methods. These advantages are particularly crucial when dealing with large-scale and complex scenario feature classification tasks. Through the integrated multi-dimensional analysis, it fully demonstrates the technical innovation and practical value of this technical solution in the construction of the Dirichlet mixture model and the generation of distance feature data, laying a solid foundation for subsequent feature complexity calculation and scene type information mapping.
[0181] In an optional implementation manner,
[0182] Based on the hyperbolic space manifold features constructed from the optimized feature vectors and the projection position information, enhanced feature data and a space mapping matrix are obtained, and the preset parameter mapping sub-network is used to generate scene rendering parameters including:
[0183] Input the optimized feature vectors into a feature decomposition network to obtain multi-dimensional feature tensors, and perform tensor decomposition operations on the multi-dimensional feature tensors to obtain feature component data; input the projection position information into a space transformation network to obtain position feature vectors, and perform manifold mapping operations on the position feature vectors to obtain a space mapping matrix;
[0184] Map the feature component data to the hyperbolic space to obtain initial feature data, perform Poisson manifold embedding operations on the initial feature data to construct a curvature weight function, and perform non-linear transformation on the initial feature data based on the curvature weight function to obtain enhanced feature data;
[0185] Input the enhanced feature data and the spatial mapping matrix into the parameter mapping sub-network. Construct a fine-grained decomposer in the parameter mapping sub-network to obtain a multi-scale feature tensor, and establish a feature pyramid network for the multi-scale feature tensor to obtain hierarchical feature data;
[0186] Perform self-attention mechanism operation on the hierarchical feature data to obtain an attention weight matrix, construct a recurrent neural network based on the attention weight matrix to obtain time series feature data, and input the time series feature data into a generative adversarial network to obtain an adversarial feature vector;
[0187] Input the adversarial feature vector into a variational autoencoder for feature reconstruction to obtain reconstructed feature data, perform residual optimization on the reconstructed feature data to obtain a residual feature vector, and generate scene rendering parameters based on the residual feature vector.
[0188] Input the optimized feature vector into a feature decomposition network to obtain a multi-dimensional feature tensor. Pass the 256-dimensional optimized feature vector through a feature decomposition network composed of 3 fully connected layers. Each layer uses LeakyReLU as the activation function, and outputs a multi-dimensional feature tensor with a shape of 16×16×16. Perform tensor decomposition operation on this multi-dimensional feature tensor. Use the Tucker decomposition method to decompose it into 3 low-dimensional core tensors and a factor matrix to obtain feature component data. For example, decompose the 16×16×16 tensor into an 8×8×8 core tensor and 3 factor matrices of size 16×8 through Tucker decomposition, with a compression rate reaching 75% while retaining more than 90% of the data information.
[0189] Input the projection position information into the spatial transformation network to obtain a position feature vector. The projection position information refers to the projection coordinates of an object in a three-dimensional scene, usually expressed in the form of (x, y, z). The spatial transformation network includes a spatial sampling layer and a feature mapping layer. The spatial sampling layer normalizes the input position information and scales the coordinate range to the [-1, 1] interval; the feature mapping layer consists of 4 convolutional layers, extracting spatial features layer by layer and outputting a 128-dimensional position feature vector. Perform a manifold mapping operation on the position feature vector, using the Riemannian metric transformation to map the Euclidean space features to the hyperbolic space to obtain a spatial mapping matrix. The dimension of this matrix is 128×128, which is used for subsequent feature fusion and transformation.
[0190] Map the feature component data to the hyperbolic space to obtain the initial feature data. Use the Poincaré ball model as the representation form of the hyperbolic space, and transform the feature component data in the Euclidean space to the hyperbolic space with a curvature of -1 through the exponential mapping function to obtain the initial feature data. Perform Poisson manifold embedding operation on the initial feature data to construct a curvature weight function. Poisson manifold embedding calculates the harmonic field of the feature on the manifold by solving the Poisson equation, and generates an adaptive curvature weight function. This weight function can dynamically adjust the weight distribution according to the intrinsic geometric structure of the data, and assigns higher weights to high-curvature regions. Perform a non-linear transformation on the initial feature data based on the curvature weight function, and use hyperbolic trigonometric function transformation to enhance the feature expression ability to obtain enhanced feature data.
[0191] Input the enhanced feature data and the spatial mapping matrix into the parameter mapping sub-network. Construct a fine-grained decomposer in the parameter mapping sub-network. This decomposer consists of a multi-layer convolutional neural network with residual connections between each layer, and outputs a multi-scale feature tensor. Specifically, extract features in parallel through three different-sized convolutional kernels of 3×3, 5×5, and 7×7, and generate feature maps of sizes 32×32, 16×16, and 8×8 respectively to form a multi-scale feature tensor. Establish a feature pyramid network for the multi-scale feature tensor, and generate hierarchical feature data through top-down feature fusion and bottom-up feature enhancement. The feature pyramid network contains a 5-layer structure, and each layer is connected through upsampling or downsampling operations to achieve information interaction of features at different scales.
[0192] Perform self-attention mechanism operation on the hierarchical feature data to obtain an attention weight matrix. In specific implementation, calculate the similarity between features to generate a self-attention weight matrix with a dimension of N×N, where N is the number of feature vectors. For example, when the number of features is 1024, a 1024×1024 attention weight matrix is generated to capture the long-range dependence relationship between features. Construct a recurrent neural network based on the attention weight matrix, use LSTM units to process temporal information, set the input gate threshold to 0.5, the forget gate threshold to 0.4, the output gate threshold to 0.6, and the number of hidden layer nodes to 512 to obtain time series feature data. Input the time series feature data into a generative adversarial network, which consists of a generator and a discriminator. The generator adopts a 5-layer transposed convolution structure, and the discriminator adopts a 4-layer convolution structure to obtain an adversarial feature vector.
[0193] The adversarial feature vectors are input into a variational autoencoder for feature reconstruction. The encoder compresses the 256-dimensional adversarial feature vectors into 64-dimensional latent space representations, and the decoder reconstructs the latent space representations into features of the original dimension through a three-layer fully connected network to obtain reconstructed feature data. Residual optimization is performed on the reconstructed feature data to calculate the difference between the reconstructed features and the original features, generating residual feature vectors. Based on the residual feature vectors, scene rendering parameters are generated through a feature mapping network, including lighting parameters (ambient light intensity, directional light intensity, light source position), material parameters (diffuse reflectivity, specular intensity, roughness), and camera parameters (position, direction, field of view angle), etc.
[0194] In this embodiment, the feature decomposition network can capture the key components of features. Tensor decomposition significantly reduces the computational complexity while maintaining the integrity of data information, laying a good foundation for subsequent processing. By mapping Euclidean space features to hyperbolic space and combining Poisson manifold embedding technology, more accurate modeling of the spatial structure is achieved, enhancing the geometric invariance of feature expression. The design of the hyperbolic space mapping combined with the curvature weight function significantly improves the expression ability of features. Poisson manifold embedding enables the system to better maintain the internal structure of data through an adaptive weight allocation mechanism, improving the accuracy and robustness of feature expression.
[0195] Figure 4 It is a schematic structural diagram of the scene adaptive adjustment system for the intelligent Internet of Things and the virtual-real fusion of the metaverse in the embodiment of the present invention. As Figure 4 shown, the system includes:
[0196] The first unit acquires scene image data and environmental parameter data; performs resolution decomposition and block division on the scene image data, and conducts scene type recognition through a preset scene recognition model to obtain scene type information;
[0197] The second unit constructs a multi-dimensional environmental parameter tensor based on the environmental parameter data, establishes an environmental digital twin model, and maps the environmental digital twin model to the virtual space using an adaptive coordinate mapping network;
[0198] The third unit is used to construct multiple scene feature vector subspaces based on the scene image data, each scene feature vector subspace corresponding to a type of scene type information. According to the scene type information, the corresponding scene feature vector subspace is selected, and the scene image data is mapped to the selected scene feature vector subspace to generate target scene feature vectors;
[0199] The fourth unit is used to obtain the projection position information of virtual objects in the virtual space in the physical space based on the environmental digital twin model; input the target scene feature vectors into a preset feature optimization sub-network for iterative optimization to obtain optimized feature vectors;
[0200] The fifth unit is used to obtain enhanced feature data and a spatial mapping matrix based on the hyperbolic space manifold features constructed from the optimized feature vectors and the projection position information, and generate scene rendering parameters by using a preset parameter mapping sub-network; perform rendering processing on the virtual object based on the scene rendering parameters, and superimpose and display the rendered virtual object on the entity space display interface.
[0201] In a third aspect of the embodiments of the present invention, there is provided an electronic device, including:
[0202] A processor; a memory for storing instructions executable by the processor; wherein the processor is configured to call the instructions stored in the memory to execute the method described above.
[0203] In a fourth aspect of the embodiments of the present invention, there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0204] The present invention may be a method, an apparatus, a system, and / or a computer program product. The computer program product may include a computer-readable storage medium, on which computer-readable program instructions for performing various aspects of the present invention are uploaded.
[0205] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A scene adaptive adjustment method for the virtual-real integration of intelligent Internet of Things and the metaverse, characterized in that, Including: Obtain scene image data and environmental parameter data; Perform resolution decomposition and chunking on the scene image data, and perform scene type recognition through a preset scene recognition model to obtain scene type information; Construct a multi-dimensional environmental parameter tensor based on the environmental parameter data, establish an environmental digital twin model, and use an adaptive coordinate mapping network to map the environmental digital twin model to the virtual space; Construct multiple scene feature vector subspaces based on the scene image data, each scene feature vector subspace corresponding to a kind of scene type information, select the corresponding scene feature vector subspace according to the scene type information, and map the scene image data to the selected scene feature vector subspace to generate a target scene feature vector; Obtain the projection position information of the virtual object in the virtual space in the physical space based on the environmental digital twin model; input the target scene feature vector into a preset feature optimization sub-network for iterative optimization to obtain an optimized feature vector; Based on the hyperbolic space manifold feature and the projection position information constructed by the optimized feature vector, obtain enhanced feature data and a spatial mapping matrix, and use a preset parameter mapping sub-network to generate scene rendering parameters including: Input the optimized feature vector into a feature decomposition network to obtain a multi-dimensional feature tensor, perform tensor decomposition operation on the multi-dimensional feature tensor to obtain feature component data; input the projection position information into a spatial transformation network to obtain a position feature vector, and perform a manifold mapping operation on the position feature vector to obtain a spatial mapping matrix; Map the feature component data to the hyperbolic space to obtain initial feature data, perform Poisson manifold embedding operation on the initial feature data to construct a curvature weight function, and perform a non-linear transformation on the initial feature data based on the curvature weight function to obtain enhanced feature data; Input the enhanced feature data and the spatial mapping matrix into a parameter mapping sub-network, construct a fine-grained decomposer in the parameter mapping sub-network to obtain a multi-scale feature tensor, and establish a feature pyramid network for the multi-scale feature tensor to obtain hierarchical feature data; Perform a self-attention mechanism operation on the hierarchical feature data to obtain an attention weight matrix, construct a recurrent neural network based on the attention weight matrix to obtain time series feature data, and input the time series feature data into a generative adversarial network to obtain an adversarial feature vector; Input the adversarial feature vector into a variational autoencoder for feature reconstruction to obtain reconstructed feature data, perform residual optimization on the reconstructed feature data to obtain a residual feature vector, and generate scene rendering parameters based on the residual feature vector; Perform rendering processing on the virtual object based on the scene rendering parameters, and superimpose and display the rendered virtual object on the physical space display interface.
2. The method according to claim 1, wherein Perform resolution decomposition and chunking on the scene image data, and perform scene type recognition through a preset scene recognition model to obtain scene type information including: Perform hierarchical perception on the scene image data to generate multiple scene image layers with different resolutions; Perform adaptive chunking on each scene image layer, and determine the optimal chunk size of each scene image layer based on the information entropy of the image chunks to obtain a corresponding image chunk sequence; Construct a discrete wavelet transform pyramid for each of the image patch sequences, and extract multi-scale frequency domain features; Input the multi-scale frequency domain features into a preset heterogeneous convolutional neural network. The heterogeneous convolutional neural network uses different convolutional kernel sizes for feature extraction of frequency domain features at different scales to obtain image patch feature vectors; Input the image patch feature vectors into a preset self-calibrating feature enhancement network, and iteratively optimize the feature vectors based on a feedback mechanism to generate enhanced feature vectors; Recombine the enhanced feature vectors according to adaptive attention weights to generate a multi-scale feature sequence; input the multi-scale feature sequence into a preset two-stream neural network. The spatial stream branch of the two-stream neural network extracts the spatial features of the multi-scale feature sequence, and the temporal stream branch of the two-stream neural network extracts the temporal features of the multi-scale feature sequence; fuse the spatial features and the temporal features to generate a fused feature vector; Input the fused feature vector into a preset hierarchical scene classifier. The hierarchical scene classifier performs progressive classification on the fused feature vector based on the scene semantic hierarchical relationship to obtain scene type information.
3. The method according to claim 1, characterized in that, Construct a multi-dimensional environmental parameter tensor based on environmental parameter data, establish an environmental digital twin model, and use an adaptive coordinate mapping network to map the environmental digital twin model to a virtual space, including: Align the environmental parameter data based on a spatio-temporal synchronization mechanism to construct a multi-dimensional environmental parameter tensor; Perform adaptive wavelet packet decomposition on the multi-dimensional environmental parameter tensor to generate a multi-scale environmental feature spectrum; Input the multi-scale environmental feature spectrum into a preset atlas embedding network, and extract the correlation features between parameters through a multi-head attention mechanism to generate high-dimensional environmental feature vectors; Construct a hybrid memory neural network, input the high-dimensional environmental feature vectors into the hybrid memory neural network, store long-term dependence information based on an external memory matrix, and combine short-term dynamic memory units to capture transient change features to generate environmental dynamic prediction data; Construct a recursive neural tensor network, generate a fused feature tensor based on the high-dimensional environmental feature vectors and the environmental dynamic prediction data, and input it into the recursive neural tensor network. The recursive neural tensor network uses a tensor decomposition method to perform multi-dimensional correlation analysis on the input data, establish a dynamic coupling relationship between environmental elements, and generate an environmental digital twin model; Construct an adaptive coordinate mapping network. The adaptive coordinate mapping network includes a spatial transformation sub-network and a deformation compensation sub-network. Map the topological structure of the environmental digital twin model to the virtual space coordinate system through the spatial transformation sub-network, and use the deformation compensation sub-network to dynamically correct the geometric deformation during the mapping process, and map the corrected environmental digital twin model to the rendering coordinate system of the virtual space.
4. The method according to claim 3, wherein Construct a recursive neural tensor network, generate a fused feature tensor based on the high-dimensional environmental feature vector and the environmental dynamic prediction data, and input the fused feature tensor into the recursive neural tensor network. The recursive neural tensor network uses tensor decomposition to perform multi-dimensional correlation analysis on the input data, establish the dynamic coupling relationship between environmental elements, and generate an environmental digital twin model, including: Construct a Riemannian manifold mapping for the high-dimensional environmental feature vector to generate a manifold feature, and project the environmental dynamic prediction data into the curvature tensor space to generate a tensor feature; Input the manifold feature and the tensor feature into a preset Riemannian optimization model, construct an optimization objective function based on the geodesic distance metric, and use the conjugate gradient method with an adaptive step size to iteratively solve the optimization objective function to obtain an optimized feature; Project the optimized feature into the Euclidean space through the Riemannian exponential mapping to generate a fused feature tensor; Construct a recursive neural tensor network, which includes an orthogonal tensor decomposition unit and a non-negative tensor decomposition unit. The orthogonal tensor decomposition unit performs local feature decomposition on the fused feature tensor, and the non-negative tensor decomposition unit performs global feature decomposition on the fused feature tensor; input the local feature decomposition result and the global feature decomposition result into a preset tensor recurrence unit, and establish the correlation relationship between features through recursive feature transfer; Construct a tensor reconstruction module based on spectral decomposition, reconstruct the features using nuclear norm regularization constraints to generate an environmental element correlation tensor, and generate an environmental digital twin model based on the environmental element correlation tensor.
5. The method according to claim 1, wherein Construct multiple scene feature vector subspaces based on the scene image data, each scene feature vector subspace corresponds to a type of scene type information, select the corresponding scene feature vector subspace according to the scene type information, and map the scene image data to the selected scene feature vector subspace to generate a target scene feature vector, including: Input the scene image data into a multi-layer convolutional neural network to extract multi-scale feature data, and perform an attention mechanism operation on the multi-scale feature data to obtain attention feature data; Input the attention feature data into a variational autoencoder network to obtain high-dimensional feature data, and construct the multiple scene feature vector subspaces according to the high-dimensional feature data; Perform a manifold embedding operation on the multiple scene feature vector subspaces to obtain manifold feature data, and input the manifold feature data into a Riemannian optimization network to obtain multiple feature distribution data; Input the multiple feature distribution data into a Dirichlet mixture model to obtain multiple probability parameters, construct distance feature data based on the multiple probability parameters, perform a spectral decomposition operation on the distance feature data to obtain feature complexity data, and establish the corresponding relationship between the scene type information and the multiple scene feature vector subspaces according to the feature complexity data; Select the corresponding target scene feature vector subspace from the multiple scene feature vector subspaces according to the scene type information, perform Markov chain sampling on the target scene feature vector subspace to obtain an optimal feature subspace, and generate a target scene feature vector.
6. The method according to claim 5, wherein Inputting the multiple feature distribution data into a Dirichlet mixture model to obtain multiple probability parameters, constructing distance feature data based on the multiple probability parameters, performing spectral decomposition operation on the distance feature data to obtain feature complexity data, and establishing a correspondence between the scene type information and the multiple scene feature vector subspaces according to the feature complexity data, including: Performing annealing Monte Carlo sampling on the multiple feature distribution data to obtain initial sampling data, inputting the initial sampling data into a Hamiltonian dynamics model to construct a potential energy function, solving a phase space equation system for the potential energy function to obtain phase distribution data, and performing importance sampling on the initial sampling data based on the phase distribution data to obtain sampling feature data; Inputting the sampling feature data into a Dirichlet mixture model to obtain multiple probability parameters, constructing a direction distribution matrix according to the multiple probability parameters, and calculating geodesic distance and Riemannian distance respectively for the direction distribution matrix to obtain distance feature data; Constructing a Riemannian metric tensor based on the distance feature data, mapping the Riemannian metric tensor to a dual space to obtain a dual feature tensor, performing high-order tensor decomposition on the dual feature tensor to obtain eigenvalue data and eigenvector data, and inputting the eigenvalue data and the eigenvector data into a quantum neural network to obtain feature complexity data; Constructing a potential function of a conditional random field model based on the feature complexity data, inputting the scene type information into the conditional random field model for variational inference to obtain a marginal probability distribution, performing variational approximation based on the marginal probability distribution to obtain a probability mapping matrix, and establishing a correspondence between the scene type information and the multiple scene feature vector subspaces based on the probability mapping matrix.
7. A scene adaptive adjustment system for the virtual-real integration of intelligent Internet of Things and the metaverse, which is used to implement the method described in any one of the foregoing claims 1-6, characterized in that Including: A first unit for acquiring scene image data and environmental parameter data; Performing resolution decomposition and blocking on the scene image data, and performing scene type recognition through a preset scene recognition model to obtain scene type information; A second unit for constructing a multi-dimensional environmental parameter tensor based on the environmental parameter data, establishing an environmental digital twin model, and mapping the environmental digital twin model to a virtual space by using an adaptive coordinate mapping network; A third unit for constructing multiple scene feature vector subspaces based on the scene image data, where each scene feature vector subspace corresponds to a kind of scene type information, selecting a corresponding scene feature vector subspace according to the scene type information, and mapping the scene image data to the selected scene feature vector subspace to generate a target scene feature vector; A fourth unit for acquiring projection position information of a virtual object in the virtual space in the physical space based on the environmental digital twin model; inputting the target scene feature vector into a preset feature optimization sub-network for iterative optimization to obtain an optimized feature vector; A fifth unit for obtaining enhanced feature data and a space mapping matrix based on the hyperbolic space manifold feature constructed from the optimized feature vector and the projection position information, and generating scene rendering parameters by using a preset parameter mapping sub-network; Performing rendering processing on the virtual object based on the scene rendering parameters, and superimposing and displaying the rendered virtual object on the physical space display interface.
8. An electronic device, characterized in that, Including: A processor; A memory for storing processor-executable instructions; Wherein, the processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Voice-driven virtual human posture synthesis method for sensing environment
CN113838218A
VR brain-computer interface system and method for olfaction-induced electroencephalogram
CN119536521A