Data acquisition method and system of intelligent Internet of Things sensor in universe virtual-real fusion

Through the adaptive fusion processing of intelligent IoT sensors and deep space-time correlation networks, the problem of insufficient synchronization and consistency in the fusion of virtual and real in the meta-universe is solved, and the two-way mapping and interaction of virtual and real scenes is realized, and the user experience and data collection efficiency is improved.

CN120372550AInactive Publication Date: 2025-07-25HANGZHOU MOXI TECH DEV CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510495699.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-07-25
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing intelligent IoT sensors lack deep mining of space-time features in the fusion of virtual and real meta-universe, and it is difficult to achieve adaptability, resulting in insufficient synchronization and consistency between virtual and real scenes, and the interaction is often unidirectional, lacking bidirectional mapping and feedback mechanisms.

Method used

Through intelligent IoT sensors, they collect physical scene information and perform digital transformation and preprocessing, use deep space-time correlation networks to perform adaptive fusion processing, build meta-universe virtual scenes, and realize bidirectional mapping and interaction of virtual and real scenes through intelligent perception engines, and dynamically adjust the sampling strategy to optimize sampling frequency, accuracy and data transmission.

Benefits of technology

It improves the realism and immersion of virtual scenes, enhances the synchronization and consistency between virtual scenes and the real world, improves the user's interactive experience, and realizes the deep integration and real-time interaction between the physical world and the virtual world.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372550A_ABST
    Figure CN120372550A_ABST
Patent Text Reader

Abstract

The invention provides a data acquisition method and system of an intelligent Internet of Things sensor in universe virtual-real fusion, and relates to the technical field of universe, and the method comprises the steps: collecting physical scene information through the intelligent Internet of Things sensor, and carrying out the preprocessing; adaptive fusion processing is carried out by using a deep space-time correlation network; constructing a meta universe virtual scene; bidirectional mapping and interaction of virtual and real scenes are realized through an intelligent perception engine, and a sampling strategy is optimized based on deep reinforcement learning. According to the invention, accurate mapping and real-time interaction between the physical world and the virtual world are realized, and the data acquisition efficiency and the virtual-real fusion experience are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the metaverse technology, and in particular to a data acquisition method and system for intelligent Internet of Things sensors in the virtual-real fusion of the metaverse. Background Art

[0002] With the rise of the metaverse concept, the integration of the virtual world and the real world has become an important research direction. As an important bridge connecting the virtual and the real, intelligent Internet of Things sensors play a key role in the metaverse scenario. At present, intelligent Internet of Things sensors have been widely used in the data acquisition of various physical scenarios, providing rich data sources for the construction of the metaverse virtual scenario.

[0003] However, there are still some deficiencies in the data acquisition of existing intelligent Internet of Things sensors in the virtual-real fusion of the metaverse: First of all, traditional data acquisition methods lack in-depth mining of spatio-temporal features, and it is difficult to fully capture the dynamic change features of physical scenarios, resulting in insufficient synchronization and consistency between the virtual scenario and the real scenario.

[0004] Secondly, existing data acquisition strategies are often fixed and lack adaptability, and cannot dynamically adjust sampling parameters according to user interaction requirements and scenario changes, affecting the real-time performance and accuracy of virtual-real fusion.

[0005] Finally, the interaction between the virtual scenario and the physical scenario in the existing technology is often one-way, lacking an effective two-way mapping and feedback mechanism, and it is difficult to achieve deep fusion and interaction between the virtual and real worlds. Summary of the Invention

[0006] Embodiments of the present invention provide a data acquisition method and system for intelligent Internet of Things sensors in the virtual-real fusion of the metaverse, which can solve the problems in the existing technology.

[0007] In the first aspect of the embodiments of the present invention, Collect physical scenario information through intelligent Internet of Things sensors, and perform digital conversion and preprocessing on the collected physical scenario information; Transmit the preprocessed physical scenario information to the data center server through the Internet of Things protocol, and perform adaptive fusion processing on the physical scenario information based on a deep spatio-temporal correlation network, including: using a temporal convolutional neural network to extract the temporal features of the physical scenario information, using a graph convolutional neural network to extract the spatial features of the physical scenario information, fusing the temporal features and the spatial features to obtain a scenario feature vector, and generating standardized scenario feature data according to the scenario feature vector; Construct a metaverse virtual scenario based on the standardized scenario feature data, and display the metaverse virtual scenario on a metaverse interaction terminal; Implement bidirectional mapping and interaction between virtual and real scenarios through an intelligent perception engine, including: receiving interaction instructions from the user at the metaverse interaction terminal, optimizing the adaptive sampling strategy through deep reinforcement learning according to the interaction instructions, and dynamically adjusting the sampling strategy of the intelligent IoT sensors, including adjusting the sampling frequency, sampling accuracy, and data transmission strategy; collecting updated physical scenario information based on the sampling strategy and mapping the updated physical scenario information to the metaverse virtual scenario.

[0008] Using a temporal convolutional neural network to extract the temporal features of physical scenario information includes: The temporal convolutional neural network includes a feature decomposition branch and a feature reconstruction branch, where the feature decomposition branch is used to decompose the physical scenario information into multi-scale temporal features, and the feature reconstruction branch is used to adaptively fuse the multi-scale temporal features by combining a dynamic attention mechanism to generate a temporal feature representation of the physical scenario.

[0009] Using a graph convolutional neural network to extract the spatial features of physical scenario information includes: Construct a spatial relationship graph structure of the physical scenario, the spatial relationship graph structure includes a set of sensor nodes and the spatial connection relationship between nodes, calculate the spatial correlation weight based on the Euclidean distance between sensor nodes, and the spatial correlation weight is calculated through a Gaussian kernel function; Map the heterogeneous data of different types of sensors to a unified feature space through the multi-modal feature embedding layer of the graph convolutional neural network. The multi-modal feature embedding layer includes a transformation matrix and a bias term corresponding to the sensor type, and generates node features through a non-linear activation function; Perform multi-layer graph convolution operations on the node features based on the message passing mechanism. The graph convolution operation aggregates features according to the neighborhood set and node degree value of the node to generate the spatial dependence features of the node; Introduce an adaptive spatial attention mechanism to dynamically weight the node features, calculate the attention coefficient between nodes through a learnable attention vector, and selectively aggregate the spatial dependence features of the nodes according to the attention coefficient to generate the spatial features of the physical scenario information.

[0010] Optimizing the adaptive sampling strategy through deep reinforcement learning according to the interaction instructions, and dynamically adjusting the sampling strategy of the intelligent IoT sensors includes: Convert the interaction instruction into a corresponding interaction feature vector, obtain the real-time monitoring data of the intelligent IoT sensors, and construct a scenario state representation including an environmental state matrix and a data quality matrix based on the real-time monitoring data. Among them, the environmental state matrix contains the environmental parameters collected by each sensor, and the data quality matrix contains the data reliability indicators of each sensor; The interaction feature vector and the scene state representation are combined to construct the state space of deep reinforcement learning. Based on the state space, the sampling frequency, sampling accuracy, and data transmission policy are calculated to generate the sampling policy action space. Based on the sampling policy action space, a sampling operation is performed, and the data quality score, latency performance score, and coverage performance score of the sampling operation are calculated. The data quality score, latency performance score, and coverage performance score are weighted and combined to obtain a comprehensive reward value. The comprehensive reward value is used to update the policy network parameters through the proximal policy optimization algorithm. Based on the updated policy network, the optimized sampling parameters are output and applied to sensor sampling control. The policy network is a dual structure including an Actor network for generating the action probability distribution and a Critic network for evaluating the state value.

[0011] Using the comprehensive reward value to update the policy network parameters through the proximal policy optimization algorithm includes: Based on the Actor network, the action probability distribution selected in the current state is calculated, and based on the Critic network, the value estimate of the current state is calculated. According to the action probability distribution, the probability ratio is calculated. The probability ratio is the ratio of the current policy probability to the old policy probability, and the clipped probability ratio is obtained by clipping the probability ratio. The generalized advantage estimation method is used to calculate the advantage function estimate value, which is calculated based on the discounted cumulative reward and the temporal difference error of the state value estimate. The clipped probability ratio and the advantage function estimate value are multiplied to construct the proximal policy objective function. Based on the proximal policy objective function, the parameter gradient of the Actor network is calculated, and the parameter gradient of the Critic network is calculated based on the state value estimation error. A KL divergence constraint mechanism is introduced. When the KL divergence between the new and old policies exceeds the preset divergence threshold, the parameter update is terminated in advance. The parameter gradients of the Actor network and the Critic network are respectively used to update the corresponding network parameters.

[0012] Using the generalized advantage estimation method to calculate the advantage function estimate value includes: The advantage function is calculated according to the state value function and the action value function, where the advantage function is expressed as the difference between the action value function and the state value function of the state-action pair. An exponential weighting coefficient is introduced based on the advantage function to construct a generalized advantage estimation function, where the generalized advantage estimation function includes an infinite sequence weighted sum of temporal difference errors, and the temporal difference errors are calculated from the immediate reward value, the discount factor, and the state value function of adjacent time steps; The estimated value of the generalized advantage estimation function is obtained by recursively adding the temporal difference error of the current time step to the advantage estimation value of the next time step adjusted by the discount factor and the generalized advantage estimation parameter; A state-dependent baseline function is introduced to correct the estimated value, where the baseline function is calculated based on the expected return under the policy, and the correction includes subtracting the baseline function value from the actual return; The corrected estimated value is normalized, where the normalization process includes standardizing the estimated value of the corrected generalized advantage estimation function using the mean of the advantage estimation values within the batch.

[0013] In a second aspect of the embodiments of the present invention, there is provided a data acquisition system for intelligent IoT sensors in the virtual-real fusion of the metaverse, including: A first unit for collecting physical scene information through intelligent IoT sensors and performing digital conversion and preprocessing on the collected physical scene information; A second unit for transmitting the preprocessed physical scene information to a data center server through an IoT protocol and performing adaptive fusion processing on the physical scene information based on a deep spatio-temporal correlation network, including: extracting temporal features of the physical scene information using a temporal convolutional neural network, extracting spatial features of the physical scene information using a graph convolutional neural network, fusing the temporal features and the spatial features to obtain a scene feature vector, and generating standardized scene feature data according to the scene feature vector; A third unit for constructing a metaverse virtual scene based on the standardized scene feature data and displaying the metaverse virtual scene on a metaverse interaction terminal; A fourth unit for realizing two-way mapping and interaction between virtual and real scenes through an intelligent perception engine, including: receiving an interaction instruction of a user on the metaverse interaction terminal, optimizing an adaptive sampling strategy through deep reinforcement learning according to the interaction instruction, and dynamically adjusting the sampling strategy of the intelligent IoT sensors, including adjusting the sampling frequency, sampling accuracy, and data transmission strategy; collecting updated physical scene information based on the sampling strategy and mapping the updated physical scene information to the metaverse virtual scene.

[0014] In a third aspect of the embodiments of the present invention, There is provided an electronic device, including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.

[0015] In the fourth aspect of the embodiments of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.

[0016] The beneficial effects of this application are as follows: By collecting physical scene information through intelligent Internet of Things sensors, performing digital conversion and preprocessing, and then using a deep spatio-temporal correlation network for adaptive fusion processing, the temporal and spatial features of the physical scene can be effectively extracted, generating standardized scene feature data, providing a reliable data basis for constructing a highly realistic metaverse virtual scene, and improving the realism and immersion of the virtual scene.

[0017] Based on the intelligent perception engine, two-way mapping and interaction between virtual and real scenes are realized. By optimizing the sampling strategy through deep reinforcement learning, the sampling parameters of the sensors can be dynamically adjusted according to user interaction needs, achieving real-time and accurate collection of physical scene information, enhancing the synchronization and consistency between the virtual scene and the real world, and improving the user's interaction experience.

[0018] This method combines intelligent Internet of Things technology, deep learning algorithms with the construction of metaverse virtual scenes, realizes the deep integration and real-time interaction between the physical world and the virtual world, provides key technical support for metaverse applications, and has important theoretical significance and application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 is a schematic flow chart of the data acquisition method of the intelligent Internet of Things sensor in the virtual-real fusion of the metaverse in the embodiments of the present invention; Figure 2 is a schematic diagram of the energy consumption prediction accuracy result in the embodiments of the present invention; Figure 3 is a schematic diagram comparing the cumulative reward values of different optimization algorithms in the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0021] The technical solution of the present invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in some embodiments.

[0022] Figure 1 It is a schematic flowchart of the data acquisition method of the intelligent IoT sensor in the virtual-real fusion of the metaverse in the embodiment of the present invention. As Figure 1 shown, the method includes: Collect physical scene information through an intelligent IoT sensor, and perform digital conversion and preprocessing on the collected physical scene information; Transmit the preprocessed physical scene information to the data center server through the IoT protocol, and perform adaptive fusion processing on the physical scene information based on the deep spatio-temporal correlation network, including: using a temporal convolutional neural network to extract the temporal features of the physical scene information, using a graph convolutional neural network to extract the spatial features of the physical scene information, fusing the temporal features and the spatial features to obtain a scene feature vector, and generating standardized scene feature data according to the scene feature vector; Construct a metaverse virtual scene based on the standardized scene feature data, and display the metaverse virtual scene on the metaverse interaction terminal; Realize the bidirectional mapping and interaction of virtual and real scenes through an intelligent perception engine, including: receiving the interaction instructions of the user on the metaverse interaction terminal, optimizing the adaptive sampling strategy through deep reinforcement learning according to the interaction instructions, and dynamically adjusting the sampling strategy of the intelligent IoT sensor, including adjusting the sampling frequency, sampling accuracy, and data transmission strategy; collecting the updated physical scene information based on the sampling strategy, and mapping the updated physical scene information to the metaverse virtual scene.

[0023] In an alternative embodiment, using a temporal convolutional neural network to extract the temporal features of physical scene information includes: The temporal convolutional neural network includes a feature decomposition branch and a feature reconstruction branch, where the feature decomposition branch is used to decompose the physical scene information into multi-scale temporal features, and the feature reconstruction branch is used to adaptively fuse the multi-scale temporal features by combining a dynamic attention mechanism to generate a temporal feature expression of the physical scene.

[0024] Construct the overall architecture of the temporal convolutional neural network. The network includes two main parts: a feature decomposition branch and a feature reconstruction branch. The feature decomposition branch is used to decompose the input physical scene information into temporal features of multiple scales, and the feature reconstruction branch uses a dynamic attention mechanism to adaptively fuse these multi-scale features, and finally generates a temporal feature expression of the physical scene.

[0025] In the feature decomposition branch, a multi-layer one-dimensional convolutional layer is adopted to achieve multi-scale decomposition of the input sequence. Specifically, 3 one-dimensional convolutional layers are used, with the convolutional kernel sizes being 3, 5, and 7 respectively, and the stride being 1 for all. This can capture feature patterns at different time scales. After each convolutional layer, a batch normalization layer and a ReLU activation function are connected to enhance the network's non-linear expression ability. For example, for an input sequence of length 100, after passing through 3 convolutional layers, 3 groups of feature maps can be obtained, with the sizes being 98x64, 96x64, and 94x64 respectively.

[0026] A dynamic attention mechanism is introduced in the feature reconstruction branch. First, global average pooling is applied to the 3 groups of feature maps respectively to obtain 3 64-dimensional vectors. Then, a two-layer fully connected network is used to generate attention weights. The number of neurons in the first layer is 32, and the second layer is 3. The weights are normalized through the Softmax function to obtain the fusion coefficients for the 3-scale features. For example, the obtained weights are [0.4, 0.3, 0.3].

[0027] Using the obtained attention weights, weighted summation is performed on the 3 groups of feature maps. The specific operation is to multiply each group of feature maps by the corresponding weight coefficient and then add them element by element. This step realizes the adaptive fusion of multi-scale features. The size of the fused feature map is 94x64.

[0028] To further extract temporal features, a one-dimensional max pooling operation is applied to the fused feature map, with the pooling kernel size being 2 and the stride being 2. This step can reduce the time dimension of the feature map while retaining important feature information. After pooling, the size of the obtained feature map becomes 47x64.

[0029] A fully connected layer is used to map the feature map to the target dimension to obtain the final temporal feature expression. Assuming the target dimension is 128, the input dimension of the fully connected layer is 47x64 = 3008, and the output dimension is 128. After the fully connected layer, a ReLU activation function is connected to increase the non-linear expression ability.

[0030] During the network training process, the backpropagation algorithm is adopted to optimize the network parameters. The loss function can be selected according to the specific task. For example, the mean squared error loss can be used for regression tasks, and the cross-entropy loss can be used for classification tasks. The Adam algorithm is selected as the optimizer, with the initial learning rate set to 0.001 and decaying by 10% every 50 epochs. The batch size is set to 64, and the number of training epochs is 200.

[0031] To improve the generalization ability of the model, data augmentation techniques are adopted during training. For temporal data, methods such as time window sliding, random cropping, and adding Gaussian noise can be used to expand the training samples. For example, for an original sequence of length 1000, 100 subsequences of length 100 can be randomly cropped as training samples.

[0032] In the model inference stage, the input physical scene information is preprocessed into a time series with a fixed length. For example, for original sequences of different lengths, they can be uniformly adjusted to a sequence of length 100 through zero-padding or linear interpolation. Then the preprocessed sequence is input into the trained temporal convolutional neural network, and through feature decomposition, dynamic attention fusion, pooling, and fully connected layers, a 128-dimensional temporal feature expression vector is finally obtained.

[0033] To verify the effectiveness of the proposed method, experiments can be conducted on an actual physical scene dataset. Taking the earthquake waveform recognition task as an example, the STEAD dataset is used for training and testing. This dataset contains waveform data of approximately 1 million earthquake and noise events, and each sample is a three-component time series with a length of 6000. The dataset is divided into a training set, a validation set, and a test set according to the ratio of 8:1:1.

[0034] In the training stage, the original waveform data is preprocessed into a sequence with a length of 100. Then the above-mentioned temporal convolutional neural network structure is used for feature extraction to obtain a 128-dimensional feature vector. Finally, a fully connected layer and a Softmax function are connected for binary classification (earthquake event / noise). Evaluated on the test set, the proposed method achieves a classification accuracy of 97.5%, which is significantly improved compared to the traditional short-time Fourier transform features (accuracy 95.2%) and long short-term memory network (accuracy 96.8%).

[0035] Through visual analysis, it is found that the extracted temporal features can effectively capture the key patterns of earthquake waveforms. For example, for the recognition of the arrival time of P waves, traditional methods are easily interfered by noise and produce misjudgments, while the features extracted based on the temporal convolutional neural network show better robustness. This shows that the method can adaptively focus on important features at different time scales, thereby improving the expression ability of the temporal features of physical scenes.

[0036] This embodiment proposes a method for extracting temporal features of physical scenes based on a temporal convolutional neural network. Through the feature decomposition branch, multi-scale feature extraction is achieved, and the dynamic attention mechanism is used for feature adaptive fusion, and finally high-quality temporal feature expressions are obtained. This method shows excellent performance in practical applications such as earthquake waveform recognition, providing an effective solution for the extraction of temporal features of physical scene information.

[0037] In an alternative embodiment, using a graph convolutional neural network to extract the spatial features of physical scene information includes: Construct a spatial relationship graph structure for the physical scene. The spatial relationship graph structure includes a set of sensor nodes and the spatial connection relationships between the nodes. Calculate the spatial correlation weights based on the Euclidean distances between the sensor nodes, and the spatial correlation weights are calculated through a Gaussian kernel function; Map the heterogeneous data of different types of sensors to a unified feature space through the multi-modal feature embedding layer of the graph convolutional neural network. The multi-modal feature embedding layer includes a transformation matrix and a bias term corresponding to the sensor type, and generate node features through a non-linear activation function; Perform multi-layer graph convolution operations on the node features based on the message passing mechanism. The graph convolution operation aggregates features according to the neighborhood set and node degree value of the node to generate the spatial dependence features of the node; Introduce an adaptive spatial attention mechanism to dynamically weight the node features. Calculate the attention coefficients between nodes through a learnable attention vector, and selectively aggregate the spatial dependence features of the nodes according to the attention coefficients to generate the spatial features of the physical scene information.

[0038] Construct a spatial relationship graph structure for the physical scene. This graph structure includes a set of sensor nodes and the spatial connection relationships between the nodes. Specifically, each sensor in the physical scene can be represented as a node in the graph, and connections are established between the nodes according to the spatial position relationship. For example, for an intelligent factory scenario, sensors such as temperature sensors, humidity sensors, and pressure sensors distributed throughout the workshop can be used as nodes in the graph.

[0039] Calculate the spatial correlation weights based on the Euclidean distances between the sensor nodes. Specifically, a Gaussian kernel function can be selected to calculate the weights. For example, for nodes i and j, their spatial correlation weights can be expressed as exp(-d^2 / σ^2), where d is the Euclidean distance between the two nodes, and σ is an adjustable kernel width parameter. In practical applications, an appropriate σ value can be selected according to the scene characteristics, such as σ = 10 meters. In this way, the weights between nodes with closer distances are larger, and the weights between nodes with farther distances are smaller, effectively characterizing the spatial correlation between the nodes.

[0040] Map the heterogeneous data of different types of sensors to a unified feature space through the multi-modal feature embedding layer of the graph convolutional neural network. This multi-modal feature embedding layer includes a transformation matrix and a bias term corresponding to the sensor type. For example, for a temperature sensor, a 100×1 transformation matrix and a 100-dimensional bias vector can be designed to map 1-dimensional temperature data to a 100-dimensional feature space. For an image sensor, a 100×3072 transformation matrix can be designed to map 32×32×3 RGB image data to the same 100-dimensional feature space. Further processed through a non-linear activation function (such as the ReLU function), finally generate the feature representation of each node.

[0041] Perform multi-layer graph convolution operations on node features based on the message passing mechanism. The graph convolution operation aggregates features according to the neighborhood set of nodes and the node degree value to generate the spatial dependence features of nodes. Specifically, for node i, the features of all its neighbor nodes j can be collected first, then normalized according to the degree value of node i, and finally linearly transformed through a learnable weight matrix. For example, 3 layers of graph convolution can be set, and the output dimensions of each layer are 64, 32, and 16 respectively. In this way, through multi-layer graph convolution operations, each node can effectively fuse its neighborhood information and capture local spatial dependence relationships.

[0042] Introduce an adaptive spatial attention mechanism to dynamically weight node features. Specifically, a learnable attention vector can be designed to calculate the attention coefficients between nodes. For example, the attention vector can be set to 16 dimensions, which matches the output dimension of the last layer of graph convolution. By calculating the inner product of the node features and the attention vector and then normalizing through the softmax function, the attention coefficients of each node to other nodes are obtained. According to these attention coefficients, selective aggregation is performed on the spatial dependence features of nodes, and finally the spatial features of the physical scene information are generated.

[0043] In practical applications, the above parameters can be adjusted according to specific scenarios. For example, for a large factory with 1000 sensor nodes, the output dimensions of the graph convolution layer can be increased accordingly, such as set to 256, 128, and 64. The dimension of the attention vector can also be adjusted to 64 dimensions accordingly. This can improve the expressive power of the model and better capture the spatial features in complex scenarios.

[0044] To verify the effectiveness of this method, experiments can be carried out in a specific application scenario. For example, in an intelligent factory environment, 100 temperature sensors, 100 humidity sensors, and 50 pressure sensors are arranged. First, a spatial relationship graph containing 250 nodes is constructed, and the connections between nodes are established based on a distance threshold of 10 meters. Then, the above method is used to extract spatial features, and these features are used to predict the energy consumption situation of the factory.

[0045] During the feature extraction process, the multi-modal feature embedding layer maps the temperature, humidity, and pressure data to a 100-dimensional feature space respectively. Then, through 3 layers of graph convolution (output dimensions of 64, 32, and 16) and the adaptive spatial attention mechanism, 16-dimensional spatial features of each node are generated. Finally, the features of all nodes are concatenated to obtain a 4000-dimensional scene spatial feature vector.

[0046] The experimental results show that using the spatial features extracted by this method for energy consumption prediction can reduce the prediction error by 20% compared with traditional methods (such as directly using the original sensor data). This proves the effectiveness of this method in extracting the spatial features of physical scenarios.

[0047] By visualizing the attention coefficients, it can be found that the model can automatically identify the most important sensor nodes for energy consumption prediction. For example, the temperature sensors near the main production equipment usually have higher attention weights, which is consistent with the actual production experience.

[0048] This method effectively extracts the spatial features of physical scenario information through a graph convolutional neural network and an adaptive spatial attention mechanism. This method can not only process heterogeneous sensor data but also make full use of the spatial relationships between sensors, providing valuable feature representations for subsequent analysis and decision-making tasks.

[0049] Figure 2 Schematic diagram of the energy consumption prediction accuracy results of the embodiments of the present invention: This chart shows the performance comparison of different methods in a certain task. The vertical axis shows the root mean square error (kWh), and the horizontal axis lists five different methods: "linear regression", "random forest", "LSTM", "traditional GCN", and "this method". Judging from the data performance, the error of the linear regression method is the highest, approximately between 240 - 260 kWh; the random forest comes second, with an error in the range of 180 - 220 kWh; the performance of LSTM is slightly better than that of the random forest, with an error of about 160 - 190 kWh; the performance of the traditional GCN is further improved, and the error drops to about 150 - 170 kWh; while the method proposed in this paper performs the best, with the lowest error, only between 130 - 150 kWh. Each method in the figure is represented by two different bars (diagonal filling and dot filling). The overall trend shows that from the traditional linear regression to the new method proposed in this paper, the prediction error shows an obvious downward trend, indicating that the prediction accuracy of the method is continuously improving. This comparison result clearly demonstrates the advantages and disadvantages of various methods in prediction accuracy, highlighting the superiority of the method proposed in this paper.

[0050] In an alternative embodiment, the adaptive sampling strategy is optimized through deep reinforcement learning according to the interaction instruction, and the dynamic adjustment of the sampling strategy of the intelligent IoT sensor includes: Converting the interaction instruction into a corresponding interaction feature vector, obtaining the real-time monitoring data of the intelligent IoT sensor, and constructing a scenario state representation including an environmental state matrix and a data quality matrix according to the real-time monitoring data, where the environmental state matrix includes the environmental parameters collected by each sensor, and the data quality matrix includes the data reliability indexes of each sensor; Construct the combination of the interaction feature vector and the scene state representation into the state space of deep reinforcement learning, calculate the sampling frequency, sampling accuracy, and data transmission strategy according to the state space, and generate a sampling policy action space; Execute a sampling operation based on the sampling policy action space, calculate the data quality score, delay performance score, and coverage performance score of the sampling operation, and perform a weighted combination of the data quality score, delay performance score, and coverage performance score to obtain a comprehensive reward value; Use the comprehensive reward value to update the policy network parameters through the proximal policy optimization algorithm, output optimized sampling parameters based on the updated policy network, and apply the optimized sampling parameters to sensor sampling control, where the policy network is a dual structure including an Actor network for generating an action probability distribution and a Critic network for evaluating the state value.

[0051] Convert the user's interaction instruction into an interaction feature vector. For example, when the user issues an instruction of "improve the temperature monitoring accuracy", the system parses the instruction into semantic features through natural language processing technology and maps it to a predefined instruction type code. For example, the encoding of the instruction type "accuracy adjustment" is 1, the encoding of the instruction parameter "temperature sensor" is 101, and the encoding of the instruction value "improve" is 2. These encodings are combined to form a fixed-dimensional feature vector [1, 101, 2], representing a request to improve the accuracy of the temperature sensor.

[0052] Obtain the real-time monitoring data of the intelligent IoT sensor and construct a scene state representation. This representation consists of two parts: an environmental state matrix and a data quality matrix. The environmental state matrix records the environmental parameters collected by each sensor. For example, in a system containing three sensors of temperature, humidity, and light, the environmental state matrix can be represented as [[25.3, 62.7, 850], [24.9, 63.1, 855], [25.1, 62.5, 852]], corresponding to the three parameter values at three time points respectively. The data quality matrix contains the data reliability indicators of each sensor, such as signal-to-noise ratio, packet loss rate, and battery power, and can be represented as: [[0.92, 0.01, 0.78], [0.88, 0.02, 0.77], [0.90, 0.01, 0.76]], corresponding to the three quality indicators of the three sensors respectively.

[0053] Combine the interaction feature vector with the scene state representation to construct the state space of deep reinforcement learning. For example, for the above interaction feature vector [1, 101, 2] and the scene state representation, the combined state space can be represented as a multi-dimensional vector containing interaction intent, environmental parameters, and quality metrics. Based on this state space, the system calculates the sampling policy action space through the policy network, including sampling frequency, sampling accuracy, and data transmission policy. The sampling frequency can be set to [0.1Hz, 0.5Hz, 1Hz, 2Hz, 5Hz], the sampling accuracy can be set to [8bit, 12bit, 16bit], and the data transmission policy can be set to [real-time transmission, batch transmission, conditional trigger transmission]. For the temperature sensor, the system outputs the action [2Hz, 16bit, real-time transmission], indicating sampling at a frequency of 2Hz with an accuracy of 16bit and transmitting the data in real-time.

[0054] Based on the generated sampling policy action space, the system performs sampling operations and evaluates its performance. The data quality score is calculated based on the accuracy and integrity of the collected data. For example, for the temperature sensor, if the accuracy of the data obtained by sampling with 16bit accuracy is 99.2%, a quality score of 0.992 can be assigned. The latency performance score is calculated based on the latency time from data collection to processing. If the average latency under the real-time transmission policy is 50ms, compared to the target latency of 100ms, a latency score of 0.95 can be assigned. The coverage performance score is calculated based on the monitoring coverage rate of the sensor for the target area. If the effective coverage rate of the temperature sensor under the current sampling policy is 85%, a coverage score of 0.85 can be assigned. These three scores are weighted and combined according to the weights 0.4, 0.3, and 0.3 to obtain a comprehensive reward value of 0.932.

[0055] Using the calculated comprehensive reward value, the system updates the policy network parameters through the Proximal Policy Optimization (PPO) algorithm. The policy network adopts an Actor-Critic dual structure, where the Actor network is responsible for generating the action probability distribution, and the Critic network is responsible for evaluating the state value. For example, for the sampling frequency of the temperature sensor, the Actor network outputs the probability distribution [0.05, 0.15, 0.25, 0.40, 0.15], indicating the probabilities of selecting each frequency value, while the Critic network outputs the value estimate of the current state, such as 0.85. Through the objective function of the PPO algorithm, the system calculates the policy gradient and updates the network parameters to gradually optimize the sampling decision of the policy network.

[0056] The updated policy network outputs optimized sampling parameters. For example, the sampling policy of the temperature sensor is adjusted from [2Hz, 16bit, real-time transmission] to [1Hz, 16bit, conditional trigger transmission], and the system applies these parameters to sensor sampling control. In practical applications, this method can dynamically adjust the sampling policy according to environmental changes and user needs. For example, when it detects that the environmental temperature fluctuation intensifies, it automatically increases the sampling frequency of the temperature sensor; when the battery power decreases, it reduces the sampling accuracy and frequency to extend the device's battery life.

[0057] In an alternative embodiment, updating the policy network parameters using the comprehensive reward value by the proximal policy optimization algorithm includes: Calculating the action probability distribution selected in the current state based on the Actor network, and calculating the value estimate of the current state based on the Critic network; Calculating the probability ratio according to the action probability distribution, where the probability ratio is the ratio of the current policy probability to the old policy probability, and performing clipping processing on the probability ratio to obtain the clipped probability ratio; Using the generalized advantage estimation method to calculate the advantage function estimate value, which is calculated based on the discounted cumulative reward and the temporal difference error of the state value estimate; Multiplying the clipped probability ratio by the advantage function estimate value to construct the proximal policy objective function, calculating the parameter gradient of the Actor network based on the proximal policy objective function, and calculating the parameter gradient of the Critic network based on the state value estimation error; Introducing a KL divergence constraint mechanism to terminate parameter update in advance when the KL divergence between the new and old policies exceeds the preset divergence threshold; Using the parameter gradient of the Actor network and the parameter gradient of the Critic network to update the corresponding network parameters respectively.

[0058] The system includes two main network components: the Actor network and the Critic network. The Actor network is responsible for generating action policies, that is, the probability distribution of selecting actions in a given state; the Critic network is responsible for evaluating the state value, that is, predicting the future cumulative reward that can be obtained in the current state. These two networks work together and continuously optimize the control strategy through interactive learning.

[0059] When updating the parameters of the policy network, the probability distribution P(a|s) of selecting each action a in state s is calculated based on the current Actor network. For example, in a robot control task, if state s represents the current position and speed information of the robot (x = 2.5 meters, v = 1.2 meters / second), the Actor network outputs the action probability distribution as: forward (0.7), backward (0.1), left turn (0.1), right turn (0.1). At the same time, the Critic network calculates the value estimate V (s) , such as V (s) = 8.5, indicating the discounted cumulative reward expected to be obtained starting from this state.

[0060] Calculate the probability ratio r(θ), which is the ratio of the current policy probability to the old policy probability. Suppose that under the old policy before the update, the probability of selecting the "forward" action is 0.6, then the probability ratio r(θ) = 0.7 / 0.6 = 1.167. To prevent overly large policy updates from causing training instability, the system clips the probability ratio, setting the clipping range to [0.8, 1.2], then the clipped probability ratio is min(1.167, 1.2) = 1.167.

[0061] The Generalized Advantage Estimation (GAE) method is used to calculate the advantage function estimate A(s,a). This estimate is calculated based on the discounted cumulative reward and the temporal difference error of the state value estimate. Specifically, for state s and the selected action a, if the actual immediate reward r = 2.0, the value estimate of the next state s' is V(s') = 7.8, and the discount factor γ = 0.99, then the temporal difference error is: δ = r + γV(s') - V(s) = 2.0 + 0.99×7.8 - 8.5 = 1.222.

[0062] By accumulating the multi-step temporal difference errors, the system calculates the advantage function estimate A(s,a) = 1.5, indicating how much better the selected action a is compared to the average performance.

[0063] Multiply the clipped probability ratio by the advantage function estimate to construct the proximal policy objective function L^CLIP. For the above example, L^CLIP = 1.167×1.5 = 1.7505. The system calculates the parameter gradients of the Actor network based on this objective function. At the same time, the parameter gradients of the Critic network are calculated based on the state value estimation error. If the true cumulative reward is 9.0 and the estimate of the Critic network is 8.5, then the mean squared error is (9.0 - 8.5)² = 0.25, and the system calculates the parameter gradients of the Critic network accordingly.

[0064] To prevent large policy updates from causing training instability, a KL divergence constraint mechanism is introduced. KL divergence measures the degree of difference between the new and old policy distributions. For example, if the new policy probability distribution is [0.7, 0.1, 0.1, 0.1] and the old policy is [0.6, 0.2, 0.1, 0.1], the calculated KL divergence is 0.15. Suppose the preset divergence threshold is 0.2. Since 0.15 < 0.2, the system continues with parameter updates; if the KL divergence exceeds 0.2, the system will terminate this parameter update in advance.

[0065] The parameter gradients of the Actor network and the Critic network are used to update the corresponding network parameters respectively. Assume the current value of the Actor network parameter θ is [0.25, 0.31, -0.42, 0.18], the learning rate α = 0.001, and the calculated parameter gradient is [0.12, 0.08, -0.15, 0.05]. Then the updated parameter is [0.25 + 0.001×0.12, 0.31 + 0.001×0.08, -0.42 + 0.001×(-0.15), 0.18 + 0.001×0.05] = [0.25012, 0.31008, -0.42015, 0.18005]. Similarly, the Critic network parameters are updated in a similar manner.

[0066] Generally, a large amount of state - action - reward data is collected over multiple episodes, and then the above update process is executed in batches. For example, in an autonomous driving scenario, 1000 state - action pairs are collected, each mini - batch contains 64 samples, and 5 training epochs are executed. After each update, the system evaluates the policy performance. For example, if the average episode reward increases from 320 to 385, it indicates that the policy quality has improved.

[0067] Figure 3 Schematic diagram for comparing the cumulative reward values of different optimization algorithms in the embodiments of the present invention: This figure shows the performance comparison of different reinforcement learning algorithms during the training process. The horizontal axis represents the number of training iterations (from 0 to 500), and the vertical axis represents the cumulative reward value (from 0 to 1100). Five algorithms are compared in the figure: the proposed technical solution (PPO with KL constraint), basic PPO, A2C, TRPO, and DQN. From the data of performance improvement, the proposed technical solution has increased by 15.5% compared with basic PPO, 27.2% compared with TRPO, and 66.3% compared with DQN. After 500 training iterations, the final cumulative reward values of each algorithm are: the proposed technical solution reaches 1023, basic PPO reaches 886, TRPO reaches 804, A2C reaches 678, and DQN is the lowest at 615. From the curve trend, it can be observed that all algorithms grow relatively fast in the initial stage of training (0 - 100 iterations), and then the growth rate gradually slows down. However, the proposed technical solution always maintains a leading edge and still shows a stable upward trend in the later stage of training, demonstrating excellent learning effects and stability.

[0068] In the existing reinforcement learning technologies, traditional policy optimization methods such as DQN and A2C often have problems such as low learning efficiency, slow convergence speed, and unstable performance when dealing with complex environments. This is mainly because these methods lack an effective constraint mechanism during policy update, and it is easy to have an overly large policy update step size, resulting in an unstable training process. Although basic PPO and TRPO introduce trust region constraints to improve training stability, their constraint methods are relatively simple, and it is difficult to achieve a fast learning speed while ensuring training stability.

[0069] The PPO algorithm based on KL constraint proposed in this embodiment innovatively designs an adaptive KL divergence constraint mechanism to dynamically adjust the step size of policy update, effectively balancing the relationship between exploration and exploitation. This method not only inherits the stability advantage of the PPO algorithm but also improves the learning efficiency of the algorithm through an optimized constraint form. Specifically, during the policy update process of this solution, the tightness of the KL divergence constraint is adaptively adjusted according to historical training data. While ensuring the stability of policy update, the algorithm is allowed to make larger step updates at appropriate times, thereby accelerating the learning speed.

[0070] The experimental results show that the technical solution of this embodiment has significant advantages compared with the existing technologies: First, it shows a faster learning speed and higher reward acquisition ability in the initial stage of training; second, it maintains a stable performance improvement trend throughout the training process without obvious performance fluctuations; finally, it can still continue to improve in the later stage of training and finally reach a better performance level. These improvement effects fully prove the superiority of this solution in reinforcement learning tasks and provide a more effective solution for intelligent decision-making in complex environments.

[0071] In an alternative embodiment, calculating the advantage function estimate using the Generalized Advantage Estimation (GAE) method includes: Calculating the advantage function based on the state-value function and the action-value function, where the advantage function is expressed as the difference between the action-value function and the state-value function for a state-action pair; Introducing an exponentially weighted coefficient based on the advantage function to construct the Generalized Advantage Estimation function, where the Generalized Advantage Estimation function includes the weighted sum of an infinite sequence of temporal difference errors, and the temporal difference error is calculated from the immediate reward value, the discount factor, and the state-value function at adjacent time steps; Using a recursive approach to add the temporal difference error at the current time step to the advantage estimate at the next time step adjusted by the discount factor and the Generalized Advantage Estimation parameter to obtain the estimate of the Generalized Advantage Estimation function; Introducing a state-dependent baseline function to correct the estimate, where the baseline function is calculated based on the expected return under the policy, and the correction includes subtracting the baseline function value from the actual return; Normalizing the corrected estimate, where the normalization includes standardizing the estimate of the Generalized Advantage Estimation function using the mean of the advantage estimates within a batch.

[0072] The agent obtains experience by interacting with the environment and optimizes its decision-making strategy based on this experience. In practical applications, the estimation of the advantage function plays an important role in reducing the variance of the policy gradient method. This embodiment elaborates in detail a calculation method based on Generalized Advantage Estimation.

[0073] Calculate the advantage function based on the state-value function and the action-value function. The state-value function represents the expected cumulative return that can be obtained by starting from a certain state and acting according to the current policy; the action-value function represents the expected cumulative return that can be obtained by performing a specific action in a certain state and then acting according to the current policy. The advantage function is expressed as the difference between the action-value function and the state-value function, i.e., A (s,a) = Q (s,a) - V(s), where s represents the state, a represents the action, Q represents the action-value function, and V represents the state-value function.

[0074] In a certain robot navigation task, the state s represents the position of the robot at coordinates (10, 15), and the action a represents moving north. If the value V (s) of this state is 85 points, and the action-value Q (s,a) after moving north in this state is 92 points, then the advantage function A (s,a) = 92 - 85 = 7 points, indicating that this action is 7 points better than the average performance.

[0075] Next, an exponentially weighted coefficient is introduced based on the advantage function to construct a generalized advantage estimation function. This function includes the weighted sum of an infinite sequence of temporal difference errors. The temporal difference error is calculated from the immediate reward value, the discount factor, and the state value function at adjacent time steps, denoted as δ t = r t + γV (st+1) - V (st) , where rt represents the immediate reward obtained at time step t, γ represents the discount factor (usually taking values between 0.9 and 0.99), s t and s t+1 represent the states at time step t and t + 1 respectively.

[0076] In actual calculation, the generalized advantage estimate value is calculated recursively. The generalized advantage estimate value at the current time step t is equal to the sum of the temporal difference error at the current moment and the advantage estimate value at the next time step adjusted by the discount factor and the generalized advantage estimation parameter. Specifically, it can be expressed as A t = δ t + (γλ)A t t+1, where λ is the generalized advantage estimation parameter, usually taking values between 0.9 and 0.99, and is used to control the trade-off between short-term and long-term rewards.

[0077] For example, assume that in three consecutive time steps, the temporal difference errors are δ t = 2.5, δ t t+1 = 1.8, δ t t+2= 3.2, the discount factor γ = 0.95, and the generalized advantage estimation parameter λ = 0.97. Then, starting from the last time step, calculate recursively: A t t+2 = δ t t+2 = 3.2, A t t+1 = δ t t+1 + (γλ)A t t+2 = 1.8 + (0.95 × 0.97) × 3.2 ≈4.74, A t t = δ t t + (γλ)A t t+1 = 2.5 + (0.95 × 0.97) × 4.74 ≈ 6.88.

[0078] To improve the accuracy of the estimation, a state-dependent baseline function is introduced to correct the estimated value. The baseline function is usually calculated based on the expected return under the policy, and the correction process includes subtracting the baseline function value from the actual return. In practice, the state value function V (s) is usually used as the baseline function.

[0079] For example, if the expected return based on the current policy for a certain state s is 78, and the actual cumulative return for this state in the sampled trajectory is 85, then the corrected return value is 85 - 78 = 7, indicating that the performance of this trajectory is 7 units better than expected.

[0080] Finally, normalize the corrected estimates to improve training stability. The normalization process involves standardizing the estimates of the corrected generalized advantage estimation function using the mean and standard deviation of the advantage estimates within the batch. Specifically, for each advantage estimate A within the batch, the normalized value is calculated as (A - μ) / σ, where μ represents the mean of all advantage estimates within the batch and σ represents the standard deviation.

[0081] For example, assume that in a batch containing 100 samples, the calculated average of the advantage estimates is 1.2, the standard deviation is 3.5, and the advantage estimate for a certain sample is 4.8. Then the normalized value is (4.8 - 1.2) / 3.5 = 1.03.

[0082] Collect the trajectory data of the agent's interaction with the environment, including the state sequence s0, s1, ..., s T , the action sequence a0, a1, ..., a T-1 , and the immediate reward sequence r0, r1, ..., r T-1 . Use a neural network or other function approximator to estimate the state-value function V (s) . For each time step t in the trajectory, calculate the temporal difference error δ t = rt + γV (st+1) - V (st) . If t is the terminal time step, then δ t = r t - V (st) .

[0083] Starting from the last time step, recursively calculate the generalized advantage estimates for each time step. First, set A T -1 = δ T -1, and then calculate A T -2, A T -3, ..., A0 in turn, and its calculation formula is A t = δ t + (γλ)A t+1. Based on the calculated generalized advantage estimate and the state value function as the baseline, correct the rewards for each state-action pair in the trajectory. Collect multiple trajectory data to form a batch, calculate the mean μ and standard deviation σ of all advantage estimates within the batch, and then normalize each advantage estimate to obtain the final advantage function estimate.

[0084] In the second aspect of the embodiments of the present invention, there is provided a data acquisition system for intelligent IoT sensors in the virtual-real fusion of the metaverse, including: A first unit for collecting physical scene information through intelligent IoT sensors and performing digital conversion and preprocessing on the collected physical scene information; A second unit for transmitting the preprocessed physical scene information to a data center server through an IoT protocol and performing adaptive fusion processing on the physical scene information based on a deep spatio-temporal association network, including: extracting temporal features of the physical scene information using a temporal convolutional neural network, extracting spatial features of the physical scene information using a graph convolutional neural network, fusing the temporal features and the spatial features to obtain a scene feature vector, and generating standardized scene feature data according to the scene feature vector; A third unit for constructing a metaverse virtual scene based on the standardized scene feature data and displaying the metaverse virtual scene on a metaverse interaction terminal; A fourth unit for realizing two-way mapping and interaction between virtual and real scenes through an intelligent perception engine, including: receiving an interaction instruction from a user on a metaverse interaction terminal, optimizing an adaptive sampling strategy through deep reinforcement learning according to the interaction instruction, and dynamically adjusting the sampling strategy of intelligent IoT sensors, including adjusting the sampling frequency, sampling accuracy, and data transmission strategy; collecting updated physical scene information based on the sampling strategy and mapping the updated physical scene information to the metaverse virtual scene.

[0085] In the third aspect of the embodiments of the present invention, There is provided an electronic device, including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.

[0086] In the fourth aspect of the embodiments of the present invention, There is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.

[0087] The present invention may be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for performing various aspects of the present invention.

[0088] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A data acquisition method for intelligent Internet of Things sensors in the virtual-real fusion of the metaverse, characterized in that, Including: Collect physical scene information through intelligent Internet of Things sensors, and perform digital conversion and preprocessing on the collected physical scene information; Transmit the preprocessed physical scene information to the data center server through the Internet of Things protocol, and perform adaptive fusion processing on the physical scene information based on the deep spatio-temporal correlation network, including: using a temporal convolutional neural network to extract the temporal features of the physical scene information, using a graph convolutional neural network to extract the spatial features of the physical scene information, fusing the temporal features and the spatial features to obtain a scene feature vector, and generating standardized scene feature data according to the scene feature vector; Construct a metaverse virtual scene based on the standardized scene feature data, and display the metaverse virtual scene on the metaverse interaction terminal; Realize the bidirectional mapping and interaction between the virtual and real scenes through the intelligent perception engine, including: receiving the interaction instructions of the user on the metaverse interaction terminal, optimizing the adaptive sampling strategy through deep reinforcement learning according to the interaction instructions, and dynamically adjusting the sampling strategy of the intelligent Internet of Things sensors, including adjusting the sampling frequency, sampling accuracy, and data transmission strategy; collecting the updated physical scene information based on the sampling strategy, and mapping the updated physical scene information to the metaverse virtual scene.

2. The method according to claim 1, wherein Using a temporal convolutional neural network to extract the temporal features of the physical scene information includes: The temporal convolutional neural network includes a feature decomposition branch and a feature reconstruction branch, where the feature decomposition branch is used to decompose the physical scene information into multi-scale temporal features, and the feature reconstruction branch is used to adaptively fuse the multi-scale temporal features by combining a dynamic attention mechanism to generate a temporal feature representation of the physical scene.

3. The method according to claim 1, wherein Using a graph convolutional neural network to extract the spatial features of the physical scene information includes: Construct a spatial relationship graph structure of the physical scene, where the spatial relationship graph structure includes a set of sensor nodes and the spatial connection relationships between the nodes, calculate the spatial correlation weights based on the Euclidean distance between the sensor nodes, and the spatial correlation weights are calculated through a Gaussian kernel function; Map the heterogeneous data of different types of sensors to a unified feature space through the multi-modal feature embedding layer of the graph convolutional neural network, where the multi-modal feature embedding layer includes a transformation matrix and a bias term corresponding to the sensor type, and generate node features through a non-linear activation function; Perform multi-layer graph convolutional operations on the node features based on the message passing mechanism, where the graph convolutional operation aggregates features according to the neighborhood set and node degree value of the node to generate the spatial dependence features of the node; Introduce an adaptive spatial attention mechanism to dynamically weight the node features, calculate the attention coefficients between the nodes through a learnable attention vector, and selectively aggregate the spatial dependence features of the nodes according to the attention coefficients to generate the spatial features of the physical scene information.

4. The method according to claim 1, wherein Optimizing the adaptive sampling strategy through deep reinforcement learning according to the interaction instructions, and dynamically adjusting the sampling strategy of the intelligent Internet of Things sensors includes: Convert the interaction instruction into a corresponding interaction feature vector, obtain the real-time monitoring data of the intelligent IoT sensor, and construct a scenario state representation including an environmental state matrix and a data quality matrix based on the real-time monitoring data, where the environmental state matrix includes environmental parameters collected by each sensor, and the data quality matrix includes data reliability indicators of each sensor; Combine the interaction feature vector and the scenario state representation to construct a state space for deep reinforcement learning, calculate the sampling frequency, sampling accuracy, and data transmission strategy based on the state space, and generate a sampling strategy action space; Perform a sampling operation based on the sampling strategy action space, calculate the data quality score, delay performance score, and coverage performance score of the sampling operation, and perform a weighted combination of the data quality score, delay performance score, and coverage performance score to obtain a comprehensive reward value; Use the comprehensive reward value to update the policy network parameters through the proximal policy optimization algorithm, output optimized sampling parameters based on the updated policy network, and apply the optimized sampling parameters to sensor sampling control, where the policy network is a dual structure including an Actor network for generating an action probability distribution and a Critic network for evaluating the state value.

5. The method according to claim 4, characterized in that, Using the comprehensive reward value to update the policy network parameters through the proximal policy optimization algorithm includes: Calculate the action probability distribution selected in the current state based on the Actor network, and calculate the value estimate of the current state based on the Critic network; Calculate the probability ratio according to the action probability distribution, where the probability ratio is the ratio of the current policy probability to the old policy probability, and perform a clipping process on the probability ratio to obtain the clipped probability ratio; Calculate the advantage function estimate value using the generalized advantage estimation method, where the advantage function estimate value is calculated based on the discounted cumulative reward and the temporal difference error of the state value estimate; Multiply the clipped probability ratio by the advantage function estimate value to construct a proximal policy objective function, calculate the parameter gradient of the Actor network based on the proximal policy objective function, and calculate the parameter gradient of the Critic network based on the state value estimation error; Introduce a KL divergence constraint mechanism, and terminate the parameter update in advance when the KL divergence between the new and old policies exceeds the preset divergence threshold; Use the parameter gradient of the Actor network and the parameter gradient of the Critic network to update the corresponding network parameters respectively.

6. The method according to claim 5, characterized in that Calculating the advantage function estimate value using the generalized advantage estimation method includes: Calculate the advantage function according to the state value function and the action value function, where the advantage function is expressed as the difference between the action value function and the state value function of the state-action pair; Introduce an exponential weighting coefficient based on the advantage function to construct a generalized advantage estimation function, where the generalized advantage estimation function includes an infinite sequence weighted sum of temporal difference errors, and the temporal difference error is calculated through the immediate reward value, the discount factor, and the state value function of adjacent time steps; The estimated value of the generalized advantage estimation function is obtained by recursively adding the temporal difference error at the current time step to the advantage estimation value at the next time step adjusted by the discount factor and the generalized advantage estimation parameter. A state-dependent baseline function is introduced to correct the estimated value, where the baseline function is calculated based on the expected return under the policy, and the correction includes subtracting the baseline function value from the actual return. Normalization processing is performed on the corrected estimated value, where the normalization processing includes standardizing the estimated value of the corrected generalized advantage estimation function using the mean of the advantage estimation values within the batch.

7. An intelligent IoT sensor in a data acquisition system for the virtual-real fusion in the metaverse, which is used to implement the method described in any one of claims 1-6, characterized in that, It includes: The first unit is used to collect physical scene information through intelligent IoT sensors, and perform digital conversion and preprocessing on the collected physical scene information. The second unit is used to transmit the preprocessed physical scene information to the data center server through the IoT protocol, and perform adaptive fusion processing on the physical scene information based on the deep spatio-temporal correlation network, including: extracting the temporal features of the physical scene information using a temporal convolutional neural network, extracting the spatial features of the physical scene information using a graph convolutional neural network, fusing the temporal features and the spatial features to obtain a scene feature vector, and generating standardized scene feature data according to the scene feature vector. The third unit is used to construct a metaverse virtual scene based on the standardized scene feature data, and display the metaverse virtual scene on the metaverse interaction terminal. The fourth unit is used to achieve two-way mapping and interaction between the virtual and real scenes through an intelligent perception engine, including: receiving the interaction instructions of the user on the metaverse interaction terminal, optimizing the adaptive sampling strategy through deep reinforcement learning according to the interaction instructions, and dynamically adjusting the sampling strategy of the intelligent IoT sensors, including adjusting the sampling frequency, sampling accuracy, and data transmission strategy; collecting the updated physical scene information based on the sampling strategy, and mapping the updated physical scene information to the metaverse virtual scene.

8. An electronic device, characterized in that, It includes: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, The computer program instructions, when executed by the processor, implement the method according to any one of claims 1 to 6.

Citation Information

Cited By

  • Aircraft multi-target performance constraint collaborative optimization method based on multi-dimensional virtual-real parameter mapping

    CN121806936A

  • Aircraft multi-objective performance constraint collaborative optimization method based on multi-dimensional virtual-real parameter mapping

    CN121806936B