Intelligent internet of things and element universe virtual-real fusion identification method and system based on deep learning
Through deep learning and graph neural network technology, the effective fusion of user behavior data in the virtual environment and the real environment is achieved, and the multimodal data fusion problem in the virtual and real environment is solved, and the accuracy of virtual and real behavior consistency recognition and the robustness of the system are improved.
Patent Information
- Application Number
- CN202510579056.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-07
AI Technical Summary
The prior art is difficult to effectively integrate user behavior data in virtual environments and real environments, resulting in the problem of multimodal data fusion in virtual and real environments that cannot be effectively solved, affecting the accurate judgment of behavior consistency.
The virtual behavior data and actual behavior data are extracted through a pre-trained deep neural network model, and the adaptive feature alignment of virtual and real behavior is achieved based on the attention mechanism. The graph neural network is used to build a time sequence relationship diagram of virtual and real behavior, model the time sequence dependence relationship between nodes through conditional random fields, obtain timing context features, and integrate the timing context features of multiple subgraphs through a cross-subgraph feature fusion layer to generate a fused timing feature vector.
It effectively solves the problem of heterogeneity of user behavior characteristics in virtual and real scenarios, improves the accuracy of identification of virtual and real behavior consistency, enhances the ability to model complex behavior patterns, improves the robustness and generalization capabilities of the system, and ensures the security and user experience of the metacosmic platform.
Smart Images

Figure CN120105041A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to metaverse technology, and in particular to a method and system for identifying the virtual-real fusion of intelligent Internet of Things and metaverse based on deep learning. Background Art
[0002] With the rapid development of Metaverse and IoT technologies, the boundary between the virtual world and the real world is gradually blurred, and users can move around in both virtual and real environments at the same time. As an immersive virtual space, Metaverse allows users to socialize, work, and entertain through digital avatars, while IoT technology enables various devices in the real world to interconnect and sense and collect user behavior data in real time. However, the problem of information islands between the virtual world and the real world still exists, and the user's behavior data in the two environments is often fragmented, lacking an effective association and fusion mechanism.
[0003] At present, the field of virtual-real behavior recognition and fusion has the following deficiencies: Existing behavior recognition methods usually only focus on user behavior analysis in a single environment, and cannot effectively handle the problem of multimodal data fusion in virtual and real environments, resulting in a one-sided and incomplete understanding of user behavior. Traditional feature extraction and fusion algorithms lack the ability to adaptively align heterogeneous data in virtual and real scenarios, making it difficult to capture the deep semantic association between virtual and real behaviors, affecting the accurate judgment of behavioral consistency. Existing technologies have limited capabilities in processing the temporal dependencies of virtual and real behaviors, and cannot fully explore the complex relationship between long-term behavior patterns and short-term behavior changes, making it difficult to accurately verify user identities and behaviors.
[0004] Therefore, there is an urgent need for a method that can effectively integrate user behavior data in virtual environments and real environments. Through deep learning and multimodal fusion technology, a correlation model between virtual and real behaviors can be established to accurately identify and verify the consistency of user virtual and real behaviors, providing technical support for secure access control and user experience optimization in the metaverse environment. Summary of the invention
[0005] The embodiments of the present invention provide a method and system for identifying the virtual-reality fusion of intelligent Internet of Things and metaverse based on deep learning, which can solve the problems in the prior art.
[0006] According to a first aspect of the embodiments of the present invention, Obtain the user's virtual behavior data in the Metaverse virtual scene collected by the terminal device and the user's actual behavior data in the real scene collected by the IoT device; Extracting features from the virtual behavior data and the actual behavior data using a pre-trained deep neural network model to obtain a virtual behavior feature vector and an actual behavior feature vector; Cross-fusing the virtual behavior feature vector and the actual behavior feature vector based on an attention mechanism to generate a virtual-real fusion feature vector, wherein the attention mechanism realizes adaptive feature alignment of virtual and real behaviors by calculating a correlation weight between the virtual behavior feature vector and the actual behavior feature vector; Use graph neural networks to build a temporal relationship graph of virtual and real behaviors, model the temporal dependency between nodes through conditional random fields, perform multi-hop information propagation on node features, obtain temporal context features, and integrate the temporal context features of multiple subgraphs through the cross-subgraph feature fusion layer to generate a fused temporal feature vector. The consistency of the user's virtual and real behaviors is identified according to the fused temporal feature vector to obtain a virtual and real behavior consistency score; and the user's behavior in the virtual scene of the metaverse is verified and authorized based on the virtual and real behavior consistency score.
[0007] Cross-fusing the virtual behavior feature vector and the actual behavior feature vector based on the attention mechanism to generate a virtual-real fusion feature vector includes: Calculating an attention weight matrix between the virtual behavior feature vector and the actual behavior feature vector, wherein, for each feature element in the virtual behavior feature vector, a correlation score between the feature element and each feature element in the actual behavior feature vector is calculated by a dot product operation, and the correlation score is normalized to obtain the attention weight matrix; Performing weighted aggregation on the virtual behavior feature vector and the actual behavior feature vector based on the attention weight matrix to obtain a context feature vector of the virtual behavior and a context feature vector of the actual behavior, wherein the context feature vector of the virtual behavior is obtained by matrix multiplication of the attention weight matrix and the actual behavior feature vector, and the context feature vector of the actual behavior is obtained by matrix multiplication of the transpose of the attention weight matrix and the virtual behavior feature vector; Performing feature splicing on the virtual behavior feature vector and the context feature vector of the virtual behavior, and performing feature splicing on the actual behavior feature vector and the context feature vector of the actual behavior, to obtain an enhanced virtual behavior feature vector and an enhanced actual behavior feature vector; The enhanced virtual behavior feature vector and the enhanced actual behavior feature vector are fused through nonlinear transformation to generate a virtual-real fusion feature vector.
[0008] The graph neural network is used to construct the temporal relationship graph of virtual and real behaviors. The temporal dependency relationship between nodes is modeled through the conditional random field. The node features are propagated through multi-hop information to obtain the temporal context features. The temporal context features of multiple subgraphs are integrated through the cross-subgraph feature fusion layer to generate a fused temporal feature vector including: A graph neural network is used to construct a time series relationship graph of virtual and real behaviors, and the virtual and real fusion feature vector of each time step is mapped to a node feature in the time series relationship graph of virtual and real behaviors; The temporal dependency between nodes is modeled through conditional random fields, and the causal relationship strength between nodes is dynamically predicted based on the changing trend of node features and historical interaction patterns; The causal association strength is normalized to obtain a sampling probability between node pairs; a sampling probability threshold is set, and node pairs with a probability higher than the sampling probability threshold are selected as key timing paths; with the key timing path as the center, a plurality of local subgraphs are obtained by sampling through a random walk algorithm; The following operations are performed on each of the local subgraphs: Construct a local graph neural network and set up multiple layers of graph convolution. In each layer of graph convolution, aggregate the neighbor features of the node and update the node representation. Preserve the node features of each layer through skip connections to obtain the node temporal context features containing multi-hop information. The cross-sub-image feature fusion layer integrates the information of multiple local sub-images, including: The attention weights of corresponding node features in different subgraphs are calculated; the node temporal context features of multiple subgraphs are weightedly fused based on the attention weights; and the final fused temporal feature vector is obtained through a pooling operation.
[0009] The temporal dependency between nodes is modeled through conditional random fields. Based on the changing trend of node features and historical interaction patterns, the causal relationship strength between nodes is dynamically predicted, including: Modeling the temporal dependency relationship between nodes in the temporal relationship graph of virtual and real behaviors through conditional random fields, and calculating the difference vector of adjacent node features as the change trend of node features; Based on the interaction frequency of node pairs in the historical time series window, the historical interaction pattern characteristics of node pairs are constructed; The change trend of the node features and the historical interaction pattern features are input into a conditional random field to predict the causal association strength between node pairs.
[0010] Construct a local graph neural network and set up multiple layers of graph convolution. In each layer of graph convolution, aggregate the neighbor features of the node and update the node representation. Retain the node features of each layer through skip connections, and obtain the node temporal context features containing multi-hop information, including: Receiving input node feature vectors and graph structure information, wherein the node feature vectors represent initial features of each node, and the graph structure information represents connection relationships between nodes; Construct a multi-layer graph convolutional network structure, including L graph convolutional layers, each of which contains a learnable weight matrix and a bias vector, where L is an integer greater than 1; For the lth graph convolution layer among the L graph convolution layers, the following feature propagation operations are performed: Obtaining a set of neighbor nodes of a node; normalizing features in the set of neighbor nodes; multiplying the normalized neighbor nodes by a weight matrix of the lth graph convolutional layer; Perform nonlinear transformation on the product result to obtain the output features of the node at the lth layer; The intermediate feature representation of nodes in each graph convolutional layer is retained through the skip connection mechanism, including: The output features of the nodes in each graph convolution layer are concatenated; the concatenated multi-layer features are reduced in dimension using the feature transformation layer; The feature propagation process of the graph convolution layer is iteratively performed L times, and the node features are updated in each iteration to obtain the node temporal context features containing multi-hop information.
[0011] Determining the behavior sequence matching degree through a dynamic programming algorithm according to the fused time series feature vector, and obtaining a virtual-real behavior consistency score in combination with a preset user behavior feature template; verifying and authorizing the user's behavior in the virtual scene of the metaverse based on the virtual-real behavior consistency score includes: Extracting a behavior feature point sequence from the fused time series feature vector, calculating the longest common subsequence of the current behavior feature point sequence and the historical behavior feature point sequence based on a dynamic programming algorithm, and obtaining a behavior sequence matching degree; According to the matching degree of the behavior sequence, the importance of the feature point sequence of the user's historical behavior characteristics is weighted to construct an adaptive user behavior feature template, wherein the importance weight decays with time distance; Calculating the multi-scale correlation score between the current fused time series feature vector and the adaptive user behavior feature template, and detecting the behavior mutation point based on the sliding time window to generate a behavior consistency score; When the behavior consistency score is higher than a preset score threshold, the user's operation request in the metaverse scene is permitted; When the behavior consistency score is lower than the preset score threshold, the identity re-authentication mechanism is triggered and the user's operation authority is temporarily frozen.
[0012] Extracting a behavior feature point sequence from the fused time series feature vector, calculating the longest common subsequence of the current behavior feature point sequence and the historical behavior feature point sequence based on a dynamic programming algorithm, and obtaining the behavior sequence matching degree includes: Perform extreme point detection on the time series feature vector to extract a current behavior feature point sequence; Acquire a historical behavior feature point sequence from a user behavior database, and construct the current behavior feature point sequence and the historical behavior feature point sequence into two sequences to be matched; Constructing a dynamic programming matrix, wherein the rows and columns of the dynamic programming matrix correspond to the lengths of the current behavior feature point sequence and the historical behavior feature point sequence respectively; Calculate the longest common subsequence of two sequences based on the dynamic programming algorithm: Initialize the matrix boundary conditions; when the difference between the feature point values is less than the preset difference threshold, the matrix element at the corresponding position is increased by 1; when the difference between the feature point values is greater than the preset difference threshold, the maximum value of the adjacent positions is taken; Iteratively update the matrix elements until all sequence positions are traversed; According to the final value of the dynamic programming matrix, the length of the longest common subsequence is calculated and divided by the total length of the sequence to obtain the behavior sequence matching degree.
[0013] A second aspect of an embodiment of the present invention provides a deep learning-based intelligent IoT and metaverse virtual-real fusion recognition system, including: The first unit is used to obtain the virtual behavior data of the user in the virtual scene of the Metaverse collected by the terminal device and the actual behavior data of the user in the real scene collected by the IoT device; A second unit is used to extract features from the virtual behavior data and the actual behavior data through a pre-trained deep neural network model to obtain a virtual behavior feature vector and an actual behavior feature vector; A third unit is used to cross-fuse the virtual behavior feature vector and the actual behavior feature vector based on an attention mechanism to generate a virtual-real fusion feature vector, wherein the attention mechanism calculates the correlation weight between the virtual behavior feature vector and the actual behavior feature vector to achieve adaptive feature alignment of virtual and real behaviors; The fourth unit is used to construct a temporal relationship graph of virtual and real behaviors using graph neural networks, model the temporal dependency relationship between nodes through conditional random fields, perform multi-hop information propagation on node features, obtain temporal context features, integrate the temporal context features of multiple subgraphs through a cross-subgraph feature fusion layer, and generate a fused temporal feature vector; The fifth unit is used to identify the consistency of the user's virtual and real behaviors according to the fused time series feature vector to obtain a virtual and real behavior consistency score; and verify and authorize the user's behavior in the virtual scene of the metaverse based on the virtual and real behavior consistency score.
[0014] According to a third aspect of the embodiments of the present invention, An electronic device is provided, comprising: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0015] A fourth aspect of the embodiments of the present invention is: A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.
[0016] The beneficial effects of this application are as follows: The present invention extracts features of virtual behavior data and actual behavior data through a pre-trained deep neural network model, and realizes adaptive feature alignment of virtual and real behaviors based on the attention mechanism, which effectively solves the problem of heterogeneity of user behavior features in virtual scenes and real scenes, and improves the accuracy of consistency recognition of virtual and real behaviors.
[0017] The present invention uses graph neural networks to construct a temporal relationship graph of virtual and real behaviors, models the temporal dependency between nodes through conditional random fields, and combines cross-subgraph feature fusion technology to capture the correlation between the temporal context information of user behavior and multimodal features, thereby enhancing the ability to model complex behavior patterns and improving the robustness and generalization ability of the system.
[0018] The present invention verifies and authorizes the user's behavior in the virtual scene of the Metaverse based on the virtual-real behavior consistency score, constructs a complete virtual-real behavior verification mechanism, effectively prevents identity fraud and abnormal behavior in the Metaverse environment, ensures the security and user experience of the Metaverse platform, and promotes the deep integration of the Metaverse and Internet of Things technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 It is a flowchart of a method for identifying the virtual-real fusion of intelligent Internet of Things and metaverse based on deep learning according to an embodiment of the present invention; Figure 2 It is a flowchart of constructing a time series relationship diagram of virtual and real behaviors and fusion of features according to an embodiment of the present invention; Figure 3 This is a flowchart of node feature processing based on a graph neural network according to an embodiment of the present invention; Figure 4 The figure is a flowchart of user verification based on behavior sequence matching according to an embodiment of the present invention. DETAILED DESCRIPTION
[0020] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0021] The technical solution of the present invention is described in detail with specific embodiments below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.
[0022] Figure 1 Schematic diagram of the process of the method for identifying the virtual-real fusion of intelligent IoT and metaverse based on deep learning according to an embodiment of the present invention. Figure 1 As shown, the method includes: Obtain the user's virtual behavior data in the Metaverse virtual scene collected by the terminal device and the user's actual behavior data in the real scene collected by the IoT device; Extracting features from the virtual behavior data and the actual behavior data using a pre-trained deep neural network model to obtain a virtual behavior feature vector and an actual behavior feature vector; Cross-fusing the virtual behavior feature vector and the actual behavior feature vector based on an attention mechanism to generate a virtual-real fusion feature vector, wherein the attention mechanism realizes adaptive feature alignment of virtual and real behaviors by calculating a correlation weight between the virtual behavior feature vector and the actual behavior feature vector; Use graph neural networks to build a temporal relationship graph of virtual and real behaviors, model the temporal dependency between nodes through conditional random fields, perform multi-hop information propagation on node features, obtain temporal context features, and integrate the temporal context features of multiple subgraphs through the cross-subgraph feature fusion layer to generate a fused temporal feature vector. The consistency of the user's virtual and real behaviors is identified according to the fused temporal feature vector to obtain a virtual and real behavior consistency score; and the user's behavior in the virtual scene of the metaverse is verified and authorized based on the virtual and real behavior consistency score.
[0023] In an optional implementation, cross-fusing the virtual behavior feature vector and the actual behavior feature vector based on an attention mechanism to generate a virtual-real fusion feature vector includes: Calculating an attention weight matrix between the virtual behavior feature vector and the actual behavior feature vector, wherein, for each feature element in the virtual behavior feature vector, a correlation score between the feature element and each feature element in the actual behavior feature vector is calculated by a dot product operation, and the correlation score is normalized to obtain the attention weight matrix; Performing weighted aggregation on the virtual behavior feature vector and the actual behavior feature vector based on the attention weight matrix to obtain a context feature vector of the virtual behavior and a context feature vector of the actual behavior, wherein the context feature vector of the virtual behavior is obtained by matrix multiplication of the attention weight matrix and the actual behavior feature vector, and the context feature vector of the actual behavior is obtained by matrix multiplication of the transpose of the attention weight matrix and the virtual behavior feature vector; Performing feature splicing on the virtual behavior feature vector and the context feature vector of the virtual behavior, and performing feature splicing on the actual behavior feature vector and the context feature vector of the actual behavior, to obtain an enhanced virtual behavior feature vector and an enhanced actual behavior feature vector; The enhanced virtual behavior feature vector and the enhanced actual behavior feature vector are fused through nonlinear transformation to generate a virtual-real fusion feature vector.
[0024] Obtain the virtual behavior feature vector and actual behavior feature vector to be processed. Assume that the virtual behavior feature vector is represented by V, with a dimension of d×1, where d represents the feature dimension, for example, d=128; the actual behavior feature vector is represented by R, with a dimension of d×1. In practical applications, these feature vectors can be extracted from the original data through a deep neural network.
[0025] Calculate the attention weight matrix between the virtual behavior feature vector V and the actual behavior feature vector R. Specifically, for each feature element v_i in the virtual behavior feature vector V, calculate its correlation score s_ij with each feature element r_j in the actual behavior feature vector R through a dot product operation. For example, if the value of the first element in V is 0.5 and the value of the first element in R is 0.7, then their dot product correlation score is 0.5×0.7=0.35.
[0026] For the relevance scores of all feature element pairs, the softmax function is used to normalize them to obtain the attention weight matrix A. Normalization ensures that the sum of the weights of each row is 1, making the attention distribution more reasonable. For example, if the original relevance score is [0.35, 0.42, 0.28], the weights may be [0.33, 0.41, 0.26] after normalization.
[0027] The virtual behavior feature vector V and the actual behavior feature vector R are weighted and aggregated based on the attention weight matrix A. Specifically, the context feature vector CV of the virtual behavior is obtained by matrix multiplication of the attention weight matrix A and the actual behavior feature vector R. In actual implementation, if A is a d×d dimensional matrix and R is a d×1 dimensional vector, then CV=A×R, and the result is a d×1 dimensional vector.
[0028] The context feature vector CR of the actual behavior is obtained by transposing the attention weight matrix A and multiplying the matrix of the virtual behavior feature vector V. That is, CR = A^T×V, where A^T represents the transposed matrix of A, and the result is also a d×1-dimensional vector.
[0029] This cross-attention mechanism allows the virtual behavior features and the actual behavior features to influence each other, thereby capturing the correlation information between them. For example, if an operation in the virtual scene is highly correlated with a specific action in the actual scene, a strong attention connection will be established between them.
[0030] The virtual behavior feature vector V is concatenated with the context feature vector CV of the virtual behavior to obtain an enhanced virtual behavior feature vector EV. In implementation, the concatenation operation can be expressed as EV=[V;CV], where ";" represents a vector concatenation operation, and the result dimension is 2d×1.
[0031] The actual behavior feature vector R is concatenated with the context feature vector CR of the actual behavior to obtain the enhanced actual behavior feature vector ER, that is, ER=[R;CR], with the same dimension of 2d×1.
[0032] The enhanced virtual behavior feature vector EV and the enhanced actual behavior feature vector ER are fused through nonlinear transformation to generate a virtual-real fusion feature vector F. The nonlinear transformation can be implemented by a multi-layer perceptron (MLP), which includes weight matrices W1 and W2 and bias items b1 and b2.
[0033] EV and ER are concatenated to obtain a 4d×1 joint vector [EV;ER], which is then transformed through two layers of MLP: Intermediate representation H = ReLU(W1×[EV;ER] + b1); final fused vector F = W2×H + b2.
[0034] ReLU stands for the rectified linear unit activation function, which is used to introduce nonlinear transformation capabilities. Assuming that the dimension of W1 is k×4d and the dimension of b1 is k×1, the dimension of H is k×1; the dimension of W2 is m×k and the dimension of b2 is m×1, then the dimension of the final fusion vector F is m×1, where k can be set to 256 and m can be set to 128.
[0035] Assume that the original virtual behavior feature vector V and the actual behavior feature vector R are both 128-dimensional vectors. Through the above cross-attention fusion method, the 128×128-dimensional attention weight matrix A is first calculated, and then the 128-dimensional context feature vectors CV and CR are obtained. After feature concatenation, EV and ER are both 256-dimensional vectors. Finally, after MLP transformation, a 128-dimensional virtual-real fusion feature vector F is obtained.
[0036] Batch processing is used to process multiple samples simultaneously. For example, if a batch of data contains 32 samples, the input V and R can be represented as a 32×128 matrix, where each row represents the feature vector of a sample. Accordingly, attention weight calculation and subsequent operations are also performed in batch processing.
[0037] In order to improve the expressiveness of fused features, a multi-head attention mechanism can be introduced on the basis of the attention mechanism. For example, using 8 attention heads, each head generates 16-dimensional context features (a total of 128 dimensions), and then the outputs of all heads are concatenated to form the final context features. This multi-head mechanism enables the model to learn information from different representation subspaces, further enhancing the effect of feature fusion.
[0038] Through the above-mentioned virtual-real behavior feature fusion method based on the attention mechanism, the correlation information between virtual behavior and actual behavior can be effectively captured, and a virtual-real fusion feature vector with rich semantics can be generated, providing strong support for subsequent analysis and decision-making tasks.
[0039] In an optional implementation, a graph neural network is used to construct a temporal relationship graph of virtual and real behaviors, the temporal dependency relationship between nodes is modeled through a conditional random field, multi-hop information propagation is performed on node features, temporal context features are obtained, and the temporal context features of multiple subgraphs are integrated through a cross-subgraph feature fusion layer to generate a fused temporal feature vector including: A graph neural network is used to construct a time series relationship graph of virtual and real behaviors, and the virtual and real fusion feature vector of each time step is mapped to a node feature in the time series relationship graph of virtual and real behaviors; The temporal dependency between nodes is modeled through conditional random fields, and the causal relationship strength between nodes is dynamically predicted based on the changing trend of node features and historical interaction patterns; The causal association strength is normalized to obtain a sampling probability between node pairs; a sampling probability threshold is set, and node pairs with a probability higher than the sampling probability threshold are selected as key timing paths; with the key timing path as the center, a plurality of local subgraphs are obtained by sampling through a random walk algorithm; The following operations are performed on each of the local subgraphs: Construct a local graph neural network and set up multiple layers of graph convolution. In each layer of graph convolution, aggregate the neighbor features of the node and update the node representation. Preserve the node features of each layer through skip connections to obtain the node temporal context features containing multi-hop information. The cross-sub-image feature fusion layer integrates the information of multiple local sub-images, including: The attention weights of corresponding node features in different subgraphs are calculated; the node temporal context features of multiple subgraphs are weightedly fused based on the attention weights; and the final fused temporal feature vector is obtained through a pooling operation.
[0040] A graph neural network (GNN) is used to construct a temporal relationship graph of virtual and real behaviors. The nodes of the graph represent different behavior states, and the edges represent the temporal relationship between these behaviors. The virtual and real fusion feature vector of each time step will be mapped to the node features in the graph. The construction of the feature vector can be done by preprocessing the behavior data to extract key temporal features and context information to ensure the validity of the node features.
[0041] The temporal dependencies between nodes are modeled through conditional random fields (CRFs). This process involves analyzing the changing trends of node features and historical interaction patterns to dynamically predict the strength of causal associations between nodes. Specifically, the historical behavior data of the nodes is first collected to construct feature vectors, and then these features are used to evaluate the causal relationship between nodes. For example, if the feature value of node A at time t changes significantly, and the feature value of node B at time t+1 also changes accordingly, it can be inferred that there is a causal relationship between A and B.
[0042] The predicted causal association strength is normalized to obtain the sampling probability between node pairs. This process ensures that the sampling probability of all node pairs is between 0 and 1. By setting a sampling probability threshold, node pairs above the threshold are screened as key timing paths. These key paths will serve as the focus of subsequent analysis to ensure that the model focuses on important timing relationships.
[0043] Taking the selected key timing path as the center, the random walk algorithm is used for sampling to obtain multiple local subgraphs. The random walk process starts from the key node and randomly selects adjacent nodes for traversal until the preset number of steps is reached or a specific node is traversed. Each local subgraph will contain nodes and their features related to the key path, providing a basis for subsequent graph neural network training.
[0044] For each local subgraph, a local graph neural network is constructed and multiple layers of graph convolution are set. In each layer of graph convolution, the neighbor features of the node are aggregated and the node representation is updated. The node features of each layer are retained through jump connections, so that each node can combine information from different layers to obtain node temporal context features containing multi-hop information. This process can enhance the expressiveness of node features through multiple iterations.
[0045] The information of multiple local subgraphs is integrated through the cross-subgraph feature fusion layer. First, the attention weights of the corresponding node features in different subgraphs are calculated. The weights reflect the importance of the node features in each subgraph. Then, based on these attention weights, the node temporal context features of multiple subgraphs are weighted fused. Finally, the fused temporal feature vector is obtained through the pooling operation, which will be used for subsequent analysis and prediction tasks.
[0046] Assume that there is a set of user behavior data, including the user's click, browse and purchase behaviors at different time points. Through preprocessing, the feature vector of each user at each time point is extracted, such as click-through rate, browsing time and purchase amount. Using these features to build a time series relationship diagram of virtual and real behaviors, the model can effectively identify the user's behavior pattern in a specific time period.
[0047] When modeling the temporal dependency between nodes, assuming that the click rate of user A at time t is 0.8, and the purchase rate of user B at time t+1 is 0.5, the model can infer that A's behavior may affect B's purchase decision. By setting the sampling probability threshold, key temporal paths are screened out, such as the relationship between A's click behavior and B's purchase behavior.
[0048] By constructing a local graph neural network and fusing cross-subgraph features, the model can generate an efficient time series feature vector to provide support for subsequent user behavior prediction.
[0049] Figure 2 This is a flowchart of the construction of the virtual-real behavior time series relationship diagram and feature fusion in an embodiment of the present invention: This flowchart shows a complete node temporal feature update and fusion process. Construct the execution party temporal relationship graph, which is used to map the node features in the virtual-real fusion feature vector. Model the node time interval dependency of the conditional random field and predict the relationship strength between nodes. Obtain the similarity probability between nodes through normalization. Determine the key temporal path based on the sampling probability graph, and use the random walk algorithm to divide the samples into multiple local subgraphs. Perform graph neural network operations on each local subgraph, including feature aggregation and update. Construct a local graph neural network and set up multiple layers of graph convolution layers to achieve information interaction through node feature propagation. The aggregated node features are fused with the original features to update the node temporal representation. The node temporal features of each layer are retained through jump connections to obtain the node context features of multi-hop information. The cross-subgraph node features are weighted fused based on the attention weight to obtain the final node temporal feature representation. The entire process effectively captures the temporal dependency between nodes through multi-level feature extraction and fusion, and generates node representations containing rich context information.
[0050] In an optional implementation, the temporal dependency between nodes is modeled by a conditional random field, and based on the change trend of node features and the historical interaction pattern, the causal relationship strength between nodes is dynamically predicted, including: Modeling the temporal dependency relationship between nodes in the temporal relationship graph of virtual and real behaviors through conditional random fields, and calculating the difference vector of adjacent node features as the change trend of node features; Based on the interaction frequency of node pairs in the historical time series window, the historical interaction pattern characteristics of node pairs are constructed; The change trend of the node features and the historical interaction pattern features are input into a conditional random field to predict the causal association strength between node pairs.
[0051] In the virtual-real behavior time series relationship graph, each node represents a behavior or event, and the edges between nodes represent the possible causal relationship between them. Each node has a series of feature values that change over time. The conditional random field constructed by this method predicts the strength of the causal relationship between node pairs by analyzing the changes in node features at adjacent time points and the historical interaction patterns between nodes.
[0052] When calculating the change trend of node features, for each pair of adjacent time points t and t+1 in the time series, obtain the feature vectors Fi,t and Fi,t+1 of node i at these two time points respectively. By calculating the difference between these two feature vectors, we get the change trend vector ΔFi,t = Fi,t+1 - Fi,t. This difference vector reflects the change direction and magnitude of the node features in this time period.
[0053] Assume that node i represents the browsing behavior of the user, and its feature vector contains three dimensions: browsing time, number of clicks, and page dwell time. At time point t, its feature vector Fi,t is [30, 5, 120], indicating that the browsing time is 30 seconds, the number of clicks is 5, and the page dwell time is 120 seconds. At time point t+1, its feature vector Fi,t+1 is [45, 8, 180]. Then the calculated change trend vector ΔFi,t is [15, 3, 60], indicating that the browsing time increases by 15 seconds, the number of clicks increases by 3, and the page dwell time increases by 60 seconds.
[0054] To model the relationship between two nodes i and j, it is necessary to consider the change trend vectors of these two nodes at the same time, and they can be concatenated to form a joint change trend feature vector [ΔFi,t, ΔFj,t].
[0055] Construct the historical interaction pattern features of node pairs. Select a fixed-size time window W and count the interaction frequency and pattern between node i and node j in the window. The interaction pattern can be characterized from multiple dimensions, including interaction frequency, interaction time interval distribution, interaction direction ratio, etc.
[0056] For example, suppose that in the past 10 time units, there are 8 interactions between node i and node j, 6 of which are from i to j and 2 from j to i. The time intervals of these interactions are [1, 2, 1, 3, 1, 2, 1, 2] time units respectively.
[0057] The changing trend of node features and historical interaction pattern features are combined to form the input feature vector Xi,j = [[ΔFi,t, ΔFj,t], Hi,j] of the conditional random field. The conditional random field predicts the causal association strength Si,j,t+1 between nodes i and j at the next time point t+1 based on these features.
[0058] Conditional random fields are a probabilistic graphical model, the core of which is to define a set of feature functions and corresponding weights. Each feature function evaluates a certain aspect of the input feature, and the weight reflects the degree of influence of the feature function on the final prediction result.
[0059] The weights of the conditional random field are learned using annotated historical data. These annotated data contain the actual causal strength of the node pairs at each time point. Through maximum likelihood estimation or other optimization methods, a set of optimal weight values is found so that the model's prediction of the training data is as close as possible to the actual annotation.
[0060] Given the node features and historical interaction patterns at the current time point t, the conditional random field calculates the probability distribution of each possible causal association strength value and selects the value with the highest probability as the prediction result.
[0061] The historical interaction pattern features are continuously updated using a sliding window. As new interaction data is generated, the window slides forward and maintains a fixed size, ensuring that the historical interaction pattern features always reflect the most recent interaction.
[0062] Different feature extraction methods are set according to different types of nodes and interactions. For example, for nodes representing user behaviors, we can focus on the frequency, duration, and intensity changes of behaviors; for nodes representing system events, we can focus on the changes in the triggering conditions and impact range of the events.
[0063] Suppose there is a user behavior analysis system for an e-commerce platform, where nodes can represent various user behaviors, such as browsing products, adding to shopping carts, viewing reviews, etc. The system needs to analyze the causal relationship between these behaviors in order to better understand the user's shopping decision process.
[0064] Taking browsing products (node A) and adding to shopping cart (node B) as examples, the system collects user behavior data at 10 consecutive time points. The characteristics of node A include browsing time, page switching times, and product detail expansion times; the characteristics of node B include the number of products added to the shopping cart, the total amount added to the shopping cart, and the continued browsing time after adding to the shopping cart.
[0065] By calculating the changing trend of node features and constructing historical interaction pattern features, the conditional random field predicts that the causal correlation strength of node A to node B is 0.75, indicating that there is a strong causal correlation between browsing product behavior and subsequent adding to shopping cart behavior. This prediction result can help the platform optimize product display strategies and improve user shopping conversion rates.
[0066] The temporal dependency between nodes is modeled through conditional random fields. Based on the changing trend of node characteristics and the characteristics of historical interaction patterns, the causal relationship strength between nodes can be effectively predicted dynamically, providing a powerful tool for analyzing the causal relationship between various elements in complex systems.
[0067] In an optional implementation, a local graph neural network is constructed, and a multi-layer graph convolution layer is set; in each layer of graph convolution, neighbor features of nodes are aggregated and node representations are updated; node features of each layer are retained through skip connections, and node temporal context features containing multi-hop information are obtained, including: Receiving input node feature vectors and graph structure information, wherein the node feature vectors represent initial features of each node, and the graph structure information represents connection relationships between nodes; Construct a multi-layer graph convolutional network structure, including L graph convolutional layers, each of which contains a learnable weight matrix and a bias vector, where L is an integer greater than 1; For the lth graph convolution layer among the L graph convolution layers, the following feature propagation operations are performed: Obtaining a set of neighbor nodes of a node; normalizing features in the set of neighbor nodes; multiplying the normalized neighbor nodes by a weight matrix of the lth graph convolutional layer; Perform nonlinear transformation on the product result to obtain the output features of the node at the lth layer; The intermediate feature representation of nodes in each graph convolutional layer is retained through the skip connection mechanism, including: The output features of the nodes in each graph convolution layer are concatenated; the concatenated multi-layer features are reduced in dimension using the feature transformation layer; The feature propagation process of the graph convolution layer is iteratively performed L times, and the node features are updated in each iteration to obtain the node temporal context features containing multi-hop information.
[0068] Receive input node feature vectors and graph structure information. The node feature vector represents the initial features of each node, which can be the node's attribute information, historical behavior data, or other related features. The graph structure information represents the connection relationship between nodes, usually provided in the form of an adjacency matrix or an edge list. For example, in a social network, a node can be a user, and the node features can include information such as the user's age, gender, interests, and hobbies, while the graph structure information represents the friendship relationship between users.
[0069] Construct a multi-layer graph convolutional network structure, including L graph convolutional layers. In this embodiment, L is 3, that is, a 3-layer graph convolutional network is constructed. Each graph convolutional layer contains a learnable weight matrix and a bias vector for feature conversion and information transfer. The dimension of the weight matrix depends on the input feature dimension and the output feature dimension. For example, the weight matrix dimension of the first layer is 64×32, the second layer is 32×32, and the third layer is 32×32.
[0070] For the lth graph convolution layer, first obtain the set of neighbor nodes for each node. For example, for node A, its neighbor nodes may include nodes B, C, and D. Then, the features in the set of neighbor nodes are normalized to balance the differences in the number of neighbors of different nodes. Normalization can be achieved by dividing the features of each neighbor node by the square root of the degree of the node (the number of neighbor nodes), and then by the square root of the degree of the center node.
[0071] Multiply the neighbor node features with the weight matrix of the lth graph convolutional layer. Specifically, assume that the feature vectors of neighbor nodes B, C, and D are [0.3, 0.5, 0.2, ...], [0.1, 0.7, 0.4, ...], [0.6, 0.2, 0.8, ...], respectively. The feature vectors obtained after normalization are [0.2, 0.33, 0.13, ...], [0.07, 0.47,0.27, ...], [0.4, 0.13, 0.53, ...], respectively. Multiply these normalized feature vectors with the weight matrix to obtain the transformed feature representation.
[0072] The product result is transformed nonlinearly to obtain the output features of the node in the lth layer. The nonlinear transformation usually uses the ReLU function, which retains positive values and sets negative values to zero to enhance the network's expressiveness. For example, for node A, the features after the first layer of graph convolution may be [0.25, 0.41, 0, 0.32, ...].
[0073] A skip connection mechanism is adopted. Specifically, the output features of the node in each layer of graph convolution are concatenated. For example, for node A, assuming that its features after three layers of graph convolution are [0.25, 0.41, 0, 0.32, ...] (first layer), [0.35, 0, 0.22, 0.18, ...] (second layer) and [0.15, 0.27, 0.39, 0, ...] (third layer), the concatenated features are [0.25, 0.41, 0, 0.32, ..., 0.35, 0, 0.22, 0.18, ..., 0.15,0.27, 0.39, 0, ...].
[0074] The feature transformation layer is used to reduce the dimension of the concatenated multi-layer features. The feature transformation layer can be a fully connected layer that maps high-dimensional features to a lower-dimensional space. In this embodiment, the feature transformation layer reduces the concatenated 96-dimensional features (assuming that each layer outputs 32 dimensions) to 64 dimensions to obtain the final node representation.
[0075] By iteratively executing the feature propagation process of the graph convolution layer L times, the node features are updated each time, and finally the node temporal context features containing multi-hop information are obtained. These features integrate the node's own information and the information of neighboring nodes with different hops, and can fully characterize the position and contextual relationship of the node in the graph structure.
[0076] In specific application scenarios, such as recommendation systems, users and items can be modeled as nodes in a graph, and user interactions with items can be modeled as edges. By extracting the temporal context features of user nodes through the above method, it is possible to capture the evolution of user interests and social influence, thereby improving the accuracy of recommendations. In actual tests, the recommendation system using this method has a 15.3% increase in click-through rate and a 12.7% increase in conversion rate compared to traditional methods.
[0077] In the anomaly detection scenario, network devices can be modeled as nodes and inter-device communications can be modeled as edges. By extracting the temporal context features of device nodes, abnormal communication patterns can be identified and network intrusion behaviors can be discovered in a timely manner. Experiments show that the anomaly detection accuracy of this method reaches 92.8%, which is 7.5 percentage points higher than the baseline method.
[0078] The method provided in this embodiment effectively extracts node temporal context features containing multi-hop information through a multi-layer graph convolutional network and a skip connection mechanism, providing a powerful tool for graph data analysis and mining.
[0079] Figure 3 This is a flowchart of node feature processing based on a graph neural network according to an embodiment of the present invention: The figure shows a node feature processing flow based on a graph neural network. The method receives the input node feature vector and graph structure information, where the node feature vector is used to characterize the initial feature attributes of each node, and the graph structure information describes the topological connection relationship between nodes. On this basis, a multi-layer graph convolutional network structure containing L layers is constructed. Each layer is equipped with a learnable weight matrix and bias vector parameter. L is an integer value greater than 1, which is used to control the depth of the network. For each layer of graph convolution operation, a specific feature propagation process is performed: first, the set of neighbor nodes of the target node is obtained, the features of these neighbor nodes are standardized, and then the processed features are multiplied with the weight matrix of the layer. Finally, the output feature representation of the node in the layer is obtained through a nonlinear activation function. In particular, the intermediate features of the node in each convolutional layer are retained through a skip connection mechanism. Specifically, the output features of each layer are spliced, and the feature conversion layer is used to reduce the dimension. Finally, a residual connection is added to fuse the reduced-dimensional features with the initial node features. This process is iterated L times, and each iteration updates the feature representation of the node, and finally a complete node temporal context feature containing multi-hop information is obtained. This design can effectively capture the structural information and semantic features of nodes at different scales, and improve the expressiveness of feature representation.
[0080] In an optional implementation, the behavior sequence matching degree is determined by a dynamic programming algorithm according to the fused time series feature vector, and a virtual-real behavior consistency score is obtained in combination with a preset user behavior feature template; and the user's behavior in the metaverse virtual scene is verified and authorized based on the virtual-real behavior consistency score, including: Extracting a behavior feature point sequence from the fused time series feature vector, calculating the longest common subsequence of the current behavior feature point sequence and the historical behavior feature point sequence based on a dynamic programming algorithm, and obtaining a behavior sequence matching degree; According to the matching degree of the behavior sequence, the importance of the feature point sequence of the user's historical behavior characteristics is weighted to construct an adaptive user behavior feature template, wherein the importance weight decays with time distance; Calculating the multi-scale correlation score between the current fused time series feature vector and the adaptive user behavior feature template, and detecting the behavior mutation point based on the sliding time window to generate a behavior consistency score; When the behavior consistency score is higher than a preset score threshold, the user's operation request in the metaverse scene is permitted; When the behavior consistency score is lower than the preset score threshold, the identity re-authentication mechanism is triggered and the user's operation authority is temporarily frozen.
[0081] Extract the behavioral feature point sequence from the fused time series feature vector. The time series feature vector obtained by the system contains multi-dimensional data, such as the user's operation actions, interaction mode, movement trajectory and other information in the virtual scene. Through the feature point detection algorithm, key behavioral feature points are extracted from these time series feature vectors to form a behavioral feature point sequence. For example, for a user in a virtual scene, a behavioral feature point sequence such as "login-browse-interact-purchase-exit" may be extracted.
[0082] Based on the dynamic programming algorithm, the longest common subsequence of the current behavior feature point sequence and the historical behavior feature point sequence is calculated to obtain the behavior sequence matching degree. In the specific implementation, let the currently extracted behavior feature point sequence be A, with a length of m; the historical record of the user's behavior feature point sequence is B, with a length of n. The system establishes a two-dimensional table of (m+1)×(n+1) and calculates the length of the longest common subsequence by filling in the table. The final behavior sequence matching degree can be expressed as the ratio of the longest common subsequence length to the total sequence length, ranging from 0 to 1. For example, if the current sequence is "login-browse-interact-purchase-exit" and the historical sequence is "login-interact-browse-purchase-exit", the longest common subsequence calculated by dynamic programming is "login-interact-purchase-exit", with a length of 4, then the matching degree is 4 / 5=0.8.
[0083] According to the matching degree of the behavior sequence, the importance of the feature point sequence of the user's historical behavior characteristics is weighted to construct an adaptive user behavior feature template. The system assigns different weights to each feature point in the historical behavior sequence according to its frequency of occurrence and time distance. The time distance decay function adopts an exponential decay form, and the behavior feature point closer to the current time has a higher weight. For example, for a feature point, if it has appeared 5 times in the last 7 days, its basic weight can be set to 0.5; if it appeared 3 days ago, the time decay factor can be set to 0.9^3≈0.73, and the final weight of the feature point is 0.5×0.73≈0.365. The system integrates all feature points and their weights to form a user behavior feature template.
[0084] Calculate the multi-scale correlation score between the current fused time series feature vector and the adaptive user behavior feature template. The system calculates the correlation score at different time window scales. For example, three time scales can be set: 5 minutes, 30 minutes, and 2 hours. In each time scale, the system calculates the similarity between the current feature vector and the template, and performs a weighted average of the scores at different scales to obtain an overall correlation score. The system also detects behavioral mutation points based on a sliding time window, that is, when the behavior pattern in a specific time window changes significantly compared to the previous time window, it is marked as a behavioral mutation point. The final behavioral consistency score is a combination of the correlation score and the mutation point detection results, and the value range is 0 to 100. For example, if the multi-scale correlation score is 85 points, and 1 moderate behavioral mutation point is detected (deducting 10 points), the final behavioral consistency score is 75 points.
[0085] When the behavior consistency score is higher than the preset score threshold, the system allows the user's operation request in the metaverse scenario. The preset score threshold can be set to different values according to the security level of the operation, and can generally be set to 70 points. For example, if a user requests to browse products in a virtual mall, the calculated behavior consistency score is 85 points, which is higher than the preset threshold of 70 points, and the system allows the operation request.
[0086] When the behavior consistency score is lower than the preset score threshold, the system triggers the identity re-authentication mechanism and temporarily freezes the user's operating permissions. For example, a user requests to conduct a large virtual asset transaction in the virtual world, and the calculated behavior consistency score is 65 points, which is lower than the preset threshold of 70 points. The system will immediately require the user to re-authenticate, which may include secondary password verification, biometric verification, or mobile phone SMS verification. Before the user completes the re-authentication, the system temporarily freezes the user's relevant operating permissions to prevent potential account theft risks.
[0087] The parameters will be adjusted dynamically for different users and scenarios. For example, for newly registered users, due to less historical behavior data, the system may lower the preset scoring threshold to 60 points; and for high-value virtual asset trading operations, the system may increase the preset scoring threshold to 85 points to provide stricter security.
[0088] An adaptive update mechanism for templates is also implemented. As the time a user spends in the Metaverse environment increases, the system continues to collect user behavior data and regularly updates the user's behavior feature template. The update frequency can be set to once a week or once a day to ensure that the template can accurately reflect the user's latest behavior features. For users who have been inactive for a long time, the weight in their behavior feature template will gradually decay. When the user becomes active again, the system will be more cautious in evaluating the consistency of behavior.
[0089] Through the comprehensive application of the above technical means, the virtual-real behavior consistency verification method provided in this implementation can effectively identify abnormal operation behaviors, ensure the safety of users in the metaverse environment, and provide a smooth user experience.
[0090] Figure 4 This is a flowchart of user verification based on behavior sequence matching according to an embodiment of the present invention: The figure shows a user verification process based on behavior sequence matching, which can be integrated as follows: The method first extracts the behavior feature point sequence from the fused time series feature vector, and uses a dynamic programming algorithm to calculate the longest common subsequence between the current behavior feature point sequence and the historical behavior feature point sequence, thereby obtaining the matching degree of the behavior sequence. Based on this matching degree, the system performs importance weighting on the user's historical behavior features and constructs an adaptive user behavior feature template, in which the importance weight of the feature decays with the increase of time distance. After that, the system calculates the multi-dimensional correlation score between the current fused time series feature vector and the adaptive behavior feature template, and detects the behavior mutation points by sliding the time window to comprehensively generate a behavior consistency score. This score is compared with the preset threshold: when the behavior consistency score is higher than the preset threshold, the system will authorize the user's operation request in the metaverse scenario and synchronously update the user's behavior feature template; conversely, when the score is lower than the preset threshold, the system will trigger the identity re-authentication mechanism and take security measures such as temporarily freezing the user's high-risk operation permissions. This verification mechanism based on behavior sequence matching can dynamically evaluate the consistency of user behavior and effectively prevent potential security risks.
[0091] In an optional implementation, a behavior feature point sequence is extracted from the fused time series feature vector, and the longest common subsequence of the current behavior feature point sequence and the historical behavior feature point sequence is calculated based on a dynamic programming algorithm to obtain a behavior sequence matching degree, which includes: Perform extreme point detection on the time series feature vector to extract a current behavior feature point sequence; Acquire a historical behavior feature point sequence from a user behavior database, and construct the current behavior feature point sequence and the historical behavior feature point sequence into two sequences to be matched; Constructing a dynamic programming matrix, wherein the rows and columns of the dynamic programming matrix correspond to the lengths of the current behavior feature point sequence and the historical behavior feature point sequence respectively; Calculate the longest common subsequence of two sequences based on the dynamic programming algorithm: Initialize the matrix boundary conditions; when the difference between the feature point values is less than the preset difference threshold, the matrix element at the corresponding position is increased by 1; when the difference between the feature point values is greater than the preset difference threshold, the maximum value of the adjacent positions is taken; Iteratively update the matrix elements until all sequence positions are traversed; According to the final value of the dynamic programming matrix, the length of the longest common subsequence is calculated and divided by the total length of the sequence to obtain the behavior sequence matching degree.
[0092] A behavior feature point sequence is extracted from the fused time series feature vector, which may be obtained by fusion of multimodal sensor data. For example, the data of an acceleration sensor, a gyroscope sensor, and a pressure sensor may be fused to form a fused time series feature vector after filtering, normalization, and time domain feature extraction.
[0093] The time series feature vector is subjected to extreme point detection to extract the current behavior feature point sequence. The extreme point detection adopts the sliding window method with a window size of 5 sampling points. When the value of the center point is greater than the values of other points in the window, it is determined to be a maximum point; when the value of the center point is less than the values of other points in the window, it is determined to be a minimum point. Maximum points and minimum points are collectively referred to as extreme points, which constitute the current behavior feature point sequence in chronological order.
[0094] For example, for a time series feature vector [0.2, 0.3, 0.5, 0.4, 0.2, 0.1, 0.3, 0.6,0.7, 0.4], after extreme point detection, it can be determined that the third point (value is 0.5) is the maximum point, the sixth point (value is 0.1) is the minimum point, and the ninth point (value is 0.7) is the maximum point. Therefore, the current behavior feature point sequence is [(3,0.5), (6,0.1), (9,0.7)], where each element represents (position, value).
[0095] Get the historical behavior feature point sequence from the user behavior database. The database stores the feature point sequence of the user's historical behavior, and each sequence corresponds to a specific type of behavior. For example, the historical behavior feature point sequence obtained from the database is [(2,0.48), (5,0.15), (8,0.68)].
[0096] The current behavior feature point sequence and the historical behavior feature point sequence are constructed into two sequences to be matched. In this embodiment, the two sequences to be matched are: Sequence A: [(3,0.5), (6,0.1), (9,0.7)]; Sequence B: [(2,0.48), (5,0.15), (8,0.68)]; Construct a dynamic programming matrix DP, where the rows and columns of the matrix correspond to the lengths of the current behavior feature point sequence and the historical behavior feature point sequence, respectively. For the above example, the size of the constructed dynamic programming matrix DP is 4×4, where DP[0][0] is the initial position and DP[3][3] is the final position.
[0097] Initialize the matrix boundary conditions: Initialize the first row and the first column of the DP matrix to 0, that is, DP[0][j]=0(j=0,1,2,3) and DP[i][0]=0(i=0,1,2,3).
[0098] Iteratively calculate the matrix element values: Starting from DP[1][1], calculate the value of each element in row priority order. For position DP[i][j], compare the difference between the i-th feature point A[i] of sequence A and the j-th feature point B[j] of sequence B. The preset difference threshold is 0.1. When |A[i].value - B[j].value|<0.1, the two feature points are considered to match, and the matrix element DP[i][j] at the corresponding position is DP[i-1][j-1]+ 1; when the difference in the feature point value is greater than or equal to the preset difference threshold, take the maximum value of the adjacent positions, that is, DP[i][j]= max(DP[i-1][j], DP[i][j-1]).
[0099] For DP[1][1], compare A[1]=(3,0.5) and B[1]=(2,0.48), |0.5-0.48|=0.02<0.1, so DP[1][1]=DP[0][0]+1=1.
[0100] For DP[1][2], compare A[1]=(3,0.5) and B[2]=(5,0.15), |0.5-0.15|=0.35>0.1, so DP[1][2]=max(DP[0][2], DP[1][1])=1.
[0101] According to the final value of the dynamic programming matrix DP[3][3]=2, the length of the longest common subsequence is 2, indicating that there are 2 feature points matching in the two sequences.
[0102] Calculate the behavior sequence matching degree: Divide the length of the longest common subsequence by the smaller value of the total length of the sequence to get the behavior sequence matching degree. In this example, the length of both sequences is 3, so the behavior sequence matching degree = 2 / 3≈0.67, which is a 67% match.
[0103] When the behavior sequence matching degree exceeds the preset threshold (such as 0.6), the current behavior is judged to match the historical behavior; otherwise, it is judged to be mismatched. In this example, the matching degree is 0.67, which exceeds the preset threshold 0.6, so the current behavior is judged to match the historical behavior.
[0104] In order to improve the matching accuracy, the position information of the feature points can be considered at the same time. When comparing whether two feature points match, in addition to comparing the numerical difference, the position difference can also be limited to a certain range, such as |A[i].position -B[j].position|<2. This ensures that the matched feature points are not only similar in numerical value, but also in relative position in the time series.
[0105] The behavior feature matching method of the present invention is applicable to various scenarios, such as motion recognition of smart watches, user behavior recognition of smart homes, abnormal behavior detection of security systems, etc. By accurately calculating the matching degree of behavior sequences, accurate recognition and analysis of user behaviors can be achieved, thereby improving user experience and system security.
[0106] In actual applications, the preset difference threshold and matching threshold can be adjusted according to specific scenarios to balance the recognition accuracy and recall rate. For example, in a security system, the matching threshold can be lowered to increase the detection rate of abnormal behavior; in daily motion recognition, the matching threshold can be increased to reduce the false recognition rate.
[0107] A second aspect of an embodiment of the present invention provides a deep learning-based intelligent IoT and metaverse virtual-real fusion recognition system, including: The first unit is used to obtain the virtual behavior data of the user in the virtual scene of the Metaverse collected by the terminal device and the actual behavior data of the user in the real scene collected by the IoT device; A second unit is used to extract features from the virtual behavior data and the actual behavior data through a pre-trained deep neural network model to obtain a virtual behavior feature vector and an actual behavior feature vector; A third unit is used to cross-fuse the virtual behavior feature vector and the actual behavior feature vector based on an attention mechanism to generate a virtual-real fusion feature vector, wherein the attention mechanism calculates the correlation weight between the virtual behavior feature vector and the actual behavior feature vector to achieve adaptive feature alignment of virtual and real behaviors; The fourth unit is used to construct a temporal relationship graph of virtual and real behaviors using graph neural networks, model the temporal dependency relationship between nodes through conditional random fields, perform multi-hop information propagation on node features, obtain temporal context features, integrate the temporal context features of multiple subgraphs through a cross-subgraph feature fusion layer, and generate a fused temporal feature vector; The fifth unit is used to identify the consistency of the user's virtual and real behaviors according to the fused time series feature vector to obtain a virtual and real behavior consistency score; and verify and authorize the user's behavior in the virtual scene of the metaverse based on the virtual and real behavior consistency score.
[0108] According to a third aspect of the embodiments of the present invention, An electronic device is provided, comprising: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0109] A fourth aspect of the embodiments of the present invention is: A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.
[0110] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.
[0111] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A deep learning-based intelligent IoT and metaverse virtual-real fusion recognition method, characterized in that: include: Obtain the user's virtual behavior data in the Metaverse virtual scene collected by the terminal device and the user's actual behavior data in the real scene collected by the IoT device; Extracting features from the virtual behavior data and the actual behavior data using a pre-trained deep neural network model to obtain a virtual behavior feature vector and an actual behavior feature vector; Cross-fusing the virtual behavior feature vector and the actual behavior feature vector based on an attention mechanism to generate a virtual-real fusion feature vector, wherein the attention mechanism realizes adaptive feature alignment of virtual and real behaviors by calculating a correlation weight between the virtual behavior feature vector and the actual behavior feature vector; Use graph neural networks to build a temporal relationship graph of virtual and real behaviors, model the temporal dependency between nodes through conditional random fields, perform multi-hop information propagation on node features, obtain temporal context features, and integrate the temporal context features of multiple subgraphs through the cross-subgraph feature fusion layer to generate a fused temporal feature vector. The consistency of the user's virtual and real behaviors is identified according to the fused temporal feature vector to obtain a virtual and real behavior consistency score; and the user's behavior in the virtual scene of the metaverse is verified and authorized based on the virtual and real behavior consistency score.
2. The method according to claim 1, characterized in that Cross-fusing the virtual behavior feature vector and the actual behavior feature vector based on the attention mechanism to generate a virtual-real fusion feature vector includes: Calculating an attention weight matrix between the virtual behavior feature vector and the actual behavior feature vector, wherein, for each feature element in the virtual behavior feature vector, a correlation score between the feature element and each feature element in the actual behavior feature vector is calculated by a dot product operation, and the correlation score is normalized to obtain the attention weight matrix; Performing weighted aggregation on the virtual behavior feature vector and the actual behavior feature vector based on the attention weight matrix to obtain a context feature vector of the virtual behavior and a context feature vector of the actual behavior, wherein the context feature vector of the virtual behavior is obtained by matrix multiplication of the attention weight matrix and the actual behavior feature vector, and the context feature vector of the actual behavior is obtained by matrix multiplication of the transpose of the attention weight matrix and the virtual behavior feature vector; Performing feature splicing on the virtual behavior feature vector and the context feature vector of the virtual behavior, and performing feature splicing on the actual behavior feature vector and the context feature vector of the actual behavior, to obtain an enhanced virtual behavior feature vector and an enhanced actual behavior feature vector; The enhanced virtual behavior feature vector and the enhanced actual behavior feature vector are fused through nonlinear transformation to generate a virtual-real fusion feature vector.
3. The method according to claim 1, characterized in that The graph neural network is used to construct the temporal relationship graph of virtual and real behaviors. The temporal dependency relationship between nodes is modeled through the conditional random field. The node features are propagated through multi-hop information to obtain the temporal context features. The temporal context features of multiple subgraphs are integrated through the cross-subgraph feature fusion layer to generate a fused temporal feature vector including: A graph neural network is used to construct a time series relationship graph of virtual and real behaviors, and the virtual and real fusion feature vector of each time step is mapped to a node feature in the time series relationship graph of virtual and real behaviors; The temporal dependency between nodes is modeled through conditional random fields, and the causal relationship strength between nodes is dynamically predicted based on the changing trend of node features and historical interaction patterns; The causal association strength is normalized to obtain a sampling probability between node pairs; a sampling probability threshold is set, and node pairs with a probability higher than the sampling probability threshold are selected as key timing paths; and multiple local subgraphs are obtained by sampling with a random walk algorithm with the key timing path as the center; The following operations are performed on each of the local subgraphs: Construct a local graph neural network and set up multiple layers of graph convolution. In each layer of graph convolution, aggregate the neighbor features of the node and update the node representation. Preserve the node features of each layer through skip connections to obtain the node temporal context features containing multi-hop information. The cross-sub-image feature fusion layer integrates the information of multiple local sub-images, including: The attention weights of corresponding node features in different subgraphs are calculated; the node temporal context features of multiple subgraphs are weightedly fused based on the attention weights; and the final fused temporal feature vector is obtained through a pooling operation.
4. The method according to claim 3, characterized in that The temporal dependency between nodes is modeled through conditional random fields. Based on the changing trend of node features and historical interaction patterns, the causal relationship strength between nodes is dynamically predicted, including: Modeling the temporal dependency relationship between nodes in the temporal relationship graph of virtual and real behaviors through conditional random fields, and calculating the difference vector of adjacent node features as the change trend of node features; Based on the interaction frequency of node pairs in the historical time series window, the historical interaction pattern characteristics of node pairs are constructed; The change trend of the node features and the historical interaction pattern features are input into a conditional random field to predict the causal association strength between node pairs.
5. The method according to claim 3, characterized in that: Construct a local graph neural network and set up multiple layers of graph convolution. In each layer of graph convolution, aggregate the neighbor features of the node and update the node representation. Retain the node features of each layer through skip connections, and obtain the node temporal context features containing multi-hop information, including: Receiving input node feature vectors and graph structure information, wherein the node feature vectors represent initial features of each node, and the graph structure information represents connection relationships between nodes; Construct a multi-layer graph convolutional network structure, including L graph convolutional layers, each of which contains a learnable weight matrix and a bias vector, where L is an integer greater than 1; For the lth graph convolution layer among the L graph convolution layers, the following feature propagation operations are performed: Obtaining a set of neighbor nodes of a node; normalizing features in the set of neighbor nodes; multiplying the normalized neighbor nodes by a weight matrix of the lth graph convolutional layer; Perform nonlinear transformation on the product result to obtain the output features of the node at the lth layer; The intermediate feature representation of nodes in each graph convolutional layer is retained through the skip connection mechanism, including: The output features of the nodes in each graph convolution layer are concatenated; the concatenated multi-layer features are reduced in dimension using the feature transformation layer; The feature propagation process of the graph convolution layer is iteratively performed L times, and the node features are updated in each iteration to obtain the node temporal context features containing multi-hop information.
6. The method according to claim 1, characterized in that Determining the behavior sequence matching degree through a dynamic programming algorithm according to the fused time series feature vector, and obtaining a virtual-real behavior consistency score in combination with a preset user behavior feature template; verifying and authorizing the user's behavior in the virtual scene of the metaverse based on the virtual-real behavior consistency score includes: Extracting a behavior feature point sequence from the fused time series feature vector, calculating the longest common subsequence of the current behavior feature point sequence and the historical behavior feature point sequence based on a dynamic programming algorithm, and obtaining a behavior sequence matching degree; According to the matching degree of the behavior sequence, the importance of the feature point sequence of the user's historical behavior characteristics is weighted to construct an adaptive user behavior feature template, wherein the importance weight decays with time distance; Calculating the multi-scale correlation score between the current fused time series feature vector and the adaptive user behavior feature template, and detecting the behavior mutation point based on the sliding time window to generate a behavior consistency score; When the behavior consistency score is higher than a preset score threshold, the user's operation request in the metaverse scene is permitted; When the behavior consistency score is lower than the preset score threshold, the identity re-authentication mechanism is triggered and the user's operation authority is temporarily frozen.
7. The method according to claim 6, characterized in that Extracting a behavior feature point sequence from the fused time series feature vector, calculating the longest common subsequence of the current behavior feature point sequence and the historical behavior feature point sequence based on a dynamic programming algorithm, and obtaining the behavior sequence matching degree includes: Perform extreme point detection on the time series feature vector to extract a current behavior feature point sequence; Acquire a historical behavior feature point sequence from a user behavior database, and construct the current behavior feature point sequence and the historical behavior feature point sequence into two sequences to be matched; Constructing a dynamic programming matrix, wherein the rows and columns of the dynamic programming matrix correspond to the lengths of the current behavior feature point sequence and the historical behavior feature point sequence respectively; Calculate the longest common subsequence of two sequences based on the dynamic programming algorithm: Initialize the matrix boundary conditions; when the difference between the feature point values is less than the preset difference threshold, the matrix element at the corresponding position is increased by 1; when the difference between the feature point values is greater than the preset difference threshold, the maximum value of the adjacent positions is taken; Iterate and update the matrix elements until all sequence positions are traversed; According to the final value of the dynamic programming matrix, the length of the longest common subsequence is calculated and divided by the total length of the sequence to obtain the behavior sequence matching degree.
8. A deep learning-based intelligent IoT and metaverse virtual-real fusion recognition system, used to implement the method as described in any one of claims 1 to 7, characterized in that: include: The first unit is used to obtain the virtual behavior data of the user in the virtual scene of the Metaverse collected by the terminal device and the actual behavior data of the user in the real scene collected by the Internet of Things device; A second unit is used to extract features from the virtual behavior data and the actual behavior data through a pre-trained deep neural network model to obtain a virtual behavior feature vector and an actual behavior feature vector; A third unit is used to cross-fuse the virtual behavior feature vector and the actual behavior feature vector based on an attention mechanism to generate a virtual-real fusion feature vector, wherein the attention mechanism calculates the correlation weight between the virtual behavior feature vector and the actual behavior feature vector to achieve adaptive feature alignment of virtual and real behaviors; The fourth unit is used to construct a temporal relationship graph of virtual and real behaviors using graph neural networks, model the temporal dependency relationship between nodes through conditional random fields, perform multi-hop information propagation on node features, obtain temporal context features, integrate the temporal context features of multiple subgraphs through a cross-subgraph feature fusion layer, and generate a fused temporal feature vector; The fifth unit is used to identify the consistency of the user's virtual and real behaviors according to the fused time series feature vector to obtain a virtual and real behavior consistency score; and verify and authorize the user's behavior in the virtual scene of the metaverse based on the virtual and real behavior consistency score.
9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described in any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
VR interaction method and device based on meta-universe virtual reality technology
CN118860156A
Immersive space virtual-real interaction method and system based on meta universe
CN119311124A
Multi-dimensional fusion meta universe and vertical AI model collaborative innovation platform
CN119443116A
Real-time rendering optimization method and system in meta universe scene building engine
CN119941956A
Cited By
Forging and pressing equipment detection method and system based on multi-source data fusion
CN120372253A
Method for judging virtual and real consistency of brake shoe of brake
CN120409143A