Intelligent Internet of Things and Metaverse Virtual-Reality Fusion Recognition Method and System Based on Deep Learning

Through deep learning and graph neural networks, the virtual and real behavior timing relationship diagram is constructed, which solves the problem of fusion of behavior data in virtual and real environments, and achieves high-accurate virtual and real behavior consistency recognition and verification, improving the security and user experience of the meta-universe environment.

CN120105041BActive Publication Date: 2025-07-18HANGZHOU MOXI TECH DEV CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510579056.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-07-18
Estimated Expiration
2045-05-07

AI Technical Summary

Technical Problem

The prior art is difficult to effectively integrate user behavior data in the virtual environment and the real environment, resulting in inconsistency between virtual and real behavior recognition, affecting identity authentication and secure access control.

Method used

Through deep learning and graph neural networks, virtual and real behavior timing relationship diagrams are built, feature alignment and cross-fusion are used to use attention mechanisms, and time-dependent dependencies are modeled in combination with conditional random fields to generate virtual and real behavior consistency scores for verification and authorization.

Benefits of technology

It improves the accuracy of identification of the consistency between virtual and real behavior, enhances the security and user experience of the metacosmic environment, and realizes accurate verification of user identity and behavior.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120105041B_ABST
    Figure CN120105041B_ABST
Patent Text Reader

Abstract

The present invention provides an intelligent IoT and metaverse virtual-real fusion recognition method and system based on deep learning, which relates to the technical field of the metaverse. It includes obtaining virtual and actual behavior data, extracting feature vectors using a deep neural network model, performing cross-fusion by combining an attention mechanism, and constructing a temporal relationship graph through a graph neural network, and finally realizing the scoring and verification of the consistency of the user's virtual and real behaviors. This method can effectively improve the accuracy of virtual-real behavior recognition and enhance the security and interaction experience in the metaverse scenario.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to metaverse technology, and particularly to an intelligent Internet of Things and metaverse virtual-real fusion recognition method and system based on deep learning. Background Art

[0002] With the rapid development of metaverse and Internet of Things technologies, the boundary between the virtual world and the real world is gradually blurred, and users can move in both virtual and real environments simultaneously. As an immersive virtual space, the metaverse allows users to conduct activities such as socializing, working, and entertaining through digital avatars, while Internet of Things technology enables various devices in the real world to be interconnected and to sense and collect users' behavioral data in real time. However, the problem of information silos between the virtual world and the real world still exists. The behavioral data of users in the two environments are often fragmented, lacking an effective association and fusion mechanism.

[0003] Currently, there are the following deficiencies in the field of virtual-real behavior recognition and fusion: existing behavior recognition methods usually only focus on the analysis of users' behaviors in a single environment and cannot effectively handle the problem of multi-modal data fusion in virtual and real environments, resulting in a one-sided and incomplete understanding of users' behaviors. Traditional feature extraction and fusion algorithms lack the ability of adaptive alignment for heterogeneous data in virtual-real scenarios and are difficult to capture the deep semantic associations between virtual behaviors and actual behaviors, affecting the accurate judgment of behavior consistency. Existing technologies have limited capabilities in dealing with the temporal dependence relationships of virtual-real behaviors and cannot fully explore the complex relationships between long-term behavior patterns and short-term behavior changes, thus making it difficult to achieve accurate verification of users' identities and behaviors.

[0004] Therefore, there is an urgent need for a method that can effectively fuse users' behavioral data in virtual and real environments, establish an association model between virtual-real behaviors through deep learning and multi-modal fusion technologies, and achieve accurate recognition and verification of the consistency of users' virtual-real behaviors, providing technical support for secure access control and user experience optimization in the metaverse environment. Summary of the Invention

[0005] Embodiments of the present invention provide an intelligent Internet of Things and metaverse virtual-real fusion recognition method and system based on deep learning, which can solve the problems in the prior art.

[0006] In the first aspect of the embodiments of the present invention,

[0007] Obtain the virtual behavior data of the user in the metaverse virtual scenario collected by the terminal device and the actual behavior data of the user in the real scenario collected by the Internet of Things device;

[0008] Extract features from the virtual behavior data and the actual behavior data through a pre-trained deep neural network model to obtain a virtual behavior feature vector and an actual behavior feature vector;

[0009] Cross - fuse the virtual behavior feature vector and the actual behavior feature vector based on the attention mechanism to generate a virtual - real fusion feature vector. Among them, the attention mechanism realizes the adaptive feature alignment of virtual - real behaviors by calculating the correlation weight between the virtual behavior feature vector and the actual behavior feature vector;

[0010] Use a graph neural network to construct a virtual - real behavior temporal relationship graph, model the temporal dependence relationship between nodes through a conditional random field, perform multi - hop information propagation on node features to obtain temporal context features, and integrate the temporal context features of multiple sub - graphs through a cross - sub - graph feature fusion layer to generate a fused temporal feature vector;

[0011] Identify the virtual - real behavior consistency of the user according to the fused temporal feature vector to obtain a virtual - real behavior consistency score; verify and authorize the user's behavior in the meta - universe virtual scene based on the virtual - real behavior consistency score.

[0012] Cross - fusing the virtual behavior feature vector and the actual behavior feature vector based on the attention mechanism to generate a virtual - real fusion feature vector includes:

[0013] Calculate the attention weight matrix between the virtual behavior feature vector and the actual behavior feature vector. Among them, for each feature element in the virtual behavior feature vector, calculate its correlation score with each feature element in the actual behavior feature vector through dot - product operation, and normalize the correlation score to obtain the attention weight matrix;

[0014] Based on the attention weight matrix, perform weighted aggregation on the virtual behavior feature vector and the actual behavior feature vector to obtain the context feature vector of the virtual behavior and the context feature vector of the actual behavior. Among them, the context feature vector of the virtual behavior is obtained by multiplying the attention weight matrix with the matrix of the actual behavior feature vector, and the context feature vector of the actual behavior is obtained by multiplying the transpose of the attention weight matrix with the matrix of the virtual behavior feature vector;

[0015] Concatenate the virtual behavior feature vector with the context feature vector of the virtual behavior, and concatenate the actual behavior feature vector with the context feature vector of the actual behavior to obtain an enhanced virtual behavior feature vector and an enhanced actual behavior feature vector;

[0016] Fuse the enhanced virtual behavior feature vector and the enhanced actual behavior feature vector through a non - linear transformation to generate a virtual - real fusion feature vector.

[0017] Construct a virtual-real behavior temporal relationship graph using a graph neural network, model the temporal dependence relationship between nodes through a conditional random field, perform multi-hop information propagation on node features to obtain temporal context features, and integrate the temporal context features of multiple subgraphs through a cross-subgraph feature fusion layer to generate a fused temporal feature vector, including:

[0018] Construct a virtual-real behavior temporal relationship graph using a graph neural network, and map the virtual-real fusion feature vector at each time step to the node features in the virtual-real behavior temporal relationship graph;

[0019] Model the temporal dependence relationship between nodes through a conditional random field, and dynamically predict the causal association strength between nodes based on the change trend of node features and the historical interaction pattern;

[0020] Normalize the causal association strength to obtain the sampling probability between node pairs; set a sampling probability threshold, and filter out node pairs higher than the sampling probability threshold as key temporal paths; centered on the key temporal paths, sample multiple local subgraphs through a random walk algorithm;

[0021] Perform the following operations on each of the local subgraphs:

[0022] Construct a local graph neural network and set multiple graph convolutional layers; in each graph convolution, aggregate the neighbor features of nodes and update the node representations; retain the node features of each layer through skip connections to obtain node temporal context features containing multi-hop information;

[0023] Integrate the information of multiple local subgraphs through a cross-subgraph feature fusion layer, specifically including:

[0024] Calculate the attention weights of the corresponding node features in different subgraphs; perform weighted fusion on the node temporal context features of multiple subgraphs based on the attention weights; obtain the final fused temporal feature vector through a pooling operation.

[0025] Model the temporal dependence relationship between nodes through a conditional random field, and dynamically predict the causal association strength between nodes based on the change trend of node features and the historical interaction pattern, including:

[0026] Model the temporal dependence relationship between nodes in the virtual-real behavior temporal relationship graph through a conditional random field, and calculate the difference vector of adjacent node features as the change trend of node features;

[0027] Construct the historical interaction pattern features of node pairs based on the interaction frequency of node pairs within the historical temporal window;

[0028] Input the change trend of the node features and the historical interaction pattern features into the conditional random field to predict the causal association strength between node pairs.

[0029] Construct a local graph neural network and set up multiple graph convolutional layers; in each graph convolution layer, aggregate the neighbor features of the nodes and update the node representations; through skip connections, retain the node features of each layer to obtain the node temporal context features containing multi-hop information, including:

[0030] Receive the input node feature vector and graph structure information, where the node feature vector represents the initial features of each node, and the graph structure information represents the connection relationship between nodes;

[0031] Construct a multi-layer graph convolutional network structure, including L graph convolutional layers, and each graph convolutional layer contains a learnable weight matrix and a bias vector, where L is an integer greater than 1;

[0032] For the l-th graph convolutional layer among the L graph convolutional layers, perform the following feature propagation operations:

[0033] Obtain the set of neighbor nodes of the node; normalize the features in the set of neighbor nodes; multiply the normalized neighbor nodes by the weight matrix of the l-th graph convolutional layer;

[0034] Perform a non-linear transformation on the product result to obtain the output features of the node at the l-th layer;

[0035] Retain the intermediate feature representations of the nodes in each graph convolutional layer through the skip connection mechanism, specifically including:

[0036] Concatenate the output features of the node in each graph convolutional layer; use a feature transformation layer to perform dimensionality reduction on the concatenated multi-layer features;

[0037] Iteratively execute the feature propagation process of the graph convolutional layer L times, and update the node features each time to obtain the node temporal context features containing multi-hop information.

[0038] Determine the behavior sequence matching degree through a dynamic programming algorithm according to the fused temporal feature vector, and obtain the virtual-real behavior consistency score in combination with a preset user behavior feature template; verify and authorize the user's behavior in the metaverse virtual scene based on the virtual-real behavior consistency score, including:

[0039] Extract the behavior feature point sequence from the fused temporal feature vector, calculate the longest common subsequence of the current behavior feature point sequence and the historical behavior feature point sequence based on the dynamic programming algorithm to obtain the behavior sequence matching degree;

[0040] Perform importance weighting on the historical behavior feature point sequence of the user according to the behavior sequence matching degree, and construct an adaptive user behavior feature template, where the importance weight decays with the time distance;

[0041] Calculate the multi-scale correlation score between the current fused temporal feature vector and the adaptive user behavior feature template, detect behavior mutation points based on a sliding time window, and generate a behavior consistency score;

[0042] When the behavior consistency score is higher than a preset score threshold, approve the operation request of the user in the metaverse scenario;

[0043] When the behavior consistency score is lower than the preset score threshold, trigger the identity re-authentication mechanism and temporarily freeze the user's operation permissions.

[0044] Extract the behavior feature point sequence from the fused temporal feature vector, calculate the longest common subsequence between the current behavior feature point sequence and the historical behavior feature point sequence based on the dynamic programming algorithm, and obtain the behavior sequence matching degree, including:

[0045] Perform extreme point detection on the temporal feature vector to extract the current behavior feature point sequence;

[0046] Obtain the historical behavior feature point sequence from the user behavior database, and construct the current behavior feature point sequence and the historical behavior feature point sequence into two sequences to be matched;

[0047] Construct a dynamic programming matrix, where the rows and columns of the dynamic programming matrix correspond to the lengths of the current behavior feature point sequence and the historical behavior feature point sequence respectively;

[0048] Calculate the longest common subsequence of the two sequences based on the dynamic programming algorithm:

[0049] Initialize the matrix boundary conditions; when the difference in feature point values is less than the preset difference threshold, add 1 to the corresponding matrix element; when the difference in feature point values is greater than the preset difference threshold, take the maximum value of the adjacent positions;

[0050] Iteratively update the matrix elements until all sequence positions are traversed;

[0051] According to the final value of the dynamic programming matrix, calculate the length of the longest common subsequence, and divide it by the total length of the sequence to obtain the behavior sequence matching degree.

[0052] In the second aspect of the embodiments of the present invention, a deep learning-based intelligent Internet of Things and metaverse virtual-real fusion recognition system is provided, including:

[0053] The first unit is used to obtain the virtual behavior data of the user in the metaverse virtual scene collected by the terminal device and the actual behavior data of the user in the real scene collected by the Internet of Things device;

[0054] A second unit, configured to extract features from the virtual behavior data and the actual behavior data through a pre-trained deep neural network model to obtain a virtual behavior feature vector and an actual behavior feature vector;

[0055] A third unit, configured to cross-fuse the virtual behavior feature vector and the actual behavior feature vector based on an attention mechanism to generate a virtual-real fusion feature vector, wherein the attention mechanism realizes adaptive feature alignment of virtual-real behaviors by calculating the correlation weight between the virtual behavior feature vector and the actual behavior feature vector;

[0056] A fourth unit, configured to construct a virtual-real behavior temporal relationship graph by using a graph neural network, model the temporal dependence relationship between nodes through a conditional random field, perform multi-hop information propagation on node features to obtain temporal context features, and integrate the temporal context features of multiple subgraphs through a cross-subgraph feature fusion layer to generate a fused temporal feature vector;

[0057] A fifth unit, configured to identify the virtual-real behavior consistency of a user according to the fused temporal feature vector to obtain a virtual-real behavior consistency score; and verify and authorize the behavior of the user in the metaverse virtual scene based on the virtual-real behavior consistency score.

[0058] In a third aspect of the embodiments of the present invention,

[0059] There is provided an electronic device, including:

[0060] A processor;

[0061] A memory for storing instructions executable by the processor;

[0062] Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.

[0063] In a fourth aspect of the embodiments of the present invention,

[0064] There is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.

[0065] The beneficial effects of the present application are as follows:

[0066] The present invention extracts features from virtual behavior data and actual behavior data through a pre-trained deep neural network model, and realizes adaptive feature alignment of virtual-real behaviors based on an attention mechanism, effectively solving the problem of heterogeneity of user behavior features in virtual scenes and real scenes, and improving the accuracy of virtual-real behavior consistency recognition.

[0067] The present invention uses graph neural networks to construct a temporal relationship graph of virtual and real behaviors, models the temporal dependency between nodes through conditional random fields, and combines cross-subgraph feature fusion technology to capture the correlation between the temporal context information of user behavior and multimodal features, thereby enhancing the ability to model complex behavior patterns and improving the robustness and generalization ability of the system.

[0068] The present invention verifies and authorizes the user's behavior in the virtual scene of the Metaverse based on the virtual-real behavior consistency score, constructs a complete virtual-real behavior verification mechanism, effectively prevents identity fraud and abnormal behavior in the Metaverse environment, ensures the security and user experience of the Metaverse platform, and promotes the deep integration of the Metaverse and Internet of Things technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] Figure 1 It is a flowchart of a method for identifying the virtual-real fusion of intelligent Internet of Things and metaverse based on deep learning according to an embodiment of the present invention;

[0070] Figure 2 It is a flowchart of constructing a time series relationship diagram of virtual and real behaviors and fusion of features according to an embodiment of the present invention;

[0071] Figure 3 This is a flowchart of node feature processing based on a graph neural network according to an embodiment of the present invention;

[0072] Figure 4 The figure is a flowchart of user verification based on behavior sequence matching according to an embodiment of the present invention. DETAILED DESCRIPTION

[0073] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0074] The technical solution of the present invention is described in detail with specific embodiments below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.

[0075] Figure 1 Schematic diagram of the process of the method for identifying the virtual-real fusion of intelligent IoT and metaverse based on deep learning according to an embodiment of the present invention. Figure 1 As shown, the method includes:

[0076] Obtain the virtual behavior data of the user in the metaverse virtual scene collected by the terminal device and the actual behavior data of the user in the real scene collected by the Internet of Things devices;

[0077] Extract features from the virtual behavior data and the actual behavior data through a pre-trained deep neural network model to obtain a virtual behavior feature vector and an actual behavior feature vector;

[0078] Based on the attention mechanism, cross-fuse the virtual behavior feature vector and the actual behavior feature vector to generate a virtual-real fusion feature vector, where the attention mechanism realizes the adaptive feature alignment of virtual-real behaviors by calculating the correlation weight between the virtual behavior feature vector and the actual behavior feature vector;

[0079] Use a graph neural network to construct a virtual-real behavior temporal relationship graph, model the temporal dependence relationship between nodes through a conditional random field, perform multi-hop information propagation on node features to obtain temporal context features, and integrate the temporal context features of multiple subgraphs through a cross-subgraph feature fusion layer to generate a fused temporal feature vector;

[0080] Identify the virtual-real behavior consistency of the user according to the fused temporal feature vector to obtain a virtual-real behavior consistency score; verify and authorize the user's behavior in the metaverse virtual scene based on the virtual-real behavior consistency score.

[0081] In an alternative embodiment, cross-fusing the virtual behavior feature vector and the actual behavior feature vector based on the attention mechanism to generate a virtual-real fusion feature vector includes:

[0082] Calculate the attention weight matrix between the virtual behavior feature vector and the actual behavior feature vector. For each feature element in the virtual behavior feature vector, calculate its correlation score with each feature element in the actual behavior feature vector through a dot product operation, and normalize the correlation score to obtain the attention weight matrix;

[0083] Based on the attention weight matrix, perform weighted aggregation on the virtual behavior feature vector and the actual behavior feature vector to obtain a context feature vector of virtual behavior and a context feature vector of actual behavior, where the context feature vector of virtual behavior is obtained by multiplying the attention weight matrix with the matrix of the actual behavior feature vector, and the context feature vector of actual behavior is obtained by multiplying the transpose of the attention weight matrix with the matrix of the virtual behavior feature vector;

[0084] Concatenate the virtual behavior feature vector and the context feature vector of the virtual behavior, and concatenate the actual behavior feature vector and the context feature vector of the actual behavior to obtain an enhanced virtual behavior feature vector and an enhanced actual behavior feature vector;

[0085] Fuse the enhanced virtual behavior feature vector and the enhanced actual behavior feature vector through non - linear transformation to generate a virtual - real fusion feature vector.

[0086] Obtain the virtual behavior feature vector and the actual behavior feature vector to be processed. Assume the virtual behavior feature vector is denoted as V, with a dimension of d×1, where d represents the feature dimension, for example, d = 128; the actual behavior feature vector is denoted as R, with the same dimension of d×1. In practical applications, these feature vectors can be extracted from the original data through a deep neural network.

[0087] Calculate the attention weight matrix between the virtual behavior feature vector V and the actual behavior feature vector R. Specifically, for each feature element v_i in the virtual behavior feature vector V, calculate its correlation score s_ij with each feature element r_j in the actual behavior feature vector R through dot - product operation. For example, if the first element value in V is 0.5 and the first element value in R is 0.7, then their dot - product correlation score is 0.5×0.7 = 0.35.

[0088] For the correlation scores of all feature element pairs, perform normalization processing through the softmax function to obtain the attention weight matrix A. The normalization processing ensures that the sum of weights in each row is 1, making the attention allocation more reasonable. For example, if the original correlation scores are [0.35, 0.42, 0.28], the normalized weights may be [0.33, 0.41, 0.26].

[0089] Perform weighted aggregation on the virtual behavior feature vector V and the actual behavior feature vector R based on the attention weight matrix A. Specifically, the context feature vector CV of the virtual behavior is obtained by multiplying the attention weight matrix A with the actual behavior feature vector R. In actual implementation, if A is a d×d - dimensional matrix and R is a d×1 - dimensional vector, then CV = A×R, and the result is a d×1 - dimensional vector.

[0090] The context feature vector CR of the actual behavior is obtained by multiplying the transpose of the attention weight matrix A with the virtual behavior feature vector V. That is, CR = A^T×V, where A^T represents the transpose matrix of A, and the result is also a d×1 - dimensional vector.

[0091] This cross-attention mechanism allows virtual behavior features and actual behavior features to influence each other, thereby capturing the correlation information between them. For example, if an operation in the virtual scenario is highly correlated with a specific action in the actual scenario, a strong attention connection will be established between them.

[0092] Feature concatenation is performed on the virtual behavior feature vector V and the context feature vector CV of the virtual behavior to obtain the enhanced virtual behavior feature vector EV. In implementation, the concatenation operation can be expressed as EV = [V; CV], where ";" represents the vector concatenation operation, and the resulting dimension is 2d×1.

[0093] Feature concatenation is performed on the actual behavior feature vector R and the context feature vector CR of the actual behavior to obtain the enhanced actual behavior feature vector ER. That is, ER = [R; CR], and the dimension is also 2d×1.

[0094] The enhanced virtual behavior feature vector EV and the enhanced actual behavior feature vector ER are fused through a non-linear transformation to generate the virtual-real fusion feature vector F. This non-linear transformation can be implemented through a multi-layer perceptron (MLP), which includes weight matrices W1 and W2 and bias terms b1 and b2.

[0095] EV and ER are concatenated to obtain a joint vector [EV; ER] of 4d×1, and then it is transformed through a two-layer MLP:

[0096] The intermediate representation H = ReLU(W1×[EV; ER] + b1); the final fusion vector F = W2×H + b2.

[0097] ReLU represents the rectified linear unit activation function, which is used to introduce the non-linear transformation ability. Assume that the dimension of W1 is k×4d, the dimension of b1 is k×1, then the dimension of H is k×1; the dimension of W2 is m×k, the dimension of b2 is m×1, then the dimension of the final fusion vector F is m×1, where k can be set to 256 and m can be set to 128.

[0098] Assume that the original virtual behavior feature vector V and the actual behavior feature vector R are both 128-dimensional vectors. Through the above cross-attention fusion method, first, the 128×128-dimensional attention weight matrix A is calculated, and then the 128-dimensional context feature vectors CV and CR are obtained. After feature concatenation, both EV and ER are 256-dimensional vectors. Finally, through the MLP transformation, the 128-dimensional virtual-real fusion feature vector F is obtained.

[0099] Multiple samples are processed simultaneously in a batch processing manner. For example, if a batch of data contains 32 samples, the input V and R can be represented as a 32×128 matrix, where each row represents the feature vector of a sample. Correspondingly, the attention weight calculation and subsequent operations will also be performed in a batch processing manner.

[0100] To improve the expressive ability of the fused features, a multi-head attention mechanism can be introduced based on the attention mechanism. For example, 8 attention heads are used, and each head generates 16-dimensional context features (a total of 128 dimensions), and then the outputs of all heads are concatenated to form the final context features. This multi-head mechanism enables the model to learn information from different representation subspaces and further enhances the effect of feature fusion.

[0101] Through the above-mentioned method for fusing virtual and real behavior features based on the attention mechanism, the correlation information between virtual behaviors and actual behaviors can be effectively captured, generating a virtual-real fusion feature vector with rich semantics, providing strong support for subsequent analysis and decision-making tasks.

[0102] In an alternative implementation, a graph neural network is used to construct a virtual-real behavior temporal relationship graph, and a conditional random field is used to model the temporal dependence relationship between nodes, perform multi-hop information propagation on node features, obtain temporal context features, and integrate the temporal context features of multiple subgraphs through a cross-subgraph feature fusion layer to generate a fused temporal feature vector, including:

[0103] Use a graph neural network to construct a virtual-real behavior temporal relationship graph, and map the virtual-real fusion feature vector at each time step to the node features in the virtual-real behavior temporal relationship graph;

[0104] Model the temporal dependence relationship between nodes through a conditional random field, and dynamically predict the causal association strength between nodes based on the change trend of node features and historical interaction patterns;

[0105] Normalize the causal association strength to obtain the sampling probability between node pairs; set a sampling probability threshold, and filter out node pairs higher than the sampling probability threshold as key temporal paths; centered on the key temporal paths, sample multiple local subgraphs through a random walk algorithm;

[0106] Perform the following operations on each of the local subgraphs:

[0107] Construct a local graph neural network and set multiple graph convolutional layers; in each graph convolution, aggregate the neighbor features of nodes and update the node representations; retain the node features of each layer through skip connections to obtain node temporal context features containing multi-hop information;

[0108] Integrate the information of multiple local subgraphs through a cross-subgraph feature fusion layer, specifically including:

[0109] Calculate the attention weights for the corresponding node features in different sub - graphs; weighted - fuse the node temporal context features of multiple sub - graphs based on the attention weights; and obtain the final fused temporal feature vector through a pooling operation.

[0110] Construct a virtual - real behavior temporal relationship graph using a graph neural network (GNN). The nodes of this graph represent different behavior states, and the edges represent the temporal relationships between these behaviors. The virtual - real fusion feature vector at each time step will be mapped as the node feature in this graph. The construction of the feature vector can pre - process the behavior data to extract key temporal features and context information to ensure the effectiveness of the node features.

[0111] Model the temporal dependence relationship between nodes through a conditional random field (CRF). This process includes analyzing the change trends and historical interaction patterns of node features to dynamically predict the causal association strength between nodes. Specifically, first collect the historical behavior data of nodes to construct feature vectors, and then use these features to evaluate the causal relationship between nodes. For example, if the feature value of node A changes significantly at time t and the feature value of node B also changes at time t + 1, it can be inferred that there is a causal association between A and B.

[0112] Normalize the predicted causal association strength to obtain the sampling probability between node pairs. This process ensures that the sampling probabilities of all node pairs are between 0 and 1. By setting a sampling probability threshold, select the node pairs above this threshold as the key temporal paths. These key paths will be the focus of subsequent analysis to ensure that the model pays attention to important temporal relationships.

[0113] Centered on the selected key temporal paths, sample multiple local sub - graphs through a random - walk algorithm. The random - walk process includes starting from a key node, randomly selecting adjacent nodes for traversal until a preset number of steps is reached or a specific node is traversed. Each local sub - graph will contain the nodes related to the key path and their features, providing a basis for the subsequent training of the graph neural network.

[0114] For each local sub - graph, construct a local graph neural network and set multiple graph convolutional layers. In each layer of graph convolution, aggregate the neighbor features of nodes and update the node representations. Retain the node features of each layer through skip connections, enabling each node to combine information from different layers, thereby obtaining the node temporal context features containing multi - hop information. This process can enhance the expressive ability of node features through multiple iterations.

[0115] Integrate the information of multiple local subgraphs through a cross-subgraph feature fusion layer. First, calculate the attention weights of the corresponding node features in different subgraphs. These weights reflect the importance of the node features in each subgraph. Then, based on these attention weights, perform weighted fusion on the node temporal context features of multiple subgraphs. Finally, obtain the fused temporal feature vector through a pooling operation, which will be used for subsequent analysis and prediction tasks.

[0116] Suppose there is a set of user behavior data, including users' click, browse, and purchase behaviors at different time points. Through preprocessing, extract the feature vectors of each user at each time point, such as click-through rate, browse duration, and purchase amount. Use these features to construct a virtual-real behavior temporal relationship graph, and the model can effectively identify the behavior patterns of users within a specific time period.

[0117] When modeling the temporal dependence relationship between nodes, assume that the click-through rate of user A at time t is 0.8, while the purchase rate of user B at time t + 1 is 0.5. The model can infer that A's behavior may affect B's purchase decision. By setting a sampling probability threshold, filter out the key temporal paths, such as the relationship between A's click behavior and B's purchase behavior.

[0118] By constructing a local graph neural network and cross-subgraph feature fusion, the model can generate an efficient temporal feature vector to support subsequent user behavior prediction.

[0119] Figure 2 The following is the flowchart for constructing and fusing the virtual-real behavior temporal relationship graph in the embodiments of the present invention:

[0120] This flowchart shows a complete process of updating and fusing node temporal features. Construct a virtual-real temporal relationship graph, which is used to map the node features in the virtual-real fusion feature vector. Model the node time interval dependence of the conditional random field to predict the relationship strength between nodes. Obtain the similarity probability between nodes through normalization. Determine the key temporal paths based on the sampling probability graph, and use the random walk algorithm to divide the samples into multiple local subgraphs. Perform graph neural network operations on each local subgraph, including feature aggregation and update. Construct a local graph neural network and set multiple graph convolutional layers to achieve information interaction through node feature propagation. Fuse the aggregated node features with the original features to update the temporal representation of the nodes. Retain the node temporal features of each layer through skip connections to obtain the node context features of multi-hop information. Perform weighted fusion on the cross-subgraph node features based on the attention weights to obtain the final node temporal feature representation. Through multi-level feature extraction and fusion, the entire process effectively captures the temporal dependence relationship between nodes and generates node representations containing rich context information.

[0121] In an alternative embodiment, the temporal dependence relationship between nodes is modeled by a conditional random field. Based on the change trend of node features and historical interaction patterns, dynamically predicting the causal association strength between nodes includes:

[0122] Model the temporal dependence relationship between nodes in the virtual-real behavior temporal relationship graph through a conditional random field, and calculate the difference vector of adjacent node features as the change trend of node features;

[0123] Construct the historical interaction pattern features of node pairs based on the interaction frequency of node pairs within the historical time window;

[0124] Input the change trend of the node features and the historical interaction pattern features into the conditional random field to predict the causal association strength between node pairs.

[0125] In the virtual-real behavior temporal relationship graph, each node represents a behavior or event, and the edges between nodes indicate the possible causal associations between them. Each node has a series of feature values that change over time. The conditional random field constructed by this method predicts the causal association strength between node pairs by analyzing the changes in node features at adjacent time points and the historical interaction patterns between nodes.

[0126] When calculating the change trend of node features, for each pair of adjacent time points t and t + 1 in the time series, respectively obtain the feature vectors Fi,t and Fi,t+1 of node i at these two time points. Calculate the difference between these two feature vectors to obtain the change trend vector ΔFi,t = Fi,t+1 - Fi,t. This difference vector reflects the change direction and magnitude of node features within this time period.

[0127] Suppose node i represents the browsing behavior of a user, and its feature vector includes three dimensions: browsing duration, number of clicks, and page stay time. At time point t, its feature vector Fi,t is [30, 5, 120], indicating a browsing duration of 30 seconds, 5 clicks, and a page stay time of 120 seconds. At time point t + 1, its feature vector Fi,t+1 is [45, 8, 180]. Then the calculated change trend vector ΔFi,t is [15, 3, 60], indicating an increase in browsing duration by 15 seconds, an increase in the number of clicks by 3 times, and an increase in page stay time by 60 seconds.

[0128] For modeling the relationship between two nodes i and j, it is necessary to consider the change trend vectors of these two nodes simultaneously. They can be concatenated to form a joint change trend feature vector [ΔFi,t, ΔFj,t].

[0129] Construct the historical interaction pattern features of node pairs. Select a time window W of a fixed size, and count the interaction frequency and pattern between node i and node j within this window. The interaction pattern can be characterized from multiple dimensions, including interaction frequency, interaction time interval distribution, interaction direction ratio, etc.

[0130] For example, assume that within the past 10 time units, there are a total of 8 interactions between node i and node j, among which 6 are from i to j and 2 are from j to i. The time intervals of these interactions are [1, 2, 1, 3, 1, 2, 1, 2] time units respectively.

[0131] Combine the change trend of node features and the historical interaction pattern features to form the input feature vector Xi,j of the conditional random field = [[ΔFi,t, ΔFj,t], Hi,j]. The conditional random field predicts the causal association strength Si,j,t+1 between node i and node j at the next time point t+1 based on these features.

[0132] The conditional random field is a probabilistic graphical model, and its core is to define a set of feature functions and corresponding weights. Each feature function evaluates a certain aspect of the input features, and the weights reflect the influence degree of this feature function on the final prediction result.

[0133] Use the labeled historical data to learn the weights of the conditional random field. These labeled data contain the actual causal association strength of node pairs at each time point. Through maximum likelihood estimation or other optimization methods, find a set of optimal weight values to make the prediction of the model for the training data as close as possible to the actual labels.

[0134] Given the node features and historical interaction patterns at the current time point t, the conditional random field calculates the probability distribution of various possible causal association strength values and selects the value with the highest probability as the prediction result.

[0135] Adopt a sliding window method to continuously update the historical interaction pattern features. As new interaction data is generated, the window slides forward, maintaining a fixed size to ensure that the historical interaction pattern features always reflect the most recent interaction situation.

[0136] Set different feature extraction methods according to different types of nodes and interactions. For example, for nodes representing user behavior, the frequency, duration, and intensity changes of the behavior can be concerned; for nodes representing system events, the trigger conditions and changes in the influence range of the events can be concerned.

[0137] Suppose there is a user behavior analysis system for an e-commerce platform, where nodes can represent various user behaviors, such as browsing products, adding items to the shopping cart, and viewing reviews. The system needs to analyze the causal associations between these behaviors to better understand the user's shopping decision-making process.

[0138] Taking browsing products (node A) and adding items to the shopping cart (node B) as an example, the system has collected the behavior data of users at 10 consecutive time points. The features of node A include browsing duration, page switching times, and product details expansion times; the features of node B include the number of products added to the shopping cart, the total amount added to the shopping cart, and the continued browsing time after adding items to the shopping cart.

[0139] By calculating the change trends of node features and constructing historical interaction pattern features, the conditional random field predicts that the causal association strength of node A to node B is 0.75, indicating a strong causal association between the behavior of browsing products and the subsequent behavior of adding items to the shopping cart. This prediction result can help the platform optimize the product display strategy and improve the user's shopping conversion rate.

[0140] By modeling the temporal dependence relationship between nodes through the conditional random field, based on the change trends of node features and historical interaction pattern features, it is possible to effectively and dynamically predict the causal association strength between nodes, providing a powerful tool for analyzing the causal relationships between various elements in a complex system.

[0141] In an alternative implementation, a local graph neural network is constructed, and multiple graph convolutional layers are set; in each graph convolution, the neighbor features of the nodes are aggregated and the node representations are updated; through skip connections, the node features of each layer are retained, and the node temporal context features containing multi-hop information are obtained, including:

[0142] Receive the input node feature vector and graph structure information, where the node feature vector represents the initial features of each node, and the graph structure information represents the connection relationship between nodes;

[0143] Construct a multi-layer graph convolutional network structure, including L graph convolutional layers, each graph convolutional layer containing a learnable weight matrix and a bias vector, where L is an integer greater than 1;

[0144] For the l-th graph convolutional layer among the L graph convolutional layers, perform the following feature propagation operations:

[0145] Obtain the set of neighbor nodes of the node; normalize the features in the set of neighbor nodes; multiply the normalized neighbor nodes by the weight matrix of the l-th graph convolutional layer;

[0146] Perform a non-linear transformation on the product result to obtain the output features of the node at the l-th layer;

[0147] The intermediate feature representations of nodes in each graph convolutional layer are retained through a skip connection mechanism, specifically including:

[0148] Concatenate the output features of nodes in each graph convolutional layer; use a feature transformation layer to perform dimensionality reduction on the concatenated multi-layer features;

[0149] Iteratively execute the feature propagation process of the graph convolutional layer L times, updating the node features each time to obtain the node temporal context features containing multi-hop information.

[0150] Receive the input node feature vector and graph structure information. The node feature vector represents the initial features of each node, which can be the attribute information of the node, historical behavior data, or other relevant features. The graph structure information represents the connection relationship between nodes, usually provided in the form of an adjacency matrix or an edge list. For example, in a social network, nodes can be users, and node features can include information such as the age, gender, and hobbies of the users, while the graph structure information represents the friendship relationship between users.

[0151] Construct a multi-layer graph convolutional network structure, including L graph convolutional layers. In this embodiment, L is set to 3, that is, a 3-layer graph convolutional network is constructed. Each graph convolutional layer contains a learnable weight matrix and a bias vector for feature transformation and information transmission. The dimension of the weight matrix depends on the input feature dimension and the output feature dimension. For example, the dimension of the first-layer weight matrix is 64×32, the second layer is 32×32, and the third layer is 32×32.

[0152] For the l-th graph convolutional layer, first obtain the set of neighbor nodes of each node. For example, for node A, its neighbor nodes may include nodes B, C, and D. Then, normalize the features in the set of neighbor nodes to balance the differences in the number of neighbor nodes of different nodes. The normalization process can be achieved by dividing the features of each neighbor node by the square root of the degree (the number of neighbor nodes) of that node and then dividing by the square root of the degree of the central node.

[0153] Multiply the neighbor node features by the weight matrix of the l-th graph convolutional layer. Specifically, assume that the feature vectors of neighbor nodes B, C, and D are [0.3, 0.5, 0.2,...], [0.1, 0.7, 0.4,...], [0.6, 0.2, 0.8,...] respectively. After normalization, the obtained feature vectors are [0.2, 0.33, 0.13,...], [0.07, 0.47, 0.27,...], [0.4, 0.13, 0.53,...] respectively. Multiply these normalized feature vectors by the weight matrix to obtain the transformed feature representation.

[0154] Perform a non - linear transformation on the product result to obtain the output feature of the node at the l - th layer. The non - linear transformation usually uses the ReLU function, which preserves positive values and sets negative values to zero, enhancing the expressive power of the network. For example, for node A, the feature after the first - layer graph convolution may be [0.25, 0.41, 0, 0.32, ...].

[0155] Adopt the skip - connection mechanism. Specifically, concatenate the output features of the node at each layer of the graph convolution layer. For example, for node A, assuming its features after three - layer graph convolution are [0.25, 0.41, 0, 0.32, ...] (the first layer), [0.35, 0, 0.22, 0.18, ...] (the second layer), and [0.15, 0.27, 0.39, 0, ...] (the third layer), then the concatenated feature is [0.25, 0.41, 0, 0.32, ..., 0.35, 0, 0.22, 0.18, ..., 0.15, 0.27, 0.39, 0, ...].

[0156] Use the feature transformation layer to perform dimensionality reduction on the concatenated multi - layer features. The feature transformation layer can be a fully - connected layer, which maps high - dimensional features to a lower - dimensional space. In this embodiment, the feature transformation layer reduces the concatenated 96 - dimensional feature (assuming 32 dimensions are output for each layer) to 64 dimensions to obtain the final node representation.

[0157] By iteratively executing the feature propagation process of the graph convolution layer L times, updating the node features each time, finally obtain the node temporal context features containing multi - hop information. These features fuse the node's own information and the information of neighbor nodes with different hop counts, and can comprehensively characterize the position and context relationship of the node in the graph structure.

[0158] In specific application scenarios, such as a recommendation system, users and items can be modeled as nodes in a graph, and the interaction behavior of users with items can be modeled as edges. By extracting the temporal context features of user nodes using the above - mentioned method, the evolution of user interests and social influence can be captured, thereby improving the recommendation accuracy. In actual tests, compared with traditional methods, the click - through rate of the recommendation system using this method has increased by 15.3%, and the conversion rate has increased by 12.7%.

[0159] In the anomaly detection scenario, network devices can be modeled as nodes, and the communication between devices can be modeled as edges. By extracting the temporal context features of device nodes, abnormal communication patterns can be identified, and network intrusion behaviors can be detected in a timely manner. Experiments show that the anomaly detection accuracy of this method reaches 92.8%, which is 7.5 percentage points higher than the baseline method.

[0160] The method provided in this embodiment effectively extracts the node temporal context features containing multi-hop information through a multi-layer graph convolutional network and a skip connection mechanism, providing a powerful tool for graph data analysis and mining.

[0161] Figure 3 This is the flowchart of node feature processing based on the graph neural network in the embodiment of the present invention:

[0162] This figure shows a node feature processing flow based on a graph neural network. This method receives the input node feature vector and graph structure information, where the node feature vector is used to represent the initial feature attributes of each node, and the graph structure information describes the topological connection relationship between nodes. On this basis, a multi-layer graph convolutional network structure with L layers is constructed, and each layer is equipped with learnable weight matrix and bias vector parameters. L is an integer greater than 1 and is used to control the depth of the network. For each layer of graph convolution operation, a specific feature propagation process is performed: first, obtain the set of neighbor nodes of the target node, normalize the features of these neighbor nodes, then multiply the processed features by the weight matrix of this layer, and finally obtain the output feature representation of the node at this layer through a non-linear activation function. In particular, the skip connection mechanism is used to retain the intermediate features of the node in each convolutional layer. The specific approach is to splice the output features of each layer, use a feature transformation layer for dimensionality reduction, and finally add a residual connection to fuse the dimensionality-reduced features with the initial node features. This process is iteratively executed L times, and the feature representation of the node is updated each time, and finally the complete node temporal context features containing multi-hop information are obtained. This design can effectively capture the structural information and semantic features of nodes at different scales and improve the expressive ability of feature representation.

[0163] In an alternative embodiment, the behavior sequence matching degree is determined through a dynamic programming algorithm according to the fused temporal feature vector, and the virtual-real behavior consistency score is obtained by combining a preset user behavior feature template; verifying and authorizing the user's behavior in the metaverse virtual scene based on the virtual-real behavior consistency score includes:

[0164] Extract the behavior feature point sequence from the fused temporal feature vector, calculate the longest common subsequence of the current behavior feature point sequence and the historical behavior feature point sequence based on the dynamic programming algorithm, and obtain the behavior sequence matching degree;

[0165] Weight the importance of the historical behavior feature point sequence of the user according to the behavior sequence matching degree, and construct an adaptive user behavior feature template, where the importance weight decays with the time distance;

[0166] Calculate the multi-scale correlation score between the current fused temporal feature vector and the adaptive user behavior feature template, detect behavior mutation points based on a sliding time window, and generate a behavior consistency score;

[0167] When the behavior consistency score is higher than a preset score threshold, approve the user's operation request in the metaverse scenario;

[0168] When the behavior consistency score is lower than a preset score threshold, trigger the identity re-authentication mechanism and temporarily freeze the user's operation permissions.

[0169] Extract the sequence of behavior feature points from the fused temporal feature vector. The temporal feature vector obtained by the system contains multi-dimensional data, such as information on the user's operation actions, interaction patterns, movement trajectories, etc. in the virtual scenario. Through a feature point detection algorithm, key behavior feature points are extracted from these temporal feature vectors to form a sequence of behavior feature points. For example, for a user in a virtual scenario, a sequence of behavior feature points such as "login - browse - interact - purchase - logout" may be extracted.

[0170] Calculate the longest common subsequence between the current sequence of behavior feature points and the historical sequence of behavior feature points based on the dynamic programming algorithm to obtain the behavior sequence matching degree. In specific implementation, let the currently extracted sequence of behavior feature points be A with a length of m, and the historical sequence of the user's behavior feature points be B with a length of n. The system constructs a two-dimensional table of size (m + 1)×(n + 1) and calculates the length of the longest common subsequence by filling the table. The final behavior sequence matching degree can be expressed as the ratio of the length of the longest common subsequence to the total length of the sequence, with a value range from 0 to 1. For example, if the current sequence is "login - browse - interact - purchase - logout" and the historical sequence is "login - interact - browse - purchase - logout", the longest common subsequence calculated by dynamic programming is "login - interact - purchase - logout" with a length of 4, then the matching degree is 4 / 5 = 0.8.

[0171] Perform importance weighting on the historical sequence of the user's behavior feature points according to the behavior sequence matching degree to construct an adaptive user behavior feature template. The system assigns different weights to each feature point in the historical behavior sequence according to its occurrence frequency and time distance. The time distance decay function adopts an exponential decay form, and the behavior feature points closer to the current time obtain higher weights. For example, for a feature point, if it has occurred 5 times in the last 7 days, its basic weight can be set to 0.5; if its occurrence time is 3 days ago and the time decay factor is 0.9^3≈0.73, then the final weight of this feature point is 0.5×0.73≈0.365. The system integrates all feature points and their weights to form a user behavior feature template.

[0172] Calculate the multi-scale correlation scores between the current fused temporal feature vector and the adaptive user behavior feature template. The system calculates the correlation scores at different time window scales respectively, for example, three time scales of 5 minutes, 30 minutes, and 2 hours can be set. Within each time scale, the system calculates the similarity between the current feature vector and the template, and performs a weighted average on the scores of different scales to obtain the overall correlation score. The system also detects behavior mutation points based on a sliding time window, that is, when the behavior pattern within a specific time window changes significantly compared to the previous time window, it is marked as a behavior mutation point. The final behavior consistency score is a combination of the correlation score and the mutation point detection result, with a value range of 0 to 100. For example, if the multi-scale correlation score is 85 points and 1 medium-level behavior mutation point is detected (deduct 10 points), the final behavior consistency score is 75 points.

[0173] When the behavior consistency score is higher than the preset score threshold, the system approves the user's operation request in the metaverse scenario. The preset score threshold can be set to different values according to the security level of the operation. Generally, it can be set to 70 points. For example, when the user requests to browse goods in the virtual mall and the calculated behavior consistency score is 85 points, which is higher than the preset threshold of 70 points, the system approves the operation request.

[0174] When the behavior consistency score is lower than the preset score threshold, the system triggers the identity re-authentication mechanism and temporarily freezes the user's operation permissions. For example, when the user requests to conduct a large virtual asset transaction in the virtual world and the calculated behavior consistency score is 65 points, which is lower than the preset threshold of 70 points, the system will immediately require the user to conduct identity re-authentication, which may include methods such as secondary password verification, biometric verification, or mobile phone SMS verification. Before the user completes the re-authentication, the system temporarily freezes the relevant operation permissions of the user to prevent potential account theft risks.

[0175] The system will dynamically adjust each parameter according to different users and different scenarios. For example, for newly registered users, due to less historical behavior data, the system may lower the preset score threshold to 60 points; while for transactions involving high-value virtual assets, the system may raise the preset score threshold to 85 points to provide more stringent security protection.

[0176] The system also implements an adaptive update mechanism for the template. As the user's activity time in the metaverse environment increases, the system continuously collects the user's behavior data and regularly updates the user's behavior feature template. The update frequency can be set to once a week or once a day to ensure that the template can accurately reflect the user's latest behavior characteristics. For users who have not been active for a long time, the weights in their behavior feature templates will gradually decay. When such users become active again, the system will conduct a more cautious behavior consistency assessment.

[0177] Through the comprehensive application of the above technical means, the virtual-real behavior consistency verification method provided by this embodiment can effectively identify abnormal operation behaviors, ensure the security of users in the metaverse environment, and at the same time provide a smooth user experience.

[0178] Figure 4 The following is the user verification flowchart based on behavior sequence matching in the embodiments of the present invention:

[0179] This figure shows a user verification process based on behavior sequence matching, which can be integrated and described as follows:

[0180] This method first extracts a sequence of behavior feature points from the fused temporal feature vectors, and uses the dynamic programming algorithm to calculate the longest common subsequence between the current sequence of behavior feature points and the historical sequence of behavior feature points, so as to obtain the matching degree of the behavior sequence. Based on this matching degree, the system performs importance weighting on the historical behavior features of the user to construct an adaptive user behavior feature template, where the importance weight of the feature decays as the time distance increases. Then, the system calculates the multi-dimensional correlation score between the current fused temporal feature vector and the adaptive behavior feature template, and at the same time detects behavior mutation points by sliding a time window, and comprehensively generates a behavior consistency score. This score is compared with a preset threshold: when the behavior consistency score is higher than the preset threshold, the system authorizes the user's operation request in the metaverse scenario and synchronously updates the user's behavior feature template; on the contrary, when the score is lower than the preset threshold, the system triggers an identity re-verification mechanism and takes security measures such as temporarily freezing the high-risk operation permissions of the user. This verification mechanism based on behavior sequence matching can dynamically evaluate the consistency of user behavior and effectively prevent potential security risks.

[0181] In an alternative embodiment, extracting a sequence of behavior feature points from the fused temporal feature vectors and calculating the longest common subsequence between the current sequence of behavior feature points and the historical sequence of behavior feature points based on the dynamic programming algorithm to obtain the behavior sequence matching degree includes:

[0182] Performing extreme point detection on the temporal feature vectors to extract the current sequence of behavior feature points;

[0183] Obtaining the historical sequence of behavior feature points from the user behavior database, and constructing the current sequence of behavior feature points and the historical sequence of behavior feature points into two sequences to be matched;

[0184] Constructing a dynamic programming matrix, where the rows and columns of the dynamic programming matrix correspond to the lengths of the current sequence of behavior feature points and the historical sequence of behavior feature points respectively;

[0185] Calculating the longest common subsequence of the two sequences based on the dynamic programming algorithm:

[0186] Initialize the matrix boundary conditions; when the difference in the values of the feature points is less than the preset difference threshold, add 1 to the matrix element at the corresponding position; when the difference in the values of the feature points is greater than the preset difference threshold, take the maximum value of the adjacent positions.

[0187] Iteratively update the matrix elements until all sequence positions are traversed.

[0188] According to the final value of the dynamic programming matrix, calculate the length of the longest common subsequence and divide it by the total length of the sequence to obtain the behavior sequence matching degree.

[0189] Extract the behavior feature point sequence from the fused time-series feature vector, which can be obtained by fusing multi-modal sensor data. For example, the data of the acceleration sensor, gyroscope sensor, and pressure sensor can be fused, and after filtering, normalization, and time-domain feature extraction, a fused time-series feature vector is formed.

[0190] Perform extreme point detection on the time-series feature vector to extract the current behavior feature point sequence. The extreme point detection uses the sliding window method with a window size of 5 sampling points. When the value of the center point is greater than the values of other points in the window, it is determined as a maximum point; when the value of the center point is less than the values of other points in the window, it is determined as a minimum point. The maximum points and minimum points are collectively referred to as extreme points, and these extreme points form the current behavior feature point sequence in chronological order.

[0191] For example, for a time-series feature vector [0.2, 0.3, 0.5, 0.4, 0.2, 0.1, 0.3, 0.6, 0.7, 0.4], after extreme point detection, it can be determined that the 3rd point (value 0.5) is a maximum point, the 6th point (value 0.1) is a minimum point, and the 9th point (value 0.7) is a maximum point. Therefore, the current behavior feature point sequence is [(3,0.5), (6,0.1), (9,0.7)], where each element represents (position, value).

[0192] Obtain the historical behavior feature point sequence from the user behavior database. This database stores the feature point sequences of the user's historical behaviors, and each sequence corresponds to a specific type of behavior. For example, the historical behavior feature point sequence obtained from the database is [(2,0.48), (5,0.15), (8,0.68)].

[0193] Construct the current behavior feature point sequence and the historical behavior feature point sequence into two sequences to be matched. In this embodiment, the two sequences to be matched are respectively:

[0194] Sequence A: [(3,0.5), (6,0.1), (9,0.7)];

[0195] Sequence B: [(2, 0.48), (5, 0.15), (8, 0.68)];

[0196] Construct a dynamic programming matrix DP. The rows and columns of the matrix correspond to the lengths of the current behavior feature point sequence and the historical behavior feature point sequence respectively. For the above example, the size of the constructed dynamic programming matrix DP is 4×4, where DP[0][0] is the initial position and DP[3][3] is the final position.

[0197] Initialize the matrix boundary conditions: Initialize the first row and the first column of the DP matrix to 0, that is, DP[0][j] = 0 (j = 0, 1, 2, 3) and DP[i][0] = 0 (i = 0, 1, 2, 3).

[0198] Iteratively calculate the values of the matrix elements: Starting from DP[1][1], calculate the values of each element in row-major order. For the position DP[i][j], compare the difference between the i-th feature point A[i] of sequence A and the j-th feature point B[j] of sequence B. Assume the preset difference threshold is 0.1. When |A[i].value - B[j].value| < 0.1, it is considered that the two feature points match, and the matrix element DP[i][j] at the corresponding position = DP[i - 1][j - 1] + 1; when the difference between the feature point values is greater than or equal to the preset difference threshold, take the maximum value of the adjacent positions, that is, DP[i][j] = max(DP[i - 1][j], DP[i][j - 1]).

[0199] For DP[1][1], compare A[1] = (3, 0.5) and B[1] = (2, 0.48), |0.5 - 0.48| = 0.02 < 0.1, so DP[1][1] = DP[0][0] + 1 = 1.

[0200] For DP[1][2], compare A[1] = (3, 0.5) and B[2] = (5, 0.15), |0.5 - 0.15| = 0.35 > 0.1, so DP[1][2] = max(DP[0][2], DP[1][1]) = 1.

[0201] According to the final value of the dynamic programming matrix DP[3][3] = 2, the length of the longest common subsequence is 2, indicating that there are 2 feature points that match in these two sequences.

[0202] Calculate the matching degree of the behavior sequences: Divide the length of the longest common subsequence by the smaller value of the total lengths of the sequences to obtain the matching degree of the behavior sequences. In this example, the lengths of both sequences are 3, so the matching degree of the behavior sequences = 2 / 3 ≈ 0.67, that is, a matching degree of 67%.

[0203] When the matching degree of the behavior sequence exceeds a preset threshold (such as 0.6), it is determined that the current behavior matches the historical behavior; otherwise, it is determined as a mismatch. In this example, the matching degree is 0.67, which exceeds the preset threshold of 0.6, so it is determined that the current behavior matches the historical behavior.

[0204] To improve the matching accuracy, the position information of the feature points can be considered simultaneously. When comparing whether two feature points match, in addition to comparing the numerical differences, the position difference can also be limited to not exceed a certain range, such as |A[i].position - B[j].position| < 2. This can ensure that the matching feature points are not only close in value but also close in relative position in the time series.

[0205] The behavior feature matching method of the present invention is applicable to various scenarios, such as motion recognition of smart watches, user behavior recognition of smart homes, and abnormal behavior detection of security systems. By accurately calculating the matching degree of the behavior sequence, accurate recognition and analysis of user behavior can be achieved, improving the user experience and system security.

[0206] In practical applications, the preset difference threshold and matching degree threshold can be adjusted according to the specific scenario to balance the accuracy and recall rate of recognition. For example, in a security system, the matching degree threshold can be reduced to improve the detection rate of abnormal behaviors; in daily motion recognition, the matching degree threshold can be increased to reduce the false recognition rate.

[0207] In the second aspect of the embodiments of the present invention, a smart IoT and metaverse virtual-real fusion recognition system based on deep learning is provided, including:

[0208] A first unit for obtaining virtual behavior data of a user in a metaverse virtual scenario collected by a terminal device and actual behavior data of the user in a real scenario collected by an IoT device;

[0209] A second unit for extracting features from the virtual behavior data and the actual behavior data through a pre-trained deep neural network model to obtain a virtual behavior feature vector and an actual behavior feature vector;

[0210] A third unit for cross-fusing the virtual behavior feature vector and the actual behavior feature vector based on an attention mechanism to generate a virtual-real fusion feature vector, where the attention mechanism realizes adaptive feature alignment of virtual and real behaviors by calculating the correlation weight between the virtual behavior feature vector and the actual behavior feature vector;

[0211] The fourth unit is used to construct a virtual-real behavior time-series relationship graph by using a graph neural network, model the time-series dependence relationship between nodes through a conditional random field, perform multi-hop information propagation on node features to obtain time-series context features, and integrate the time-series context features of multiple subgraphs through a cross-subgraph feature fusion layer to generate a fused time-series feature vector;

[0212] The fifth unit is used to identify the virtual-real behavior consistency of the user according to the fused time-series feature vector to obtain a virtual-real behavior consistency score; and verify and authorize the user's behavior in the metaverse virtual scene based on the virtual-real behavior consistency score.

[0213] In the third aspect of the embodiments of the present invention,

[0214] A kind of electronic device is provided, including:

[0215] A processor;

[0216] A memory for storing instructions executable by the processor;

[0217] Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.

[0218] In the fourth aspect of the embodiments of the present invention,

[0219] A computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.

[0220] The present invention can be a method, a device, a system and / or a computer program product. The computer program product may include a computer-readable storage medium, on which computer-readable program instructions for executing various aspects of the present invention are uploaded.

[0221] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: it is still possible to modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. An intelligent Internet of Things and metaverse virtual-real fusion recognition method based on deep learning, characterized in that Including: Obtaining virtual behavior data of a user in a metaverse virtual scene collected by a terminal device and actual behavior data of the user in the real scene collected by an Internet of Things device; Performing feature extraction on the virtual behavior data and the actual behavior data through a pre-trained deep neural network model to obtain a virtual behavior feature vector and an actual behavior feature vector; Based on an attention mechanism, cross-fusing the virtual behavior feature vector and the actual behavior feature vector to generate a virtual-real fusion feature vector, wherein the attention mechanism realizes adaptive feature alignment of virtual and real behaviors by calculating the correlation weight between the virtual behavior feature vector and the actual behavior feature vector; Using a graph neural network to construct a virtual-real behavior temporal relationship graph, modeling the temporal dependence relationship between nodes through a conditional random field, performing multi-hop information propagation on node features to obtain temporal context features, and integrating the temporal context features of multiple subgraphs through a cross-subgraph feature fusion layer to generate a fused temporal feature vector; Identifying the virtual-real behavior consistency of the user according to the fused temporal feature vector to obtain a virtual-real behavior consistency score; verifying and authorizing the behavior of the user in the metaverse virtual scene based on the virtual-real behavior consistency score; Calculating an attention weight matrix between the virtual behavior feature vector and the actual behavior feature vector, wherein for each feature element in the virtual behavior feature vector, calculating its correlation score with each feature element in the actual behavior feature vector through a dot product operation, and normalizing the correlation score to obtain the attention weight matrix; Based on the attention weight matrix, performing weighted aggregation on the virtual behavior feature vector and the actual behavior feature vector to obtain a context feature vector of the virtual behavior and a context feature vector of the actual behavior, wherein the context feature vector of the virtual behavior is obtained by multiplying the attention weight matrix with the matrix of the actual behavior feature vector, and the context feature vector of the actual behavior is obtained by multiplying the transpose of the attention weight matrix with the matrix of the virtual behavior feature vector; Performing feature splicing on the virtual behavior feature vector and the context feature vector of the virtual behavior, and performing feature splicing on the actual behavior feature vector and the context feature vector of the actual behavior to obtain an enhanced virtual behavior feature vector and an enhanced actual behavior feature vector; Fusing the enhanced virtual behavior feature vector and the enhanced actual behavior feature vector through a non-linear transformation to generate a virtual-real fusion feature vector; Using a graph neural network to construct a virtual-real behavior temporal relationship graph, and mapping the virtual-real fusion feature vector at each time step to the node features in the virtual-real behavior temporal relationship graph; Modeling the temporal dependence relationship between nodes through a conditional random field, and dynamically predicting the causal association strength between nodes based on the change trend of node features and the historical interaction pattern; Normalize the causal association strength to obtain the sampling probability between node pairs; set a sampling probability threshold, and filter out the node pairs higher than the sampling probability threshold as the key timing paths; centering on the key timing paths, sample multiple local subgraphs through the random walk algorithm; Perform the following operations on each of the local subgraphs: Construct a local graph neural network and set multiple graph convolutional layers; in each layer of graph convolution, aggregate the neighbor features of the nodes and update the node representations; retain the node features of each layer through skip connections to obtain the node timing context features containing multi-hop information; Integrate the information of multiple local subgraphs through a cross-subgraph feature fusion layer, specifically including: Calculate the attention weights of the corresponding node features in different subgraphs; perform weighted fusion on the node timing context features of multiple subgraphs based on the attention weights; obtain the final fused timing feature vector through a pooling operation.

2. The method according to claim 1, wherein Model the temporal dependence relationship between nodes through a conditional random field. Based on the change trend of node features and the historical interaction pattern, dynamically predict the causal association strength between nodes, including: Model the temporal dependence relationship between nodes in the virtual-real behavior temporal relationship graph through a conditional random field, and calculate the difference vector of adjacent node features as the change trend of node features; Construct the historical interaction pattern features of node pairs based on the interaction frequency of node pairs within the historical time window; Input the change trend of the node features and the historical interaction pattern features into the conditional random field to predict the causal association strength between node pairs.

3. The method according to claim 1, wherein Construct a local graph neural network and set multiple graph convolutional layers; in each layer of graph convolution, aggregate the neighbor features of the nodes and update the node representations; retain the node features of each layer through skip connections to obtain the node timing context features containing multi-hop information, including: Receive the input node feature vector and graph structure information, where the node feature vector represents the initial features of each node, and the graph structure information represents the connection relationship between nodes; Construct a multi-layer graph convolutional network structure, including L graph convolutional layers, each graph convolutional layer containing a learnable weight matrix and a bias vector, where L is an integer greater than 1; For the l-th graph convolutional layer in the L graph convolutional layers, perform the following feature propagation operations: Obtain the set of neighbor nodes of the node; normalize the features in the set of neighbor nodes; multiply the normalized neighbor nodes by the weight matrix of the l-th graph convolutional layer; Perform a non-linear transformation on the product result to obtain the output feature of the node at the l-th layer; Retain the intermediate feature representations of the nodes in each graph convolutional layer through a skip connection mechanism, specifically including: Concatenate the output features of the node in each layer of graph convolution; use a feature transformation layer to perform dimensionality reduction on the concatenated multi-layer features; Iteratively execute the feature propagation process of the graph convolutional layer L times, updating the node features each time to obtain the node timing context features containing multi-hop information.

4. The method according to claim 1, wherein Determine the behavior sequence matching degree through the dynamic programming algorithm according to the fused temporal feature vector, and obtain the virtual-real behavior consistency score in combination with the preset user behavior feature template; Verifying and authorizing the user's behavior in the metaverse virtual scene based on the virtual-real behavior consistency score includes: Extract the behavior feature point sequence from the fused temporal feature vector, and calculate the longest common subsequence of the current behavior feature point sequence and the historical behavior feature point sequence based on the dynamic programming algorithm to obtain the behavior sequence matching degree; Perform importance weighting on the user's historical behavior feature point sequence according to the behavior sequence matching degree, and construct an adaptive user behavior feature template, where the importance weight decays with time distance; Calculate the multi-scale correlation score between the current fused temporal feature vector and the adaptive user behavior feature template, and detect behavior mutation points based on a sliding time window to generate a behavior consistency score; When the behavior consistency score is higher than the preset score threshold, approve the user's operation request in the metaverse scenario; When the behavior consistency score is lower than the preset score threshold, trigger the identity re-authentication mechanism and temporarily freeze the user's operation permissions.

5. The method according to claim 4, wherein Extracting the behavior feature point sequence from the fused temporal feature vector, and calculating the longest common subsequence of the current behavior feature point sequence and the historical behavior feature point sequence based on the dynamic programming algorithm to obtain the behavior sequence matching degree includes: Perform extreme point detection on the temporal feature vector to extract the current behavior feature point sequence; Obtain the historical behavior feature point sequence from the user behavior database, and construct the current behavior feature point sequence and the historical behavior feature point sequence into two sequences to be matched; Construct a dynamic programming matrix, where the rows and columns of the dynamic programming matrix correspond to the lengths of the current behavior feature point sequence and the historical behavior feature point sequence respectively; Calculate the longest common subsequence of the two sequences based on the dynamic programming algorithm: Initialize the matrix boundary conditions; when the difference in feature point values is less than the preset difference threshold, add 1 to the matrix element at the corresponding position; when the difference in feature point values is greater than the preset difference threshold, take the maximum value of the adjacent positions; Iteratively update the matrix elements until all sequence positions are traversed; According to the final value of the dynamic programming matrix, calculate the length of the longest common subsequence, and divide it by the total length of the sequence to obtain the behavior sequence matching degree.

6. An intelligent Internet of Things and metaverse virtual-real fusion recognition system based on deep learning, for implementing the method described in any one of claims 1-5, characterized in that, Including: The first unit is used to obtain the virtual behavior data of the user in the metaverse virtual scene collected by the terminal device and the actual behavior data of the user in the real scene collected by the Internet of Things device; The second unit is used to extract features from the virtual behavior data and the actual behavior data through a pre-trained deep neural network model to obtain a virtual behavior feature vector and an actual behavior feature vector; The third unit is used to cross-fuse the virtual behavior feature vector and the actual behavior feature vector based on the attention mechanism to generate a virtual-real fusion feature vector, where the attention mechanism realizes adaptive feature alignment of virtual-real behaviors by calculating the correlation weight between the virtual behavior feature vector and the actual behavior feature vector; The fourth unit is used to construct a virtual-real behavior time-series relationship graph by using a graph neural network, model the time-series dependence relationship between nodes through a conditional random field, perform multi-hop information propagation on node features to obtain time-series context features, and integrate the time-series context features of multiple subgraphs through a cross-subgraph feature fusion layer to generate a fused time-series feature vector; The fifth unit is used to identify the virtual-real behavior consistency of the user according to the fused time-series feature vector to obtain a virtual-real behavior consistency score; and verify and authorize the user's behavior in the metaverse virtual scene based on the virtual-real behavior consistency score.

7. An electronic device, characterized in that, Comprising: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 5.

8. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Immersive space virtual-real interaction method and system based on meta universe

    CN119311124A

  • Multi-dimensional fusion meta universe and vertical AI model collaborative innovation platform

    CN119443116A