Multi-modal anomaly detection method and system based on neighborhood context
By combining local structural features and neighborhood semantic features in graph neural networks, a multimodal anomaly detection method is proposed to solve the problem of difficult identification of disguised fraudulent behavior, achieve efficient detection of abnormal behavior in complex graphs, and improve detection accuracy and robustness.
Patent Information
- Application Number
- CN202511022582.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-24
- Publication Date
- 2025-11-18
AI Technical Summary
Existing graph neural network-based methods are ineffective in identifying disguised fraudulent behavior. Traditional detection methods fail when faced with low-frequency, covert, and collaborative fraud, and single-modal detection strategies are unable to capture the inconsistencies of nodes under different modalities.
A multimodal anomaly detection method based on neighborhood context is adopted. By jointly modeling the local structural features and neighborhood semantic features of nodes, and using the alignment learning of graph embedding and text embedding, the consistency break phenomenon of nodes under different modalities is captured. The model is optimized by combining multi-relation attention mechanism for weighted aggregation and modal alignment training to improve detection accuracy.
It improves the accuracy and robustness of detecting disguised anomaly nodes, reduces feature engineering costs, and enhances the ability to identify anomalous behaviors in complex graphs.
Smart Images

Figure CN120973991A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of anomaly detection, and in particular to a multi-modal anomaly detection method and system based on neighborhood context. BACKGROUND
[0002] In real life, anomaly detection is one of the basic technologies to ensure the health of the platform ecosystem, and is widely used in e-commerce platforms, social networks, public opinion control and other key scenarios. Traditional fraudulent behavior usually shows abnormal operation, such as malicious accounts posting a large number of posts in a short time, concentrating on scoring, publishing false links, etc. This kind of behavior is easy to be identified by existing detection methods based on graph structure because of its obvious characteristics.
[0003] However, in recent years, fraudulent behavior has gradually evolved towards low frequency, concealment and collaboration, giving birth to a more difficult to prevent form - disguised fraud. Instead of creating strong conflict signals, these fraudsters actively mimic the behavior patterns and social paths of normal users by adjusting their node features (feature disguise) and carefully designing neighbor relationships (relationship disguise), significantly improving the concealment of fraudulent behavior. The key purpose is to prolong the survival time and evade detection, making traditional detection methods based on node feature differences or structural anomalies ineffective.
[0004] At the same time, existing graph neural network (GNN) methods also have obvious limitations in the face of disguised fraud. GNN relies on the aggregation of node original features and neighbor information, while disguised fraud weakens the model's discrimination ability through feature imitation and neighbor assimilation mechanism, and may even exacerbate the disguise effect through the training process.
[0005] Further observation shows that although disguised fraud nodes can be highly similar to normal nodes in structure, there are still significant differences in text content, semantic expression and other modalities. Therefore, a single modality detection strategy cannot fully identify disguised fraud, and an anomaly detection method that integrates multi-modal features and mines structural and semantic inconsistencies is needed to improve the accuracy and robustness of detection. SUMMARY
[0006] In order to improve the accuracy and robustness of disguised anomaly node detection, the present application proposes a multi-modal anomaly detection method based on neighborhood context. By jointly modeling the local structural features of nodes and the semantic features of neighbors, and modeling the context alignment degree between graph embedding and text embedding, the consistency breakdown phenomenon of nodes in different modalities is effectively captured, so as to identify abnormal behavior hidden in a large number of normal nodes.
[0007] To solve the above technical problems, the present application adopts the following technical solutions: A multi-modal anomaly detection method based on neighborhood context, comprising the following steps: Step 1: Collecting graph modality feature input and text modality feature input of each node in the multi-relation heterogeneous graph; Step 2: Modality encoding is performed on the graph modality feature input and the text modality feature input respectively, and the graph modality feature input and the text modality feature input are mapped to a low-dimensional embedding space of a unified dimension to obtain structure embedding and text embedding respectively; Step 3: Based on the obtained structure embedding or text embedding, for the relationship type between nodes, the neighbor set of nodes belonging to the same relationship type is obtained, and the neighbor node representation under different relationships is weighted and aggregated to obtain neighborhood graph embedding and neighborhood text embedding; Step 4: Modality alignment training based on structure embedding or text embedding, neighborhood graph embedding and neighborhood text embedding; Step 5: Based on the consistency degree of nodes and neighborhoods under different modalities, an abnormal score is obtained to obtain a node fraud probability; Step 6: Based on the node fraud probability, a weighted binary cross-entropy loss is used to minimize the error with the true label, and the model is optimized.
[0008] Further, the multi-relation heterogeneous graph is:
[0009] wherein, denotes a node set, denotes an edge set, each edge is composed of a pair of nodes and their relationship type Each relationship corresponds to an adjacency matrix ; each node contains structure modality features and text modality features.
[0010] Further, the graph modality feature input in step 1 includes user behavior statistics information, denoted as , wherein is the feature dimension of the graph modality; the text modality feature input is denoted as , which is the semantic description of the node, is the text modality feature dimension.
[0011] Further, the structure embedding in step 2 is:
[0012] The text embedding is:
[0013] wherein, is the original graph structure, denotes a graph encoder with neighbor aggregation mechanism, Encoding function for text.
[0014] Further, the step 3 comprises: Based on the obtained structure embedding or text embedding, for each relationship type between nodes, a neighbor set belonging to the same relationship type is obtained; Based on the relationship-aware attention aggregation function, the neighbor node representations under different relationships are weighted and aggregated; Each field of each node under different relationship types is traversed, and neighborhood graph embedding and neighborhood text embedding are obtained respectively.
[0015] Further, the relationship-aware attention aggregation function is:
[0016] Wherein, j is the neighbor node index, is the relationship of the neighbor set of the node i, wherein in turn represents the definition m of a plurality of different relationships in the multi-relational graph, is the input feature vector of the node j , including structure embedding or text embedding ; is a learnable query vector of the relationship r , used to measure the relative importance of the node j to the node i in the relationship; exp(•) is an exponential mapping; u is a normalized traversal index, is the input feature vector of the node u; The neighborhood graph embedding and the neighborhood text embedding are respectively:
[0017]
[0018] Wherein, , respectively as the neighborhood graph embedding and the neighborhood text embedding.
[0019] Further, the modal alignment training in the step 4 comprises node self-alignment, neighborhood context alignment and joint training; The node self-alignment comprises: minimizing the contrastive loss function between the node itself in the two modal representations :
[0020] Wherein n is the number of nodes, is the temperature coefficient, is the text embedding vector of node j; The neighborhood context alignment includes: (a) Graph embedding - loss function of text neighborhood alignment :
[0021] (b) Text embedding - loss function of graph neighborhood alignment :
[0022] (3) Joint training: the overall loss function of the cross-modal pre-training stage is as follows:
[0023] wherein, is the neighborhood consistency weight parameter. When , the framework degenerates into pure CLIP itself alignment.
[0024] Further, the node fraud probability in step 5 is:
[0025] is the Sigmod mapping function, which compresses any real number to the interval (0,1), is the fraud probability of node , is the node-level consistency difference score:
[0026] wherein, is the learnable relationship weight, is the bidirectional residual weight, is the graph→text residual weight, and the larger the value is, the more attention is paid to the degree of deviation of the graph modality from the text modality. is the text→graph residual weight; and respectively, for each relationship r , the graph-to-text residual and the text-to-graph residual under the node i are calculated.
[0027] Further, the weighted binary cross-entropy loss in step 6 is:
[0028] wherein, is the true label; is the focal factor, Weighted binary cross-entropy loss.
[0029] In another aspect, the present application provides a multi-modal anomaly detection system based on neighborhood context, comprising: A modal feature input collection module is configured to collect graph modal feature input and text modal feature input of each node in a multi-relation heterogeneous graph. A structure embedding and text embedding acquisition module is configured to respectively encode the graph modal feature input and the text modal feature input, map the graph modal feature input and the text modal feature input to a low-dimensional embedding space of a unified dimension, and acquire structure embedding and text embedding. A neighborhood graph embedding and neighborhood text embedding acquisition module is configured to acquire a neighbor set of nodes belonging to the same type of relationship based on the acquired structure embedding or text embedding, and aggregate the neighbor node representations of different relationships to acquire neighborhood graph embedding and neighborhood text embedding. A modal alignment training module is configured to perform modal alignment training based on the structure embedding or text embedding, the neighborhood graph embedding and the neighborhood text embedding. A node fraud probability acquisition module is configured to acquire a node fraud probability based on the consistency degree of the node and the neighborhood in different modalities. A model optimization module is configured to minimize the error of the node fraud probability and the true label based on the weighted binary cross-entropy loss, and optimize the model.
[0030] Compared with the prior art, the present application has the following beneficial effects: The present application proposes a camouflage type anomaly node detection method based on cross-modal consistency and relationship perception aggregation mechanism, which uses structure modal and text modal alignment learning to improve the model's ability to recognize abnormal behavior in complex graphs. By designing node alignment and neighborhood alignment tasks, combining multi-relation attention mechanism to weight and fuse neighbor information, the potential inconsistency between modalities is effectively captured, thereby achieving efficient anomaly detection, reducing feature engineering cost and improving detection accuracy. BRIEF DESCRIPTION OF DRAWINGS
[0031] In order to more clearly illustrate the technical solutions in the present application or prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor.
[0032] Figure 1 is a flowchart of an embodiment of the present application. DETAILED DESCRIPTION
[0033] In order to make the above objects, features and advantages of the present application more clear and easily understood, the specific embodiments of the present application are described in detail below with reference to the drawings. In the following description, a large number of specific details are set forth in order to provide a thorough understanding of the present application. However, the present application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the spirit of the present application, so the present application is not limited to the specific embodiments disclosed below. The present application provides a multi-modal anomaly detection method based on neighborhood context consistency detection, which is represented as , wherein represents a node set, represents an edge set, each edge is composed of a pair of nodes and their relationship type , each relationship corresponds to an adjacency matrix . Each node contains input information of structural modal and text modal, which are represented as structural embedding and text embedding , respectively, and the label is , wherein represents an abnormal node.
[0034] The method of the present application mainly includes the following steps: Step 1: Constructing graph feature and text feature input. First, for each node , collect its graph modal feature and text modal feature. The graph modal feature usually includes user behavior statistical information, such as the number of comments, rating distribution, active period, etc. These features can effectively describe the overall activity pattern of the node in the platform, which can be denoted as , wherein is the feature dimension of the graph modal; while the text modal feature is derived from the comment text published by the user, which is encoded into a fixed length vector as the semantic description of the node, is the feature dimension of the text modal. This step lays the foundation for subsequent multi-modal feature fusion.
[0035] Step 2: Modal encoding. For the two kinds of input features of the node, a special encoder is used to process them, and the original features are projected into embedding representation with unified dimension . The graph modal input is encoded by graph neural network (such as GCN or GraphSAGE): , wherein, represents the graph modal input feature of the node , and is the original graph structure, The step represents a graph encoder with neighbor aggregation mechanism. This step extracts the representation of a node in the structural neighborhood, which can effectively model the local structural properties of the node.
[0036] Similarly, the text modality embedding is processed by a text encoding network (such as MLP or Transformer encoder):
[0037] where, is the text modality input of node , is the text encoding function, which can use a pre-trained language model or a lightweight encoder. Through this process, the original high-dimensional heterogeneous features can be mapped to a low-dimensional embedding space of uniform dimension, facilitating consistent modeling between different modalities.
[0038] Step 3: Adaptive neighborhood fusion. After obtaining the modality representation of the node itself, the invention further models the neighbor set of the node for each relationship type . Let be the neighbor set of node i under relationship , where represents the m different relationships defined in the multi-relational graph in turn. In other words, collects all nodes directly connected to node i through relationship r.
[0039] This method uses a relationship-aware multi-head attention mechanism to weight and aggregate the neighbor node representations under different relationships, in order to fully exploit the behavior characteristics of the node under various contextual relationships. The relationship-aware attention aggregation function Attn( ) is defined as follows:
[0040] where j is the neighbor node index. is the input feature vector of node j. This formula does not limit the modality, and can take structural embedding or text embedding . is the learnable query vector of relationship r, which measures the relative importance of node j to node i under that relationship. exp(•) is the exponential mapping.
[0041] u is the normalized traversal index. When calculating the denominator, it is used to traverse each neighbor of node i under relationship r, and j belongs to . is the input feature vector of node u. From this formula, we get and , , These are used as neighborhood graph embeddings and neighborhood text embeddings, respectively. In this way, nodes not only consider their own local features, but also dynamically perceive the distribution of modal features in different neighborhoods, thereby improving the ability to detect complex camouflage behaviors.
[0042] Step 4: Modality Alignment Training Objective Construction. To capture the potential inconsistencies between structural and textual modalities, this invention designs a cross-modal contrastive learning task. Specifically, on the one hand, through a node self-alignment task, the representation of nodes is encouraged to remain as consistent as possible across the two modalities, improving the model's robustness in modeling normal nodes; on the other hand, neighborhood context alignment is introduced, forcing nodes to maintain cross-modal consistency even within the local context formed by their neighbors. This dual alignment mechanism can effectively reveal the feature deviations of disguised anomalous nodes across different modalities. Specifically, the following two sub-tasks are constructed: (1) Node self-alignment. This subtask aims to maximize the consistency of each node's representation in both the graph structure modality and the text semantic modality. Specifically, this method minimizes the contrast loss function of a node's own representation between the two modalities. :
[0043] Where n is the number of nodes. This is the temperature coefficient. The graph modality embedding vector is obtained for node i. Embed the neighborhood text of node i. is the text embedding vector of node j, used in the denominator as a contrast "negative sample".
[0044] (2) Neighborhood context alignment. In addition to aligning the modal representation of the node itself, this method further designs a neighbor alignment task to capture the semantic association between the node and its context.
[0045] (a) Loss function for graph embedding-text neighborhood alignment :
[0046] (b) Text embedding - loss function for graph neighborhood alignment :
[0047] (3) Joint Training Objective. Finally, the overall loss function for the cross-modal pre-training stage is defined as follows:
[0048] in This is the neighborhood consistency weight parameter. When... The time frame degenerates to pure CLIP itself alignment.
[0049] Step 5: Consistency-driven anomaly scoring. After completing the cross-modal pre-training, the present application performs anomaly scoring based on the consistency degree of nodes and neighborhoods in different modalities. By comparing the residual differences of nodes after aggregating neighborhood information in the structure modality and the text modality, the degree of deviation of the nodes in different modalities is quantified. To prevent bias introduced by a single relationship, a relationship weighting mechanism is introduced to adaptively adjust the importance of different relationships. At the same time, bidirectional residual adaptive balances the anomaly signals of the structure-to-text and text-to-structure two paths. Finally, the consistency difference score of the node is obtained as the basis for judging abnormal nodes.
[0050] First, for each relationship r, calculate the graph-to-text residual of node i and the text-to-graph residual :
[0051]
[0052] where , denote the neighborhood text embedding and graph embedding obtained by converging only in the relationship .
[0053] If the node and the neighborhood remain consistent after pre-training, then ; once there is structural or semantic camouflage, the residual will increase significantly. Next, to avoid the dominance of single relationship noise, a learnable query vector is used to give the residual weight:
[0054] where denotes the normalized attention coefficient of node i in relationship r. At the same time, the bidirectional coefficient in the Softmax form is used to adaptively balance . Therefore, the node-level consistency difference score is defined as
[0055] where is the learnable relationship weight, is the bidirectional residual weight. is the graph→text residual weight, and the larger the value, the more attention is paid to the deviation of the graph modality from the text modality. is the text→graph residual weight, and the larger the value, the more attention is paid to the deviation of the text modality from the graph modality. The node fraud probability is directly mapped by Sigmoid:
[0056] for the Sigmod mapping function, any real number is compressed to the interval (0,1), for the fraud probability of node . for the node-level consistency differential score.
[0057] Step 6: Optimize the objective function. In real-world graph anomaly detection scenarios, fraudulent nodes often account for only a small fraction (e.g. ), and there is a serious class imbalance problem. In this setting, the model is more likely to be dominated by mainstream (normal) nodes, leading to a decline in the ability to identify disguised anomalies. In addition, the disguising behavior of some nodes may be ambiguous, and unstable or incorrect pseudo-label predictions may occur early in training, misleading the optimization process.
[0058] Therefore, the present application introduces a Focal weighted binary cross-entropy loss in the optimization phase to strengthen the model's attention to difficult-to-classify nodes and reduce the interference of pseudo-label noise on the training process. By giving higher loss weights to nodes that are difficult to align or uncertain, the model's ability to identify disguised abnormal nodes in an extremely imbalanced data environment is improved, significantly improving overall detection performance:
[0059] wherein is the true label; is the focal factor, used to enhance attention to difficult-to-align samples.
[0060] Embodiment 1 The embodiment of the present application is based on the Yelp restaurant review dataset to build a multi-modal graph structure model for disguised abnormal node detection. The Yelp dataset contains a total of 45,954 review nodes, of which 14.5% are fraudulent nodes. According to the edge type, define multiple relationships, including "R-U-R (review-user-review)", "R-T-R (review-time-review)", and "R-S-R (review-semantic-review)", a total of 3,846,979 edges are generated. According to the ratio of 4:2:4, the training set, validation set and test set are divided, ensuring that the proportion of abnormal nodes in each subset remains consistent, and are used for model training, tuning and final evaluation, respectively.
[0061] As shown in the specific embodiment Figure 1 , the present application comprises the following steps: Step 1: Node input feature processing. For each node, its structural modal features and text modal features are extracted respectively. The structural features include the node's behavior statistics such as the number of comments, rating behavior, time activity, and other behavior features, with a dimension of 32; the text features are derived from the user's comments on the restaurant content, which are normalized by TF-IDF to generate a 128-dimensional vector representation.
[0062] Among them, the 32-dimensional graph modal vector includes the following features: Activity (2 dimensions): total number of comments published by the user, character length of the user's nickname Rating distribution (16 dimensions): absolute number and proportion of 1-5 star comments, proportion of good comments, proportion of bad comments, minimum / median / maximum / average star rating, and rating entropy Voting feedback (10 dimensions): total number of useful votes / useless votes, proportion, and extreme value statistics Time behavior (3 dimensions): number of days between the first and last comments, entropy value after counting the number of comments by year, whether it is a same-day attack (1 if the first and last comments fall on the same day, otherwise 0) Text overview (1 dimension): average comment length Step 2: Modal encoder construction. The structural modal features and text modal features of the node are input into independently designed encoding modules. The structural features are processed through 2-layer GCN to model the local connections and context characteristics of the node in the graph structure, with a hidden layer dimension of 128; the text features are processed through the BERT pre-trained language model (bert-base-uncased), with a maximum length of 128 tokens, to obtain text embedding with a dimension of 128.
[0063] Step 3: Neighbor feature aggregation based on relationship type. A relationship-aware multi-head attention mechanism is used for neighbor feature weighting aggregation, with a total of 4 attention heads for aggregation. After aggregation, the node's graph modal neighborhood representation and text modal neighborhood representation are obtained.
[0064] Step 4: Modal alignment task training. In the training phase, a cross-modal alignment target is set, including the consistency alignment between the node's own representation and the neighborhood aggregation representation. The contrast learning temperature coefficient is set to 0.1; is taken . The model training batch size is set to 64; the total number of training rounds is set to 300.
[0065] Step 5: Abnormal node scoring and reasoning. In the reasoning stage, for each node, the consistency residual between its structure modal and neighborhood aggregation representation under the text modal is calculated to obtain the consistency difference score. Finally, the node anomaly probability is output through the Sigmoid activation function, and the determination threshold is set to 0.5. Nodes higher than 0.5 are determined as pseudo abnormal nodes.
[0066] Step 6: In order to deal with the problem of extremely unbalanced positive and negative samples, weighted binary cross entropy loss is introduced in the training process, and Focal Loss is combined to further enhance the learning ability of difficult samples. The focal factor of Focal Loss is set to 2.
[0067] Embodiment 2 The embodiment provides a multi-modal anomaly detection system based on neighborhood context, comprising: A modal feature input collection module is configured to collect graph modal feature input and text modal feature input of each node in a multi-relation heterogeneous graph. A structure embedding and text embedding acquisition module is configured to respectively encode the graph modal feature input and the text modal feature input, map the graph modal feature input and the text modal feature input to a low-dimensional embedding space of a unified dimension, and acquire structure embedding and text embedding. A neighborhood graph embedding and neighborhood text embedding acquisition module is configured to acquire a neighbor set of nodes belonging to the same type of relationship based on the acquired structure embedding or text embedding, and acquire neighborhood graph embedding and neighborhood text embedding by weighted aggregation of neighbor node representations under different relationships. A modal alignment training module is configured to perform modal alignment training based on the structure embedding or the text embedding, the neighborhood graph embedding and the neighborhood text embedding. A node fraud probability acquisition module is configured to acquire a node fraud probability based on the consistency degree of the node and the neighborhood under different modalities. A model optimization module is configured to minimize the error between the node fraud probability and the true label by using a weighted binary cross entropy loss, and perform model optimization.
[0068] The above is only the preferred specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any changes or replacements within the technical range disclosed in the present application can be easily thought of by those skilled in the art, and should be covered within the protection scope of the present application.
[0069] It should be understood that parts not elaborated in the specification are all prior art.
[0070] It should be understood that the above description is merely a detailed explanation of the preferred embodiments and is not intended to limit the patent protection scope of the present application. Any modification or alternation made by those skilled in the art without departing from the scope of the present application shall fall within the patent protection scope of the present application. The patent protection scope of the present application shall be subject to the appended claims.
Claims
1. A multimodal anomaly detection method based on neighborhood context, characterized in that, Includes the following steps: Step 1: Collect graph modal feature input and text modal feature input of each node in the multi-relation heterogeneous graph; Step 2: Perform modal encoding on the graph modal feature input and the text modal feature input respectively, and map the graph modal feature input and the text modal feature input to a low-dimensional embedding space of the same dimension to obtain the structural embedding and the text embedding respectively; Step 3: Based on the obtained structural embedding or text embedding, for the relationship type between each node, obtain the neighbor set of nodes belonging to the same relationship type, and perform weighted aggregation on the neighbor node representations under different relationships to obtain the neighborhood graph embedding and neighborhood text embedding; Step 4: Perform modality alignment training based on structural embedding or text embedding, neighborhood graph embedding, and neighborhood text embedding; Step 5: Obtain the probability of node fraud by scoring anomalies based on the consistency between the node and its neighborhood under different modalities; Step 6: Based on the node fraud probability, the model is optimized by minimizing the error between the node and the true label using a weighted binary cross-entropy loss.
2. The multimodal anomaly detection method based on neighborhood context according to claim 1, characterized in that, The multi-relationship heterogeneous graph is as follows: in, Represents a set of nodes. Represents a set of edges, where each edge consists of a pair of nodes and their relation type. Composition, each relationship Corresponding to an adjacency matrix Each node It includes structural modal features and textual modal features.
3. The multimodal anomaly detection method based on neighborhood context according to claim 2, characterized in that, In step 1, the graph modal feature input includes user behavior statistics, denoted as... ,in The feature dimension of the graph modality; the feature input of the text modality is represented as... As a semantic description of a node, This represents the dimension of text modality features.
4. The multimodal anomaly detection method based on neighborhood context according to claim 3, characterized in that, The structural embedding in step 2 is as follows: Text embedding is: in, For the original graph structure, This represents a graph encoder with a neighbor aggregation mechanism. This is a text encoding function.
5. The multimodal anomaly detection method based on neighborhood context according to claim 4, characterized in that, Step 3 includes: Based on the obtained structural or textual embeddings, and considering the relationship type between nodes, obtain the set of neighbors of nodes belonging to the same relationship type; Based on the relation-aware attention aggregation function, the neighbor node representations under different relations are weighted and aggregated. Iterate through each node in each domain under different relation types, and obtain the neighborhood graph embedding and neighborhood text embedding respectively.
6. The multimodal anomaly detection method based on neighborhood context according to claim 5, characterized in that, The relation-aware attention aggregation function is: in, j Indexing neighboring nodes, For relationship The set of neighbors of the next node i, where These represent the definitions in a multi-relationship graph, respectively. m Different kinds of relationships, For nodes j The input feature vector, including structural embeddings Or text embedding ; It is a relationship r Learnable query vectors, used to measure nodes j For nodes i The relative importance of this relationship; exp(•) is the exponential mapping; u It is a normalized traversal index. It is the input feature vector of node u; Neighborhood graph embedding and neighborhood text embedding are respectively: in, , They are used as neighborhood graph embeddings and neighborhood text embeddings, respectively.
7. The multimodal anomaly detection method based on neighborhood context according to claim 6, characterized in that, The modality alignment training in step 4 includes node self-alignment, neighborhood context alignment, and joint training. Node self-alignment includes minimizing the contrast loss function of a node itself between two modal representations. : in n For the number of nodes, For temperature coefficient, Let be the text embedding vector of node j; Neighborhood context alignment includes: (a) Loss function for graph embedding-text neighborhood alignment : (b) Text embedding - loss function for graph neighborhood alignment : (3) Joint training: The overall loss function for the cross-modal pre-training stage is as follows: in, The neighborhood consistency weight parameter; when The time frame degenerates into pure CLIP self-alignment.
8. The multimodal anomaly detection method based on neighborhood context according to claim 7, characterized in that, The node fraud probability in step 5 is: The Sigmoid mapping function compresses any real number into the interval (0,1). For nodes The probability of fraud, Node-level consistency differential score: in, For learnable relation weights, For two-way residual weights, The graph-to-text residual weights are determined by the degree to which the graph modality deviates from the text modality. A larger value indicates greater attention is paid to the degree to which the graph modality deviates from the text modality. Text → Graph residual weights; and For each relation r compute nodes i The graph-to-text residual and the text-to-graph residual are shown below.
9. The multimodal anomaly detection method based on neighborhood context according to claim 8, characterized in that, The weighted binary cross-entropy loss in step 6 is: in, This is a real label; As the focal factor, This is the weighted binary cross-entropy loss.
10. A multimodal anomaly detection system based on neighborhood context, characterized in that, include: Modal feature input acquisition module: It is used to acquire graph modal feature input and text modal feature input of each node in a multi-relation heterogeneous graph; The structural embedding and text embedding acquisition module is used to perform modal encoding on the graph modal feature input and the text modal feature input respectively, map the graph modal feature input and the text modal feature input to a low-dimensional embedding space of the same dimension, and acquire the structural embedding and the text embedding respectively. The neighborhood graph embedding and neighborhood text embedding acquisition module is used to obtain the neighborhood graph embedding and neighborhood text embedding based on the acquired structural embedding or text embedding and the relationship type between each node. It also performs weighted aggregation on the neighbor node representations under different relationships to obtain the neighborhood graph embedding and neighborhood text embedding. Modality alignment training module: It is used for modality alignment training based on structural embedding or text embedding, neighborhood graph embedding and neighborhood text embedding; Node fraud probability acquisition module: It is used to obtain the node fraud probability by performing anomaly scoring based on the consistency degree between the node and its neighbors under different modalities; Model optimization module: It is used to optimize the model by minimizing the error between the node and the true label using a weighted binary cross-entropy loss based on the node fraud probability. The neighborhood context-based multimodal anomaly detection system is used to perform the steps in the neighborhood context-based multimodal anomaly detection method according to any one of claims 1-9.
Citation Information
Cited By
Internet of Things equipment anomaly detection method, anomaly detection model training method and system
CN121486241A