Log anomaly detection method, medium and program product based on bidirectional knowledge distillation
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-07
- Publication Date
- 2026-08-07
AI Technical Summary
[0002]相关的日志异常检测方法需要依赖专家规则或统计模型,难以应对复杂攻击和高维数据
[0007] A fourth aspect of this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described method.
Smart Images

Figure CN122533864A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of network security, log analysis and anomaly detection technology, and specifically to a log anomaly detection method, medium and program product based on bidirectional knowledge distillation. Background Technology
[0002] Existing log anomaly detection methods rely on expert rules or statistical models, making them ill-suited for complex attacks and high-dimensional data. While deep learning methods are widely used, they suffer from limitations such as single-view modeling, sensitivity to noise, and insufficient generalization ability. Therefore, there is an urgent need for a detection method that can comprehensively capture the spatiotemporal characteristics of system behavior and possesses strong robustness and high generalization ability. Summary of the Invention
[0003] In view of the above problems, this application provides a log anomaly detection method, medium and program product based on bidirectional knowledge distillation, and also provides a log anomaly detection device and equipment based on bidirectional knowledge distillation.
[0004] According to the first aspect of this application, a log anomaly detection method based on bidirectional knowledge distillation is provided, comprising: acquiring semantic graph data and structural graph data related to the log to be detected, wherein nodes in the semantic graph data and structural graph data represent at least one of the logical entities or physical entities of the system, edge weights in the semantic graph structural data represent the similarity between the node behavior sequences corresponding to multiple nodes, edge weights in the structural graph structural data represent the degree of interaction between nodes connected by edges, and node behavior sequences represent the event type sequences in which nodes participate; processing the semantic graph data and structural graph data using a log anomaly detection model to obtain log anomaly detection results related to the log to be detected; the log anomaly detection model is obtained by training a deep learning model based on bidirectional distillation loss values, wherein the bidirectional distillation loss values are obtained based on the bidirectional distillation loss function, according to the sample semantic category distribution and sample structural category distribution of the sample nodes in the sample semantic graph data and sample structural graph data, and the bidirectional distillation loss function value is used to characterize the bidirectional transfer state between the sample structural graph data and the sample semantic graph data based on the sample semantic category distribution and sample structural category distribution, wherein the bidirectional transfer state reflects the degree of mutual transfer of cross-modal knowledge, the degree of heterogeneous feature fusion, and the degree of collaborative optimization under graph structural constraints.
[0005] The second aspect of this application provides a log anomaly detection device based on bidirectional knowledge distillation, comprising: an acquisition module for acquiring semantic graph data and structural graph data related to the log to be detected, wherein nodes in the semantic graph data and structural graph data represent at least one of the logical entities or physical entities of the system, edge weights in the semantic graph structural data represent the similarity between the node behavior sequences corresponding to multiple nodes, edge weights in the structural graph structural data represent the degree of interaction between nodes connected by edges, and node behavior sequences represent the event type sequences in which nodes participate; and a log anomaly detection result module for processing the semantic graph data and structural graph data using a log anomaly detection model to obtain log anomaly detection results related to the log to be detected; the log anomaly detection model is obtained by training a deep learning model based on bidirectional distillation loss values, wherein the bidirectional distillation loss values are obtained based on the bidirectional distillation loss function, according to the sample semantic category distribution and sample structural category distribution of the sample nodes in the sample semantic graph data and sample structural graph data, and the bidirectional distillation loss function value is used to characterize the bidirectional transfer state between the sample structural graph data and the sample semantic graph data based on the sample semantic category distribution and sample structural category distribution, wherein the bidirectional transfer state reflects the degree of mutual transfer of cross-modal knowledge, the degree of heterogeneous feature fusion, and the degree of collaborative optimization under graph structural constraints.
[0006] A third aspect of this application provides an electronic device comprising: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the method described above.
[0007] A fourth aspect of this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described method.
[0008] The fifth aspect of this application also provides a computer program product, including a computer program or instructions that, when executed by a processor, implement the steps of the above-described method.
[0009] According to embodiments of this application, semantic graph data and structural graph data related to the log to be detected are obtained. Nodes in the semantic graph data and structural graph data can represent at least one of the system's logical entities or physical entities. Edge weights in the semantic graph data can represent the similarity between the behavior sequences of multiple nodes, while edge weights in the structural graph data represent the degree of interaction between nodes connected by the edges. Based on a dual-graph modeling framework using both structural graph data and semantic graph data, system behavior can be comprehensively characterized from two complementary dimensions: topological structure and semantic association. By processing the semantic graph data and structural graph data using the log anomaly detection model, log anomaly detection results related to the log to be detected can be obtained. This overcomes the limitations of single-view modeling in related methods, enabling a more comprehensive capture of the spatiotemporal characteristics and semantic connotations of system behavior, and improving the detection capability against complex and covert attacks. During the training of the log anomaly detection model, based on the bidirectional distillation loss function, the bidirectional distillation loss value is obtained according to the sample semantic category distribution and sample structural category distribution of each sample node in the sample semantic graph data and sample structural graph data. The bidirectional distillation loss function value is used to characterize the bidirectional transfer state between the sample structural graph data and the sample semantic graph data based on the sample semantic category distribution and sample structural category distribution. Through symmetrical knowledge transfer, the structural branch and semantic branch are optimized collaboratively during the training process, dynamically evaluating the sample prediction quality, prioritizing the learning of reliable knowledge, suppressing the influence of noise, and significantly improving the robustness and stability of the model. Attached Figure Description
[0010] The above-mentioned contents, other objects, features and advantages of this application will become clearer from the following description of embodiments of this application with reference to the accompanying drawings.
[0011] Figure 1 The diagram illustrates an application scenario of the log anomaly detection method and apparatus based on bidirectional knowledge distillation according to an embodiment of this application.
[0012] Figure 2 A flowchart of a log anomaly detection method based on bidirectional knowledge distillation according to an embodiment of this application is shown.
[0013] Figure 3 A schematic diagram of semantic graph data and structural graph data according to an embodiment of this application is shown.
[0014] Figure 4 A schematic diagram of a reliability estimation mechanism according to an embodiment of this application is shown.
[0015] Figure 5 A schematic diagram of bidirectional distillation according to an embodiment of this application is shown.
[0016] Figure 6 An architecture diagram of a log anomaly detection method based on bidirectional knowledge distillation according to an embodiment of this application is shown.
[0017] Figure 7 A structural block diagram of a log anomaly detection device based on bidirectional knowledge distillation according to an embodiment of this application is shown.
[0018] Figure 8 A block diagram of an electronic device suitable for implementing a log anomaly detection method based on bidirectional knowledge distillation, according to an embodiment of this application, is shown. Detailed Implementation
[0019] The embodiments of this application will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of this application. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of this application for ease of explanation. However, it will be apparent that one or more embodiments may be implemented without these specific details. Furthermore, descriptions of well-known structures and technologies are omitted in the following description to avoid unnecessarily obscuring the concepts of this application.
[0020] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0021] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.
[0022] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).
[0023] With the increasing complexity of information systems and the diversification of cyberattacks, related log anomaly detection methods face numerous challenges. Rule-based methods require expert knowledge to build rule bases, making it difficult to cope with new attacks and unknown threats; statistical methods are prone to high false positive rates when dealing with high-dimensional sparse data and complex behavioral patterns.
[0024] Deep learning technology has been widely used in log anomaly detection, but related methods still have significant shortcomings. They typically model log data from a single perspective, making it difficult to comprehensively capture the complex spatiotemporal characteristics of system behavior; most methods use simple feature concatenation for multi-feature fusion, lacking the ability to model deep correlations between different features; in complex system environments, the boundary between normal and abnormal logs is blurred, and related methods are quite sensitive to noisy data; furthermore, in cross-system and cross-scenario log anomaly detection, generalization performance needs improvement. To address these issues, a log anomaly detection method that can comprehensively model system behavior and possesses strong robustness and high generalization ability is needed.
[0025] In view of this, this application provides a log anomaly detection method, medium, and program product based on bidirectional knowledge distillation. The method includes: acquiring semantic graph data and structural graph data related to the log to be detected; in the semantic graph data and structural graph data, nodes represent at least one of the logical entities or physical entities of the system; in the semantic graph structural data, edge weights represent the similarity between the node behavior sequences corresponding to multiple nodes; in the structural graph structural data, edge weights represent the degree of interaction between nodes connected by edges; and node behavior sequences represent the sequence of event types in which nodes participate. The method further involves processing the semantic graph data and structural graph data using a log anomaly detection model to obtain log anomaly detection results related to the log to be detected. The log anomaly detection model is obtained by training a deep learning model based on bidirectional distillation loss values. The bidirectional distillation loss values are obtained based on the bidirectional distillation loss function, according to the sample semantic category distribution and sample structural category distribution of the sample nodes in the sample semantic graph data and sample structural graph data. The bidirectional distillation loss function value is used to characterize the bidirectional transfer state between the sample structural graph data and sample semantic graph data based on the sample semantic category distribution and sample structural category distribution. The bidirectional transfer state reflects the degree of cross-modal knowledge transfer, the degree of heterogeneous feature fusion, and the degree of collaborative optimization under graph structural constraints.
[0026] In the technical solution of this application, the user information (including but not limited to user personal information, user image information, user device information, such as location information) and data (including but not limited to data used for analysis, stored data, and displayed data) involved are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of related data all comply with relevant laws, regulations, and standards, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entry points for users to choose to authorize or refuse.
[0027] In scenarios involving automated decision-making using personal information, the methods, devices, and systems provided in this application all offer users corresponding entry points for choosing to agree to or reject the automated decision-making results. If the user chooses to reject, the process proceeds to the expert decision-making stage. Here, "automated decision-making" refers to the activity of automatically analyzing and evaluating an individual's behavioral habits, interests, or economic, health, and credit status through computer programs, and then making a decision. Here, "expert decision-making" refers to the activity of making decisions by personnel who specialize in a particular field, possess specialized experience, knowledge, and skills, and have reached a certain level of professional expertise.
[0028] Figure 1 The diagram illustrates an application scenario of the log anomaly detection method and apparatus based on bidirectional knowledge distillation according to an embodiment of this application.
[0029] like Figure 1 As shown, the application scenario according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 serves as a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.
[0030] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 via the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social media platform software, etc. (for example only).
[0031] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.
[0032] Server 105 can be a server that provides various services, such as a backend management server that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (this is just an example). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.
[0033] It should be noted that the log anomaly detection method provided in this application embodiment can generally be executed by server 105. Correspondingly, the log anomaly detection device provided in this application embodiment can generally be located in server 105. The log anomaly detection method provided in this application embodiment can also be executed by a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105. Correspondingly, the log anomaly detection device provided in this application embodiment can also be located in a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105.
[0034] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0035] Figure 2 A flowchart of a log anomaly detection method based on bidirectional knowledge distillation according to an embodiment of this application is shown.
[0036] like Figure 2 As shown, the log anomaly detection method based on bidirectional knowledge distillation in this embodiment includes operations S210 to S220, and this log anomaly detection method can be executed by an electronic device.
[0037] In operation S210, semantic graph data and structural graph data related to the log to be detected are obtained.
[0038] In operation S220, the log anomaly detection model is used to process semantic graph data and structural graph data to obtain log anomaly detection results related to the log to be detected.
[0039] The logs to be tested can include the raw log data of the system under test, including system call logs, network logs, file access logs, etc. Each log entry can contain fields such as timestamp, entity identifier, operation type, and attribute information. A log parser can extract key entities (such as processes, files, and network connections) and their interactions.
[0040] In semantic graph data and structural graph data, nodes represent at least one of the following: system logical entities or system physical entities. System logical entities may include networks, processes, or files. System physical entities may include hosts, routers, or switches, etc. This application does not limit the nodes in semantic graph data and structural graph data.
[0041] It can obtain semantic graph data and structure graph data related to the log to be detected. The nodes in the semantic graph data and the structure graph data are the same; the difference lies in how the edges are constructed.
[0042] Edges in semantic graph data can be constructed based on the semantic similarity of entity attributes (such as file path similarity and process name similarity), and can also consider access co-occurrence relationships (such as two files frequently accessed by the same process). Node features in semantic graph data are encoded using a pre-trained encoding model to encode attribute text. Node features in a semantic graph can be obtained by encoding node attributes and type information.
[0043] In semantic graph structured data, edge weights represent the similarity between the behavior sequences of multiple nodes, and node behavior sequences can represent the sequence of event types in which a node participates. Each node maintains a behavior path table, which records the sequence of event types in which the node participates, containing behavioral pattern information of the entity. For example, a node behavior sequence might include node A visiting node B, or node A creating node C.
[0044] In structural graph data, nodes can represent system entities, and edges represent interaction events between entities (such as process file creation or network connection). Edges of the same type can be merged, and edge weights are calculated based on interaction frequency. Node features in structural graph data can be derived by hashing entity attributes. The type, attributes, and behavior sequence of each node can be hashed and then normalized to obtain the node features.
[0045] Structure diagram data It can be constructed based on entity interaction relationships, and can be formally represented as: , where the set of nodes Represents system entities, edge set Represents interaction events between entities, a set of relation types Different event types are represented (such as "process creates file", "process accesses network", etc.). For each log event, the source entity and the target entity are the two ends of an edge, and the event type is used as the relation label for the edge.
[0046] In the structural graph data, the edge weight represents the degree of interaction between the nodes connected by the edge. Specifically, if there are multiple interactions of the same type between the same pair of nodes, they are merged into one edge. The edge weight is calculated by the interaction frequency, as shown in formula (1).
[0047] (1);
[0048] in, Represents a node and Edge weights between them For nodes and Types of relationships Number of interactions Represents a node and Types of relationships The maximum number of interactions, where h and w represent any two distinct nodes in the structure graph data, and p represents the interaction type between nodes in the structure graph data.
[0049] The maximum number of interactions can be determined based on the following steps: First, based on the structure graph data, determine the number of interactions of different interaction types between each group of nodes. Then, determine the maximum number of interactions from the multiple interaction counts. For example, if the number of interactions of the same type between node A and node B is 50, the number of interactions of the same type between node A and node C is 120, and the number of interactions of the same type between node B and node D is 30, then the maximum number of interactions can be determined to be 120.
[0050] This application can acquire log data from the system to be detected, parse and construct sample structure graph data and sample semantic graph data; input the sample structure graph data and sample semantic graph data into the structure branch and semantic branch respectively to obtain the corresponding sample semantic category distribution and sample structure category distribution. Based on the bidirectional distillation loss function, the bidirectional distillation loss value is obtained according to the sample semantic category distribution and sample structure category distribution of each sample node in the sample semantic graph data and sample structure graph data. Based on the bidirectional distillation loss value, a deep learning model can be trained to obtain a log anomaly detection model.
[0051] The bidirectional distillation loss function value is used to characterize the bidirectional transfer state between sample structure graph data and sample semantic graph data based on the sample semantic category distribution and sample structure category distribution. The bidirectional transfer state reflects the degree of mutual transfer of cross-modal knowledge, the degree of heterogeneous feature fusion, and the degree of collaborative optimization under graph structure constraints.
[0052] According to embodiments of this application, semantic graph data and structural graph data related to the log to be detected are obtained. Nodes in the semantic graph data and structural graph data can represent at least one of the system's logical entities or physical entities. Edge weights in the semantic graph data can represent the similarity between the behavior sequences of multiple nodes, while edge weights in the structural graph data represent the degree of interaction between nodes connected by the edges. Based on a dual-graph modeling framework using both structural graph data and semantic graph data, system behavior can be comprehensively characterized from two complementary dimensions: topological structure and semantic association. By processing the semantic graph data and structural graph data using the log anomaly detection model, log anomaly detection results related to the log to be detected can be obtained. This overcomes the limitations of single-view modeling in related methods, enabling a more comprehensive capture of the spatiotemporal characteristics and semantic connotations of system behavior, and improving the detection capability against complex and covert attacks. During the training of the log anomaly detection model, based on the bidirectional distillation loss function, the bidirectional distillation loss value is obtained according to the sample semantic category distribution and sample structural category distribution of each sample node in the sample semantic graph data and sample structural graph data. The bidirectional distillation loss function value is used to characterize the bidirectional transfer state between the sample structural graph data and the sample semantic graph data based on the sample semantic category distribution and sample structural category distribution. Through symmetrical knowledge transfer, the structural branch and semantic branch are optimized collaboratively during the training process, dynamically evaluating the sample prediction quality, prioritizing the learning of reliable knowledge, suppressing the influence of noise, and significantly improving the robustness and stability of the model.
[0053] According to embodiments of this application, the edges of semantic graph data include attribute similarity edges and access co-occurrence edges. The edge weight of attribute similarity edges indicates that the similarity between the node behavior sequences corresponding to multiple nodes is greater than or equal to a predetermined similarity. The edge weight of access co-occurrence edges indicates that the number of times the same node accesses the same node is greater than or equal to a predetermined number of times.
[0054] Semantic graph data Edges in the network can be constructed based on attribute similarity and visit co-occurrence relationships, where the node set Keep it consistent with the structure diagram data.
[0055] Similar edge attributes represent pairs of nodes of the same type. The similarity of node behavior sequences is calculated as shown in formula (2).
[0056] (2);
[0057] in, Represents a node and Similarity of node behavior sequences between them This refers to node i mentioned above. This refers to node j mentioned above. and They are nodes and The sequence of node behaviors, when If the similarity is greater than or equal to a predetermined value, attribute-similar edges can be created between node pairs, with the edge weight being the similarity value. Represents the union, Indicates intersection.
[0058] Visiting co-occurrence edges represent pairs of target nodes visited by the same node. For example, node A visits both node B and node C. For any pair of nodes... If the same node is accessed simultaneously and If the number of co-occurrences is greater than or equal to a predetermined number, a co-occurrence access edge can be established, with the edge weight... As shown in formula (3).
[0059] (3);
[0060] in, This indicates that the same node was accessed simultaneously. and The number of times they co-occur.
[0061] According to an embodiment of this application, the edge weight of a hybrid edge is determined based on the edge weight of an edge with similar attributes and the edge weight of an edge that is co-occurring in access. A hybrid edge is an edge that is both an edge with similar attributes and an edge that is co-occurring in access.
[0062] When node pair When both attribute similarity and access co-occurrence conditions are met, the two edge types can be merged into a hybrid edge, with the edge weight... As shown in formula (4).
[0063] (4);
[0064] in, This represents the weight of edges with similar attributes. This indicates the weight of the accessed co-occurring edge.
[0065] According to embodiments of this application, a comprehensive characterization of system entity behavior features and interaction relationships is achieved through dual-perspective modeling using semantic graph data and structural graph data. Semantic graph data integrates attribute similarity and access co-occurrence to accurately capture behavioral pattern associations; structural graph data quantifies entity interaction intensity and reconstructs the system topology. This dual-graph collaboration enables complementary heterogeneous features, overcoming the limitations of a single perspective and improving the comprehensiveness and accuracy of complex attack detection.
[0066] Figure 3 A schematic diagram of semantic graph data and structural graph data according to an embodiment of this application is shown.
[0067] like Figure 3 As shown, the log parser can extract key entities from the logs, such as P1, P2, F1, F2, N1, and N2 in the figure. The nodes in the semantic graph data and the structural graph data are the same.
[0068] In structural graph data, edges are formed when there are multiple interactions of the same type between the same pair of nodes. The edge weight is calculated based on the interaction frequency. Figure 3 The diagram shows that there are edge relationships between P1 and N1, P1 and F1, P2 and F2, N2 and F2, and P1 and P2, respectively.
[0069] Edges in semantic graph data can include attribute-similar edges (i.e., similar edges) and visit co-occurrence edges (i.e., co-occurrence edges). The edge weight of attribute-similar edges indicates that the similarity between the behavior sequences of multiple nodes is greater than or equal to a predetermined similarity. The edge weight of visit co-occurrence edges indicates that the number of times the same node is visited is greater than or equal to a predetermined number. A hybrid edge is an edge that is both attribute-similar and visit co-occurrence. The edge weight of a hybrid edge is determined based on the edge weights of attribute-similar edges and visit co-occurrence edges. This is how edge relationships in semantic graph data are constructed. Figure 3 The diagram shows that there are edge relationships between P1 and P3, P2 and F1, P2 and N1, F1 and F2, and P1 and P2, respectively.
[0070] According to an embodiment of this application, the bidirectional distillation loss function value is determined based on a first distillation loss value and a second distillation loss value. The first distillation loss function value is obtained based on the distribution of sample semantic categories of sample nodes in the sample semantic graph data, and according to the degree of difference between the distribution of sample structure categories of sample nodes in the sample structure graph data and the distribution of sample semantic categories of sample nodes. The second distillation loss function value is obtained based on the distribution of sample structure categories of sample nodes in the sample structure graph data, and according to the degree of difference between the distribution of sample semantic categories of sample nodes in the sample semantic graph data and the distribution of sample structure categories of sample nodes.
[0071] Based on the distribution of sample semantic categories of sample nodes in the sample semantic graph data, the first distillation loss function value can be obtained according to the degree of difference between the distribution of sample structure categories of sample nodes in the sample structure graph data and the distribution of sample semantic categories of sample nodes.
[0072] Based on the distribution of sample structure categories of sample nodes in the sample structure graph data, the second distillation loss function value can be obtained according to the degree of difference between the distribution of sample semantic categories of sample nodes and the distribution of sample structure categories of sample nodes in the sample semantic graph data.
[0073] The first and second distillation loss values can be fused to determine the bidirectional distillation loss function value, which can then be used to train a deep learning model to obtain a trained log anomaly detection model.
[0074] According to embodiments of this application, the bidirectional distillation mechanism enables the mutual transfer of semantic and structural knowledge. The first distillation ensures alignment between structure and semantics, while the second distillation ensures alignment between semantics and structure, thus promoting bidirectional transfer of cross-modal knowledge. By enhancing the depth of heterogeneous feature fusion through bidirectional constraints, the model's collaborative modeling ability for complex behavioral patterns is improved, optimization consistency under graph structure constraints is strengthened, and detection robustness and generalization performance are enhanced.
[0075] According to embodiments of this application, the first distillation loss function value is obtained based on the sample semantic category distribution of sample nodes in the sample semantic graph data, according to the semantic reliability of the sample nodes, and the degree of difference between the sample structural category distribution of sample nodes in the sample structure graph data and the sample semantic category distribution of sample nodes. The semantic reliability characterizes the reliability of the sample semantic category distribution of sample nodes. The second distillation loss function value is obtained based on the sample structural category distribution of sample nodes in the sample structure graph data, according to the structural reliability of the sample nodes, and the degree of difference between the sample semantic category distribution of sample nodes in the sample semantic graph data and the sample structural category distribution of sample nodes. The structural reliability characterizes the reliability of the sample structural category distribution of sample nodes.
[0076] By utilizing reliability-weighted symmetric knowledge transfer, the two branches act as teacher and student to each other. Semantic reliability characterizes the reliability of the sample semantic category distribution of sample nodes. Based on semantic reliability, the reliability of sample nodes can be determined, and the sample structure graph data can be updated based on sample nodes with higher than preset semantic reliability. Using the sample semantic category distribution of sample nodes in the sample semantic graph data as a benchmark, the first distillation loss function value can be obtained based on the semantic reliability of sample nodes and the degree of difference between the sample structure category distribution of sample nodes in the sample structure graph data and the sample semantic category distribution of sample nodes. As shown in formula (5).
[0077] (5);
[0078] in, This represents the semantic category distribution of the samples obtained from the semantic branch. This represents the distribution of sample structure categories obtained from structural branching. This represents the semantic reliability of the i-th node. The distillation temperature. Represents the divergence function. Indicates the number of nodes.
[0079] Structural reliability characterizes the reliability of the sample node's sample structure category distribution. Based on structural reliability, the reliability of a sample node can be determined, and the sample semantic graph data can be updated based on sample nodes with a higher than preset structural reliability. Using the sample structure category distribution of sample nodes in the sample structure graph data as a benchmark, and based on the structural reliability of the sample nodes and the degree of difference between the sample semantic category distribution of sample nodes in the sample semantic graph data and the sample structure category distribution of the sample nodes, the second distillation loss function value is obtained. As shown in formula (6).
[0080] (6);
[0081] in, This represents the semantic category distribution of the samples obtained from the semantic branch. This represents the distribution of sample structure categories obtained from structural branching. Indicates the structural reliability of the i-th node. The distillation temperature. Represents the divergence function. Indicates the number of nodes.
[0082] Highly reliable sample nodes can serve as more trustworthy teachers and are given higher weight in distillation loss, thereby enabling the transfer of reliable knowledge and avoiding unreliable samples from misleading the learning of another branch.
[0083] The value of the two-way distillation loss function is determined based on the first distillation loss value and the second distillation loss value, as shown in formula (7).
[0084] (7);
[0085] in, This represents the value of the two-way distillation loss function.
[0086] According to embodiments of this application, semantic reliability and structural reliability are introduced as weighting coefficients to achieve adaptive bidirectional distillation. The high-reliability distribution dominates the knowledge transfer direction, while the low-reliability distribution receives the correction signal, suppressing interference from noisy data and uncertain predictions. This reliability-weighted mechanism enhances the model's utilization of high-quality knowledge, improves the accuracy and stability of cross-modal knowledge transfer, and strengthens detection robustness in complex scenarios.
[0087] According to embodiments of this application, semantic reliability and structural reliability are determined as follows: For any sample node in the sample semantic graph data or any sample node in the sample structural graph data, sample category entropy and sample noise category entropy are obtained based on the sample category distribution and sample noise category distribution of the sample node. The sample noise category distribution is obtained based on the sample noise features, which are obtained by adding noise to the sample features. The sample features are the sample structural features or sample semantic features of the sample node, and the sample noise features are the sample structural noise features or sample semantic noise features of the sample node. The sample category entropy is either the sample semantic category entropy or the sample structural category entropy, and the sample noise category entropy is either the sample semantic noise category entropy or the sample structural noise category entropy. Based on the sample category entropy and the sample noise category entropy, the reliability of the sample node is obtained, which is either semantic reliability or structural reliability.
[0088] Figure 4 A schematic diagram of a reliability estimation mechanism according to an embodiment of this application is shown.
[0089] like Figure 4 As shown, a reliability estimation mechanism based on noise perturbation invariance is demonstrated. The reliability of the sample prediction is evaluated by comparing the entropy changes of the original prediction and the prediction after noise perturbation.
[0090] Taking the determination of semantic reliability for any sample node in the sample semantic graph data as an example, the process of calculating structural reliability is similar to the process of calculating semantic reliability, and will not be described in detail here.
[0091] For any sample node in the sample semantic graph data, you can add sample semantic features to that sample node. The sample noise features are obtained by perturbing with sub-Gaussian noise. Forward propagation is then performed on the sample semantic features and sample noise feature distributions respectively to obtain the corresponding sample class distribution and sample noise class distribution. Based on the information entropy calculation method, the sample class entropy and sample noise class entropy are obtained from the sample class distribution and sample noise class distribution of the sample nodes. The sample class entropy and sample noise class entropy are then... The reliability is obtained by averaging the entropy differences and performing scale normalization. (i.e., semantic initial reliability). Based on the semantic initial reliability, the semantic reliability of each sample can be calculated using a power-law distribution mapping (i.e.,...). Figure 4 (Adoption probability in the data). The semantic reliability of each node in the structure graph data and semantic graph data is calculated separately to facilitate subsequent bidirectional knowledge distillation.
[0092] For any sample node in the sample semantic graph data, you can add sample semantic features to that sample node. The noise characteristics of the sample are obtained by perturbing with sub-Gaussian noise, as shown in formula (8).
[0093] (8);
[0094] in, This represents the sample noise feature of the i-th sample node. Indicates the first Gaussian noise This represents the semantic features of the i-th sample node.
[0095] By performing forward propagation on the distributions of sample semantic features and sample noise features respectively, the corresponding sample class distributions can be obtained. and sample noise category distribution As shown in formula (9).
[0096] (9);
[0097] in, This represents the activation function.
[0098] The reliability estimation mechanism assesses sample prediction quality based on the invariance of information entropy to noise disturbances. The sample class entropy is obtained from the sample node's sample class distribution and the sample noise class distribution. and sample noise category entropy As shown in formulas (10) and (11).
[0099] (10);
[0100] (11);
[0101] Where I represents the total number of sample nodes. The number of categories is 2, i.e., logs are normal or abnormal, and the initial semantic reliability is... Defined as The average value of the sub-entropy difference is calculated and scaled, as shown in formula (12).
[0102] (12);
[0103] in, This represents the scaling factor, and K represents the entropy difference across K iterations. The smaller the value, the less sensitive the prediction is to noise disturbances, meaning the more stable and reliable the model's prediction for that sample. Conversely, The larger the value, the more susceptible the prediction is to noise, and the lower the initial semantic reliability.
[0104] The semantic reliability of each sample is calculated using a power-law distribution mapping based on the initial semantic reliability, as shown in formula (13). For the first sample within a batch... One sample:
[0105] (13);
[0106] in Let be the semantic reliability of the i-th sample node. The maximum initial semantic reliability within the current batch. The power-law parameter indicates that the greater the semantic reliability, the more reliable the sample is, and it is given higher weight during the distillation process.
[0107] According to embodiments of this application, by comparing the sample category entropy with the noise category entropy to quantify node reliability, it is possible to identify high-confidence nodes that predict stably and low-confidence nodes that are susceptible to noise interference. The entropy difference mechanism adaptively evaluates distribution quality, suppresses the negative impact of noisy samples on knowledge transfer, enhances the model's ability to perceive data uncertainty, and improves the accuracy of the distillation process and the overall robustness of the model.
[0108] According to embodiments of this application, a log anomaly detection model is used to process semantic graph data and structure graph data to obtain log anomaly detection results related to the log to be detected. This includes: processing the type attributes of nodes in the structure graph data using a first convolutional neural network to obtain structural features; processing the type attributes of nodes in the semantic graph data using a second convolutional neural network to obtain semantic features; processing the structural features based on a first classifier to obtain structural category results; processing the semantic features based on a second classifier to obtain semantic category results; and processing the structural category results and semantic category results using an activation function to obtain log anomaly detection results related to the log to be detected.
[0109] The first convolutional neural network is used to process the type attributes of nodes in the structure graph data to obtain structural features. The update rule of the structural features is shown in formula (14).
[0110] (14);
[0111] Where R represents the set of real numbers, This represents the feature vector of node j in the l-th layer. Let i represent the feature vector of node i in the l-th layer. Represents the sample node of the (l+1)th layer The corresponding structural feature vector, For sample nodes In relationship The following is a collection of neighbors. The normalization constant is For relationship A specific weight matrix, This is the self-connection weight matrix. This is the ReLU activation function.
[0112] Structural features are obtained by fusing averaging and max pooling. As shown in formula (15).
[0113] (15);
[0114] Here, READOUT represents an aggregate function.
[0115] Semantic features can be obtained by processing the type attributes of nodes in semantic graph data using a second convolutional neural network. Specifically, the initial node embedding can be generated by a text encoder, as shown in formula (16).
[0116] (16);
[0117] in, Let represent the initial semantic features of sample node i, and BERT represent the text encoder. This represents the raw data of sample node i before it was processed by the encoder.
[0118] Representation learning is performed using a BERT encoder and a heterogeneous graph convolutional network. Heterogeneous graph convolution is performed on the semantic graph data. Since the semantic graph contains three edge types (similar attributes, access co-occurrence, and mixed edges), each edge type uses an independent convolutional kernel, and the sample semantic feature update rule is shown in formula (17).
[0119] (17);
[0120] in, This represents the activation function. It is a relationship The corresponding normalized adjacency matrix, For relationship A specific weight matrix, This is the self-connected weight matrix. After L layers of graph convolution, the node embeddings incorporate semantic information from multi-hop neighborhoods. The final sample semantic features can be obtained through the same readout mechanism as the structural branching. .
[0121] Semantic branches can capture the co-occurrence of node attributes and semantics, complementing structural branches.
[0122] Based on the structural features processed by the first classifier, the structural category result can be obtained. As shown in formula (18). Based on the semantic features processed by the second classifier, the semantic category results are obtained. As shown in formula (19).
[0123] (18);
[0124] (19);
[0125] in, and The classifier weight matrix is... and For bias.
[0126] By using activation functions to process the structural category results and semantic category results, log anomaly detection results related to the log to be detected can be obtained, as shown in formula (19).
[0127] (19);
[0128] in, The result of log anomaly detection is represented by the Softmax function, as shown in formula (20).
[0129] (20);
[0130] in, This represents the category result of the c-th system log behavior, where C represents the number of system log behavior categories. This represents the category result of the j-th system log behavior. The probability distributions for each category are used. The final label, i.e., the log anomaly detection result, is determined by the highest probability. As shown in formula (21).
[0131] (twenty one);
[0132] in, Let represent the probability distribution of the c-th system log behavior. 0 indicates normal system log behavior, and 1 indicates abnormal system log behavior.
[0133] According to embodiments of this application, the log anomaly detection model is obtained by training a deep learning model based on embedding loss, classification loss, and bidirectional distillation loss. The embedding loss is obtained based on the embedding loss function, according to the sample semantic features and sample structural features of each sample node in the sample semantic graph data and sample structure graph data. The embedding loss function is used to measure the similarity between sample structural features and sample semantic features. The classification loss is obtained based on the classification loss function, according to the sample semantic category results of the sample semantic graph data and the sample structure category results of the sample structure graph data.
[0134] Based on the embedding loss function, the embedding loss value can be obtained from the semantic features and structural features of the sample nodes in the sample semantic graph data and the sample structure graph data. As shown in formula (22).
[0135] (twenty two);
[0136] The embedding loss function measures the similarity between sample structural features and sample semantic features, constraining the consistency of the graph-level representations of the structural graph data and the semantic graph data. This embedding loss value uses cosine similarity to measure the distance between the graph-level representations of sample structural features and sample semantic features in the feature space, encouraging both branches to learn consistent representations in the feature space and preventing excessive differentiation between the two branches that could lead to knowledge distillation failure. Simultaneously, this embedding loss value also acts as a regularization factor, improving the model's generalization ability.
[0137] Based on the classification loss function, the classification loss value is obtained according to the sample semantic category results of the sample semantic graph data and the sample structure category results of the sample structure graph data. As shown in formula (23).
[0138] (twenty three);
[0139] A deep learning model is trained using embedding loss, classification loss, and bidirectional distillation loss to obtain a log anomaly detection model. The training of the log anomaly detection model employs a multi-objective joint optimization strategy. The total loss function value... As shown in formula (23).
[0140] (twenty three);
[0141] in and For distillation weight, These are category weights used to handle data imbalance.
[0142] The performance of the method in this application and other related methods on the target dataset is compared in Table 1.
[0143] Table 1: Comparison of detection performance on the target dataset
[0144] method Accuracy Recall rate accuracy F1 value Lightweight fuzzing method 0.95 0.97 0.99 0.96 Source tracing method 0.98 0.95 0.99 0.96 The method in this application 1.0 0.95 0.96 0.97
[0145] Table 1 shows the performance comparison of the method in this application with lightweight fuzzy methods, source tracing methods, etc. on the target dataset, showing that the method outperforms the relevant methods in terms of precision, recall, and F1 score.
[0146] In this application, the training set, validation set, and test set are divided into a 7:1:2 ratio for training and validation.
[0147] The optimizer can be Adam, with an initial learning rate of 0.001 and cosine annealing scheduling. The batch size can be set to 32. The evaluation metrics are precision, recall, F1 score, and false positive rate (FPR).
[0148] On the target dataset, the method in this application achieved the following performance: precision of 1.0, recall of 0.95, and F1 score of 0.97.
[0149] Performance comparisons with other related methods are as follows: the lightweight fuzzy method has a precision of 0.95, a recall of 0.97, and an F1 score of 0.96. The source tracing method has a precision of 0.98, a recall of 0.95, and an F1 score of 0.96. The method in this application has a precision of 1.0, a recall of 0.95, and an F1 score of 0.97.
[0150] To verify the contributions of each module, this application also conducted ablation experiments. These experiments were divided into three aspects: the first aspect involved experiments without distillation, with an F1 score of 0.85; the second aspect involved experiments with bidirectional distillation, with an F1 score of 0.89; and the third aspect involved experiments with bidirectional distillation based on node reliability, with an F1 score of 0.97. The experiments demonstrate that the reliability estimation mechanism has a significant effect on improving detection robustness.
[0151] The method in this application can be deployed in an enterprise SOC (Security Operations Center) or a cloud security platform to achieve real-time log anomaly detection. Real-time log access: Streaming logs are accessed via Syslog or Kafka; Online graph construction: Near-real-time structural and semantic graph data can be constructed using a sliding window approach; Model inference: Pre-trained models are loaded for real-time anomaly scoring; Alarm generation: Security alarms are generated by combining thresholds and a rule engine; Feedback learning: Manual annotation feedback is supported, enabling online model fine-tuning.
[0152] According to embodiments of this application, a multi-task collaborative training mechanism is employed. Embedded loss constrains the semantic consistency of heterogeneous features, classification loss optimizes node category discrimination capabilities, and bidirectional distillation loss promotes cross-modal knowledge transfer. The joint optimization of these three mechanisms achieves a unified approach to feature alignment, accurate classification, and knowledge fusion, comprehensively enhancing the model's ability to model complex log patterns and its detection performance.
[0153] Figure 5 A schematic diagram of bidirectional distillation according to an embodiment of this application is shown.
[0154] like Figure 5As shown, the structure of the bidirectional knowledge distillation module is illustrated, including knowledge transfer from structure graph data to semantic graph data, knowledge transfer from semantic graph data to structure graph data, and a reliability estimation mechanism.
[0155] First, structural graph data (i.e., the structure graph) and semantic graph data (i.e., the semantic graph) can be obtained based on the log data. Then, by processing the structural graph data using a first convolutional neural network (i.e., R-GCN), structural features can be obtained; by processing the semantic graph data using a second convolutional neural network (Bert-GCN), semantic features can be obtained. Processing the structural features using a first classifier (i.e., the fully connected layer corresponding to the structure graph + Softmax) yields structural category results; processing the semantic features using a second classifier (i.e., the fully connected layer corresponding to the semantic graph + Softmax) yields semantic category results. Finally, by processing the structural and semantic category results using an activation function, log anomaly detection results related to the log to be detected are obtained. .
[0156] Figure 5 The upper part of the graph represents the structural branch, and the lower part represents the semantic branch. The embedding loss value (i.e., embedding loss) can be obtained based on the embedding loss function, according to the semantic and structural features of each sample node in the sample semantic graph and sample structural graph data.
[0157] Based on the distribution of sample semantic categories of sample nodes in the sample semantic graph data, and according to the semantic reliability of sample nodes and the degree of difference between the distribution of sample structure categories of sample nodes in the sample structure graph data and the distribution of sample semantic categories of sample nodes, the value of the first distillation loss function is obtained.
[0158] Based on the distribution of sample structure categories of sample nodes in the sample structure graph data, and according to the structural reliability of sample nodes, and the degree of difference between the distribution of sample semantic categories of sample nodes and the distribution of sample structure categories of sample nodes in the sample semantic graph data, the second distillation loss function value is obtained.
[0159] Figure 5 The knowledge reliability metric in this paper quantifies the reliability of the sample semantic category distribution based on the sample node and the reliability of the sample structural category distribution based on the sample node. Bidirectional knowledge distillation is performed on the node features with semantic reliability and structural reliability exceeding a preset threshold to train the log anomaly detection model.
[0160] Figure 6 An architecture diagram of a log anomaly detection method based on bidirectional knowledge distillation according to an embodiment of this application is shown.
[0161] like Figure 6As shown, the overall framework of the log anomaly detection method based on bidirectional knowledge distillation is presented, including modules such as log data collection, dual graph construction, reliable knowledge distillation, bidirectional knowledge distillation, and log anomaly detection.
[0162] First, the raw logs can be parsed to obtain structural graph data and semantic graph data. The methods for constructing structural graph data and semantic graph data have already been described in [the original text]. Figure 3 The details are explained in the text and will not be elaborated here. Then, based on reliable knowledge distillation, the reliability of each node is determined, so that subsequent bidirectional knowledge distillation based on semantic and structural reliability can be performed. Reliable knowledge distillation is... Figure 4 Explanation will be provided, but details will not be elaborated here. Based on the acquired structural graph data and semantic graph data, bidirectional knowledge distillation is performed by combining semantic reliability and structural reliability. Figure 5 The above has been explained in detail, so it will not be repeated here. Based on the embedding loss function, the embedding loss value can be obtained according to the sample semantic graph data and sample structure graph data, respectively, the sample semantic features and sample structure features of the sample nodes. (i.e., embedding alignment). Based on the classification loss function, the classification loss value is obtained according to the sample semantic category results of the sample semantic graph data and the sample structure category results of the sample structure graph data. .
[0163] Based on the semantic category distribution of sample nodes in the sample semantic graph data, and according to the semantic reliability of the sample nodes, and the degree of difference between the structural category distribution of sample nodes in the sample structure graph data and the semantic category distribution of sample nodes, a first distillation loss function value is obtained. Based on the structural category distribution of sample nodes in the sample structure graph data, and according to the structural reliability of the sample nodes, and the degree of difference between the semantic category distribution of sample nodes in the sample semantic graph data and the structural category distribution of sample nodes, a second distillation loss function value is obtained. The bidirectional distillation loss function value is then determined based on the first and second distillation loss values. Based on embedding loss value Classification loss value and two-way distillation loss value Train a deep learning model to obtain a log anomaly detection model.
[0164] By using a trained log anomaly detection model to process log data, the probability distribution of each category can be determined, and log anomaly detection results can be obtained. Log anomaly detection results can include normal results and abnormal results.
[0165] Figure 7 A structural block diagram of a log anomaly detection device according to an embodiment of this application is shown.
[0166] like Figure 7As shown, the log anomaly detection device of this embodiment includes an acquisition module 710 and a log anomaly detection result module 720.
[0167] The acquisition module 710 is used to acquire semantic graph data and structural graph data related to the log to be detected. In the semantic graph data and structural graph data, the nodes represent at least one of the logical entities or physical entities of the system. In the semantic graph structural data, the edge weights represent the similarity between the node behavior sequences corresponding to multiple nodes. In the structural graph structural data, the edge weights represent the degree of interaction between the nodes connected by the edges. The node behavior sequence represents the sequence of event types in which the nodes participate.
[0168] The log anomaly detection result module 720 is used to process semantic graph data and structural graph data using the log anomaly detection model to obtain log anomaly detection results related to the log to be detected. The log anomaly detection model is obtained by training a deep learning model based on bidirectional distillation loss value. The bidirectional distillation loss value is based on the bidirectional distillation loss function, which is obtained according to the sample semantic category distribution and sample structural category distribution of sample nodes in the sample semantic graph data and sample structural graph data. The bidirectional distillation loss function value is used to characterize the bidirectional transfer state between the sample structural graph data and the sample semantic graph data based on the sample semantic category distribution and sample structural category distribution. The bidirectional transfer state reflects the degree of cross-modal knowledge transfer, the degree of heterogeneous feature fusion, and the degree of collaborative optimization under graph structure constraints.
[0169] According to embodiments of this application, semantic graph data and structural graph data related to the log to be detected are obtained. Nodes in the semantic graph data and structural graph data can represent at least one of the system's logical entities or physical entities. Edge weights in the semantic graph data can represent the similarity between the behavior sequences of multiple nodes, while edge weights in the structural graph data represent the degree of interaction between nodes connected by the edges. Based on a dual-graph modeling framework using both structural graph data and semantic graph data, system behavior can be comprehensively characterized from two complementary dimensions: topological structure and semantic association. By processing the semantic graph data and structural graph data using the log anomaly detection model, log anomaly detection results related to the log to be detected can be obtained. This overcomes the limitations of single-view modeling in related methods, enabling a more comprehensive capture of the spatiotemporal characteristics and semantic connotations of system behavior, and improving the detection capability against complex and covert attacks. During the training of the log anomaly detection model, based on the bidirectional distillation loss function, the bidirectional distillation loss value is obtained according to the sample semantic category distribution and sample structural category distribution of each sample node in the sample semantic graph data and sample structural graph data. The bidirectional distillation loss function value is used to characterize the bidirectional transfer state between the sample structural graph data and the sample semantic graph data based on the sample semantic category distribution and sample structural category distribution. Through symmetrical knowledge transfer, the structural branch and semantic branch are optimized collaboratively during the training process, dynamically evaluating the sample prediction quality, prioritizing the learning of reliable knowledge, suppressing the influence of noise, and significantly improving the robustness and stability of the model.
[0170] According to an embodiment of this application, the bidirectional distillation loss function value is determined based on a first distillation loss value and a second distillation loss value. The first distillation loss function value is obtained based on the distribution of sample semantic categories of sample nodes in the sample semantic graph data, and according to the degree of difference between the distribution of sample structure categories of sample nodes in the sample structure graph data and the distribution of sample semantic categories of sample nodes. The second distillation loss function value is obtained based on the distribution of sample structure categories of sample nodes in the sample structure graph data, and according to the degree of difference between the distribution of sample semantic categories of sample nodes in the sample semantic graph data and the distribution of sample structure categories of sample nodes.
[0171] According to an embodiment of this application, the first distillation loss function value is obtained based on the sample semantic category distribution of sample nodes in the sample semantic graph data, according to the semantic reliability of the sample nodes, and the degree of difference between the sample structural category distribution of sample nodes in the sample structure graph data and the sample semantic category distribution of sample nodes. The semantic reliability characterizes the reliability of the sample semantic category distribution of sample nodes. The second distillation loss function value is obtained based on the sample structural category distribution of sample nodes in the sample structure graph data, according to the structural reliability of the sample nodes, and the degree of difference between the sample semantic category distribution of sample nodes in the sample semantic graph data and the sample structural category distribution of sample nodes. The structural reliability characterizes the reliability of the sample structural category distribution of sample nodes.
[0172] According to embodiments of this application, semantic reliability and structural reliability are determined as follows: For any sample node in the sample semantic graph data or any sample node in the sample structural graph data, sample category entropy and sample noise category entropy are obtained based on the sample category distribution and sample noise category distribution of the sample node. The sample noise category distribution is obtained based on the sample noise features, which are obtained by adding noise to the sample features. The sample features are the sample structural features or sample semantic features of the sample node, and the sample noise features are the sample structural noise features or sample semantic noise features of the sample node. The sample category entropy is either the sample semantic category entropy or the sample structural category entropy, and the sample noise category entropy is either the sample semantic noise category entropy or the sample structural noise category entropy. Based on the sample category entropy and the sample noise category entropy, the reliability of the sample node is obtained, which is either semantic reliability or structural reliability.
[0173] The log anomaly detection result module 720 includes: a structural feature acquisition unit, a semantic feature acquisition unit, a structural category result acquisition unit, a semantic category result acquisition unit, and a log anomaly detection result unit.
[0174] The structural feature is obtained by using a first convolutional neural network to process the type attributes of nodes in the structural graph data to obtain structural features.
[0175] The semantic feature acquisition unit is used to process the type attributes of nodes in the semantic graph data using the second convolutional neural network to obtain semantic features.
[0176] The structural category results are used to process structural features based on the first classifier to obtain structural category results.
[0177] The semantic category result unit is used to process semantic features based on the second classifier to obtain the semantic category result.
[0178] The log anomaly detection result unit is used to process the structural category result and semantic category result using the activation function to obtain the log anomaly detection result related to the log to be detected.
[0179] According to embodiments of this application, the edges of semantic graph data include attribute similarity edges and access co-occurrence edges. The edge weight of attribute similarity edges indicates that the similarity between the node behavior sequences corresponding to multiple nodes is greater than or equal to a predetermined similarity. The edge weight of access co-occurrence edges indicates that the number of times the same node accesses the same node is greater than or equal to a predetermined number of times.
[0180] According to an embodiment of this application, the edge weight of a hybrid edge is determined based on the edge weight of an edge with similar attributes and the edge weight of an edge that is co-occurring in access. A hybrid edge is an edge that is both an edge with similar attributes and an edge that is co-occurring in access.
[0181] According to embodiments of this application, the log anomaly detection model is obtained by training a deep learning model based on embedding loss, classification loss, and bidirectional distillation loss. The embedding loss is obtained based on the embedding loss function, according to the sample semantic features and sample structural features of each sample node in the sample semantic graph data and sample structure graph data. The embedding loss function is used to measure the similarity between sample structural features and sample semantic features. The classification loss is obtained based on the classification loss function, according to the sample semantic category results of the sample semantic graph data and the sample structure category results of the sample structure graph data.
[0182] According to embodiments of this application, any plurality of modules in the acquisition module 710 and the log anomaly detection result module 720 can be merged into one module, or any one of these modules can be split into multiple modules. Alternatively, at least a portion of the functionality of one or more of these modules can be combined with at least a portion of the functionality of other modules and implemented in one module. According to embodiments of this application, at least one of the acquisition module 710 and the log anomaly detection result module 720 can be at least partially implemented as a hardware circuit, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or any other reasonable means of integrating or packaging the circuit, or implemented in software, hardware, or firmware, or in any appropriate combination of any of these three implementation methods. Alternatively, at least one of the acquisition module 710 and the log anomaly detection result module 720 can be at least partially implemented as a computer program module, which can perform corresponding functions when the computer program module is run.
[0183] Figure 8 A block diagram of an electronic device suitable for implementing a cardiac resuscitation assessment method according to an embodiment of this application is shown.
[0184] like Figure 8 As shown, an electronic device 800 according to an embodiment of this application includes a processor 801, which can perform various appropriate actions and processes according to a program stored in a ROM 802 (Read-Only Memory) or a program loaded from a storage portion 808 into a RAM 803 (Random Access Memory). The processor 801 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 801 may also include onboard memory for caching purposes. The processor 801 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of this application.
[0185] RAM 803 stores various programs and data required for the operation of electronic device 800. Processor 801, ROM 802, and RAM 803 are interconnected via bus 804. Processor 801 executes various operations of the method flow according to embodiments of this application by executing programs in ROM 802 and / or RAM 803. It should be noted that the programs may also be stored in one or more memories other than ROM 802 and RAM 803. Processor 801 may also execute various operations of the method flow according to embodiments of this application by executing programs stored in said one or more memories.
[0186] According to embodiments of this application, the electronic device 800 may further include an input / output (I / O) interface 805, which is also connected to a bus 804. The electronic device 800 may also include one or more of the following components connected to the input / output (I / O) interface 805: an input section 806 including a keyboard, mouse, etc.; an output section 807 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 808 including a hard disk, etc.; and a communication section 809 including a network interface card such as a LAN card, modem, etc. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the input / output (I / O) interface 805 as needed. A removable medium 811, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 810 as needed so that computer programs read from it can be installed into the storage section 808 as needed.
[0187] This application also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs, which, when executed, implement the method according to the embodiments of this application.
[0188] According to embodiments of this application, the computer-readable storage medium can be a non-volatile computer-readable storage medium, such as including but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this application, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to embodiments of this application, the computer-readable storage medium may include ROM 802 and / or RAM 803 and / or one or more memories other than ROM 802 and RAM 803 described above.
[0189] Embodiments of this application also include a computer program product comprising a computer program containing program code for performing the methods shown in the flowchart. When the computer program product is run on a computer system, the program code is used to cause the computer system to implement the recommended methods provided in the embodiments of this application.
[0190] When the computer program is executed by the processor 801, it performs the functions defined in the system / apparatus of this application embodiment. According to the embodiments of this application, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0191] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and may be downloaded and installed via the communication section 809, and / or installed from a removable medium 811. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.
[0192] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 809, and / or installed from the removable medium 811. When the computer program is executed by the processor 801, it performs the functions defined in the system of this application embodiment. According to the embodiments of this application, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0193] According to embodiments of this application, program code for executing the computer programs provided in the embodiments of this application can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, Python, "C", or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0194] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0195] Those skilled in the art will understand that the features described in the various embodiments of this application can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in this application. In particular, the features described in the various embodiments of this application can be combined and / or combined in various ways without departing from the spirit and teachings of this application. All such combinations and / or combinations fall within the scope of this application.
[0196] The embodiments of this application have been described above. However, these embodiments are merely illustrative and not intended to limit the scope of this application. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Without departing from the scope of this application, those skilled in the art can make various substitutions and modifications, all of which should fall within the scope of this application.
Claims
1. A log anomaly detection method based on bidirectional knowledge distillation, characterized in that, include: Obtain semantic graph data and structure graph data related to the log to be detected. In the semantic graph data and the structure graph data, the nodes represent at least one of the logical entities or physical entities of the system. In the semantic graph data, the edge weights represent the similarity between the node behavior sequences corresponding to multiple nodes. In the structure graph data, the edge weights represent the degree of interaction between the nodes connected by the edges. The node behavior sequence represents the sequence of event types in which the node participates. The semantic graph data and the structural graph data are processed using a log anomaly detection model to obtain log anomaly detection results related to the log to be detected; The log anomaly detection model is obtained by training a deep learning model based on bidirectional distillation loss. The bidirectional distillation loss is based on the bidirectional distillation loss function and is obtained according to the sample semantic category distribution and sample structural category distribution of sample nodes in the sample semantic graph data and the sample structural graph data, respectively. The bidirectional distillation loss function value is used to characterize the bidirectional transfer state between the sample structural graph data and the sample semantic graph data based on the sample semantic category distribution and the sample structural category distribution. The bidirectional transfer state reflects the degree of cross-modal knowledge transfer, the degree of heterogeneous feature fusion, and the degree of collaborative optimization under graph structure constraints.
2. The method according to claim 1, characterized in that, The bidirectional distillation loss function value is determined based on the first distillation loss value and the second distillation loss value; The first distillation loss function value is obtained based on the sample semantic category distribution of the sample nodes in the sample semantic graph data, and according to the degree of difference between the sample structure category distribution of the sample nodes in the sample structure graph data and the sample semantic category distribution of the sample nodes. The second distillation loss function value is obtained based on the sample structure category distribution of the sample nodes in the sample structure graph data, and according to the degree of difference between the sample semantic category distribution of the sample nodes in the sample semantic graph data and the sample structure category distribution of the sample nodes.
3. The method according to claim 2, characterized in that, The first distillation loss function value is based on the sample semantic category distribution of the sample nodes in the sample semantic graph data, and is obtained according to the semantic reliability of the sample nodes and the degree of difference between the sample structure category distribution of the sample nodes in the sample structure graph data and the sample semantic category distribution of the sample nodes. The semantic reliability characterizes the reliability of the sample semantic category distribution of the sample nodes. The second distillation loss function value is based on the sample structure category distribution of the sample nodes in the sample structure graph data, and is obtained according to the structural reliability of the sample nodes and the degree of difference between the sample semantic category distribution of the sample nodes and the sample structure category distribution of the sample nodes in the sample semantic graph data. The structural reliability characterizes the reliability of the sample structure category distribution of the sample nodes.
4. The method according to claim 3, characterized in that, The semantic reliability and the structural reliability are determined in the following manner: For any sample node in the sample semantic graph data or any sample node in the sample structure graph data, Based on the sample category distribution and sample noise category distribution of the sample nodes, sample category entropy and sample noise category entropy are obtained. The sample noise category distribution is obtained based on sample noise features, which are obtained by adding noise to sample features. The sample features are either the sample structure features or the sample semantic features of the sample nodes. The sample noise features are either the sample structure noise features or the sample semantic noise features of the sample nodes. The sample category entropy is either the sample semantic category entropy or the sample structure category entropy. The sample noise category entropy is either the sample semantic noise category entropy or the sample structure noise category entropy. The reliability of the sample node is obtained based on the sample category entropy and the sample noise category entropy, wherein the reliability is the semantic reliability or the structural reliability.
5. The method according to any one of claims 1 to 4, characterized in that, The process of using a log anomaly detection model to process the semantic graph data and the structure graph data to obtain log anomaly detection results related to the log to be detected includes: The first convolutional neural network is used to process the type attributes of nodes in the structure graph data to obtain structural features; The semantic graph data is processed using a second convolutional neural network to obtain semantic features; The structural features are processed based on the first classifier to obtain the structural category result; The semantic features are processed using a second classifier to obtain semantic category results; The structural category results and the semantic category results are processed using activation functions to obtain log anomaly detection results related to the log to be detected.
6. The method according to any one of claims 1 to 4, wherein, The edges of the semantic graph data include attribute similarity edges and access co-occurrence edges. The edge weight of the attribute similarity edges indicates that the similarity between the node behavior sequences corresponding to multiple nodes is greater than or equal to a predetermined similarity. The edge weight of the access co-occurrence edges indicates that the number of times the same node accesses the same node is greater than or equal to a predetermined number of times.
7. The method according to claim 6, wherein, The edge weight of the hybrid edge is determined based on the edge weight of the attribute-similar edge and the edge weight of the access co-occurrence edge, wherein the hybrid edge is an edge that is both an attribute-similar edge and an access co-occurrence edge.
8. The method according to any one of claims 1 to 4, characterized in that, The log anomaly detection model is obtained by training a deep learning model based on the embedding loss value, the classification loss value, and the bidirectional distillation loss value. The embedding loss value is obtained based on the embedding loss function, according to the sample semantic features and sample structural features of the sample nodes in the sample semantic graph data and the sample structure graph data, respectively. The embedding loss function is used to measure the similarity between the sample structural features and the sample semantic features. The classification loss value is obtained based on the classification loss function, according to the sample semantic category results of the sample semantic graph data and the sample structure category results of the sample structure graph data.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the steps of the method according to any one of claims 1 to 8.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8.