A weakly supervised event representation learning method and system for assisted decision making
By constructing multiple hypersphere normal prototypes and introducing a multi-instance learning method with weak supervision signals at the sequence level, the problem of accurate localization in anomaly detection in existing technologies is solved, achieving more efficient anomaly recognition and decision support.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NORTHERN INST OF AUTOMATIC CONTROL TECH
- Filing Date
- 2025-04-09
- Publication Date
- 2026-07-24
AI Technical Summary
Existing discrete event sequence anomaly detection technologies struggle to accurately pinpoint specific abnormal events within a sequence, and unsupervised or single-class learning methods result in high false alarm and false negative rates, failing to effectively support enterprise decision-making.
By constructing a multi-hypersphere normal prototype, patterns of different normal events are learned, and weak supervision signals at the sequence level are introduced for multi-instance learning. A sequence encoder and an anomaly discriminator are used to distinguish between normal and abnormal events.
It improves the accuracy and robustness of abnormal event detection, clarifies the boundary between normal and abnormal events, and provides more accurate auxiliary decision support.
Smart Images

Figure CN120508895B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of electronic digital data processing technology, and in particular to a weakly supervised event representation learning method and system for decision support. Background Technology
[0002] With the rapid development of information technology, modern information systems have become the core infrastructure supporting the operation of enterprises and organizations. The massive amounts of monitoring data they continuously generate, such as user behavior records, provide important evidence for security protection and risk control. For intelligent operation and maintenance systems, the Discrete Event Sequence Anomaly Detection (DESAD) technology abstracts system records into discrete events and constructs event sequences to achieve automatic identification and detection of abnormal behavior. This technology can help enterprises and organizations promptly discover potential security threats to intelligent operation and maintenance systems, assisting them in making corresponding decisions and taking effective protective measures, thereby avoiding significant losses.
[0003] However, existing research on anomaly detection for discrete event sequences has significant limitations. Most current research focuses on sequence-level anomaly detection, which involves abstracting user activities into events and aggregating them into sequences. Then, it borrows sequence models from natural language processing to capture the temporal dependencies between user activities, thereby achieving anomaly detection. While these methods can identify anomalous sequences to some extent, they struggle to precisely locate specific anomalous events within the sequence. Considering that in real-world applications, an event sequence may contain hundreds or even thousands of independent events, developing fine-grained event-level detection technologies is of significant practical importance for reducing manual screening costs, improving threat localization efficiency, assisting enterprise decision-making, and implementing effective protection.
[0004] For anomaly detection in discrete event sequences at the event level, it is impractical to provide anomaly annotations for such a large number of events due to the extreme rarity and concealment of anomalous events. Most existing studies employ unsupervised or single-class learning paradigms, identifying anomalous events that deviate from the pattern by modeling normal behavior patterns. This approach has significant shortcomings in practical applications—it cannot exhaustively enumerate all normal patterns, and the boundary between normal and anomalous is often blurred, leading to high false positive and false negative rates. Summary of the Invention
[0005] The technical problem to be solved by this invention is to provide a weakly supervised event representation learning method and system for decision support. By constructing multiple hypersphere normal prototypes, it learns the patterns of different normal events and introduces sequence-level weak supervision signals to perform multi-instance learning, distinguishing between normal and abnormal events, and effectively learning the feature representations of events, thereby providing better assistance for decision-making in intelligent operation and maintenance systems.
[0006] This invention is achieved through the following technical solution: A weakly supervised event representation learning method for decision support includes the following steps: S1: Form an event sequence from the user's actions on the operation and maintenance network, and use a sequence encoder to generate an initial event representation from the event sequence; S2: Randomly set multiple hypersphere centers to construct a normal prototype of multiple hyperspheres; S3: Using the representation of the hypersphere center in the normal prototype of the hypersphere and the representation of the normal event in the initial event representation as input, train the sequence encoder and the normal prototype of the hypersphere to obtain the sequence encoder and the normal prototype of the hypersphere after the first stage parameter optimization. S4: The sequence encoder with optimized parameters in the first stage generates the event representation after the first stage optimization. The anomaly discriminator calculates the anomaly score of each event in the event sequence based on the event representation after the first stage optimization and the representation of the center of the hypersphere in the normal prototype of the hypersphere after the first stage optimization, and obtains the prediction result of the event. S5: Perform multi-instance learning, obtain the prediction results of the event sequence based on the prediction results of the event, and introduce event sequence level labels as weak supervision signals. Use the prediction results of the event sequence and the event sequence level labels to train the sequence encoder, normal prototype of the multi-hypersphere and anomaly discriminator after parameter optimization in the first stage, and obtain the sequence encoder, normal prototype of the multi-hypersphere and anomaly discriminator after parameter optimization in the second stage. S6: The sequence encoder with optimized parameters in the second stage generates the event representation optimized in the second stage. The anomaly discriminator with optimized parameters in the second stage calculates the final anomaly score of each event in the event sequence based on the event representation optimized in the second stage and the representation of the center of the hypersphere in the normal prototype of the hypersphere optimized in the second stage, and obtains the final prediction result of the event, screens abnormal events, and assists the manual decision-making of the intelligent operation and maintenance system.
[0007] In the optimized version, in step S1, the sequence encoder generates an initial event representation from the event sequence through two processes: event embedding and sequence context encoding.
[0008] Furthermore, the sequence encoder in step S1 generates the initial event representation from the event sequence through two processes: event embedding and sequence context encoding, as follows: S11: The embedding layer of the sequence encoder projects events into the embedding space to obtain the corresponding embedding vectors; S12: The event sequence is encoded according to equation (1) using a bidirectional gated loop unit to generate the initial event representation: (1); in, Encoding representing a sequence of events, This indicates a bidirectional gated loop unit. Representing events respectively , ... Embedded vector, Indicates the number of events.
[0009] Furthermore, in step S2, the normal prototype of the multi-hypersphere is constructed using the following method: Random settings The center of a hypersphere is represented by equation (2): (2); in: Indices representing the index of a hypersphere. Represent real numbers, Representing dimension, Indicates the center of the hypersphere. Represents the space of real numbers.
[0010] Furthermore, in step S3, the sequence encoder and the normal prototype of the multi-hypersphere are trained using the following method to obtain the sequence encoder and the normal prototype of the multi-hypersphere after the first stage parameter optimization: S31: During training, the distance from the representation of each normal event in the initial event representation to the center of each hypersphere is calculated according to equation (3): (3); in: The initial event is represented in the first... Normal events up to the 1st The Euclidean distance from the center of a hypersphere The initial event is represented in the first... Representation of normal events Indicates the center of the hypersphere. Indices representing the index of a hypersphere. This indicates the calculation of Euclidean distance; S32: Based on the distance from the representation of each normal event in the initial event representation to the center of each hypersphere, calculate the index from the representation of each normal event in the initial event representation to the nearest hypersphere center according to equation (4): (4); in: This represents the index of each normal event in the initial event representation to the nearest hypersphere center. This indicates a minimize operation. Indicates the total number of hyperspheres; S33: The multi-hypersphere center loss is calculated according to equation (5) based on the index of each normal event representation in the initial event representation to the nearest hypersphere center. (5); in: This indicates that the loss is more than the center of the sphere. This represents a normal sequence of events within an event sequence. This indicates the number of normal event sequences in the event sequence. The distance to the event represents the nearest hypersphere center. This indicates the calculation of the square Euclidean distance; S34: Calculate the separability loss of the hypersphere according to equation (6): (6); in: This indicates the loss of divisibility of the hypersphere. This indicates the calculation of binary cross-entropy loss. This indicates an exponential operation. This represents the hypersphere center that is second closest to the representation of each normal event in the initial event representation. This represents the index of the representation of each normal event in the initial event representation to the second nearest hypersphere; S35: Based on the loss of the multiple hypersphere centers and the loss of hypersphere separability, the loss function of the first stage is calculated according to equation (7): (7); in: This represents the loss function for the first stage. The weight representing the loss of separability of the hypersphere; S36: Perform iterative training with the goal of minimizing the loss function of the first stage to complete the first stage training of the sequence encoder and the normal prototype of the multi-hypersphere, and obtain the sequence encoder and the normal prototype of the multi-hypersphere with optimized parameters in the first stage.
[0011] Furthermore, in step S4, the following method is used to generate the first-stage optimized event representation using the sequence encoder with optimized parameters from the first stage. Based on the first-stage optimized event representation and the representation of the hypersphere center in the normal prototype of the hypersphere with optimized parameters from the first stage, the anomaly score of each event in the event sequence is calculated to obtain the event prediction result: S41: The embedding layer of the sequence encoder after the first-stage parameter optimization projects the events into the embedding space to obtain the first-stage optimized embedding vector; S42: The event sequence is encoded according to equation (8) using a bidirectional gated loop unit to generate the event representation after the first stage of optimization. (8); in, This represents the encoding of the event sequence after the first stage of optimization. This indicates a bidirectional gated loop unit. Representing events respectively , ... The embedding vector after the first stage of optimization Indicates the number of events.
[0012] S43: Calculate the distance from the representation of each event in the first-stage optimized event representation to the center of each hypersphere in the first-stage parameter-optimized hypersphere normal prototype according to equation (9): (9); in: This represents the first stage of optimized event representation. The event, in the first stage of parameter optimization, in the normal prototype of the hypersphere, the [number]th event... The Euclidean distance from the center of a hypersphere This represents the first stage of optimized event representation. The representation of an event, This represents the first stage of parameter optimization in the normal prototype of the multi-hypersphere. The center of a hypersphere Indices representing the index of a hypersphere. This indicates the calculation of Euclidean distance; S44: Based on the distance from the representation of each event in the first-stage optimized event representation to the center of each hypersphere in the first-stage parameter-optimized normal prototype, calculate the index of the nearest hypersphere center in the first-stage parameter-optimized normal prototype for each event representation in the first-stage optimized event representation according to Equation (10): (10); in: This represents the index of each event in the first-stage optimized event representation to the nearest hypersphere center in the first-stage parameter-optimized hypersphere normal prototype. This indicates a minimize operation. Indicates the total number of hyperspheres; S45: The anomaly discriminator calculates the deviation score of each event in the event sequence based on the nearest hypersphere center according to equation (11): (11); in: This represents the deviation score of each event in the event sequence from the nearest hypersphere center in the multi-hypersphere prototype after parameter optimization in the first stage. Represents the first event in the event sequence. One event, Represents a sequence of events. This represents the first stage of optimized event representation. The representation of an event, The event distance from the first-stage optimization indicates the center of the hypersphere in the most recent first-stage parameter-optimized hypersphere normal prototype. S46: Calculate the anomaly discriminator score for each event in the event sequence according to equation (12): (12); in: This represents the anomaly discriminator score for each event in the event sequence. This represents the activation function. Represents the weight matrix. This indicates the first stage of optimization and enhancement. This event indicates that... Indicates bias. This represents the matrix transpose operation; S47: Based on the deviation score of each event in the event sequence from the nearest hypersphere center in the multi-hypersphere prototype after parameter optimization in the first stage, and the anomaly discriminator score of each event in the event sequence, calculate the anomaly score of each event in the event sequence according to equation (13): (13); in: This represents the anomaly score for each event in the event sequence. Indicates model parameters, Indicates the dual-score balance factor; S48: Compare the abnormal score of each event in the event sequence with a set threshold. If the abnormal score of each event in the event sequence is less than the set threshold, the prediction result of the event is judged to be normal. If the abnormal score of each event in the event sequence is greater than or equal to the set threshold, the prediction result of the event is judged to be abnormal.
[0013] Furthermore, in step S5, the prediction results of the event sequence are obtained based on the prediction results of the events using the following method: S51: Arrange the events in the event sequence from highest to lowest anomaly score, and determine the preceding events according to equation (14). Index of events: (14); in: Indicates the first In the sequence of events, the first Index of events, Indicates before calculation A function for an event. Indicates the first A sequence of events Indicates the first A sequence of events Abnormal score of events; S52: Based on Center front The index of each event is obtained according to equation (15). Prediction results: (15); in: Indicates the first Prediction results for a sequence of events, This indicates taking the absolute value.
[0014] Furthermore, in step S5, event sequence-level labels are introduced as weak supervision signals using the following method. The sequence encoder, the normal prototype of the hypersphere, and the anomaly discriminator are trained using the prediction results of the event sequence and the event sequence-level labels to obtain the sequence encoder, the normal prototype of the hypersphere, and the anomaly discriminator after parameter optimization in the second stage: S53: Introduce event sequence level labels as weak supervision signals, and calculate the loss function for the second stage using the event sequence level labels and the prediction results of the event sequence according to equation (16): (16); in: This represents the loss function for the second stage. A set representing a sequence of events. Indicates the number of event sequences. Indicates the first Event sequence level labels; S54: With the goal of minimizing the loss function of the second stage, train the sequence encoder, the normal prototype of the hypersphere, and the anomaly discriminator after the parameter optimization of the first stage to obtain the sequence encoder, the normal prototype of the hypersphere, and the anomaly discriminator after the parameter optimization of the second stage.
[0015] The optimized method used in step S6 is as follows to obtain the final prediction result of the event: S61: Encode the event sequence using the sequence encoder optimized by the second-stage parameters to generate the event representation optimized by the second stage; S62: Calculate the distance from the representation of each event in the second-stage optimized event representation to the center of each hypersphere in the second-stage parameter-optimized normal prototype of the multi-hypersphere; S63: Based on the distance from the representation of each event in the second-stage optimized event representation to the center of each hypersphere in the second-stage parameter-optimized normal prototype, calculate the index of the nearest hypersphere center in the second-stage parameter-optimized normal prototype for each event representation in the second-stage optimized event representation. S64: Based on the event representation optimized in the second stage and the index of each event representation in the second stage optimized event representation to the nearest hypersphere center in the normal prototype of the multi-hypersphere after the second stage parameter optimization, calculate the final deviation score and the final anomaly discriminator score for each event based on the nearest hypersphere center, and then further calculate the final anomaly score. S65: Screen for abnormal events based on the final abnormality score.
[0016] A weakly supervised event representation learning system for decision support is provided to execute a weakly supervised event representation learning method for decision support as described in any of the above, comprising a sequence encoder, a normal prototype of a multi-hypersphere, and an anomaly discriminator. The sequence encoder is used to form an event sequence from the user's operation behavior on the operation and maintenance network, and generate an initial event representation from the event sequence; the sequence encoder after the first stage parameter optimization is used to generate an optimized event representation from the event sequence; the sequence encoder after the second stage parameter optimization is used to generate an optimized event representation from the event sequence. The normal prototype of the multi-hypersphere includes multiple randomly set hypersphere centers; the normal prototype of the multi-hypersphere after the first stage parameter optimization includes multiple hypersphere centers after the first stage optimization; the normal prototype of the multi-hypersphere after the second stage parameter optimization includes multiple hypersphere centers after the second stage optimization. The anomaly discriminator is used to calculate the anomaly score of each event in the event sequence based on the event representation optimized in the first stage and the center of the hypersphere in the normal prototype of the hypersphere after parameter optimization in the first stage, so as to obtain the prediction result of the event; the anomaly discriminator after parameter optimization in the second stage is used to calculate the final anomaly score of each event in the event sequence based on the event representation optimized in the second stage and the center of the hypersphere in the normal prototype of the hypersphere after parameter optimization in the second stage, so as to obtain the final prediction result of the event.
[0017] Beneficial effects of the invention: This invention provides a weakly supervised event representation learning method and system for decision support. It uses multiple hyperspheres to represent different normal modes of events, overcoming the limitation of single-class learning paradigms that cannot exhaustively enumerate all normal modes, and providing a good unsupervised starting point. By introducing sequence-level weakly supervised signals for multi-instance learning, the detection becomes more robust, the boundary between normal and abnormal events becomes clearer, and more accurate event representations are learned, thereby providing more accurate decision support for intelligent operation and maintenance systems. Attached Figure Description
[0018] Figure 1 This is a schematic diagram of the process of this invention.
[0019] Figure 2 This is a schematic diagram of the system structure of the present invention. Detailed Implementation
[0020] A weakly supervised event representation learning method for decision support includes the following steps, the flowchart of which is shown below. Figure 1 As shown: S1: Form an event sequence from the user's actions on the operation and maintenance network, and use a sequence encoder to generate an initial event representation from the event sequence; Specifically, the user's operations on the network mainly include: logging in, accessing files, connecting devices, accessing the network, sending emails, and receiving emails.
[0021] The sequence encoder includes an embedding layer and a bidirectional gated loop unit, which can generate an initial event representation from an event sequence through two processes: event embedding and sequence context encoding.
[0022] The specific method by which a sequence encoder generates an initial event representation from an event sequence is as follows: S11: The embedding layer of the sequence encoder projects events into the embedding space to obtain the corresponding embedding vectors; S12: The bidirectional gated loop unit using a sequence encoder encodes the event sequence according to equation (1), thereby generating the initial event representation: (1); in, Encoding representing a sequence of events, , events , ... The context indicates that This indicates a bidirectional gated loop unit. Representing events respectively , ... Embedded vector, Indicates the number of events.
[0023] Sequence encoders can capture the semantic information of events. The bidirectional gated recurrent units they use can capture long-term dependencies between events better, generate more effective initial event representations, and provide stronger feature support for subsequent analysis tasks.
[0024] S2: Randomly set multiple hypersphere centers to construct a normal prototype of multiple hyperspheres; Specifically, the normal prototype of a multi-hypersphere can be constructed using the following method: Random settings The center of a hypersphere is represented by equation (2): (2); in: Indices representing the index of a hypersphere. Represent real numbers, Representing dimension, Indicates the center of the hypersphere. Represents the space of real numbers.
[0025] Randomly setting multiple hypersphere centers, instead of a single hypersphere center, can better handle multimodal normal behavior.
[0026] S3: Using the representation of the hypersphere center in the normal prototype of the hypersphere and the representation of the normal event in the initial event representation as input, train the sequence encoder and the normal prototype of the hypersphere to obtain the sequence encoder and the normal prototype of the hypersphere after the first stage parameter optimization. Specifically, the following method can be used to train the sequence encoder and the normal prototype of the hypersphere to obtain the sequence encoder with optimized parameters in the first stage and the normal prototype of the hypersphere with optimized parameters in the first stage: S31: During training, the distance from the representation of each normal event in the initial event representation to the center of each hypersphere is calculated according to equation (3): (3); in: The initial event is represented in the first... Normal events up to the 1st The Euclidean distance from the center of a hypersphere The initial event is represented in the first... Representation of normal events Indicates the center of the hypersphere. Indices representing the index of a hypersphere. This indicates the calculation of Euclidean distance; S32: Based on the distance from the representation of each normal event in the initial event representation to the center of each hypersphere, calculate the index from the representation of each normal event in the initial event representation to the nearest hypersphere center according to equation (4): (4); in: This represents the index of each normal event in the initial event representation to the nearest hypersphere center. This indicates a minimize operation. Indicates the total number of hyperspheres; S33: The multi-hypersphere center loss is calculated according to equation (5) based on the index of each normal event representation in the initial event representation to the nearest hypersphere center. (5); in: This indicates that the loss is more than the center of the sphere. This represents a normal sequence of events within an event sequence. This indicates the number of normal event sequences in the event sequence. The distance to the event represents the nearest hypersphere center. This indicates the calculation of the square Euclidean distance; S34: Calculate the separability loss of the hypersphere according to equation (6): (6); in: This indicates the loss of divisibility of the hypersphere. This indicates the calculation of binary cross-entropy loss. This indicates an exponential operation. This represents the hypersphere center that is second closest to the representation of each normal event in the initial event representation. This represents the index of the representation of each normal event in the initial event representation to the second nearest hypersphere; S35: Based on the loss of the multiple hypersphere centers and the loss of hypersphere separability, the loss function of the first stage is calculated according to equation (7): (7); in: This represents the loss function for the first stage. The weight representing the loss of separability of the hypersphere; S36: Perform iterative training with the goal of minimizing the loss function of the first stage to complete the first stage training of the sequence encoder and the normal prototype of the multi-hypersphere, and obtain the sequence encoder and the normal prototype of the multi-hypersphere with optimized parameters in the first stage.
[0027] By learning patterns of different normal events through multiple hyperspheres, normal events cluster near the center of each hypersphere in the feature space, resulting in a multi-hypersphere center loss. This causes the detected event to be closer to the center of the nearest hypersphere, resulting in a loss of hypersphere separability. This allows the detected events to be further away from the center of the second nearest hypersphere, thus enabling the learning of better multi-hypersphere representations and optimizing the event representation.
[0028] S4: The sequence encoder with optimized parameters in the first stage generates the event representation after the first stage optimization. The anomaly discriminator calculates the anomaly score of each event in the event sequence based on the event representation after the first stage optimization and the representation of the center of the hypersphere in the normal prototype of the hypersphere after the first stage optimization, and obtains the prediction result of the event. The anomaly discriminator consists of a self-attention layer and a fully connected layer. The self-attention layer further enhances the representation of event features, while the fully connected layer is used to calculate the final anomaly discriminator score.
[0029] Specifically, the following method can be used to generate the first-stage optimized event representation using the sequence encoder with optimized parameters from the first stage. Based on the first-stage optimized event representation and the representation of the hypersphere center in the normal prototype of the hypersphere with optimized parameters from the first stage, the anomaly score of each event in the event sequence can be calculated to obtain the event prediction result: S41: The embedding layer of the sequence encoder after the first-stage parameter optimization projects the events into the embedding space to obtain the first-stage optimized embedding vector; S42: The event sequence is encoded according to equation (8) using a bidirectional gated loop unit to generate the event representation after the first stage of optimization. (8); in, This represents the encoding of the event sequence after the first stage of optimization. This indicates a bidirectional gated loop unit. Representing events respectively , ... The embedding vector after the first stage of optimization Indicates the number of events.
[0030] S43: Calculate the distance from the representation of each event in the first-stage optimized event representation to the center of each hypersphere in the first-stage parameter-optimized hypersphere normal prototype according to equation (9): (9); in: This represents the event representation after the first stage of optimization. The event, in the first stage of parameter optimization, in the normal prototype of the hypersphere, the [number]th event... The Euclidean distance from the center of a hypersphere This represents the event representation after the first stage of optimization. The representation of an event, This represents the first stage of parameter optimization in the normal prototype of the multi-hypersphere. The center of a hypersphere Indices representing the index of a hypersphere. This indicates the calculation of Euclidean distance; S44: Based on the distance from the representation of each event in the first-stage optimized event representation to the center of each hypersphere in the first-stage parameter-optimized normal prototype, calculate the index of the nearest hypersphere center in the first-stage parameter-optimized normal prototype for each event representation in the first-stage optimized event representation according to Equation (10): (10); in: This represents the index of each event in the first-stage optimized event representation to the nearest hypersphere center in the first-stage parameter-optimized hypersphere normal prototype. This indicates a minimize operation. Indicates the total number of hyperspheres; S45: The anomaly discriminator calculates the deviation score of each event in the event sequence based on the nearest hypersphere center according to equation (11): (11); in: This represents the deviation score of each event in the event sequence from the nearest hypersphere center in the multi-hypersphere prototype after parameter optimization in the first stage. Represents the first event in the event sequence. One event, Represents a sequence of events. This represents the event representation after the first stage of optimization. The representation of an event, The event distance from the first-stage optimization indicates the center of the hypersphere in the most recent first-stage parameter-optimized hypersphere normal prototype. S46: Calculate the anomaly discriminator score for each event in the event sequence according to equation (12): (12); in: This represents the anomaly discriminator score for each event in the event sequence. This represents the activation function. Represents the weight matrix. This indicates the first stage of optimization and enhancement. This event indicates that... Indicates bias. This represents the matrix transpose operation; The self-attention layer of the anomaly discriminator passes through We can obtain the event representation after the first stage optimization and the self-attention layer enhancement. ,in This enables computation of the self-attention layer, thereby further enhancing the ability to represent event features; S47: Based on the deviation score of each event in the event sequence from the nearest hypersphere center in the multi-hypersphere prototype after parameter optimization in the first stage, and the anomaly discriminator score of each event in the event sequence, calculate the anomaly score of each event in the event sequence according to equation (13): (13); in: This represents the anomaly score for each event in the event sequence. Indicates model parameters, Indicates the dual-score balance factor; S48: Compare the abnormal score of each event in the event sequence with a set threshold. If the abnormal score of each event in the event sequence is less than the set threshold, the prediction result of the event is judged to be normal. If the abnormal score of each event in the event sequence is greater than or equal to the set threshold, the prediction result of the event is judged to be abnormal.
[0031] S5: Perform multi-instance learning, obtain the prediction results of the event sequence based on the prediction results of the event, and introduce event sequence level labels as weak supervision signals. Use the prediction results of the event sequence and the event sequence level labels to train the sequence encoder, normal prototype of the multi-hypersphere and anomaly discriminator after parameter optimization in the first stage, and obtain the sequence encoder, normal prototype of the multi-hypersphere and anomaly discriminator after parameter optimization in the second stage.
[0032] Specifically, the prediction results of the event sequence can be obtained based on the prediction results of the event itself, as follows: S51: Arrange the events in the event sequence from highest to lowest anomaly score, and determine the preceding events according to equation (14). Index of events: (14); in: Indicates the first In the sequence of events, the first Index of events, Indicates before calculation A function for an event. Indicates the first A sequence of events Indicates the first A sequence of events Abnormal score of events; S52: Based on Center front The index of each event is obtained according to equation (15). Prediction results: (15); in: Indicates the first Prediction results for a sequence of events, This indicates taking the absolute value.
[0033] Furthermore, event sequence-level labels can be introduced as weak supervision signals using the following method. The prediction results of the event sequences and the event sequence-level labels are used to train the sequence encoder, the normal prototype of the hypersphere, and the anomaly discriminator, resulting in the second-stage parameter-optimized sequence encoder, the normal prototype of the hypersphere, and the anomaly discriminator: S53: Introduce event sequence level labels as weak supervision signals, and calculate the loss function for the second stage using the event sequence level labels and the prediction results of the event sequence according to equation (16): (16); in: This represents the loss function for the second stage. A set representing a sequence of events. Indicates the number of event sequences. Indicates the first Event sequence level labels; S54: With the goal of minimizing the loss function of the second stage, train the sequence encoder, the normal prototype of the hypersphere, and the anomaly discriminator after the parameter optimization of the first stage to obtain the sequence encoder, the normal prototype of the hypersphere, and the anomaly discriminator after the parameter optimization of the second stage.
[0034] After two phases of training, the sequence encoder, the normal prototype of the multi-hypersphere, and the anomaly discriminator overcome the limitation of the single-class learning paradigm in being unable to exhaustively enumerate all normal patterns, providing a good unsupervised starting point. By introducing weak supervision signals at the sequence level and performing multi-instance learning, the detection becomes more robust, the boundary between normal and abnormal events becomes clearer, and more accurate event representations can be learned, thus providing more accurate auxiliary decision-making for intelligent operation and maintenance systems.
[0035] S6: The sequence encoder with optimized parameters in the second stage generates the event representation optimized in the second stage. The anomaly discriminator with optimized parameters in the second stage calculates the final anomaly score of each event in the event sequence based on the event representation optimized in the second stage and the representation of the center of the hypersphere in the normal prototype of the hypersphere optimized in the second stage, and obtains the final prediction result of the event, screens abnormal events, and assists the manual decision-making of the intelligent operation and maintenance system.
[0036] Specifically, the final prediction result of the event can be obtained using the following method: S61: The embedding layer of the sequence encoder after the second-stage parameter optimization projects the events into the embedding space to obtain the second-stage optimized embedding vector, and uses a bidirectional gated recurrent unit to encode the event sequence according to equation (17) to generate the second-stage optimized event representation: (17); in, This represents the encoding of the event sequence after the second stage of optimization. This indicates a bidirectional gated loop unit. Representing events respectively , ... The second-stage optimized embedding vector, Indicates the number of events.
[0037] S62: Calculate the distance from the representation of each event in the second-stage optimized event representation to the center of each hypersphere in the second-stage parameter-optimized hypersphere normal prototype according to equation (18): (18); in: This represents the second stage of optimized event representation. The event, in the second stage parameter optimization of the normal prototype of the hypersphere, is the first... The Euclidean distance from the center of a hypersphere This represents the second stage of optimized event representation. The representation of an event, This indicates the first hypersphere normal prototype after the second stage parameter optimization. The center of a hypersphere Indices representing the index of a hypersphere. This indicates the calculation of Euclidean distance; S63: Based on the distance from the representation of each event in the second-stage optimized event representation to the center of each hypersphere in the second-stage parameter-optimized normal prototype, calculate the index of the nearest hypersphere center in the second-stage parameter-optimized normal prototype for each event representation in the second-stage optimized event representation according to equation (19): (19); in: This represents the index of each event in the second-stage optimized event representation to the nearest hypersphere center in the second-stage parameter-optimized normal prototype. This indicates a minimize operation. Indicates the total number of hyperspheres; S64: The anomaly discriminator calculates the deviation score of each event in the event sequence from the nearest hypersphere center in the second-stage parameter-optimized hypersphere prototype according to equation (20): (20); in: This represents the deviation score of each event in the event sequence from the nearest hypersphere center in the second-stage parameter-optimized hypersphere prototype. Represents the first event in the event sequence. One event, Represents a sequence of events. This represents the second stage of optimized event representation. The representation of an event, The event distance from the second-stage optimized event represents the center of the hypersphere in the most recent second-stage parameter-optimized hypersphere normal prototype. The anomaly discriminator calculates the anomaly discriminator score for each event in the event sequence according to equation (21): (twenty one); in: This represents the final anomaly discriminant score for each event in the event sequence. This represents the activation function. This represents the optimized weight matrix. This indicates the second phase of optimization and enhancement. This event indicates that... This represents the optimized bias. This represents the matrix transpose operation; The self-attention layer of the anomaly discriminator passes through We can obtain the event representation after the second-stage optimization and the self-attention layer enhancement. ,in The computation of the self-attention layer further enhances the ability to represent event features.
[0038] Based on the deviation score of each event in the event sequence from the nearest hypersphere center in the second-stage parameter-optimized multi-hypersphere prototype and the final anomaly discriminator score of each event in the event sequence, the final anomaly score of each event in the event sequence is calculated according to equation (22): (twenty two); in: This represents the final anomaly score for each event in the event sequence. This represents the optimized model parameters. Indicates the dual-score balance factor; S65: Compare the final abnormal score of each event in the event sequence with a set threshold. If the final abnormal score of each event in the event sequence is less than the set threshold, the final prediction result of the event is judged to be normal. If the final abnormal score of each event in the event sequence is greater than or equal to the set threshold, the final prediction result of the event is judged to be abnormal.
[0039] The event representation after the second stage of optimization is more accurate, and the representation of the hypersphere center in the normal prototype of the hypersphere after the second stage of parameter optimization is more accurate. Therefore, screening abnormal events by using the event representation after the second stage of optimization and the representation of the hypersphere center in the normal prototype of the hypersphere after the second stage of parameter optimization yields more accurate screening results, which can more effectively assist the manual decision-making of the intelligent operation and maintenance system.
[0040] This invention validates the effectiveness of a weakly supervised event representation learning method for decision support by evaluating its method on the CER r4.2 and CERT r5.2 datasets, which are widely used in the field of internal threat detection. Detailed information about the datasets is shown in Table 1. Table 1
[0041] This invention uses AUC (Area Under the ROC Curve), DR (Detection Rate), and FPR (False Positive Rate) as evaluation metrics to assess the performance of the proposed method. To ensure fair comparisons between methods, the same data partitioning is used.
[0042] Specifically, the experiments conducted in this invention include the following aspects: 1. Event-Level Discrete Event Sequence Anomaly Detection: Comparison with 16 state-of-the-art methods, including DeepLog, TIRESIAS, FMLP, RNN, GRU, Transformer, RWKV, DIEN, BST, m-RNN, m-GRU, m-LSTM, m-Transformer, mFMLP, ITDBERT, and OC4Seq. DeepLog and TIRESIAS are two classic methods that fit normal event sequences by learning to predict the next event in a given event sequence context. LSTM, RNN, GRU, Transformer, and RWKV are several widely popular sequence modeling architectures. DIEN and BST are two classic user behavior modeling methods. FMLP predicts future user behavior by filtering noise from historical user behavior data. m-RNN, m-GRU, m-LSTM, m-Transformer, and mFMLP learn to predict events at masked locations using bidirectional RNNs, GRUs, LSTMs, Transformers, and FMLP. ITDBERT is an attention-based event-level detection method. OC4Seq embeds normal events into a hypersphere and detects anomalies by predicting how close the events are to the center of the hypersphere.
[0043] Specific event-level discrete event sequence anomaly detection data are shown in Table 2: Table 2
[0044] The results show that the method of the present invention outperforms other methods on both datasets, with improved AUC and DR indices and reduced FPR indices.
[0045] 2. Component Analysis Experiment: Used to evaluate the contribution of each component in this invention. The specific experimental results are shown in Table 3. Table 3
[0046] Experimental results show that compared to training solely on normal events based on multiple hyperspheres, the proposed method achieves significant improvements across all metrics, demonstrating the effectiveness of multi-instance learning. This indicates that introducing weakly supervised sequence-level annotations can effectively help the model distinguish between normal and abnormal events. Furthermore, the proposed method outperforms methods that rely solely on multi-instance learning, further proving the effectiveness of training on normal events based on multiple hyperspheres.
[0047] A weakly supervised event representation learning system for decision support is provided, which executes a weakly supervised event representation learning method for decision support as described in any of the above-mentioned methods. Its system structure diagram is shown below. Figure 2As shown, Includes a sequence encoder, a normal prototype of a multi-hypersphere, and an anomaly detector; A sequence encoder is used to form an event sequence from user operations on the operation and maintenance network and generate an initial event representation from the event sequence; a sequence encoder after the first stage of parameter optimization is used to generate an optimized event representation from the event sequence; a sequence encoder after the second stage of parameter optimization is used to generate an optimized event representation from the event sequence. The normal prototype of the multi-hypersphere contains multiple randomly set hypersphere centers; the normal prototype of the multi-hypersphere after the first stage of parameter optimization contains multiple hypersphere centers after the first stage of optimization; the normal prototype of the multi-hypersphere after the second stage of parameter optimization contains multiple hypersphere centers after the second stage of optimization. An anomaly discriminator is used to calculate the anomaly score of each event in the event sequence based on the event representation optimized in the first stage and the center of the hypersphere in the normal prototype of the hypersphere after parameter optimization in the first stage, and obtain the prediction result of the event. An anomaly discriminator after parameter optimization in the second stage is used to calculate the final anomaly score of each event in the event sequence based on the event representation optimized in the second stage and the center of the hypersphere in the normal prototype of the hypersphere after parameter optimization in the second stage, and obtain the final prediction result of the event.
[0048] In summary, the present invention provides a weakly supervised event representation learning method and system for decision support. By constructing multiple hypersphere normal prototypes, training on different normal event patterns, and introducing sequence-level weak supervision signals for multi-instance learning, the method distinguishes between normal and abnormal events in intelligent operation and maintenance systems, effectively learns the feature representations of intelligent operation and maintenance events, and thus provides better assistance for decision-making in intelligent operation and maintenance systems.
[0049] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A weakly supervised event representation learning method for decision support, characterized in that: Includes the following steps: S1: Form an event sequence from the user's actions on the operation and maintenance network, and use a sequence encoder to generate an initial event representation from the event sequence; S2: Randomly set multiple hypersphere centers to construct a normal prototype of multiple hyperspheres; S3: Using the representation of the hypersphere center in the normal prototype of the hypersphere and the representation of the normal event in the initial event representation as input, train the sequence encoder and the normal prototype of the hypersphere to obtain the sequence encoder and the normal prototype of the hypersphere after the first stage parameter optimization. S4: The sequence encoder with optimized parameters in the first stage generates the event representation after the first stage optimization. The anomaly discriminator obtains the event prediction result based on the event representation after the first stage optimization and the representation of the center of the hypersphere in the normal prototype of the hypersphere after the first stage parameter optimization. S5: Perform multi-instance learning, obtain the prediction results of the event sequence based on the prediction results of the event, and introduce event sequence level labels as weak supervision signals. Use the prediction results of the event sequence and event sequence level labels to train the sequence encoder after parameter optimization in the first stage, the normal prototype of the multi-hypersphere after parameter optimization in the first stage, and the anomaly discriminator to obtain the sequence encoder, the normal prototype of the multi-hypersphere, and the anomaly discriminator after parameter optimization in the second stage. S6: The sequence encoder with optimized parameters in the second stage generates the event representation optimized in the second stage. The anomaly discriminator with optimized parameters in the second stage obtains the final prediction result of the event based on the event representation optimized in the second stage and the representation of the center of the hypersphere in the normal prototype of the hypersphere optimized in the second stage, and screens abnormal events to assist the manual decision-making of the intelligent operation and maintenance system.
2. The weakly supervised event representation learning method for decision support according to claim 1, characterized in that: In step S1, the sequence encoder generates an initial event representation from the event sequence through two processes: event embedding and sequence context encoding.
3. The weakly supervised event representation learning method for decision support according to claim 2, characterized in that: The method for generating the initial event representation using a sequence encoder in step S1 is as follows: S11: The embedding layer of the sequence encoder projects events into the embedding space to obtain the corresponding embedding vectors; S12: The event sequence is encoded according to equation (1) using a bidirectional gated loop unit to generate the initial event representation: (1); in, Encoding representing a sequence of events, This indicates a bidirectional gated loop unit. Representing events respectively , ... Embedded vector, Indicates the number of events.
4. The weakly supervised event representation learning method for decision support according to claim 1, characterized in that: In step S2, the normal prototype of the multi-hypersphere is constructed using the following method: Random settings The center of a hypersphere is represented by equation (2): (2); in: Indices representing the index of a hypersphere. Represent real numbers, Representing dimension, Indicates the center of the hypersphere. Represents the space of real numbers.
5. The weakly supervised event representation learning method for decision support according to claim 1, characterized in that: In step S3, the sequence encoder and the normal prototype of the hypersphere are trained using the following method to obtain the sequence encoder and the normal prototype of the hypersphere after the first stage parameter optimization: S31: During training, the distance from the representation of each normal event in the initial event representation to the center of each hypersphere is calculated according to equation (3): (3); in: The initial event is represented in the first... The first normal event to the first The Euclidean distance from the center of a hypersphere The initial event is represented in the first... The representation of a normal event. Indicates the center of the hypersphere. Indices representing the index of a hypersphere. This indicates the calculation of Euclidean distance; S32: Based on the distance from the representation of each normal event in the initial event representation to the center of each hypersphere, calculate the index from the representation of each normal event in the initial event representation to the nearest hypersphere center according to equation (4): (4); in: This represents the index of each normal event in the initial event representation to the nearest hypersphere center. This indicates a minimize operation. Indicates the total number of hyperspheres; S33: The multi-hypersphere center loss is calculated according to equation (5) based on the index of each normal event representation in the initial event representation to the nearest hypersphere center. (5); in: This indicates that the loss is more than the center of the sphere. This represents a normal sequence of events within an event sequence. This indicates the number of normal event sequences in the event sequence. The distance to the event represents the nearest hypersphere center. This indicates the calculation of the square Euclidean distance; S34: Calculate the separability loss of the hypersphere according to equation (6): (6); in: This indicates the loss of divisibility of the hypersphere. This indicates the calculation of binary cross-entropy loss. This indicates an exponential operation. This represents the hypersphere center that is second closest to the representation of each normal event in the initial event representation. This represents the index of the representation of each normal event in the initial event representation to the second nearest hypersphere; S35: Based on the loss of the multiple hypersphere centers and the loss of hypersphere separability, the loss function of the first stage is calculated according to equation (7): (7); in: This represents the loss function for the first stage. The weight representing the loss of separability of the hypersphere; S36: Perform iterative training with the goal of minimizing the loss function of the first stage to complete the first stage training of the sequence encoder and the normal prototype of the multi-hypersphere, and obtain the sequence encoder and the normal prototype of the multi-hypersphere with optimized parameters in the first stage.
6. The weakly supervised event representation learning method for decision support according to claim 1, characterized in that: In step S4, the prediction result of the event is obtained using the following method: S41: The embedding layer of the sequence encoder after the first-stage parameter optimization projects the events into the embedding space to obtain the first-stage optimized embedding vector; S42: The event sequence is encoded according to equation (8) using a bidirectional gated loop unit to generate the event representation after the first stage of optimization. (8); in, This represents the encoding of the event sequence after the first stage of optimization. This indicates a bidirectional gated loop unit. Representing events respectively , ... The embedding vector after the first stage of optimization Indicates the number of events; S43: Calculate the distance from the representation of each event in the first-stage optimized event representation to the center of each hypersphere in the first-stage parameter-optimized hypersphere normal prototype according to equation (9): (9); in: This represents the first stage of optimized event representation. The event, in the first stage of parameter optimization, in the normal prototype of the hypersphere, the [number]th event... The Euclidean distance from the center of a hypersphere This represents the event representation after the first stage of optimization. The representation of an event, This represents the first stage of parameter optimization in the normal prototype of the multi-hypersphere. The center of a hypersphere This indicates the calculation of Euclidean distance; S44: Calculate the index of the nearest hypersphere center in the normal prototype of the multi-hypersphere after the first-stage parameter optimization for each event representation in the first-stage optimized event representation according to equation (10): (10); in: This represents the index of each event in the first-stage optimized event representation to the nearest hypersphere center in the first-stage parameter-optimized hypersphere normal prototype. This indicates a minimize operation. Indicates the total number of hyperspheres; S45: The anomaly discriminator calculates the deviation score of each event in the event sequence based on the nearest hypersphere center according to equation (11): (11); in: This represents the deviation score of each event in the event sequence from the nearest hypersphere center in the multi-hypersphere prototype after parameter optimization in the first stage. Represents the first event in the event sequence. One event, Represents a sequence of events. This represents the event representation after the first stage of optimization. The representation of an event, The event distance from the first-stage optimization indicates the center of the hypersphere in the most recent first-stage parameter-optimized hypersphere normal prototype. S46: Calculate the anomaly discriminator score for each event in the event sequence according to equation (12): (12); in: This represents the anomaly discriminator score for each event in the event sequence. This represents the activation function. Represents the weight matrix. This indicates the first stage of optimization and enhancement. This event indicates that... Indicates bias. This represents the matrix transpose operation; S47: Calculate the anomaly score for each event in the event sequence according to equation (13): (13); in: This represents the anomaly score for each event in the event sequence. Indicates model parameters, Indicates the dual-score balance factor; S48: Compare the abnormal score of each event in the event sequence with a set threshold. If the abnormal score of each event in the event sequence is less than the set threshold, the prediction result of the event is judged to be normal. If the abnormal score of each event in the event sequence is greater than or equal to the set threshold, the prediction result of the event is judged to be abnormal.
7. The weakly supervised event representation learning method for decision support according to claim 1, characterized in that: In step S5, the prediction results of the event sequence are obtained based on the prediction results of the events using the following method: S51: Arrange the events in the event sequence from highest to lowest anomaly score, and determine the preceding events according to equation (14). Index of events: (14); in: Indicates the first In the sequence of events, the first Index of events, Indicates before calculation A function for an event. Indicates the first A sequence of events Indicates the first A sequence of events Abnormal score of events; S52: Based on Center front The index of each event is obtained according to equation (15). Prediction results: (15); in: Indicates the first Prediction results for a sequence of events.
8. A weakly supervised event representation learning method for decision support according to claim 7, characterized in that: In step S5, the sequence encoder, the normal prototype of the hypersphere, and the anomaly discriminator after parameter optimization in the second stage are obtained using the following method: S53: Introduce event sequence level labels as weak supervision signals, and calculate the loss function for the second stage using the event sequence level labels and the prediction results of the event sequence according to equation (16): (16); in: This represents the loss function for the second stage. A set representing a sequence of events. Indicates the number of event sequences. Indicates the first Event sequence level labels; S54: With the goal of minimizing the loss function of the second stage, train the sequence encoder, the normal prototype of the hypersphere, and the anomaly discriminator after the parameter optimization of the first stage to obtain the sequence encoder, the normal prototype of the hypersphere, and the anomaly discriminator after the parameter optimization of the second stage.
9. A weakly supervised event representation learning method for decision support according to claim 1, characterized in that: In step S6, the final prediction result of the event is obtained using the following method: S61: Encode the event sequence using the sequence encoder optimized by the second-stage parameters to generate the event representation optimized by the second stage; S62: Calculate the distance from the representation of each event in the second-stage optimized event representation to the center of each hypersphere in the second-stage parameter-optimized normal prototype of the multi-hypersphere; S63: Based on the distance from the representation of each event in the second-stage optimized event representation to the center of each hypersphere in the second-stage parameter-optimized normal prototype, calculate the index of the nearest hypersphere center in the second-stage parameter-optimized normal prototype for each event representation in the second-stage optimized event representation. S64: Based on the event representation optimized in the second stage and the index of each event representation in the second stage optimized event representation to the nearest hypersphere center in the normal prototype of the multi-hypersphere after the second stage parameter optimization, calculate the final deviation score and the final anomaly discriminator score for each event based on the nearest hypersphere center, and then further calculate the final anomaly score. S65: Screen abnormal events based on the final anomaly score to obtain the final prediction result of the event.
10. A weakly supervised event representation learning system for decision support, characterized in that: A weakly supervised event representation learning method for decision support as described in any one of claims 1 to 9, comprising a sequence encoder, a normal prototype of a multi-hypersphere, and an anomaly discriminator; The sequence encoder is used to form an event sequence from the user's operation behavior on the operation and maintenance network, and generate an initial event representation from the event sequence. The sequence encoder after the first stage parameter optimization is used to generate an event representation after the first stage optimization from the event sequence. The sequence encoder after the second stage parameter optimization is used to generate an event representation after the second stage optimization from the event sequence. The normal prototype of the multi-hypersphere includes multiple randomly set hypersphere centers. The normal prototype of the multi-hypersphere after the first stage parameter optimization includes the multi-hypersphere centers after the first stage optimization. The normal prototype of the multi-hypersphere after the second stage parameter optimization includes the multi-hypersphere centers after the second stage optimization. The anomaly discriminator is used to calculate the anomaly score of each event in the event sequence based on the event representation optimized in the first stage and the center of the hypersphere in the normal prototype of the hypersphere after parameter optimization in the first stage, so as to obtain the prediction result of the event. The anomaly discriminator, after second-stage parameter optimization, is used to calculate the final anomaly score for each event in the event sequence based on the event representation optimized in the second stage and the center of the hypersphere in the normal prototype of the hypersphere optimized in the second stage, thus obtaining the final prediction result of the event.