Detecting anomalies in processes executing on a device based on side-channel emissions
A neural network-based method for side-channel monitoring uses trace encoder and comparison models to accurately detect anomalies in device processes with minimal training data, addressing the challenge of frequent software updates.
Patent Information
- Application Number
- PCT/IB2024/051738
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-22
- Publication Date
- 2025-08-28
AI Technical Summary
Conventional one- or few-shot techniques for side-channel monitoring have inadequate accuracy in detecting anomalies due to the difficulty in collecting large amounts of training data, especially when 'normal' device behavior changes frequently due to software updates.
A neural network-based approach involving a trace encoder model and a trace comparison model is used to map side-channel measurements into informative representations, allowing for one-shot anomaly detection by comparing traces from the same process with subsequent measurements, utilizing unsupervised learning for trace encoder training and attention-based neural networks for trace similarity analysis.
Enables robust and accurate anomaly detection without the need for extensive training data, facilitating model updates with a single trace measurement, even when device behavior changes frequently.
Smart Images

Figure IB2024051738_28082025_PF_FP_ABST
Abstract
Description
[0001] DETECTING ANOMALIES IN PROCESSES EXECUTING ON A DEVICE BASED ON SIDE-CHANNEL EMISSIONS
[0002] TECHNICAL FIELD
[0003] The present disclosure relates generally to computing and / or communication devices, and more specifically to detecting anomalies in processes execution on devices based on capturing and analyzing traces of device side channel emissions.
[0004] BACKGROUND
[0005] The increasing complexity of communication networks, including 5G networks, drives the evolution of analytics systems that support operation, optimization, and planning of these networks. This includes detecting and addressing sudden, undesired changes in network operation and / or performance (e.g., failures). These analytics systems, in turn, require collecting and processing of enormous amounts of data, particularly time series data.
[0006] In general, a time series is a sequence of data or information values, each of which has an associated time instance (e.g., when the data or information value was generated and / or collected). The data or information can be anything measurable that depends on time in some way, such as prices, humidity, or number of people. One important characteristic of a time series is frequency, which is how often the data values of the data set are recorded. Frequency is also inversely related to the period (or duration) between successive data values.
[0007] Time series analysis includes techniques that attempt to understand or contextualize time series data, such as to make forecasts or predictions of future data (or events ) using a model built from past time series data. To best facilitate such analysis, it is preferrable that the time series consists of data values measured and / or recorded with a constant frequency or period.
[0008] Time series datasets can be collected from geographic locations, such as from nodes of a communication network located in one or more geographic areas (e.g., countries, regions, provinces, cities, etc.). For example, values of performance measurement (PM) counters can be collected from the various network nodes at certain time intervals. Time series data collected in this manner can be used to analyze, predict, and / or understand user behavior patterns as well as network performance trends.
[0009] Anomaly detection involves automatically analyzing a dataset - such as a time series dataset - to detect deviating samples. This operation may also be referred to as novelty detection, particularly when detecting deviations in new samples compared to a previously collected dataset. For some applications, however, it may be difficult to collect large amounts of training data, such that only one sample or a small number of samples are available as training data. This type of anomaly detection is often referred to as “one-shot” or “few-shot”, based on the number of training samples required to build a model for detection of anomalies (or novelties).
[0010] One application of anomaly detection is side-channel monitoring. Side-channel emissions (or leakage) by devices have been maliciously exploited by attackers to extract secrets. In general, side-channel emissions are via a non-intended information channel from a device to its surrounding environment. For example, side-channel emissions may include energy consumption, electromagnetic (EM) emissions, thermal signatures, sound emissions, and optical emissions. An attacker can utilize these leakages to extract sensitive information from a device.
[0011] One example attack based on side-channel emissions is extracting cryptographic keys from cryptographic algorithms implemented in software executing on subscriber identity modules (SIMs) and processors. Another example attack is on hardware implementations of cryptographic algorithms, such as Advanced Encryption Standard (AES) and Post-Quantum Cryptography (PQC) candidates such as Saber and KYBER. Side-channel emissions can also be used to reverse engineer software running on a processor.
[0012] In side-channel monitoring, an external monitor registers side-channel emissions from the device and concludes if the device behaves normally according to pre-defined criteria. In some solutions, the monitor is oblivious to the internal state of the device and only determines whether side-channel emissions are “normal” or “abnormal”. In other solutions, the monitor is aware of certain states or operations of the device that correspond to certain side-channel emissions. In this case, the monitor may also detect “illegal” state transitions where the execution flow of the device is abnormal.
[0013] SUMMARY
[0014] Even so, there are various challenges for side channel monitoring of device emissions for anomaly detection. For example, it may be difficult to collect a large amount of side-channel training data needed to build a robust monitoring model, particularly when “normal” device behavior changes frequently due to regular software updates. Thus, it is desirable to employ one- or few-shot techniques for side-channel monitoring. However, conventional one- or few-shot techniques have inadequate accuracy for detecting anomalies (or novelties) in side-channel monitoring.
[0015] An object of embodiments of the present disclosure is to provide accurate and reliable techniques for one- or few-shot side-channel monitoring of device emissions, thereby improving ability to detect attacks on and / or unauthorized (or malicious) operation of devices.
[0016] Some embodiments include methods (e.g., procedures) for detecting anomalies in processes executing on a device. Such methods can be performed, for example, by a monitoring device. For example, the monitoring device may be internal to the device being monitored, external but connected to the device being monitored, or external and not connected to the device being monitored.
[0017] These exemplary methods include capturing first and second traces of side channel information. The first trace is associated with execution of a first process and the second trace is associated with execution of a process on the device. These exemplary methods also include, using a trained side-channel encoder, determining encoded representations of the first and second traces. These exemplary methods also include, using a trained trace comparison model, determining whether the encoded representations of the first and second traces are associated with execution of a same process.
[0018] In some embodiments, these exemplary methods also includes the following operations:
[0019] • training the side-channel encoder based on a first plurality of traces of side-channel information;
[0020] • using the trained side-channel encoder, determining encoded representations of a second plurality of traces of side-channel information associated with the device or with other devices; and
[0021] • training the trace comparison model based on the encoded representations of the second plurality of traces.
[0022] In some of these embodiments, the second plurality of traces comprises:
[0023] • a plurality of first trace pairs, wherein each first trace pair includes two traces associated with execution of the first process; and
[0024] • a plurality of second trace pairs, wherein each second trace pair includes one trace associated with execution of the first process and one trace associated with execution of a different process than the first process.
[0025] In some of these embodiments, training the side-channel encoder in block 910 is performed using unsupervised machine learning and the traces of the first plurality have no known association with execution of any specific processes (i.e., unlabelled). In other of these embodiments, training the side-channel encoder is performed using supervised machine learning and each encoded representation of a trace of the second plurality is labelled with an associated process executing on the device. In some embodiments, an encoded representation is a more compact trace representation that indicates information-carrying characteristics of the side channel from the device. In some embodiments, the trace comparison model is an attention-based neural network.
[0026] In some embodiments, the encoded representations of the first and second traces are respective first and second vectors (vl, v2), each comprising a plurality (L) of data elements. In such case, determining whether the encoded representations of the first and second traces are associated with execution of a same process in block 970 includes the following operations:
[0027] • determining a plurality (L) of first difference vectors based on the second vector and respective data elements of the first vector;
[0028] • determining a plurality (L) of second difference vectors based on the first vector and respective data elements of the second vector; and
[0029] • applying the plurality (L) of first difference vectors to a first multi-head attention function; and
[0030] • applying the plurality (L) of second difference vectors to a second multi-head attention function.
[0031] In some of these embodiments, each of the first and second multi-head attention functions comprises a plurality (L) of single-head attention functions.
[0032] In some embodiments, a plurality of the first traces of side channel information associated with execution of the first process are captured and a corresponding plurality of encoded representations of the first traces are determined. In such case, determining whether the encoded representations of the first and second traces are associated with execution of a same process is based on the plurality of encoded representations of the first traces.
[0033] In other embodiments, a plurality of the second traces of side channel information associated with the device are captured and a corresponding plurality of encoded representations of the second traces are determined. In such case, determining whether the encoded representations of the first and second traces are associated with execution of a same process is based on the plurality of encoded representations of the second traces.
[0034] Other embodiments include monitoring devices (e.g., user equipment, circuit modules, loT devices, computing devices, etc.) configured to perform operations corresponding to any of the exemplary methods described herein. Other embodiments include non-transitory, computer- readable media storing program instructions that, when executed by processing circuitry, configure such monitoring devices to perform operations corresponding to any of the exemplary methods described herein.
[0035] These and other embodiments described herein may provide various benefits and / or advantages. For example, embodiments may provide robust and accurate anomaly detection without a need to collect large amounts of training data, which may be beneficial for side-channel monitoring of devices whose “normal” device behavior changes frequently due to regular software updates. Instead, embodiments may facilitate updating a model for emission behavior with a single trace measurement after new device functionality (e.g., update) has been deployed.
[0036] These and other objects, features, and advantages of embodiments of the present disclosure will become apparent upon reading the following Detailed Description in view of the Drawings briefly described below.
[0037] BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 shows an exemplary training procedure for a side channel trace encoder model, according to some embodiments of the present disclosure.
[0039] Figure 2 shows an exemplary training procedure for a dual trace comparison model, according to some embodiments.
[0040] Figure 3 shows a high-level view of a neural network architecture, according to embodiments of the present disclosure.
[0041] Figures 4-5 show certain features of the architecture of Figure 3 in more detail, according to some embodiments of the present disclosure.
[0042] Figure 6 depicts an exemplary register and comparison phase, according to some embodiments.
[0043] Figures 7A-B show experimental results that illustrate performance of some embodiments of the present disclosure.
[0044] Figures 8A-C illustrate different arrangements of a monitoring device, according to various embodiments of the present disclosure.
[0045] Figure 9 shows an exemplary method (e.g., procedure) for detecting anomalies in processes executing on a device, according to various embodiments of the present disclosure.
[0046] Figure 10 shows a monitoring device according to various embodiments of the present disclosure.
[0047] DETAILED DESCRIPTION
[0048] Some of the embodiments contemplated herein will now be described more fully with reference to the accompanying drawings. Other embodiments, however, are contained within the scope of the subject matter disclosed herein, the disclosed subject matter should not be construed as limited to only the embodiments set forth herein; rather, these embodiments are provided by way of example to convey the scope of the subject matter to those skilled in the art.
[0049] In general, all terms used herein are to be interpreted according to their ordinary meaning to a person of ordinary skill in the relevant technical field, unless a different meaning is expressly defined and / or implied from the context of use. All references to a / an / the element, apparatus, component, means, step, etc. are to be interpreted openly as referring to at least one instance of the element, apparatus, component, means, step, etc., unless explicitly stated otherwise or clearly implied from the context of use. The operations of any methods and / or procedures disclosed herein do not have to be performed in the exact order disclosed, unless an operation is explicitly described as following or preceding another operation and / or where it is implicit that an operation must follow or precede another operation. Any feature of any embodiment disclosed herein can apply to any other disclosed embodiment, as appropriate. Likewise, any advantage of any embodiment described herein can apply to any other disclosed embodiment, as appropriate.
[0050] Note that some aspects of the description given herein may focus on a 3GPP cellular communications system and, as such, 3GPP terminology or terminology similar to 3GPP terminology may be used. However, the concepts disclosed herein are not limited to a 3GPP system and can be applied to any communication system that may benefit from them.
[0051] As briefly mentioned above, side-channel emissions (or leakage) by devices have been maliciously exploited by attackers to extract secrets. In general, side-channel emissions are via a non-intended information channel from a device to its surrounding environment. For example, side-channel emissions may include energy consumption, EM emissions, thermal signatures, sound emissions, and optical emissions.
[0052] An attacker can utilize these leakages to extract sensitive information from a device. Side-channel emissions can also be used to reverse engineer software running on the device. Some side-channel leakage types, including EM, thermal, and optical emissions, are sensitive to probe placement. In such case, it is important to measure in a consistent position relative to the device to get correlated data from two different measurement suites.
[0053] In side-channel monitoring, an external monitor registers side-channel emissions from the device and concludes if the device behaves “normally” according to pre-defined criteria. In some solutions, the monitor is oblivious to the internal state of the device and only determines whether side-channel emissions are “normal” or “abnormal”. In other solutions, the monitor is aware of certain states or operations of the device that correspond to certain side-channel emissions. In this case, the monitor may also detect “illegal” state transitions where the execution flow of the device is abnormal. For example, such solutions can be used for detecting malicious or unauthorized software execution on a device.
[0054] Advantages of side-channel monitoring include that it is very hard to avoid or intentionally shape side-channel emissions, and that a monitor may be physically separated from the device. This makes it very hard for an attacker to remain undetected, as an attack on a device will unavoidably change the side channel leakage. This is beneficial in both high-security environments and as a complement to “classic” monitoring solutions. Low-cost generic hardware for side channel measurements is now available, such as the NewAE’s ChipWhisperer series.
[0055] Most commercially available devices are “closed source”, such that the owner or user has little or no possibility to install software-based intrusion detection capabilities. In such case, external side channel monitoring may be used with the device as a “black box” that has no malware detection capabilities and is unaware of the side-channel monitoring. Side channel monitoring may be used to detect various types of malware on a device. For example, a device infected by malware can act normally towards a controller by detaching the internal measurements with the status data supplied to the controller. Stuxnet is a well-known example of this type of sophisticated malware containing code that faked sensor signals so that a targeted system would not shut down due to abnormal behavior. This specific malware may have been designed to target Uranium enrichment facilities. By forcing the system sensors to report normal behavior, the controller failed to safeguard the system from operations that were detrimental to the hardware, in this case centrifuges used to separate nuclear material to such an extent that the centrifuges broke. Up to 1000 centrifuges were reportedly destroyed by the malware during a period of a few years.
[0056] Since Stuxnet and similar malware install root kits, they affect side channel emission patterns from the attacked component. Creating malware that manages to perform their original target task without creating detectable changes in the side channel emissions would be an extremely complex problem, and while some limited academic progress has been made in this field, it is unlikely that such attacks will appear in the near future.
[0057] Another threat, especially Internet of Things (loT) devices, is that a device may contain a software vulnerability or have a weak authentication setup. This may be exploited by an adversary to infect the device with malware that, unbeknownst to the owner, makes the device part of a botnet. This type of attack utilizes the device’s compute resources to perform malicious activities such denial-of- service (DoS) attacks and malware distribution to other devices. The Mirai botnet is an example of such malware that infected approximately 400000 devices. This type of malware infection is invisible to a controller (e.g., smart home hub) since the device seems to function normally, albeit with additional network traffic and some reduced performance.
[0058] As briefly mentioned above, anomaly detection involves automatically analyzing a dataset - such as a time series dataset - to detect deviating samples. This operation may also be referred to as novelty detection, particularly when detecting deviations in new samples compared to a previously collected dataset. Side-channel monitoring is one use case for novelty detection, in which newly obtained side channel measurements can be compared to previously obtained side channel measurements (or a model trained on such measurements) to detect significant and unexpected deviations.
[0059] At a high level, anomaly detection techniques can be classified as one of three types. Proximity-based techniques use a “proximity function” to compare two sequences, and then use pairwise comparisons to determine if a given sequence is similar to sequences in a training set. Prediction-based techniques learn a time series distribution from which future data points are predicted, then determine deviations between predicted and actual collected data points to determine if a sequence is anomalous.
[0060] Reconstruction-based techniques typically using neural networks to train a model to reconstruct samples from the training dataset, and then finds anomalies by comparing the reconstructed samples to actual collected samples. Examples of reconstruction-based techniques include autoencoders and Generative Adversarial Networks (GANs). Further description of GAN-based techniques can be found in “Tadgan: Time series anomaly detection using generative adversarial networks" by Geiger, et al., published in Proceedings of 2020 IEEE International Conference on Big Data.
[0061] Unsupervised learning is used in well-known large generative models for language, image, and audio data. One aspect of these large models is the ability to learn powerful data representations. In contrastive learning, the aim is to embed similar samples close to each other while dissimilar samples become further apart. This requires the ability to create positive and negative pairs. “Stochastic pairing for contrastive anomaly detection on time series” by Chambaret, et al. (published in Proceedings of Pattern Recognition and Artificial Intelligence - part II, pp. 306-317, 2022) describes data augmentation techniques that obtain positive pairs with an original and an augmented time-series sample and negative pairs by distance comparison, choosing the most distant samples according to a dynamic time warping-based function.
[0062] For some applications, however, it may be difficult to collect large amounts of training data, such that only one sample or a small number of samples are available as training data. This type of anomaly detection is often referred to as “one-shot” or “few-shot”, based on the number of training samples required to build a model for detection of anomalies (or novelties). For example, “FewSOME: One-Class Few Shot Anomaly Detection with Siamese Networks Niamh Belton” by Hagos, et al. (published in Computer Vision and Patter Recognition, 2023) describes a technique in which 30-60 samples are used to learn a representation that provides a small Euclidean distance between these samples of the normal class.
[0063] For practical reasons, one- or few-shot techniques may be needed for anomaly detection during side channel monitoring of device emissions. For example, it may be difficult to collect a large amount of side-channel training data needed to build a robust monitoring model, particularly when “normal” device behavior changes frequently due to regular software updates. However, conventional one- or few-shot techniques have inadequate accuracy for detecting anomalies (or novelties) in side-channel monitoring.
[0064] Embodiments of the present disclosure address these and other problems, issues, and / or difficulties by novel, flexible, and efficient techniques and architecture for a neural network trained to be a trace-similarity function. The proposed architecture and training techniques facilitate one-shot modelling so that a single side-channel trace of a process can be compared with a subsequent measurement to reliably detect unexpected deviations in side-channel traces from what is supposedly the same process.
[0065] At a high level, embodiments of the disclosed techniques involve three phases. First, embodiments utilize a large amount of un-labeled side-channel traces from many different processes to learn an embedded representation using methods similar to those used for learning pre-trained representations for audio, images or language models. Second, embodiments utilize a large number of side-channel traces from different processes with more than one trace from each process to learn a similarity function. Third, embodiments utilize the learned similarity function in one-shot anomaly detection by comparing a first obtained trace (representing current version of a process) to a subsequent trace from the same process, and output an indication of whether the compared traces are measurements of the same process or of different processes.
[0066] Embodiments may provide various benefits and / or advantages. For example, embodiments provide robust and accurate anomaly detection without a need to collect large amounts of training data, which is beneficial for side-channel monitoring of devices whose “normal” device behavior changes frequently due to regular software updates. Rather, embodiments facilitate updating a model for emission behavior with a single trace measurement after new device functionality (e.g., update) has been deployed. These advantages are facilitated by a similarity function representing how multiple measurements of the same process relate to each other compared to measurements from two different processes.
[0067] At a high level, some embodiments include the following:
[0068] • Trace encoder model: trained to map the side-channel measurements into a representation highlighting the information-carrying characteristics of the side-channel;
[0069] • Trace comparison model (also referred to as “dual trace comparison model”): trained using the trace encoder model as pre-processing, to determine whether a side-channel measurements is collected from a specific process; and
[0070] • Monitoring system configured with the trace encoder model and the trace comparison model, and including collection equipment (e.g., probe) arranged to capture a new trace for input to the dual trace comparison model.
[0071] As briefly mentioned above, embodiments of the techniques involve three phases, which will be referred to as trace encoder model training phase, trace comparison model training phase, and register and comparison phase. Alternately, these three phases can be thought of as functional blocks of a monitoring system or monitoring device.
[0072] In some embodiments, the trace encoder model can be trained to map a side-channel measurement into a more meaningful representation, using unsupervised learning training techniques that do not require data labels. Unsupervised learning may facilitate use of much larger quantities of data, since labelling of data can be costly and / or time consuming. More specifically, training based on unsupervised learning means that side-channel measurements can be used without knowing the corresponding process that generated the trace. Furthermore, , training based on unsupervised training may use data captured by different hardware setups or data generated by a larger variety of processes than of interest for anomaly detection.
[0073] One example of training based on unsupervised learning is the data2vec framework described in "data2Vec: A General Framework for Self-Supervised Learning in Speech, Vision, and Language” by Baevski, et al., published in Proceedings of 39tnInternational Conference on Machine Learning, 2022. At a high level, in the data2vec framework, samples are randomly masked and the model is trained to predict the missing (masked) samples.
[0074] More specifically, the model is trained by predicting the model representations of the full input data given a partial view of the input. First, a masked version of the training sample (“student mode”) is encoded and then training targets are constructed by encoding the unmasked version of the input data with the same model but when parameterized as an exponentially moving average of the model weights (“teacher mode”). The target representations encode all information in the training sample and the learning task is for the student to predict these representations given a partial view of the input.
[0075] In some embodiments, the trace encoder model may be jointly trained with the trace comparison model to obtain an encoding that is more optimal for distinguishing between multiple traces from the same process and multiple traces from different and / or modified processes.
[0076] Figure 1 depicts an exemplary training procedure for a side channel trace encoder model, according to some embodiments of the present disclosure. As noted above, the training is based on unlabeled data, such that the device process(es) (110) that generate a particular trace (120) captured by the monitoring equipment is(are) unknown. The captured trace is input to the sidechannel trace unsupervised pre-training module (130), which outputs the trace encoder model (140) trained to map the side-channel measurements into a representation of the informationcarrying characteristics of the side-channel.
[0077] In some embodiments, the trace comparison model training data includes positive pairs of traces captured from the same underlying process and / or negative pairs from what are considered to be different underlying processes. In some embodiments, depending on the desired sensitivity, the negative pairs can also be traces from execution of the same underlying process but one of the traces also involves execution of other code (e.g., another process).
[0078] The training data for the trace comparison model is labeled with two classes: a first class (e.g., zeroes) for negative pairs and a second class (e.g., ones) for positive pairs. The trace comparison model can thus be trained with a conventional loss function used for classification problems, such as a cross-entropy loss function of the log-probability of the correct class. This loss function gives a high loss when the correct class has a low probability. The output of this training will be a binary model that indicates, for a pair of traces, whether the pair are obtained from measurements of the same process or from measurements of different processes.
[0079] Figure 2 depicts an exemplary training procedure for a dual trace comparison model, according to some embodiments. As noted above, the training is based on labeled data. Different processes (210) are run to generate trace pairs (220) captured by the monitoring equipment, while two runs of the same process (230) generate other trace pairs (240) captured by the monitoring equipment. The labelled traces are input to the trained trace encoder model (140), which outputs encoded trace pairs with labels (i.e., same / different) to be used for dual trace comparison model training (250).
[0080] In some embodiments, the trace comparison model includes or is based on an attentionbased neural network (NN) architecture that processes encoded representations output by the trace encoder model trained according to the process described above.
[0081] Self-attention (or “intra-attention”) relates different positions (or entries) of a single sequence in order to compute a representation of the sequence. Self-attention has been used successfully in a variety of tasks including reading comprehension, abstractive summarization, textual entailment, and learning task-independent sentence representations. An “attention function” maps a query and a set of key-value pairs to an output, where the query, keys, values, and output are all vectors. The output is computed as a weighted sum of the values, where the weight assigned to each value is computed by a compatibility function of the query with the corresponding key. More details about attention functions can be found in “Attention is All You Need” by Vaswani, et al., published in Advances in Neural Information Processing Systems, 2017.
[0082] Figure 3 shows a high-level view of a NN architecture (300) according to embodiments of the present disclosure. The architecture expects as input two encoded sequences (xl, x2) of a fixed length L, but this does not imply that the underlying measurements need to be of the same length. Similarly to common practice in classification of audio or text sequences of varying length, an attention mask can be applied to one or both of the sequences to generate equal-length subsequences for comparison.
[0083] In the same way that attention enables a network to learn and quantify importance of different contexts, embodiments utilize attention based on differences in subtraction between different elements in two representations. Initially, the two encoded sequences (xl, x2) are input to two identical CompareSequenceB locks (310, 320), in different orders: (xl, x2) to block 310, (x2, xl) to block 320. Each CompareSequenceBlock has L CrossSequenceAttention heads (e.g., 321) for the respective elements of the sequence. The vector input to the i-th CrossSequenceAttention head is given by: xl(j)-x2(i), j = 1...L, for block 310; and x2(j)-xl(i), j = 1 . . .L, for block 320.
[0084] This creates a representation in which the differences between each element are encoded into one output vector representation that is input to a corresponding CrossSequenceAttention head. Figure 4 further illustrates operation of CompareSequenceBlock (x2, xl) and the corresponding i-th CrossSequenceAttention head (i.e., blocks 320 and 340) from Figure 3. The dimension-E input vector {x2(j)-xl(i), j = 1...L} is mapped into dimension-d query (Q), key (K), and value (V) vectors using projection matrices MQ, MK, and Mv, respectively. The structures of these matrices may be specified according to implementation requirements, while the matrix contents (referred to as “weights”) are learned or adapted by back-propagation during the model training phase. The i-th CrossSequenceAttention head output is computed by: which is a dimension-d attention vector. Note that the softmax function takes as input a vector of d real numbers and normalizes it into a probability distribution consisting of d probabilities that are proportional to the exponentials of the input numbers.
[0085] Figure 5 further illustrates operation of the multihead attention (330, 340) and concatenation (“concat”, 350) blocks of Figure 3. As noted above, each of the L CrossSequenceAttention heads generates a dimension-d attention vector, which can be thought of as a projection into the internal model dimension d. The outputs of the L CrossSequenceAttention heads are concatenated into a dimension-(d-L) vector, before being projected back to a vector of the original input dimension L.
[0086] Returning to Figure 3, the output of the concatenation block is projected through fully connected layers (360, 370) to an output dimensionality of two, corresponding to classes SAME and DIFFERENT. In this case, the softmax function in block 370 takes as input a vector of two real numbers and normalizes it into a probability distribution consisting of two probabilities that are proportional to the exponentials of the input numbers. Here, the probabilities are of the two possible outcomes of SAME and DIFFERENT.
[0087] In some embodiments, during the register and comparison phase, the trained trace comparison model is used to compare two side-channel traces from the same process to monitor if the process has changed. In some variants, the trace comparison model resides in a base unit that can be connected to or embedded in the trace-capturing monitoring device. In such variants, if no trace is available that matches a trace from the current process running on the device being monitored, then the non-matching trace is simply registered as a model for the process.
[0088] Figure 6 depicts an exemplary register and comparison phase, according to some embodiments. In operation 610, an initial run of process A is performed, resulting in a captured trace (620) that is processed by the trace encoder model (140). In operation 630, the output of the trace encoder model is registered for process A (e.g., as a representation of “normal” process behavior). In operation 640, a subsequent run of what is believed to be process A is performed, resulting in a captured trace (650) that is processed by the trace encoder model (140).
[0089] This output of the trace encoder model and the encoded trace registered for process A in operation 630 are input to the dual trace comparison model (250), which compares the two traces and generates probabilities of the two possible outcomes, i.e., SAME and DIFFERENT (e.g., different processes, original and modified processes). Alternatively, the trace comparison model may only indicate SAME or DIFFERENT, e.g., based on which of the outcomes are determined to have higher probability of occurrence.
[0090] In some embodiments, there may be multiple encoded traces that represent “normal” process behavior, also referred to as “golden samples.” The monitoring system may handle these multiple golden samples in different ways, described below.
[0091] In some variants, the monitoring system may compare an encoded captured trace to all golden samples, with the trace comparison model providing an output for each captured trace / golden sample pair. For example, the monitoring system may generate an alarm based on a threshold number of comparisons indicating DIFFERENT, where the threshold number may range from one to the number of comparisons, according to implementation requirements or preferences. As a more specific example, the monitoring system may use a “voting strategy” in which the outcome of the majority of comparisons is deemed to be the final result.
[0092] In other variants, the monitoring system may average the multiple golden samples, with the average provided as input to the trace comparison model. In other variants, the monitoring system may identify “stable” parts that are relatively constant or consistent across all of the golden samples, and weight these stable parts more heavily in the comparison than other less stable parts.
[0093] The following describes an experiment that illustrates performance of embodiments of the present disclosures. In a first experiment, 1000 traces were collected for each of three different processes, with each trace containing 400 measurement values. All 3000 traces where used for unsupervised trace encoder model training, using the data2vec training framework (mentioned above) with measurement trace inputs rather than audio inputs used previously. The output representation was an eight-element sequence of 32-dimensional points. In other words, the initial 400-element sequence was subdivided into eight 50-element windows, with each window being embedded into a 32-dimensional point. The training ran for 400 epochs and the best model on the validation set was chosen as the final trace encoder model.
[0094] For the training of the dual trace comparison model, new datasets of trace -pairs were collected. First, the original traces were randomly split into a training set (70%) and a test set (30%), with each of these sets being independently and randomly augmented into trace-pair datasets. For the negative class, a cut and paste strategy was used in which two traces were randomly chosen from one process and one trace was randomly chosen from another process. One of the traces from the same process was then modified with parts of traces from the other process, using randomly selected parts of the trace from the other process. The training set consisted of 20000 trace pairs constructed in this manner. For the positive class, 20000 pairs of traces from the same process were chosen at random for the training set.
[0095] The test set included trace pairs selected randomly from the test set (with no overlapping traces), including 2500 positive pairs and 2500 negative pairs respectively. After the trace comparison model was trained with conventional cross-entropy loss function for 250 epochs, an detection accuracy of 98.13 % was observed from running the trained model on the test set.
[0096] To further illustrate performance of the trace comparison model in this experiment, Figure 7 A shows a comparison of 400 measurement values from two test set traces that the model determined to be SAME (i.e., generated by the same process). In contrast, Figure 7B shows a comparison of 400 measurement values from two test set traces that the model determined to be DIFFERENT (i.e., generated by processes that differ in some manner). In both figures, the two test set traces are represented by solid and dashed line, respectively.
[0097] Figures 8A-C shows various arrangements of a monitoring device (820) configured to detect unexpected deviations in processes executing on a device (810), using the various techniques described above. In some embodiments, as illustrated in Figure 8A, the monitoring device may be external but connected to the device. In these embodiments, the monitoring device may capture traces of side channel information via the connection (e.g., wire) to the device. In other embodiments, as illustrated in Figure 8B, the monitoring device may be external and not connected to the device. In these embodiments, the monitoring device may capture traces of side channel information emitted by the device, such as by using an monitoring receiver capable of receiving electromagnetic emissions by the device at one or more frequencies.
[0098] In other embodiments, as illustrated in Figure 8C, the monitoring device may be internal to the device, such as a module or circuitry. In these embodiments, the monitoring device may capture traces of side channel information via internal connection to other circuitry of the device, or by receiving internal electromagnetic emissions by the device at one or more frequencies. Although not shown, Various features of the embodiments described above correspond to various operations illustrated in Figure 9, which depicts an exemplary method (e.g., procedure) for detecting anomalies in processes executing on a device, according to various embodiments of the present disclosure. In other words, various features of the operations shown in Figure 9 and described below correspond to various embodiments described above. Although Figure 9 shows specific blocks in a particular order, the operations of the exemplary method can be performed in a different order than shown and can be combined and / or divided into blocks having different functionality than shown. Optional blocks or operations are indicated by dashed lines.
[0099] The following description is based on the exemplary method being performed by a monitoring device. For example, as discussed above, the monitoring device may be internal to the device being monitored (e.g., a module or circuit of the device), external but connected to the device being monitored (e.g., via wired or wireless connection), or external and not connected to the device being monitored.
[0100] The exemplary method includes the operations of block 940, where the monitoring device captures first and second traces of side channel information. The first trace is associated with execution of a first process and the second trace is associated with execution of a process on the device. The exemplary method also includes the operations of block 950, where using a trained side-channel encoder, the monitoring device determines encoded representations of the first and second traces. The exemplary method also includes the operations of block 970, where using a trained trace comparison model, the monitoring device determines whether the encoded representations of the first and second traces are associated with execution of a same process.
[0101] In some embodiments, the exemplary method also includes the following operations, labelled with corresponding block numbers:
[0102] • (910) training the side-channel encoder based on a first plurality of traces of side-channel information;
[0103] • (920) using the trained side-channel encoder, determining encoded representations of a second plurality of traces of side-channel information associated with the device or with other devices; and
[0104] • (930) training the trace comparison model based on the encoded representations of the second plurality of traces.
[0105] In some of these embodiments, the second plurality of traces comprises:
[0106] • a plurality of first trace pairs, wherein each first trace pair includes two traces associated with execution of the first process; and
[0107] • a plurality of second trace pairs, wherein each second trace pair includes one trace associated with execution of the first process and one trace associated with execution of a different process than the first process.
[0108] In some of these embodiments, training the side-channel encoder in block 910 is performed using unsupervised machine learning (ML), and the traces of the first plurality have no known association with execution of any specific processes (i.e., unlabelled). In some of these embodiments, training the side channel encoder in block 910 includes, for each trace of the first plurality, the following operations labelled with corresponding sub-block numbers:
[0109] • (911) determining one or more partially masked versions of the trace; and
[0110] • (912) training the side channel encoder to predict the trace from the one or more partially masked versions of the trace.
[0111] In other of these embodiments, training the side-channel encoder is performed using supervised machine learning (ML), and each encoded representation of a trace of the second plurality is labelled with an associated process executing on the device.
[0112] In some embodiments, an encoded representation is a more compact trace representation that indicates information-carrying characteristics of the side channel from the device. In some embodiments, the trace comparison model is an attention-based neural network.
[0113] In some embodiments, the encoded representations of the first and second traces are respective first and second vectors (vl, v2), each comprising a plurality (L) of data elements. In such case, determining whether the encoded representations of the first and second traces are associated with execution of a same process in block 970 includes the following operations, labelled with corresponding block numbers:
[0114] • (971) determining a plurality (L) of first difference vectors based on the second vector and respective data elements of the first vector;
[0115] • (972) determining a plurality (L) of second difference vectors based on the first vector and respective data elements of the second vector; and
[0116] • (973) applying the plurality (L) of first difference vectors to a first multi-head attention function; and
[0117] • (974) applying the plurality (L) of second difference vectors to a second multi-head attention function.
[0118] In some of these embodiments, each of the first and second multi -head attention functions comprises a plurality (L) of single-head attention functions. In such case, applying the plurality (L) of first difference vectors to the first multi-head attention function in sub-block 973 includes, for each first difference vector, mapping the first difference vector to query (Q), key (K), and value (V) vectors of a corresponding single -head attention function. Also, applying the plurality (L) of second difference vectors to the second multi-head attention function in sub-block 974 includes, for each second difference vector, mapping the second difference vector to query (Q), key (K), and value (V) vectors of a corresponding single-head attention function.
[0119] In some variants of these embodiments, applying the plurality (L) of first difference vectors to the first multi-head attention function in sub-block 973 also includes, for each first difference vector, calculating an output of the corresponding single-head attention function based on a function of the mapped query (Q), key (K), and value (V) vectors. Also, applying the plurality (L) of second difference vectors to the second multi-head attention function in subblock 974 includes, for each second difference vector, calculating an output of the corresponding single -head attention function based on the function of the mapped query (Q), key (K), and value (V) vectors. In some further variants, the function of the mapped query (Q), key (K), and value (V) vectors is: wherein d is a dimension of each of the mapped query (Q), key (K), and value (V) vectors.
[0120] In some further variants, determining whether the encoded representations of the first and second traces are associated with execution of a same process in block 970 also includes the following operations, labelled with corresponding sub-block numbers:
[0121] • (975) concatenating the outputs of the respective single -head attention functions in a concatenated vector; and
[0122] • (976) applying the concatenated vector to one or more fully connected layers of a neural network, which outputs at least one of the following: a first probability that the first and second traces are associated with execution of the same process, and a second probability that the first and second traces are associated with execution of different processes.
[0123] For example, the neural network may only output the first probability, and the second probability may be determined by subtracting the first probability from one.
[0124] In some embodiments, a plurality of the first traces of side channel information associated with execution of the first process are captured (e.g., in block 940) and a corresponding plurality of encoded representations of the first traces are determined (e.g., in block 950). In such case, determining whether the encoded representations of the first and second traces are associated with execution of a same process is based on the plurality of encoded representations of the first traces. These plurality of encoded representations of the first traces can be considered “golden samples”, as discussed above.
[0125] In some of these embodiments, determining whether the encoded representations of the first and second traces are associated with execution of a same process executing on the device in block 970 includes the following operations, labelled with corresponding sub-block numbers:
[0126] • (977) determining, for each of the plurality of first traces, an indication of whether the encoded representation of the second trace is associated with execution of a same process as the encoded representation of the first trace; and
[0127] • (978) determining whether the second trace is associated with execution of a same process as the plurality of first traces based on the plurality of indications.
[0128] In some variants of these embodiments, it is determined in sub-block 978 that second trace is associated with execution of a same process as the plurality of first traces when one of the following indicate that the second trace is associated with execution of a same process as the plurality of first traces: one indication, all indications, a majority of indications, a threshold number of indications that is greater than one but less than all.
[0129] In other of these embodiments, the exemplary method also includes the operations of block 960, where the monitoring device combines the plurality of encoded representations of the first traces into a combined first trace representation. In such case, determining whether the encoded representations of the first and second traces are associated with execution of a same process is based on the combined first trace representation.
[0130] In some variants of these embodiments, the plurality of encoded representations of the first traces are combined into the combined first trace representation based on element-by- element averaging, and determining whether the encoded representations of the first and second traces are associated with execution of a same process is based on one of the following weightings for elements of the combined first trace representation:
[0131] • equal weighting for all elements; or
[0132] • element weighting based on similarity of elements from the plurality of encoded representations of the first traces that were averaged to determine the element.
[0133] In other embodiments, a plurality of the second traces of side channel information associated with the device are captured and a corresponding plurality of encoded representations of the second traces are determined. In such case, determining whether the encoded representations of the first and second traces are associated with execution of a same process in block 970 is based on the plurality of encoded representations of the second traces.
[0134] Although various embodiments are described herein above in terms of methods, apparatus, devices, computer-readable medium and receivers, the person of ordinary skill will readily comprehend that such methods can be embodied by various combinations of hardware and software in various systems, communication devices, computing devices, control devices, apparatuses, non-transitory computer-readable media, etc.
[0135] Figure 10 shows a monitoring device 1000 in accordance with some embodiments of the present disclosure. Examples of a monitoring device include, but are not limited to, a smart phone, mobile phone, cell phone, voice over IP (VoIP) phone, wireless local loop phone, desktop computer, personal digital assistant (PDA), wireless cameras, gaming console or device, music storage device, playback appliance, wearable terminal device, wireless endpoint, mobile station, tablet, laptop, laptop-embedded equipment (LEE), laptop-mounted equipment (LME), smart device, wireless customer-premise equipment (CPE), vehicle, vehicle-mounted or vehicle embedded / integrated wireless device, etc. Other examples include any user equipment (UE) identified by 3GPP, including a narrow band internet of things (NB-IoT) UE, a machine type communication (MTC) UE, and / or an enhanced MTC (eMTC) UE.
[0136] A monitoring device may support device-to-device (D2D) communication, for example by implementing a 3GPP standard for sidelink communication, Dedicated Short-Range Communication (DSRC), vehicle-to-vehicle (V2V), vehicle-to-infrastructure (V2I), or vehicle-to- everything (V2X). In other examples, a monitoring device may not necessarily have a user in the sense of a human user who owns and / or operates the relevant device. Instead, a monitoring device may represent a device that is intended for sale to, or operation by, a human user but which may not, or which may not initially, be associated with a specific human user. Alternatively, a monitoring device may represent a device that is not intended for sale to, or operation by, an end user but which may be associated with or operated for the benefit of a user.
[0137] Monitoring device 1000 includes processing circuitry 1002 that is operatively coupled via a bus 1004 to an input / output interface 1006, a power source 1008, a memory 1010, a communication interface 1012, and / or any other component, or any combination thereof. Certain monitoring devices may utilize all or a subset of the components shown in Figure 10. The level of integration between the components may vary from one monitoring device to another monitoring device. Further, certain monitoring devices may contain multiple instances of a component, such as multiple processors, memories, transceivers, transmitters, receivers, etc.
[0138] Processing circuitry 1002 is configured to process instructions and data and may be configured to implement any sequential state machine operative to execute instructions stored as machine-readable computer programs in memory 1010. Processing circuitry 1002 may be implemented as one or more hardware-implemented state machines (e.g., in discrete logic, field- programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), etc.); programmable logic together with appropriate firmware; one or more stored computer programs, general-purpose processors, such as a microprocessor or digital signal processor (DSP), together with appropriate software; or any combination of the above. For example, processing circuitry 1002 may include multiple central processing units (CPUs).
[0139] In the example, input / output interface 1006 may be configured to provide an interface or interfaces to an input device, output device, or one or more input and / or output devices. Examples of an output device include a speaker, a sound card, a video card, a display, a monitor, a printer, an actuator, an emitter, a smartcard, another output device, or any combination thereof. An input device may allow a user to capture information into monitoring device 1000. Examples of an input device include a touch-sensitive or presence- sensitive display, a camera (e.g., a digital camera, a digital video camera, a web camera, etc.), a microphone, a sensor, a mouse, a trackball, a directional pad, a trackpad, a scroll wheel, a smartcard, and the like. The presence-sensitive display may include a capacitive or resistive touch sensor to sense input from a user. A sensor may be, for instance, an accelerometer, a gyroscope, a tilt sensor, a force sensor, a magnetometer, an optical sensor, a proximity sensor, a biometric sensor, etc., or any combination thereof. An output device may use the same type of interface port as an input device. For example, a Universal Serial Bus (USB) port may be used to provide an input device and an output device.
[0140] In some embodiments, power source 1008 is structured as a battery or battery pack. Other types of power sources, such as an external power source (e.g., an electricity outlet), photovoltaic device, or power cell, may be used. Power source 1008 may further include power circuitry for delivering power from power source 1008 itself, and / or an external power source, to the various parts of monitoring device 1000 via input circuitry or an interface such as an electrical power cable. Delivering power may be, for example, for charging of power source 1008. Power circuitry may perform any formatting, converting, or other modification to the power from power source 1008 to make the power suitable for the respective components of monitoring device 1000 to which power is supplied.
[0141] Memory 1010 may be or be configured to include memory such as random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic disks, optical disks, hard disks, removable cartridges, flash drives, and so forth. In one example, memory 1010 includes one or more application programs 1014, such as an operating system, web browser application, a widget, gadget engine, or other application, and corresponding data 1016. Memory 1010 may store, for use by monitoring device 1000, any of a variety of various operating systems or combinations of operating systems.
[0142] Memory 1010 may be configured to include a number of physical drive units, such as redundant array of independent disks (RAID), flash memory, USB flash drive, external hard disk drive, thumb drive, pen drive, key drive, high-density digital versatile disc (HD-DVD) optical disc drive, internal hard disk drive, Blu-Ray optical disc drive, holographic digital data storage (HDDS) optical disc drive, external mini-dual in-line memory module (DIMM), synchronous dynamic random access memory (SDRAM), external micro-DIMM SDRAM, smartcard memory such as tamper resistant module in the form of a universal integrated circuit card (UICC) including one or more subscriber identity modules (SIMs), such as a USIM and / or ISIM, other memory, or any combination thereof. The UICC may for example be an embedded UICC (eUICC), integrated UICC (iUICC) or a removable UICC commonly known as ‘SIM card.’ Memory 1010 may allow monitoring device 1000 to access instructions, application programs and the like, stored on transitory or non-transitory memory media, to off-load data, or to upload data. An article of manufacture, such as one utilizing a communication system may be tangibly embodied as or in memory 1010, which may be or comprise a device-readable storage medium.
[0143] Processing circuitry 1002 may be configured to communicate with an access network or other network using communication interface 1012. Communication interface 1012 may comprise one or more communication subsystems and, in some cases, may include or be communicatively coupled to an antenna 1022. Communication interface 1012 may include one or more transceivers used to communicate with other devices or nodes, such as another monitoring device or a network node in an access network. Each transceiver may include a transmitter 1018 and / or a receiver 1020 appropriate to provide network communications (e.g., optical, electrical, frequency allocations, and so forth). Moreover, transmitter 1018 and receiver 1020 may be coupled to one or more antennas (e.g., antenna 1022) and may share circuit components, software, or firmware, or alternatively be implemented separately.
[0144] In the illustrated embodiment, communication functions of communication interface 1012 may include cellular communication, Wi-Fi communication, LPWAN communication, data communication, voice communication, multimedia communication, short-range communications such as Bluetooth, near-field communication, location-based communication such as the use of the global positioning system (GPS) to determine a location, another like communication function, or any combination thereof. Communications may be implemented in according to one or more communication protocols and / or standards, such as IEEE 802.11, Code Division Multiplexing Access (CDMA), Wideband Code Division Multiple Access (WCDMA), GSM, LTE, New Radio (NR), UMTS, WiMax, Ethernet, transmission control protocol / internet protocol (TCP / IP), synchronous optical networking (SONET), Asynchronous Transfer Mode (ATM), QUIC, Hypertext Transfer Protocol (HTTP), and so forth.
[0145] Monitoring device 1010 may provide an output of its monitoring results via communication interface 1012, which facilitates a wired or wireless connection to a network node or another monitoring device. The output may be periodic (e.g., once every 15 minutes), random (e.g., to even out the reporting traffic from several monitoring devices), in response to a triggering event (e.g., when an execution anomaly is detected), in response to a request (e.g., a supervisory request), or a continuous stream of results (e.g., of each comparison performed).
[0146] Communication interface 1012 may also include side-channel receiver circuitry capable of capturing side-channel information emitted by a device being monitored and / or by other devices. The captured side-channel information may include one or more reference traces and other traces, from the device being monitored, that are compared against the reference trace(s) using various techniques described above. For example, the side channel receiver circuitry may be capable of capturing and / or detecting side-channel information such as device energy consumption, device electromagnetic (EM) emissions, device thermal emissions, device sound emissions, and / or device optical emissions.
[0147] A monitoring device may be intended and / or configured for use in one or more application domains. As such, the monitoring device may include circuitry and / or software that supports operation in the intended application domain(s), in addition to other components described above for monitoring device 1000 shown in Figure 10.
[0148] As a more specific example, a monitoring device may represent a machine or other device that performs monitoring and / or measurements, and transmits the results of such monitoring and / or measurements to another monitoring device and / or a network node. The monitoring device may in this case be an M2M device, which may in a 3GPP context be referred to as an MTC device. As one particular example, the monitoring device may implement the 3GPP NB-IoT standard. In other scenarios, a monitoring device may be included in a vehicle (e.g., car, bus, truck, ship, airplane, etc.) or other equipment that is capable of monitoring and / or reporting on its operational status or other functions associated with its operation.
[0149] As mentioned above, in some embodiments, the monitoring device may be internal to the device being monitored. In such embodiments, some of the components shown in Figure 10 (e.g., power supply 1008) may be shared with the device being monitored while other components (e.g., processing circuitry 1002 and memory 1010) may be separate from the device being monitored.
[0150] The foregoing merely illustrates the principles of the disclosure. Various modifications and alterations to the described embodiments will be apparent to those skilled in the art in view of the teachings herein. It will thus be appreciated that those skilled in the art will be able to devise numerous systems, arrangements, and procedures that, although not explicitly shown or described herein, embody the principles of the disclosure and can be thus within the spirit and scope of the disclosure. Various embodiments can be used together with one another, as well as interchangeably therewith, as should be understood by those having ordinary skill in the art.
[0151] The term unit, as used herein, can have conventional meaning in the field of electronics, electrical devices and / or electronic devices and can include, for example, electrical and / or electronic circuitry, devices, modules, processors, memories, logic solid state and / or discrete devices, computer programs or instructions for carrying out respective tasks, procedures, computations, outputs, and / or displaying functions, and so on, as such as those that are described herein. Any appropriate steps, methods, features, functions, or benefits disclosed herein may be performed through one or more functional units or modules of one or more virtual apparatuses. Each virtual apparatus may comprise a number of these functional units. These functional units may be implemented via processing circuitry, which may include one or more microprocessor or microcontrollers, as well as other digital hardware, which may include Digital Signal Processor (DSPs), special-purpose digital logic, and the like. The processing circuitry may be configured to execute program code stored in memory, which may include one or several types of memory such as Read Only Memory (ROM), Random Access Memory (RAM), cache memory, flash memory devices, optical storage devices, etc. Program code stored in memory includes program instructions for executing one or more telecommunications and / or data communications protocols as well as instructions for carrying out one or more of the techniques described herein. In some implementations, the processing circuitry may be used to cause the respective functional unit to perform corresponding functions according to one or more embodiments of the present disclosure.
[0152] As described herein, device and / or apparatus can be represented by a semiconductor chip, a chipset, or a (hardware) module comprising such chip or chipset; this, however, does not exclude the possibility that a functionality of a device or apparatus, instead of being hardware implemented, be implemented as a software module such as a computer program or a computer program product comprising executable software code portions for execution or being run on a processor. Furthermore, functionality of a device or apparatus can be implemented by any combination of hardware and software. A device or apparatus can also be regarded as an assembly of multiple devices and / or apparatuses, whether functionally in cooperation with or independently of each other. Moreover, devices and apparatuses can be implemented in a distributed fashion throughout a system, so long as the functionality of the device or apparatus is preserved. Such and similar principles are considered as known to a skilled person.
[0153] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. It will be further understood that terms used herein should be interpreted as having a meaning that is consistent with their meaning in the context of this specification and the relevant art and will not be interpreted in an idealized or overly formal sense unless expressly so defined herein.
[0154] In addition, certain terms used in the present disclosure, including the specification and drawings, can be used synonymously in certain instances (e.g., “data” and “information”). It should be understood, that although these terms (and / or other terms that can be synonymous to one another) can be used synonymously herein, there can be instances when such words can be intended to not be used synonymously.
Claims
CLAIMS1. A computer-implemented method for detecting anomalies in processes executing on a device, the method comprising: capturing (940) first and second traces of side channel information, wherein the first trace is associated with execution of a first process and the second trace is associated with execution of a process on the device; using a trained side-channel encoder, determining (950) encoded representations of the first and second traces; and using a trained trace comparison model, determining (970) whether the encoded representations of the first and second traces are associated with execution of a same process.
2. The method of claim 1, further comprising: training (910) the side-channel encoder based on a first plurality of traces of side-channel information; using the trained side-channel encoder, determining (920) encoded representations of a second plurality of traces of side-channel information associated with the device or with other devices; and training (930) the trace comparison model based on the encoded representations of the second plurality of traces.
3. The method of claim 2, wherein the second plurality of traces comprises: a plurality of first trace pairs, wherein each first trace pair includes two traces associated with execution of the first process; and a plurality of second trace pairs, wherein each second trace pair includes one trace associated with execution of the first process and one trace associated with execution of a different process than the first process.
4. The method of any of claims 2-3, wherein training (910) the side-channel encoder is performed using unsupervised machine learning and the traces of the first plurality have no known association with execution of any specific processes.
5. The method of any of claims 2-4, wherein training (910) the side channel encoder comprises, for each trace of the first plurality:determining (911) one or more partially masked versions of the trace; and training (912) the side channel encoder to predict the trace from the one or more partially masked versions of the trace.
6. The method of any of claims 2-3, wherein training (910) the side-channel encoder is performed using supervised machine learning and each encoded representation of a trace of the second plurality is labelled with an associated process executing on the device.
7. The method of any of claims 1-6, wherein an encoded representation is a more compact trace representation that indicates information-carrying characteristics of the side channel from the device.
8. The method of any of claims 1-7, wherein the trace comparison model is an attentionbased neural network.
9. The method of any of claims 1-8, wherein: the encoded representations of the first and second traces are respective first and second vectors (vl, v2), each comprising a plurality (L) of data elements; and determining (970) whether the encoded representations of the first and second traces are associated with execution of a same process comprises: determining (971) a plurality (L) of first difference vectors based on the second vector and respective data elements of the first vector; determining (972) a plurality (L) of second difference vectors based on the first vector and respective data elements of the second vector; and applying (973) the plurality (L) of first difference vectors to a first multi-head attention function; and applying (974) the plurality (L) of second difference vectors to a second multihead attention function.
10. The method of claim 9, wherein: each of the first and second multi-head attention functions comprises a plurality (L) of single -head attention functions; and applying (973) the plurality (L) of first difference vectors to the first multi-head attention function comprises, for each first difference vector, mapping the first difference vector to query (Q), key (K), and value (V) vectors of a correspondingsingle -head attention function; and applying (974) the plurality (L) of second difference vectors to the second multi-head attention function comprises, for each second difference vector, mapping the second difference vector to query (Q), key (K), and value (V) vectors of a corresponding single -head attention function.
11. The method of claim 10, wherein applying (973) the plurality (L) of first difference vectors to the first multi-head attention function further comprises, for each first difference vector, calculating an output of the corresponding single -head attention function based on a function of the mapped query (Q), key (K), and value (V) vectors; and applying (974) the plurality (L) of second difference vectors to the second multi-head attention function comprises, for each second difference vector, calculating an output of the corresponding single-head attention function based on the function of the mapped query (Q), key (K), and value (V) vectors.
12. The method of claim 11, wherein the function of the mapped query (Q), key (K), and value (V) vectors is:wherein d is a dimension of each of the mapped query (Q), key (K), and value (V) vectors.
13. The method of any of claims 11-12, wherein determining (970) whether the encoded representations of the first and second traces are associated with execution of a same process further comprises: concatenating (975) the outputs of the respective single -head attention functions in a concatenated vector; and applying (976) the concatenated vector to one or more fully connected layers of a neural network, which outputs at least one of the following: a first probability that the first and second traces are associated with execution of the same process; and a second probability that the first and second traces are associated with execution of different processes.
14. The method of any of claims 1-13, wherein:a plurality of the first traces of side channel information associated with execution of the first process are captured, a corresponding plurality of encoded representations of the first traces are determined; and determining (970) whether the encoded representations of the first and second traces are associated with execution of a same process is based on the plurality of encoded representations of the first traces.
15. The method of claim 14, wherein determining (970) whether the encoded representations of the first and second traces are associated with execution of a same process executing on the device comprises: determining (977), for each of the plurality of first traces, an indication of whether the encoded representation of the second trace is associated with a same process as the encoded representation of the first trace; and determining (978) whether the second trace is associated with execution of a same process as the plurality of first traces based on the plurality of indications.
16. The method of claim 15, wherein it is determined that second trace is associated with a same process as the plurality of first traces when one of the following indicate that the second trace is associated with a same process as the plurality of first traces: one indication, all indications, a majority of indications, a threshold number of indications that is greater than one but less than all.
17. The method of claim 14, further comprising combining (960) the plurality of encoded representations of the first traces into a combined first trace representation, wherein determining (970) whether the encoded representations of the first and second traces are associated with execution of a same process is based on the combined first trace representation.
18. The method of claim 17, wherein the plurality of encoded representations of the first traces are combined into the combined first trace representation based on element-by-element averaging, and determining (970) whether the encoded representations of the first and second traces are associated with execution of a same process is based on one of the following weightings for elements of the combined first trace representation: equal weighting for all elements; or element weighting based on similarity of elements from the plurality of encodedrepresentations of the first traces that were averaged to determine the element.
19. The method of any of claims 1-13, wherein: a plurality of the second traces of side channel information associated with the device are captured; a corresponding plurality of encoded representations of the second traces are determined; and determining (970) whether the encoded representations of the first and second traces are associated with execution of a same process is based on the plurality of encoded representations of the second traces.
20. The method of any of claims 1-19, wherein the method is performed by one of the following: a monitoring device that is internal to the device, a monitoring device that is external but connected to the device, or a monitoring device that is external and not connected to the device.
21. A monitoring device (820, 1000) configured to detect anomalies in processes executing on a device (810), the monitoring device comprising: communication interface circuitry (1012) configured to capture traces of side channel information emitted by at least the device; and processing circuitry (1002) operably coupled to the communication interface circuitry, whereby the processing circuitry and the communication interface circuitry are configured to: capture first and second traces of side channel information, wherein the first trace is associated with execution of a first process and the second trace is associated with execution of a process on the device; using a trained side-channel encoder, determine encoded representations of the first and second traces; and using a trained trace comparison model, determine whether the encoded representations of the first and second traces are associated with execution of a same process.
22. The monitoring device of claim 21, wherein the processing circuitry and the communication interface circuitry are further configured to perform operations corresponding to any of the methods of claims 2-20.
23. A monitoring device (820, 1000) configured to detect anomalies in processes executing on a device (810), the monitoring device being further configured to: capture first and second traces of side channel information, wherein the first trace is associated with execution of a first process and the second trace is associated with execution of a process on the device; using a trained side-channel encoder, determine encoded representations of the first and second traces; and using a trained trace comparison model, determine whether the encoded representations of the first and second traces are associated with execution of a same process.
24. The monitoring device of claim 23, being further configured to perform operations corresponding to any of the methods of claims 2-20.
25. A non-transitory, computer-readable medium (1010) storing computer-executable instructions that, when executed by processing circuitry (1002) of a monitoring device (820, 1000) configured to detect anomalies in processes executing on a device (810), configure the monitoring device to perform operations corresponding to any of the methods of claims 1-20.
26. A computer program product (1014) comprising computer-executable instructions that, when executed by processing circuitry (1002) of a monitoring device (820, 1000) configured to detect anomalies in processes executing on a device (810), configure the monitoring device to perform operations corresponding to any of the methods of claims 1-20.
Citation Information
Patent Citations
Microcontroller program instruction execution fingerprinting and intrusion detection
US20210294893A1