Multifactor continuous authentication

The Hierarchical Reinforcement Learning-based method continuously authenticates and defends IoT devices by analyzing hardware and network data to maintain network security, addressing vulnerabilities and ensuring secure operation.

WO2025140840A1PCT designated stage expired Publication Date: 2025-07-03BRITISH TELECOM PLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2024/085242
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-29
Filing Date
2024-12-09
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Existing Internet of Things (IoT) devices face challenges in continuous authentication due to their resource constraints and vulnerabilities, lacking standard protocols, making them susceptible to cyberattacks and potential breaches that can propagate across networks.

Method used

A computer-implemented method using a machine learning model, specifically Hierarchical Reinforcement Learning (HRL), continuously authenticates IoT devices by analyzing hardware and network data to assign health scores, detecting anomalies, and triggering responses to maintain network security.

Benefits of technology

The system provides robust, continuous authentication and defense mechanisms, ensuring the security and trustworthiness of IoT devices by dynamically adapting to threats and maintaining a high Global Health Score, thereby preventing unauthorized access and enhancing network security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024085242_03072025_PF_FP_ABST
    Figure EP2024085242_03072025_PF_FP_ABST
Patent Text Reader

Abstract

A system to verify the correctness of equipment installation configured to perform the steps of: accessing a digital model of the equipment; receiving a representation of the equipment indicating a configuration of the equipment, Accessing installation-specific metadata associated with the installation of the equipment; Predicting at least one behaviour of the equipment based on metadata associated with installation of the equipment; Accessing service information comprising at least one measured behaviour of the equipment; Comparing the representation with the model to determine a first degree of conformity of the equipment with the model of the equipment; if the first degree of conformity is within a first threshold degree of conformity, comparing the at least one predicted behaviour with the at least one measured behaviour to determine a second degree of conformity of the equipment with the service information; and if the second degree of conformity is within a threshold second degree of conformity, verifying to an operator that the equipment is correctly installed.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Multifactor Continuous Authentication

[0002] The present invention relates to continuous authentication of networked devices.

[0003] Authentication is the process of proving the identity of a system user (human or machine). This is most observed through the use of passwords or PINs entered into a login screen by human users. With machine-to-machine communication, this is more likely to be token, or public key based. A weakness in traditional authentication is that this authentication only occurs once, typically at the start of a session.

[0004] Continuous authentication seeks to improve security by assessing ongoing user characteristics. In human authentication scenarios, that may be facial recognition, fingerprint, gait, etc., also known as “Type 3” authentication factors “something you are” (i.e. biometrics). These authentication factors are observable through the user’s interaction with the system. loT devices are widely opted for their low cost-low energy effective solution for automation and digitalisation of processes in various industries accounting to 10 billion-plus loT devices connected to the internet, this is expected to reach 25 billion devices by 2030. The loT devices usually have less computational power and have major vulnerabilities as most of the devices operate for life with no further patches or security updates which may make them a target for attackers.

[0005] Many organisations suffer from attempted cyberattacks on loT devices. The security of loT devices is crucial for enabling growth and sustaining the operations, any breach could have severe consequences that could jeopardise the operations and services that depend on the communication and data relayed by these devices. Further, the attackers could easily propagate to other devices and systems by compromising a single loT device on a network.

[0006] Due to the low SWaP (Size, Weight & Power) of the loT devices, it makes it challenging to secure them indefinitely, especially with continuously evolving threats, it’s not feasible to patch these devices.

[0007] There are no standard Continuous authentication protocols / methods for loT devices. Existing Continuous authentication methods focus on user-based authentication to authenticate the device. The only device-based authentication (non-continuous) that is widely used happens just once after registration, i.e., static authentication. The present invention accordingly provides, in a first aspect, a computer-implemented method to continuously authenticate a device in a network of devices using a machine learning model, the method comprising the steps of: initially analysing hardware data associated with the device and network data associated with the network, wherein the network data is associated with all other devices in the network; assigning a local health score to the device based on the hardware data and a global health score to the network based on the network data; continuously authenticating the device at intervals, wherein the frequency of the intervals is determined based on the global health score and / or the local health score; at one interval of the intervals, requesting hardware data and a network device data to ascertain if there is a change in the local health score and / or global health score; if there is a change in the local health score and / or global health score, failing to authenticate the device; subsequent to failing to authenticate the device, identifying and / or removing a threat in the network; subsequent to identifying and / or removing the threat, triggering a response to reintroduce the device to the network.

[0008] In one arrangement of the method, wherein the change is larger than a threshold.

[0009] In one arrangement of the method, wherein the threshold is predetermined.

[0010] In one arrangement of the method, wherein the threshold is based on the hardware data and / or the network data.

[0011] In one arrangement of the method, wherein the initial analysing of the hardware data and network data is performed contemporaneously.

[0012] In one arrangement of the method, wherein the machine learning model is a reinforcement learning model.

[0013] In one arrangement of the method, wherein the reinforcement model is a hierarchical reinforcement learning model.

[0014] In one arrangement of the method, wherein the device is an Internet of Things, loT, device.

[0015] In one arrangement of the method, wherein the response to reintroduce the device comprises increasing the local and / or global health score above a minimum threshold.

[0016] The present invention accordingly provides, in a second aspect, a data processing system comprising a processor configured to carry to perform the steps of: continuously authenticate a device in a network of devices using a machine learning model, the method comprising the steps of: initially analysing hardware data associated with the device and network data associated with the network, wherein the network data is associated with all other devices in the network; assigning a local health score to the device based on the hardware data and a global health score to the network based on the network data; continuously authenticating the device at intervals, wherein the frequency of the intervals is determined based on the global health score and / or the local health score; at one interval of the intervals, requesting hardware data and a network device data to ascertain if there is a change in the local health score and / or global health score; if there is a change in the local health score and / or global health score, failing to authenticate the device; subsequent to failing to authenticate the device, identifying and / or removing a threat in the network; subsequent to identifying and / or removing the threat, triggering a response to reintroduce the device to the network.

[0017] The present invention accordingly provides, in a third aspect, a computer program element comprising computer program code to, when loaded into a computer system and executed thereon, cause the computer to perform the steps of a method described above.

[0018] Exemplary arrangements of the present invention will now be described, by way of example only, with reference to the accompanying drawings, in which:

[0019] Figure 1 is a block diagram showing reinfoicement learning in the operation of embodiments of the present invention;

[0020] Figure 2 is a flowchart of an exemplary method in accordance with an exemplary arrangement of the present invention;

[0021] Figure 3 is a block diagram of an exemplary method in accordance with an exemplary arrangement of the present invention;

[0022] Figure 4 is a block diagram of primary and secondary actions in accordance with an exemplary arrangement of the present invention;

[0023] Figure 5 is a block diagram of the sub-protocols in accordance with an exemplary arrangement of the present invention;

[0024] Figure 6 is a block diagram of the continuous authentication and active defence and response system in accordance with an exemplary arrangement of the present invention.

[0025] Authentication is a critical aspect of the Internet of Things (loT), which involves verifying and validating the identity of devices and users within an loT ecosystem. With a vast network of interconnected devices exchanging data to provide various services and functionalities, it is important to ensure the security and trustworthiness of the device and the data it generates. This is crucial in preventing unauthorised access, data breaches, and potentially malicious activities. Traditional authentication systems focus on verifying the identity of human users who access computer systems, applications, or online services. In contrast, loT authentication involves verifying the identity of interconnected loT ecosystem devices, such as sensors, actuators, and / or intelligent appliances. This process is developed in a decentralized or distributed manner, as each device needs to be authenticated independently to participate in the network. Since various manufacturers produce these devices with diverse communication protocols like Zigbee, Bluetooth, etc., loT networks often consist of many instruments. Scalability is a significant concern as many loT devices are resource- constrained, thereby making the implementation of traditional authentication systems a challenge. Therefore, flexible and interoperable authentication solutions are required.

[0026] The authentication process of loT devices encounters distinct security challenges including device impersonation, firmware tampering, and the requirement for secure device onboarding procedures. Popular authentication protocols such as OAuth, OpenlD Connect, and SAML are widely utilised to ensure secure access to different services. However, due to the nature of the devices, loT authentication requires different solutions to ensure the security and trustworthiness of interconnected devices.

[0027] Continuous authentication is a security method that verifies the identity of devices or users throughout an interaction or session rather than simply at the initial login phase (known as static authentication). This ongoing process of continuous authentication ensures that the device or user remains authenticated, and it quickly detects and prevents unauthorised access. In contrast, static authentication is the traditional method where devices are only authenticated at the beginning of a session and not revalidated during the interaction. Once successfully authenticated, the devices are trusted for the duration of the session.

[0028] In device-to-device communication, continuous authentication involves monitoring the communication devices' behaviour, characteristics, or other unique traits during the interaction. These characteristics may include the device's location, error, noise, channels, hardware imperfection and / or some persistent features. If a continuous authentication system detects any deviation from expected behaviour or signs of potential compromise, it may be configured to trigger additional security measures. These measures include requesting reauthentication and / or terminating the connection.

[0029] Although static authentication is more straightforward to implement and requires fewer resources than continuous authentication, it may be less effective in identifying potential security breaches that could happen during a communication session. If an intruder manages to enter the authorised devices during the session, they could carry out unauthorised activities. Combining continuous and static authentication methods can improve the devices and network security.

[0030] In some examples, continuous device authentication (CDA) monitors a specific device’s behaviour patterns or the specific device's feature patterns. Continuous device authentication monitors continuously after the device registration or static authentication process. These patterns are unique for each device and can be observed in the device. Observing these unique patterns is called device fingerprinting (DFP), and the unique pattern is known as a device fingerprint.

[0031] Condition monitoring is the continuous process of monitoring machinery, equipment, and applications to predict critical events such as malfunctions, failures, or glitches. Condition monitoring is closely related to device fingerprinting (DFP) methods since it utilises the unique characteristics of the device, machine, or system for any given application. This similarity enables direct inference of the device’s condition from fingerprint data. Conditioning monitoring is an important component in the loT 4.0 industry, and it can be used to monitor the system health, behaviour and performance of the appliances, infrastructure, heavy machines, and other equipment. loT devices do not have the same observable biometrics as humans. Applied to machines, type 3 authentication may include features such as network utilisation and hardware characteristics. Many characteristics of a device are unchangeable given a specific configuration. For example, clock skew is a phenomenon that results in distinct variations between devices. Some characteristics result from hardware configuration and even manufacturing characteristics that are unlikely to change significantly (environment and degradation of components may be a factor). Other characteristics may be the result of software and may be predictable if the code is known.

[0032] This means that the characteristics can be a reliable indicator of a device being modified (through software or hardware modification). Different sources of features are required to develop or generate DFP based on the four-axis Taxonomy. These features are based on Device physics, operation, behaviour, and network. In some examples, these device characteristics directly or indirectly are measured for DFP and continuous authentication systems. The four-axis taxonomy is: Intrinsic characteristic, Behavioural characteristics, operational characteristic and Network Characteristics. In some examples, the quality and reliability of DFP are based on these evaluation techniques during a feature selection process. Reinforcement learning (RL), as shown in Figure 1 , is an approach to machine learning that does not require prior knowledge. In RL, agents 1 make use of exploration to drive a sequence of actions 2 within an environment 3. A reward function 4 is used to reinforce actions that provide a beneficial result. The purpose is for agents 2 to learn an optimal policy that maximises reward 5, taking the actions that are most likely to achieve a satisfactory state 6.

[0033] In some examples, the basic process of reinforcement learning is modelled as a Markov Decision Process. A sequence of states is described as Markov if the probability of moving from State Stto St+1depends on all previous states Si to St-1.

[0034] A Markov process is a tuple constructed of S a finite set of states and P a state transition probability matrix:

[0035] PSS' = P[St+i= s'l i = s]

[0036] The incorporation of action and reward to a Markov process transforms it to a Markov Decision Process. In this process, the movement from Stto St+1relies not only on the current state but also the action Attaken at the current state:

[0037] In some examples, the present invention proposes an automated continuous authentication and defence of loT devices on a network using Reinforcement Learning (RL). In some examples, the present invention uses device-specific behaviour hardware features and network metadata to continuously authenticate the device. If the authentication of the device fails, the system is configured to analyse the device data to detect and / or identify anomalies / threats. In some examples, once a threat is identified, an automated response is prompted. The loT device-based continuous authentication can be applied on the gateway or server. Depending on the detected authenticity of the individual loT devices in a network, these devices are placed in a sub-cluster network with appropriate protocols. The continuous monitoring of the device may also indicate the potential to detect device degradation and failure.

[0038] In one example, an approach of the present invention is a multifactor continuous authentication and defence system for loT devices. In this example, the present invention employs a unique hybrid approach for automated continuous authentication of the loT devices in regular time intervals. The hybrid approach consists of utilising the hardware device specific data and network metadata (such as NetFlow data) to authenticate a loT device. The multifactor authentication is the ability of the system to perform a combination of continuous and static authentications at pre-determined time intervals depending on the status of individual loT device.

[0039] The active defence component of the system is triggered for an “authentication failed” loT device, wherein the hardware & network metadata data are further analysed to detect and / or identify potential threats / anomalies on a network. The combination of the continuous authentication and active defence makes the loT devices highly secure throughout its operation. The different network topology states and clustering of the loT devices on sub networks based on the individual device status (Health score, Risk factor etc.) enables customised network protocols for sub network which further enhances the network & loT device security.

[0040] Figure 2 illustrates the architecture of the Citokines system, the different components and 4 stages of operation.

[0041] Stage-1 101 : RL agent Action

[0042] The Reinforcement learning model receives the loT device data device behaviour data- dhnand device network metadata- dnn, where n => number of loT devices on a network. The RL agent, which is external to the environment e takes an action atwhich affects the environment, as guided by the pre-set policy. The environment in this case is the network hardware and all connected loT devices. The action atchanges the state st+1of the current environment and depending on the outcome of changes, i.e., whether the action has positive effect on the environment or not, the RL agent gets rewarded accordingly. The stage-1 101 process in the present invention only involves the RL agent’s new action based on its current exposure to network environment through the loT device data and governing policy of the whole system, which is Global Health Score (GHS) of the total network.

[0043] Staoe-2 102: Network Environment Chances

[0044] The RL agent continuously monitors and analyses the loT device data to trigger a primary action which is to either authenticate or not authenticate the device based on the device behaviour data- dhnand device network metadata- dnn. This action affects the state of the network environment, where the number of authenticated devices on network at time t0could be different to the number of authenticated devices at time tr. The authenticated devices and authentication failed devices require secondary actions. The authenticated devices need to be updated with network protocols and the authentication failed devices needs to action on defence measures, also known as active defence mechanism. Stage-3 103: Network Clustering & Protocol

[0045] The authenticated devices are assigned with Local Health Score (LHS) which places the device on the appropriate risk factor (FF) which determines the sub-network cluster on which the individual loT devices needs to be placed. Each sub-network stipulates the authentication and network protocol that individual loT devices on that cluster needs to follow. The authentication failed devices have secondary actions that could involve a series of abstracted primitive actions to focus on active defence process that will detect, identify, and respond to the cause of the failure of authentication. In some examples, the primary assumption is that the failure is due to security incident, therefore the model will aim to identify and respond to the threat instantly without affecting the secure functioning and operation of other loT devices connected to the same network.

[0046] Stage-4 104: RL agent Reward

[0047] The goal of the RL agent is to maintain / increase the Global Health score of the network as high as possible, therefore the policies will be based on this, and the reward will be proportional to the change is Global health score(AGHS). The RL agent will be aiming for maximum reward, therefore taking actions that will maintain / increase the Global health score of the network, keeping the devices securely operational throughout.

[0048] There are many Reinforcements Learning (RL) algorithms and most of them follow valuebased methods or policy-based methods. The linear RL models are not always capable of mapping real world complex state space representations and therefore the Deep Learning (DL) based Reinforcements Learning (RL) can be employed to deal with complex scenarios with high dimensions. The continuous authentication and defence of loT devices requires different levels of actions and in different sequences which makes Hierarchical Reinforcement Learning (HRL) preferable in some scenarios. Hierarchical Reinforcement Learning (HRL) decomposes a reinforcement learning problem into a hierarchy of subproblems or subtasks such that higher-level parent-tasks invoke lower-level child tasks as if they were primitive actions. A decomposition may have multiple levels of hierarchy.

[0049] There are two major characteristics of the Hierarchical Reinforcement Learning (HRL): (a) temporal abstraction and (b) state space representation & transfer approach. The Temporal Abstraction is a process of breaking down the big task into several smaller tasks. These temporal abstractions can be utilised by a learning agent in order to make decisions on a higher level of abstraction. The sequence of actions that make up a temporally extended action can be fixed or governed by a policy. The Hierarchy allows different levels of policy to be set at different levels when the problem is broken down into smaller tasks. The generation of functional approximation function for a given problem is optimised and solved more efficiently as the RL agent in HRL breaks down the large problem to be solved into a smaller problem, avoiding the approach to solve the problem in one large higher dimensional space. The State space representation & transfer approach has an advantage of breaking down tasks to allow a higher state space representation to be transformed to smaller more meaningful state space abstractions. The focus of the RL agent at a given time will only be on solving a particular small state space problem, therefore, the RL agent will be dealing with fewer parameters. In the context of solving the bigger problem, the smaller problems are independent of each other as a first smaller problem doesn’t need any information about solving a second smaller problem. The combination of temporal abstraction and state space introduces an efficient learning process and an efficient exploration for solving tasks. The transfer approach enables reusability of optimised smaller task solutions that were explored during the learning process for the future tasks.

[0050] A hierarchical reinforcement learning (HRL) model is trained to continuously authenticate and defend loT devices. A governing factor known as a “Health score” is defined to determine the security status of the loT device. The lower the score, the higher the risk. There are global and / oca / health scores. The health score places the corresponding device with an appropriate “Risk factor” level, i.e., low, medium, or high. The higher the risk, the lower the authentication time intervals, i.e., high frequency of authentication will be required until the device progresses to a lower risk level (by having a higher health score. There are global and local Risk factor levels.

[0051] The RL agent operates to continuously authenticate the loT devices on a network by verifying the device using its hardware and network features at time intervals defined by a health score matrix. If the device is authenticated, it is assumed that there are no threats, and therefore network connectivity will remain unchanged. The authenticated devices are clustered into sub networks based on individual (i.e., local) health scores (LHS) and this will amend the network configurations accordingly. If the authentication of the device fails, the RL agent will utilise the network metadata to detect and / or identify an anomaly. Once the anomaly is Identified, the RL agent triggers an appropriate response to defend the network and, in some examples, reintroduce the device back to the network after cleaning (to ensure the device is benign) said device.

[0052] One aim of Continuous authentication and defence is achieved by developing the most efficient functional approximation function that keeps the Global Health Score (GHS) as high as possible. Therefore, the RL agent will explore and exploit, in a balanced manner, to take an optimal action that would be aimed at maintaining or increasing the Global Health Score (GHS) of the network.

[0053] Where, n * 0 &n e R

[0054] The Local Health Score (LHS) can be defined in a few ways depending on the learning approach applied to the system.

[0055] LHSm= (Cs- Me) where, LHSm=

[0056] LHS for Meta Learning architectures combining different ML leanring approaches,

[0057] Cs= Confidence Metric of ML model output, and

[0058] Me = Change is device behavior characteristics.

[0059] LHSrl= (1 - Me) where, LHSrl= LHS for Reinforcement Learning architectures, and

[0060] Me = Change is device behavior characteristics.

[0061] Mbc + Mbc Me = -

[0062] 2 where, Mbc = Change is hardware device behavior characteristics, and

[0063] Mbc = Change is network device behavior characteristics.

[0064] The present invention is not limited to any specific Reinforcement algorithm or any specific Machine Learning approach. The Hierarchical Reinforcement Learning (HRL) is mentioned to illustrate the concept but any machine learning approach may be employed. An alternative Meta learning architecture is illustrated in figure 3. In this alternative approach, a supervised ML model makes use of loT device data to perform continuous authentication, the authenticated devices and authentication failed devices further gets analysed by a RL model to perform network clustering and defence.

[0065] The present invention, when powered by Hierarchical Reinforcement learning (HRL), requires the RL agent to take actions for supporting the active operation of continuous authentication and defence. There are primary and secondary actions (as shown in figure 4) with different policies associated with them. The primary action is focused on continuous authentication and the secondary actions are based on the outcomes of the continuous authentication of the loT device on a network.

[0066] A primitive set of secondary actions are primarily designed to be taken on the authentication-failed devices to trigger active defence, where the network metadata will be used to detect and identify potential threats / anomalies, so that an appropriate response can be triggered by the RL agent. The other set of secondary actions involve network configuration management and clustering. The primary goal of these primary and secondary actions is to ensure that the Global Health Score (GHS) is maintained or optimised by which the continuous authentication and defence are performed / achieved.

[0067] Figure 4 illustrates how the different primary actions and secondary actions taken by the RL agent affects the state (Sn) of a RL environment, which dictates the number of loT devices connected to the network and its network configurations. The actions taken by the RL agent could change the state of the present RL environment in an intended (positive) way or negative way. Through a series of learning processes of multiple state steps, the RL agent improves by taking more appropriate actions. The reward-based feedback system alongside policies guides the RL agent to maximise the reward which will maximise the Global Health Score (GHS) of the loT network.

[0068] Network Clustering and protocols continuously monitor the individual loT devices and cluster them in appropriate sub-networks, as shown in Figure 5, based on its authenticity / Local Health Score (LHS). The sub-network clustering is crucial to ensure that devices with similar risk factors and health scores are grouped in same sub-networks, i.e., for example the loT devices with medium risk and Local Health scores ranging between, for example, a score of 50-80 would be placed in sub-network-2. In one example, there are 4 sub-networks in total, of which 3 are allocated for authenticated devices and one sub-networks for authentication failed devices.

[0069] In some examples, the sub-networks also provide customised network configurations for devices. This is to ensure that devices on high-risk sub-networks have different protocols compared to devices in low-risk sub network. For example, the time frequency for continuous authentication of loT devices in a low-risk sub-network is 60 seconds, compared to 5 seconds in a high-risk sub network cluster. Similarly, the access type, static authentication and other protocols vary across each sub-network (favouring the low-risk sub-networks). The higher the number of devices on a high-risk sub-network, the lower the Global Health Score, and vice versa. Similarly, when the device authentication fails, a separate sub-network comprising an active defence system (involving detection, identification and response) is triggered.

[0070] Figure 6 illustrates an examplary high-level data flow and gateway-edge device deployment architecture of the present invention. A goal of an RL agent is to optimise (increase) and / or maintain the global health score (GHS) of the network. This is expected to reduce the global risk factor (GRF) which ensures that a maximum number of devices are connected to the network. The RL agent learns to perform continuous authentication and defence at optimal levels to ensure a maximum of the total number of devices available are connected to network. The Global health score is calculated by summation of individual health scores of devices divided by total number of devices.

[0071] In some examples, the present invention comprises two subsystems that operate in tandem to ensure the safety and security of data. The first is a Continuous Authentication subsystem which works to gather and process information to facilitate authentication of devices. The second is an Active Defence and Response subsystem which is responsible for identification, detection and response against potential threats and anomalies accordingly.

[0072] In one example, an operational scenario includes a gateway (for example, but not limited to, a home router), edge devices (such as loT or CPS) and a configured continuous authentication and defence system. However, the deployment is not limited to the gateway. The deployment is configured to operate in a variety of ways (such as cloud / server-based deployment or a combination of techniques).

[0073] In some examples, the process begins by initiating the device registration process and static authentication protocol. Once the registration and / or static authentication is successful, the gateway proceeds to request device hardware features (data) from the edge node. This is facilitated by a Feature Requestor and Receiver module which actively collects device features (data) that are requested and subsequently sends the data to the gateway. In some examples, owing to the hybrid approach to include both hardware and network data, the network metadata of the devices can be extracted in the gateway simultaneously. Prior to transmission of the device data to the Continuous Authentication System, both data and features undergo an aggregation and pre-processing. This ensures that the transmitted information is received at the Continuous authentication module with the required data structure and type.

[0074] Various features / parameters can be identified for continuous authentication including, but not limited to, sensor defects, calibration matrices, IC defects, PUF, memory defects, magnetic sensor, phase offset, l / Q imbalance, Carrier frequency offset, Turn on / off transient signals, Clock skewness, Clock Drift, Clock Offset, temperature, Electromagnetic Radiation, RF emission, specific pattern charge depletion, instantaneous frequency error, phase error, amplitude error, Statics of Extracted Noise, signal strength, frequency responses, channel state information, external collected behavioural source, device design and settings, recourse utilization, event log, device history, RF-DNA, multivariate distribution, handshake delay, application specific delay, network delay, queueing Delay, propagation delay, transmission delay, processing delay, NOMA protocol with Time Slot, and / or NetFlow.

[0075] In some examples, the continuous Authentication module is configured to learn normal device behaviour. When deployed, the hybrid data features from the device are analysed by the RL agent in real-time to access the device’s authenticity. Once the authenticity of the device is evaluated by analysing the hybrid data features, current data is compared with a baseline, and the RL agent decides whether to authenticate the device if the current data is within a threshold deviation from the baseline, wherein the baseline defines normal device behaviour.

[0076] In some examples, the primary action of the RL agent is continuous authentication, once this process is completed, there are a series of primitive secondary actions that are triggered by the RL agent based on the outcome of the continuous authentication. The RL agent is primarily governed by the Global Health Score (GHS), although predefined primary and secondary level policies will guide the RL agent to facilitate optimal action is taken by the agent for performing the continuous authentication and defence. Further, the reward-based system provides instant feedback on actions during the training and deployment, where the RL agent will aim to maximise the reward by taking appropriate actions. In this scenario the reward is proportional to the Global Health Score (GHS) such that a higher Global Health Score corresponds to a higher the reward and vice versa. This increases security across the entire network throughout its operation.

[0077] Once the continuous authentication protocol is complete, the network clustering & protocol process is initiated where both authentication-failed and authenticated devices are clustered in respective sub-networks with custom network configurations. The devices that are authenticated by the continuous authentication will be primarily clustered into 3 sub-networks depending on the health score and risk factors. This ensures the operation of the device and its connectivity remains unchanged.

[0078] Similarly, the authentication failed devices will be clustered into a separate sub-network, where the Active Defence and Response (ADR) is triggered. The user will be promptly notified of this situation, and subsequently, the Active Defence and Response system will request for the network metadata of the authentication failed devices. The network metadata will be analysed to detect and identify the anomaly or threats. Once the identification of the anomaly is complete, the response process is initiated whereby the RL agent will trigger an optimal action to respond to the identified threat to mitigate the effect of the threat, and to return the device to a normal status for it to be reconnected to the network for normal operation. It is important to note that the device will remain in cluster four until it can be successfully continually authenticated, ensuring maximum security and safety for the device / network.

[0079] In some examples, in the event of any suspicious activity or security breach, the response system acts promptly to alert a user and follows the Cluster Four protocol to swiftly address the issue. This involves notifying the network or CPS / loT controller to isolate any affected devices and continuously monitor their behaviour for any anomalies. Additionally, the system can carry out customized actions as per the administrator's instructions, going beyond just isolating the device.

[0080] Insofar as embodiments of the invention described are implementable, at least in part, using a software-controlled programmable processing device, such as a microprocessor, digital signal processor or other processing device, data processing apparatus or system, it will be appreciated that a computer program for configuring a programmable device, apparatus or system to implement the foregoing described methods is envisaged as an aspect of the present invention. The computer program may be embodied as source code or undergo compilation for implementation on a processing device, apparatus or system or may be embodied as object code, for example.

[0081] Suitably, such a computer program is stored on a carrier medium in machine or device readable form, for example in solid-state memory, magnetic memory such as disk or tape, optically or magneto-optically readable memory such as compact disk or digital versatile disk etc., and the processing device utilises the program or a part thereof to configure it for operation. The computer program may be supplied from a remote source embodied in a communications medium such as an electronic signal, radio frequency carrier wave or optical carrier wave. Such carrier media are also envisaged as aspects of the present invention. It will be understood by those skilled in the art that, although the present invention has been described in relation to the above described example embodiments, the invention is not limited thereto and that there are many possible variations and modifications which fall within the scope of the invention. The scope of the present invention includes any novel features or combination of features disclosed herein. The applicant hereby gives notice that new claims may be formulated to such features or combination of features during prosecution of this application or of any such further applications derived therefrom. In particular with reference to the appended claims, features from dependent claims may be combined with those of the independent claims and features from respective independent claims may be combined in any appropriate manner and not merely in the specific combinations enumerated in the claims.

Claims

CLAIMS1 . A computer-implemented method to continuously authenticate a device in a network of devices using a machine learning model, the method comprising:Initially analysing hardware data associated with the device and network data associated with the network, wherein the network data is associated with all other devices in the network;- Assigning a local health score to the device based on the hardware data and a global health score to the network based on the network data;- Continuously authenticating the device at intervals, wherein the frequency of the intervals is determined based on the global health score and / or a local health score;- at one interval of the intervals, requesting hardware data and a network device data to ascertain if there is a change in the local health score and / or the global health score;If there is a change in the local health score and / or the global health score, failing to authenticate the device;- Subsequent to failing to authenticate the device, identifying and / or removing a threat in the network;- subsequent to identifying and / or removing the threat, triggering a response to reintroduce the device.

2. The computer implemented method of claim 1 , wherein the change is larger than a threshold.

3. The computer implemented method of claim 2, wherein the threshold is predetermined.

4. The computer implemented method of claim 2 or 3, wherein the threshold is based on the hardware data and / or the network data.

5. The computer implemented method of any preceding claim, wherein the machine learning model is a reinforcement learning model.

6. The computer implemented method of any preceding claim, wherein the response to reintroduce the device comprises increasing the local and / or the global health score above a minimum threshold.

7. A computer program element comprising computer program code to, when loaded into a computer system and executed thereon, cause the computer to perform the steps of a method as claimed in any of claims 1 to 6.

Citation Information

Patent Citations

  • Network sensor and method thereof for wireless network vulnerability detection

    US10498758B1

  • Continuous multifactor device authentication

    US20200244653A1