Method for operating a digital system and a digital system incorporating the method

The method uses event counters and a security matrix to detect and respond to cyber threats in digital systems, ensuring efficient and adaptive threat monitoring and response, maintaining system integrity.

JP7749006B2Active Publication Date: 2025-10-03NORTHROP GRUMMAN SYSTEMS CORP
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2023509741
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-08-31
Filing Date
2021-06-30
Publication Date
2025-10-03
Estimated Expiration
2041-06-30

AI Technical Summary

Technical Problem

Existing digital systems face challenges in efficiently monitoring for cyber threats without incurring high resource costs or limiting system complexity, and current methods for ensuring system integrity are inadequate for dynamic threat detection.

Method used

A method for operating digital systems that includes event counters, a security matrix, and a response plan module to detect and respond to non-nominal operations by comparing event counts to thresholds, implementing isolation, restriction, or corrective actions based on risk scores.

Benefits of technology

This approach enables flexible and efficient threat detection and response, maintaining system integrity by customizing monitoring to operational states, reducing potential harm from cyber threats.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007749006000001
    Figure 0007749006000001
  • Figure 0007749006000002
    Figure 0007749006000002
  • Figure 0007749006000003
    Figure 0007749006000003
Patent Text Reader

Abstract

A method for operating a digital system operable in multiple operating states having a digital resource and an event detector for detecting events in the digital resource, the method comprising: substantially determining an operational mode of at least one of the system and the digital resource during an interval; and accumulating a number of events occurring during the interval. The method further comprises comparing the accumulated events to at least one threshold associated with the operational mode; and implementing at least one of isolation, restriction, and corrective action if the comparison indicates non-nominal operation of the digital resource. A digital system incorporating the method is also disclosed.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates generally to digital systems, and more particularly to methods of operating digital systems and digital systems incorporating such methods. [Background technology]

[0002] Digital systems or subsystems, such as systems on a chip (SoC), application-specific integrated circuits (ASICs), or field-programmable gate arrays (FPGAs), as well as combinations thereof, typically implement functions such as processors, memory, controllers, interconnects, and input / output (I / O) devices. Such digital resources or components can be compromised by a hacker with unauthorized access to the digital system. Alternatively, one or more digital resources or components may malfunction, causing the digital system to operate improperly. In either case, it is important to know whether such an event has occurred. Specifically, the National Institute of Standards and Technology (NIST), the National Information Assurance Partnership (NIAP), and the Cybersecurity Maturity Model Certification (CMMF) standards already or soon will require digital subsystems to be actively monitored for dynamically emerging threats. These requirements currently apply mostly to ground networks, but in the near future they will also apply to avionics and payload systems.

[0003] Each of these resources or components is unique to it and has a particular pattern of behavior specific to the environment the system is operating in. For example, if a spacecraft subsystem is only communicating with a ground station, there may be a controller operating according to MIL-STD-1553 LAN standards and some minimal processing logic. Additionally, it is not anticipated that some other systems, such as a high-speed PCIe or SRIO link, may be operating.

[0004] Typically, each resource or component has an embedded event counter that counts the occurrence of a particular event specific to that device. For example, a processor might have an event counter that counts the number of times the processor needs to fetch data from main memory. Or, a memory controller might have an event counter that counts how many reads have occurred since reset. In either case, such counters are often used and monitored to measure performance and / or for debugging purposes. Others are added by system integrators, specifically to monitor operation at the subsystem or SoC level, for example.

[0005] Past approaches to monitoring cyberspace threats have included designs that operate quickly but incur high resource costs, or minimize resource usage but are very slow. Another approach would be to operate entirely within a closed system to ensure no rogue agents are introduced into the system. However, this approach undesirably limits the overall system's possible complexity, limiting its mission capability. Designers must also review every line of code and hardware that enters the system, ensuring the entire supply chain is free of rogue code. Summary of the Invention [Problem to be solved by the invention]

[0006] None of the above options are optimal for most applications. [Means for solving the problem]

[0007] According to one aspect, a method for operating a digital system operable in multiple operating states having a digital resource and an event detector for detecting events in the digital resource includes determining an operational mode of at least one of the system and the digital resource in effect during an interval, accumulating a number of events occurring during the interval, comparing the accumulated events to at least one threshold associated with the operational mode, and implementing at least one of isolation, restriction, and corrective action if the comparison indicates off-nominal operation of the digital resource.

[0008] According to another aspect, a digital system operable in multiple operational states includes a digital resource and an event detector configured to detect events in the digital resource. The digital system further includes a security matrix having at least one event counter configured to accumulate a number of events in the digital resource that occur when the digital system is operating in a particular operational state during a time interval and compare the accumulated events with at least one threshold associated with the particular operational state to obtain a risk score value. Still further, the threat assessment module is configured to implement at least one of isolation, restriction, and corrective action if the risk score value indicates non-nominal behavior of the digital resource.

[0009] According to yet another aspect, a digital system operable in multiple operational states comprises a digital resource and means for detecting events of the digital resource. The digital system further comprises means having at least one event counter for accumulating a number of events of the digital resource that occur when the digital system is operating in a particular operational state during a time interval, and means for comparing the accumulated events with at least one threshold associated with the particular operational state to obtain a risk score value. The digital system further includes means for implementing at least one of quarantine, restriction, and corrective action if the risk score value indicates non-nominal operation of the digital resource.

[0010] Other aspects and advantages will become apparent from consideration of the following detailed description and accompanying drawings. [Brief explanation of the drawings]

[0011] [Figure 1] This is a first example of a digital system in the form of an SoC that implements a method for operating an SoC. [Figure 2] FIG. 2 is a block diagram illustrating the state flow of the SoC of FIG. 1. [Figure 3A] FIG. 3B, when joined along similar letter lines to the right, constitutes a block diagram of the Security Matrix and Response Plan module of FIG. [Figure 3B] FIG. 3A, when joined along similar letter lines to the left, constitutes a block diagram of the Security Matrix and Response Plan module of FIG. [Figure 4] A second example of a digital system in the form of an ASIC and an FPGA, both of which include implementing the method of operating the ASIC and the FPGA. [Figure 5] 2 includes a block diagram illustrating the system of FIG. 1 in more detail. [Figure 6] 2 is a timing diagram illustrating an exemplary operation of the security matrix and response plan module of FIG. 1. DETAILED DESCRIPTION OF THE INVENTION

[0012] 1 and 2 illustrate two exemplary digital systems 10, 20, each of which may incorporate apparatus, programming, or a combination of apparatus and programming for implementing a method for operating the digital system and / or one or more subsystems and components thereof to identify and act against malicious or other threats to the operation of all or a portion of the digital system. Each digital system 10, 20 utilizes an event counter, a security matrix, and a response plan module to implement the method of operation. Specifically, according to one embodiment, when the system 10, 20, or a portion of such a system including one or more subsystems, is operating in a particular state, the event counter counts events of a device or portion thereof, such as a system / subsystem component implementing a digital (e.g., hardware) resource, and the resulting count includes one or more digital signatures of the functionality of the respective device or device portion while the entire system or subsystem remains in that state. The digital signature(s) are compared by the security matrix with the associated digital signature(s) of a nominal operational profile for the current system / subsystem state, reflecting the expected nominal behavior of the respective device or device portion, to obtain a risk score value (referred to as parameter "risk_score"). A risk score indicating unexpected behavior of the associated device or device portion results in the invocation of one or more actions by a response plan module to minimize the security and / or operational threat.

[0013] If or when a system / subsystem changes state to reflect a change in functionality, a different nominal operational profile containing one or more different nominal digital signatures can be loaded and compared with the cumulative events to derive one or more further risk scores indicating whether the functionality of the device / device portion is within the operational criteria for that state.

[0014] In an exemplary embodiment, the risk score is preferably determined on a per-device or per-subset of devices (e.g., per hardware resource in situations where a device implements multiple hardware resources), optionally on a per-function basis. Thus, for example, it may be desirable to monitor a particular function that may implement a set or subset of devices, a portion of devices, or any combination thereof. In one embodiment, the risk score reflects the cumulative activity of a particular device or portion of devices over a selected period of time and the deviation of that cumulative activity value from expected high and / or low thresholds. Each threshold may be predetermined and / or dynamically determined / updated during operation. Thus, for example, further / subsequent system states may be used to define different threshold boundaries, as described below. For example, for each hardware resource, the designer selects the events associated with monitoring and counting, the associated counting frequency, the timing and duration of the counting period and sampling interval. For example, for events that occur at almost every clock cycle in all subsystem states, accumulating such events with high precision may not be very important. Specifically, all such events may be counted, but the accuracy of the counter values ​​that cross a threshold and may trigger a potential response plan may be tracked with less precision. Thus, for example, the accuracy of tracking such events to trigger a response plan may be every 1000 or 10,000 events. Thus, the event counter of a hardware resource may pulse a signal to its associated security matrix that may reflect that an event has occurred 1000 or 10,000 times (or more or less), while counting every single event that potentially occurs at almost any clock.Thus, although the granularity at which activity is monitored in the security matrix is ​​not as great as with an event monitor, the system can still count all events that occur potentially on every clock cycle, thus enabling high-precision monitoring of activity without having large hardware counter resource requirements. In the above example, each time the event counter for the monitored resource reaches 1000 or 10000 events, the counter is reset and the security matrix counter is incremented by one.

[0015] Furthermore, it is possible to have two distinct system states that may be closely related, and may have the same threshold of 95% for nominal activity, while the remaining 5% may be used to allow for different thresholds that separate the two states. The requirement for dynamic threshold updates should be based on the response time required to detect and react to off-nominal conditions. If a hardware resource has a very large dynamic range for nominal activity and response time to a perceived threat is critical, dynamically changing activity thresholds may be preferable to changing the entire system state. Monitoring systems can be configured so that monitored resources have their own security matrix, multiple subsystem states of the resource, and their own dedicated hardware / software for formulating a response plan. That response plan can be fed into another system that monitors the entire SoC; therefore, if a subsystem's value is particularly high, consideration can be given to hierarchically cascading the monitor / response plan pair to another monitor / response plan system.

[0016] In a particular embodiment, the risk_score is determined for a particular device or portion of a device (e.g., a component implementing a hardware resource) by detecting one or more event increments, where each event increment can include a single detected event or multiple detected events, optionally applying a weighting factor to each event increment, and then, in situations where multiple weighted event increments are detected, summing all weighted events occurring during a particular interval to obtain an event sum value. The event sum value, denoted "event_sum," is then compared to at least one threshold value, e.g., high and low threshold values ​​specific to the system, subsystem, device, component, or portion thereof, and a current operating state associated with the system, subsystem, device, component, or portion thereof. If the event_sum changes from at least one threshold value, the event_sum is incremented by a particular amount. In one specific embodiment, if event_sum exceeds or falls below a respective high or low threshold (i.e., if event_sum falls outside the threshold range defined by the high and low thresholds), event_sum is incremented by an amount dependent on the absolute value indicating the degree to which the value of event_sum is outside the threshold range, as well as the current system / subsystem state. In an alternative embodiment, the value of event_sum is compared to multiple high and / or low thresholds (i.e., multiple threshold ranges), and event_sum is incremented by different amounts depending on which threshold range it falls outside, to obtain a risk_score value indicating the perceived risk of the monitored activity relative to expectations for the device or device portion and current system / subsystem state. The risk score can be used to invoke one or more actions, which can be isolation, restriction, and / or corrective in nature.

[0017] Additionally, in one embodiment, the value of event_sum may be decremented at a particular rate as a function of time during a period to obtain the risk_score, so that any high risk scores can be averaged over a period of time. In a specific embodiment, the decrement rate may be a function of a time interval, referred to as the decrement_interval, that is deemed relevant to the device or device portion being monitored as a function of a trusted clock source. When the assigned decrement_interval expires or passes, the risk_score is decremented by a value based on the value decrement_coefficient, which is a weighting factor that determines how much to subtract from the event_sum to obtain the value risk_score. In one embodiment, the value of risk_score is prevented from dropping below zero.

[0018] Thus, for example, a subsystem's operational profile may consist of risk scores across all of the subsystem's monitored hardware resources. These risk scores are aggregated in a dedicated hardware resource table and compared to the expected (i.e., nominal) values ​​for the system / subsystem, as represented by its nominal operational profile. A disappointing operational profile results in action at the subsystem or other level (e.g., system-wide) to minimize the potential harmful effects of the perceived cybersecurity threat.

[0019] Referring to FIG. 1 , one embodiment of a system-on-chip (SoC) device 10, which may be used as a spacecraft or other vehicle controller, includes one or more devices or portions of devices that implement one or more hardware resources 22 and one or more event monitors 24. There may be as many event monitors 24 as there are hardware resources 22, such that each event monitor 24 is associated with one or more events for each hardware resource 22 and monitors the occurrence of those events. Alternatively, the device 20 may include a different number of event monitors 24 than there are hardware resources 22. In either case, the output of the event monitors 24 is coupled to one or more instances of a security matrix module 26, each of which develops a risk_score for each of the one or more hardware resources. If one or more of the risk_score values ​​indicate an off-nominal operating condition, a response plan module 28 manages the operation of one or more of the hardware resources, the system, and / or the subsystems. In a preferred embodiment, as described in more detail below, response plan module 28 invokes one or more other actions associated with the hardware resource(s), system, subsystem, or portion thereof, depending on the identification of the hardware resource(s) operating outside the threshold(s) during a selected time period and the magnitude(s) of the variation of the total weighted event count(s) from the threshold(s).

[0020] Referring now to FIG. 2, SoC 10 is operable in several states, indicated by blocks 30-40 and 44-50. Blocks 30, 32, 34, 36, and 38 represent boot states operable during a boot sequence. (The acronyms SCC, TEE, I / O, HSCC, and REE stand for Security Core Complex Trusted Execution Environment, Input / Output, High Speed ​​Core Complex, and Requirements Engineering Environment, respectively, as will be apparent to those skilled in the art.) The SCC is a dedicated security partition firewalled from the rest of the SoC and includes dedicated processing elements and associated memory, as well as its own software-based execution environment, or TEE. The SCC is where high-security functions within the SoC take place, such as locking down hardware and shared resources and then configuring security policies for the rest of the higher-speed hardware and shared resources before bringing them online. The HSCC is a high-speed core complex containing high-speed processing elements, memory, and the like. It follows the security policies established by the SCC, including accessible hardware resources, memory maps, and accessible I / O. This complex performs all performance-oriented processing in accordance with security policies established and monitored by the SCC. Block 40 enforces the nominal operating state of the SoC. Block 42 implements one or more instances of security matrix 26 and response plan module 28 of FIG. 1. Block 42 also includes threat response logic in the form of hardware, software, firmware, or a combination thereof, which responds to event monitor(s) 24, develops indications of perceived threats to SoC 10, such as unauthorized intrusions, and invokes actions in the states represented by blocks 44-50 in response to the perceived threats. The states implemented by such blocks, and their general descriptions, are as follows: Audit (block 44): Perform unscheduled automated audit and self-test routines of the system, and / or subsystems, and / or devices, and / or hardware resources, and / or one or more portions thereof. Quarantine Devices (block 46): Devices and / or hardware resources that have an outside threshold number of events during a selected time period are quarantined from the rest of the SoC 10. Quiesce SoC (block 48): The SoC 10 is placed in a minimal operating state. Safe Mode (Block 50): Operation of the SoC 10 is placed in safe mode.

[0021] In the illustrated embodiment, recovery from safe mode requires external intervention (block 52), for example, by a reset signal transmitted from a ground station when SoC 10 is used on a satellite.

[0022] 3A and 3B illustrate an example of threat response logic 42, assuming that there are N event monitors monitoring N hardware resource events, that threats are evaluated by a single instance of security matrix 26 at the end of multiple consecutive intervals, and that threat response logic 42 includes one or more event counters 60-0, 60-1, ..., 60-N that accumulate counts of the N events detected by event monitors 24. Note that embodiments are not limited to this example. For example, a single hardware resource or other element can be monitored by one or more event monitors, and in any case, there need not be a one-to-one relationship between the event monitor and the monitored device or portion thereof. For example, a single event monitor can be time-multiplexed to detect events from multiple hardware resources or other elements. Furthermore, when a single hardware resource is monitored by multiple event monitors, typically different functional aspects of the single hardware resource may be monitored by the associated event monitor. Additionally, the outputs of multiple event monitors can be counted and the counts processed to derive a risk score for: (1) a single functional aspect of a single hardware resource; or (2) a portion or all of the hardware resource. or (3) all or a portion of a system, one or more subsystems or one or more portions thereof, or one or more devices or components or one or more portions thereof. Generally, but not exclusively, events occurring in a system, one or more subsystems or portions thereof, or one or more devices or components / functions may be monitored, counted, and processed to derive a risk score for the system or subsystem or portion thereof. For example, any set of one or more event monitors may be used to monitor any particular set of functions or features, provided that the hardware has the accessibility and configurability to support these various modes within a digital system. If different events are accessible and configurable to be monitored by a fixed set of event monitors, different events for different operating states may be deemed more appropriate for monitoring than other events.Therefore, some type of reconfigurability of the event monitor when system conditions are changing is a possible embodiment of the event monitor logic.

[0023] As seen in FIG. 3A, at the end of the interval, a plurality of multipliers 62-0, 62-1, ..., 62-N multiply the current values ​​of the counted events, denoted as event0_count, event1_count, eventN_count, by weighting values, denoted as event0_weight, event1_weight, ..., eventN_weight, respectively, and an adder 64 adds the multiplied values ​​to obtain a preliminary value of event_sum as follows: event_sum = event0_count*event0_weight + event1_count*event1_weight + … + eventN_count*eventN_weight …(1) The weighting values ​​reflect different perceived magnitudes of risk identified by cumulative events of the associated monitored element, such as a hardware resource.

[0024] The preliminary value of event_sum is added to the high threshold HI by adders 66 and 68, respectively. THR and low threshold Lo THR and compared with the HI of event_sum THR and LO THR. The first and second deviation values ​​are obtained, which represent the magnitude of the deviation from the event_sum. The limits 67 and 69 limit the magnitude of the first and second deviation values ​​so that they do not become equal to or less than zero. The limited deviation values ​​are obtained when the event_sum is THR and L.O. THR The event_sum value indicates whether the event_sum is outside the threshold limit represented by (x,y) and, if so, the magnitude (i.e., absolute magnitude) of the deviation of event_sum from the threshold limit. The bounded deviation value is provided to summer 70, which adds such value to the preliminary value of event_sum to obtain an updated value of event_sum.

[0025] As shown in FIG. 3B, the updated value of event_sum is compared against two upper thresholds high_threshold0 and high_threshold1 and a lower threshold low_threshold0 by blocks 72, 74, and 76, respectively, and the result of the comparison may optionally be used by block 78 to decrement the updated value of inventory_sum by the decrement_value obtained from blocks 80, 82 depending on the current interval indication in the current time period (discussed below in connection with FIG. 6) to obtain the value risk_score as follows: decrement_value = current value of decrement_interval*decrement_coefficient …(2) if((event_sum > high_threshold1) and (previous_risk_score + 2 - decrement_value > 0)) then risk_score = previous_risk_score + 2 - decrement_value; if((event_sum > high_threshold0) and (previous_risk_score + 1 - decrement_value > 0)) then risk_score = previous_risk_score + 1 - decrement_value; else if((event_sum < low_threshold0) and (previous_risk_score + 1 - decrement_value > 0)) then risk_score = previous_risk_score + 1 - decrement_value …(3) where previous_risk_score is the value of risk_score determined for the interval immediately preceding the current time period. Note that decrementing event_sum to obtain risk_score is optional and may not be performed if it is determined that the undecremented value of event_score is considered a better indicator of whether a cybersecurity event is occurring or has occurred. In this case, the updated value of event_sum is used as the value risk_score.

[0026] The magnitude of the risk_score thus determined is evaluated and classified by block 84 as falling into one of five ranges of values: "no risk," "low risk," "medium risk," "high risk," and "extreme risk." Block 86 triggers the operation of the SoC depending on the classification performed by block 84. Specifically, a threat assessment of no risk causes the SoC to operate in the nominal state of block 40 in FIG. 2. A threat assessment of low risk, medium risk, high risk, or extreme risk causes the SoC to operate in the states shown by blocks 44, 46, 48, and 50 in FIG. 2, respectively. Blocks 84 and 86 implement the response plan module 28 in FIG. 1. Meanwhile, the balance of the elements in FIGS. 3A and 3B implements the security matrix 26.

[0027] Referring again to FIGS. 3A and 6, the reset module 90 is operable at the end of each interval to reset the value accumulated by the event counter 60 to an initial value (e.g., zero) and also reset the decrement_interval value of FIG. 3B at the end of each period to an initial value (e.g., zero). Note that the latter resets need not be synchronous with respect to the operation of the various modules. Also, depending on the decrement and increment values ​​used, the risk_score value may be reset periodically (e.g., after a certain number of intervals or periods, or upon a change in the state of the system), aperiodically, or never reset and allowed to persist for part or all of the period. Thus, as seen in FIG. 6, in one embodiment, the event counter 24 counts constantly during each sampling interval, and then the accumulated value is detected at the end of the interval. The security matrix 26 updates and thus determines the impact of the risk score immediately after the end of the interval. The event monitor 24 resets after each interval to begin a new sampling interval. However, the security matrix 26 preferably maintains a persistent state so as to adequately characterize events over longer periods of time, thereby maintaining counts over multiple sampling intervals. How often the event monitor updates the counts in the security matrix, and what the counts represent, can be a function of the frequency of the events being monitored. Higher frequency events can be assigned lower fidelity per bit in the event counter, so that a single count value can represent hundreds or thousands of positive trigger events.

[0028] In general, a nominal operational profile for a system / subsystem can be constructed from nominal risk scores across all monitored hardware resources using the methodology described above, assuming that the system / subsystem is operating in a particular state and that there are no threats attempting to affect the operation of the system / subsystem. These nominal risk scores may be tabulated and stored in dedicated hardware resource tables within security matrix 26 for each system / subsystem state, and in the illustrated embodiment, are set by a threshold HI THR and L.O. THR and optionally one or more initial values, variables and / or constants for one or more of high_threshold0, high_threshold1, and low_threshold0. The subsequent value(s) of risk_score may or may not be stored in a resource table and / or may be compared to its associated nominal value for the operating state. In some embodiments, the results of the comparison may be used to determine action to be taken and / or modification of the nominal value to take into account the new operating conditions.

[0029] The response plan must be executed by a supervisor or other processor with the authority and ability to access all control mechanisms within the subsystem to bring the hardware resources, including quarantines, back into compliance. These control mechanisms can be in the form of credit configuration, device reset, forced quiesce, processor reset, and anything else related to bringing down the hardware device in either a graceful or forced manner.

[0030] The above-described method can implement any number of levels of risk_score increase, leading to different behaviors of the associated system / subsystem or other element(s) and / or portion thereof, and can use any number of high / low thresholds. For example, when the risk score reaches a certain predetermined threshold level, the security matrix 26 can generate a message sent as an interrupt to the monitoring processor agent directing appropriate action commensurate with the level of perceived risk. This can take the form of an interrupt line from the security matrix 26 with a specific interrupt vector reflecting different response plans that can bring the malfunctioning hardware resource back into compliance or isolate it. If the risk is considered relatively low, the response plan can implement increased logging in the form of enabling other monitors or software-based monitors for that hardware resource. If the perceived threat level is high, the response plan can forcefully disable the device, isolate it, and then continue the subsystem until live intervention occurs. If the threat level is critical, the response plan can shut down all processor cores except the monitoring core, forcefully halt all internal I / O and memory traffic, and enter a safe state, as described below. Therefore, the higher the risk score, the more restrictive the action can be.

[0031] If necessary, additional default response plans can be implemented for scenarios not explicitly covered by the remaining response plans. The default plan can be in the form of a system safe state in which all application processor cores are stopped and parked, all unnecessary internal and external traffic is stopped, and subsystems are only performing tasks that are mission-critical to the functioning of the spacecraft or other facility while awaiting ground or other command intervention.

[0032] 4 illustrates a digital system 20 including a combination of an ASIC 100 and an FPGA 102 that together implement a control system useful, for example, as a spacecraft or other vehicle controller. Note that ASIC 100 may be replaced with two or more ASICs, one or more FPGAs, or a combination thereof. Also, the FPGA may be replaced with two or more FPGAs or one or more additional ASICs, as desired. Various elements of the ASIC(s) and / or FPGA(s) may be implemented in software, hardware, firmware, or a combination thereof.

[0033] Response plan module 28 is implemented by ASIC 100, and security matrix 26 is implemented by FPGA 102. If desired, module 28 and security matrix 26 may alternatively be implemented by FPGA 102 and ASIC 100, respectively, or both elements 26, 28 may be implemented by either FPGA 102 or ASIC 100, or portions of one or both elements 26, 28 may be implemented by one or both of FPGA 102 and / or ASIC 100.

[0034] 5 shows system 20 in more detail. Event monitors 124a and 124b monitor events occurring during operation of DDR controller 126 and one or more high-performance CPUs 128, both of which form part of ASIC 100's high-speed core complex 130. CPU 128 communicates with AXI4 communication bus 134 via level 2 cache 132. Bus 134 also communicates with DDR controller 126, clock 136, and XAUI communication bus 138 via AXI4-to-XAUI bridge 140. DDR controller 126 further communicates with secure and non-secure RAM modules 144, 146 via DDR communication bus 142.

[0035] Secure core complex 132 of ASIC 100 includes TCM 134, which implements, among other things, functionality represented by block 136 similar to or identical to response plan module 28 of FIG. 1. One or more event monitors 124c monitor events that occur during operation of TCM 134. TCM 134 further communicates with a secure CPU 150 that is coupled to an associated boot ROM module 152. Secure CPU 150 communicates with an AXI3-to-PCIe bridge 158 via a level 2 cache module 154 and an AXI3 communication bus 156.

[0036] FPGA 102 may include a radiation-hardened FPGA called an RTG4® module manufactured and / or sold by Microsemi Corporation of Aliso Viejo, California, and includes one or more event monitors 124d-124m that monitor various events for various devices, at least some of which are coupled to one or more instances of security matrix 160. Each instance of security matrix 160 may perform similar or identical functions to security matrix 26 of FIG. 1, although different instances may perform different functions, such as comparisons with other threshold values, other increment / decrement values, and combinations thereof, and each instance of security matrix 160 may invoke the same or different threat response actions via interrupt control module 190, described below.

[0037] Specifically, event monitors 124d-124l are respectively coupled to and receive event occurrences from I / O interrupt control module 170, multiple UART modules 172, XAUI-AXI3 bridge 174 in communication with XAUI bus 138, MIL-STD-1553 LAN controller 178, SpaceWire communication network 180, firewall 182, PCIe communication bus 184, AXI3-PCIe bridge 186 in communication with AXI3-PCIe bridge 158 of ASIC 100 through PCIe communication bus 188, and interrupt control module 190. Mailbox registration module 191, bridge 174, firewall 182, bridge 186, and interrupt control module 190 communicate with each other via AXI3 communication buses 192, 194, 196, and 198. If desired, any of the communication protocols can be replaced with different protocols and associated buses and bridges as desired. Thus, for example, an AXI4 bus can be replaced with an AXI3 bus (with associated bridges as needed), a PCIe bus and associated bridges can be replaced with a Serdes bus and associated bridges, etc. Thus, one or more communication protocols can be changed without affecting the performance of various embodiments.

[0038] The UART module 172, the 1553 LAN controller 178, the SpaceWire communications network, the plurality of SRIO modules 200, and the PCIe communications bus 184 communicate with remote devices / systems (collectively referred to as the “outside world”) in any suitable manner. The SRIO module 200 may be associated with one or more of the event monitors 124m, the output(s) of which may or may not be provided to one or more instances of the security matrix 160.

[0039] The monitored event occurrences from event monitors 124d-124l, along with the monitored event occurrences from event monitors 124a-124c in ASIC 100, are provided to one or more instances of security matrix 160. A selected one or all of the monitored event occurrences are fed to each instance of security matrix 160. For each instance of security matrix 160, as described above, the counts of selected monitored event occurrences occurring during an interval are weighted, the weighted values ​​are added and compared to a threshold to obtain an event_sum value for that interval, the event_sum value is optionally decremented depending on whether the absolute magnitude of event_sum or the decremented value is a better indicator of a cybersecurity threat, and the generated risk_score value is used to determine whether to take action, as described above. [Industrial Applicability]

[0040] In summary, this approach provides a way to customize threat detection monitoring in heterogeneous digital subsystems to address the growing capabilities of external threats to national security. Alternatively, or in addition, the operational health of the system / subsystem can be verified and appropriate action can be taken to minimize potential or actual adverse effects resulting therefrom. Overall, this system enables hardware-defined monitoring to detect non-nominal behavior, triage deviations, and enact a plan to bring the system back into compliance. When compliance is not achieved, risk reduction becomes the driving force. Different operational states of a subsystem can warrant different security profiles associated with the subsystem's mode. This gives the threat monitoring system flexibility and sensitivity regarding operational context, ensuring system engineers understand what the system is doing under all circumstances.

[0041] All references herein to a "system" or "systems" include not only SoCs, ASICs, and FPGAs, but also any other device that includes one or more systems or subsystems. Also, references herein to a "system" or "systems" shall be construed as references to a "subsystem" or one or more portions thereof, respectively, and vice versa, to the extent that the features disclosed herein apply equally to the entire system or subsystem or any one or more portions thereof.

[0042] All references cited in this specification, including publications, patent applications, and patents, are herein incorporated by reference to the same extent as if individually and specifically indicated to be incorporated by reference in their entireties.

[0043] The use of the terms "a," "an," and "the" and similar references in the context of describing the present invention (particularly in the context of the claims below) should be construed to cover both the singular and the plural unless otherwise indicated herein or clearly contradicted by context. The recitation of ranges of values ​​herein is merely intended to serve as a shorthand method of referring individually to each separate value falling within the range, unless otherwise stated herein, and each separate value is incorporated herein as if it were individually listed herein. All methods described herein can be performed in any suitable order unless otherwise stated herein or clearly contradicted by context. The use of any and all examples, or exemplary language (e.g., "etc.") provided herein is intended merely to better clarify the disclosure and does not pose a limitation on the scope of the disclosure unless otherwise claimed. No language in the specification should be construed as indicating any non-claimed element as essential to the practice of the invention.

[0044] Numerous modifications to the present invention will be apparent to those skilled in the art in light of the above description. It should be understood that the illustrated embodiments are merely exemplary and should not be construed as limiting the scope of the present invention.

Claims

1. 1. A method of operating a digital system operable in a plurality of system-level operating states, the method comprising: a digital resource; and an event detector for detecting an event in the digital resource, the method comprising: determining during an interval which of the plurality of system-level operating states the digital system is in; Accumulating events of the digital resource occurring during the period; comparing the accumulated events to at least one threshold value related to the determined system-level operating conditions; If the comparison indicates non-nominal behavior of the digital resource, then implementing at least one of isolation, restriction, and corrective action.

2. The method of claim 1 , wherein the comparing comprises determining whether the digital resource is operating outside of nominal limits.

3. The method of claim 1 , wherein the taking comprises: defining a plurality of actions; and selecting one of the plurality of actions in response to the accumulated events.

4. The method of claim 3 , wherein the comparing comprises comparing the accumulated events to an upper limit and a lower limit.

5. the digital system having a plurality of further event detectors for detecting a plurality of further events; the accumulating further includes accumulating additional events, weighting the accumulated events, and summing the accumulated and weighted events to obtain an event total; The method of claim 1 , wherein the comparing comprises comparing the event sum to the at least one threshold value.

6. The method of claim 5, further comprising incrementing the event total if the event total is outside the range indicated by the at least one threshold value.

7. The method of claim 6 , further comprising decrementing the incremented event sum as a function of a time interval.

8. 7. The method of claim 6, wherein at least one of the isolation, restriction, and corrective action comprises isolating the digital resource, quiescing the digital system, operating the digital system in a safe mode, and auditing at least a portion of the digital system.

9. 1. A digital system operable in multiple system-level operating states, comprising: Digital resources and an event detector for detecting an event in the digital resource; a security matrix having at least one event counter that accumulates events of the digital resource that occur while the digital system is operating in a particular system-level operating state among the plurality of system-level operating states over a time interval, and compares the accumulated events with at least one threshold associated with the particular system-level operating state to obtain a risk score value; a threat assessment module that implements at least one of isolation, restriction, and corrective action if the risk score value indicates non-nominal behavior of the digital resource.

10. 10. The digital system of claim 9, wherein the security matrix is ​​operable to determine whether the digital resource is operating outside of nominal limits.

11. The digital system of claim 10 , wherein the security matrix is ​​operable to compare the accumulated events to upper and lower limits.

12. 11. The digital system of claim 10, wherein the digital system has a plurality of additional event detectors that detect a plurality of additional hardware events, and wherein the security matrix is ​​further operable to accumulate the additional hardware events, weight the accumulated hardware events, sum the accumulated and weighted hardware events to obtain an event total, and compare the event total to at least one threshold associated with the particular system-level operating state to obtain a risk score value.

13. 13. The digital system of claim 12, wherein the security matrix is ​​further operable to increment the event sum if the event sum is outside a range indicated by the at least one threshold.

14. 14. The digital system of claim 13, wherein the security matrix is ​​further operable to decrement the incremented event sum as a function of a time interval.

15. 10. The digital system of claim 9, wherein at least one of the isolation, restriction, and corrective action comprises quarantining the digital resource, quiescing the digital system, operating the digital system in a safe mode, and auditing at least a portion of the digital system.

16. 1. A digital system operable in multiple system-level operating states, comprising: Digital resources and detection means for detecting an event in the digital resource; an accumulation means having at least one event counter for accumulating events of the digital resource that occur when the digital system is operating in a particular system-level operating state of the plurality of system-level operating states during a time interval, and means for comparing the accumulated events with at least one threshold associated with the particular system-level operating state to obtain a risk score value; and an enforcement means for enforcing at least one of quarantine, restriction, and corrective action if the risk score value indicates non-nominal behavior of the digital resource.

17. 17. The digital system of claim 16, wherein the accumulating means includes means for comparing the accumulated events with upper and lower limits.

18. 20. The digital system of claim 17, wherein the detecting means detects a plurality of additional hardware events, the accumulating means accumulates the additional hardware events, and includes means for weighting the accumulated hardware events, and means for summing the accumulated and weighted hardware events to obtain an event total and comparing the event total with at least one threshold associated with the particular system-level operating state to obtain a risk score value.

19. 20. The digital system of claim 18, wherein the accumulating means further comprises: means for incrementing the event sum if the event sum is outside a range indicated by at least one threshold.

20. 20. The digital system of claim 19, wherein the accumulating means further includes means for decrementing the incremented event sum as a function of a time interval to obtain a risk score, and the implementing means includes means responsive to the risk score for invoking quarantine of the digital resource, quiescing the digital system, operating the digital system in a safe mode, and auditing at least a portion of the digital system.

Citation Information

Patent Citations

  • Illicit operation management device, illicit operation management module and program

    JP2008191857A

  • Information processing device and information processing method

    JP2015082191A

  • System and method of adapting patterns of dangerous behavior of programs to computer systems of users

    JP2019220132A

  • Techniques for discovering and managing application security

    JP2019510304A

  • User activity monitoring

    US20160306965A1