Oblivious Screening of Data Streams

By introducing inadvertent screening technology into the data screening system, the combination of ORAM and trusted memory is used to solve the problem of difficulty in preventing side channel attacks in data use in the prior art, achieving higher data privacy and security, and reducing storage costs.

CN112970022BActive Publication Date: 2025-05-27VISA INTERNATIONAL SERVICE ASSOCIATION
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN201980073084.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-09-28
Filing Date
2019-09-27
Publication Date
2025-05-27
Estimated Expiration
2039-09-27

AI Technical Summary

Technical Problem

The prior art is difficult to effectively protect data in use from side channel attacks, especially in trusted computing processor architectures, where attackers can infer user information by observing data access patterns.

Method used

An inadvertent screening system is used to mitigate the impact of side channel attacks by fuzzing the results of the data screening system. Specific implementations include the use of inadvertent random access memory (ORAM) combined with trusted memory, hiding memory access modes, and reducing information leakage through buffer management and virtual write operations.

Benefits of technology

Effectively mitigate the impact of side channel attacks, reduce the possibility of attackers inferring user information, improve data privacy and security, while reducing the size and usage cost of trusted memory.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112970022B_ABST
    Figure CN112970022B_ABST
Patent Text Reader

Abstract

A technique for oblivious screening may include receiving an input data stream having a plurality of input elements. For each of the received input elements, determining whether the input element satisfies a screening condition. For each of the received input elements that satisfy the screening condition, performing a write operation to store the input element in a memory subsystem. For those of the received input elements that do not satisfy the screening condition, performing at least a virtual write operation on the memory subsystem. When the memory subsystem is full, the contents of the memory subsystem may be evicted to an output data stream. The memory subsystem may include a trusted memory and an unprotected memory.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application claims priority to U.S. Provisional Application No. 62 / 738,752, filed on September 28, 2018, entitled “OBLIVIOUSFILTERING OF DATA STREAMS,” the entire contents of which are incorporated herein by reference for all purposes. Background Art

[0003] Application developers increasingly rely on shared cloud infrastructure to host and process databases that often contain sensitive user data. Even in strictly regulated industries such as healthcare and finance, production databases are often hosted in silos protected by weak defenses such as firewalls (for network isolation), monitoring tools, and role-based access controls. This puts sensitive data at risk of being compromised due to a number of reasons, including vulnerabilities in the application software stack, privilege escalation vulnerabilities in cloud infrastructure software, inadequate access controls, firewall misconfigurations, and malicious administrators. Although typical guidelines require encryption of data at rest (e.g., using authenticated encryption) and in transit (e.g., using Transport Layer Security (TLS)), such techniques are insufficient to guarantee privacy from the above threats because they do not protect data in use.

[0004] Recent advances in trusted computing processor architectures have enabled hardware-isolated execution at native processor speeds while protecting data in use. However, since applications running inside hardware enclaves interact with the host platform or external interfaces to access resources and perform I / O operations, even isolated execution environments may be exposed to side-channel observations such as access patterns, network traffic patterns, and output timing. Such side-channel observations may allow an attacker to use statistical inference to recover information about application users. For example, even in an ideal trusted execution environment, an attacker may still observe cache timing and hyperthreading leaks. As encrypted databases are increasingly commercialized and deployed, mitigation of such channels becomes increasingly important.

[0005] Embodiments of the present disclosure address these and other problems individually and collectively. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] Figure 1 A block diagram of a system according to some embodiments is shown.

[0007] Figure 2 An example of screening an encrypted data stream in accordance with some embodiments is shown.

[0008] Figure 3A system model for inadvertent screening is shown in accordance with some embodiments.

[0009] Figure 4 Graph showing buffer size and expected element level leakage according to some embodiments.

[0010] Figure 5 A system model for an oblivious random access memory (ORAM) assisted screening algorithm is shown in accordance with some embodiments.

[0011] Figure 6 The process of inadvertent screening is shown in accordance with some embodiments. DETAILED DESCRIPTION

[0012] A side channel attack may refer to a network security threat in which an unauthorized entity is able to observe data access patterns and infer certain information about the data consumer based on data access pattern analysis. Encryption of the underlying data can mitigate the risk of an attacker gaining access to the underlying data, but encryption itself does little to prevent side channel attacks because such attacks rely on data access patterns and not necessarily on the underlying data. In the context of a data filtering system, an attacker can observe at what point in time the input data stream passes through the filter. This information may allow the attacker to infer certain user information, such as user preferences, interests, access timing and frequency, etc. Therefore, the technology described herein provides an oblivious filtering system that mitigates the effects of such side channel attacks by obfuscating the results of the data filtering system.

[0013] Before discussing the embodiments of the present invention, some terms may be described in further detail.

[0014] A computing device may be a device that includes one or more electronic components (e.g., an integrated chip) that can process data and / or communicate with another device. A computing device may provide remote communication capabilities to a network and may be configured to transmit data or communications to other devices and receive data or communications from other devices. Examples of computing devices may include computers, portable computing devices (e.g., laptops, netbooks, ultrabooks, etc.), mobile devices such as mobile phones (e.g., smartphones, cellular phones, etc.), tablet computers, portable media players, personal digital assistant devices (PDAs), wearable computing devices (e.g., smart watches), electronic reader devices, etc. A computing device may also be in the form of an Internet of Things (IoT) device or vehicle (e.g., a motor vehicle such as a networked car) equipped with communication and / or network connection capabilities.

[0015] The server computer or processing server may include a powerful computer or computer cluster. For example, the server computer may be a large mainframe, a small computer cluster, or a group of servers that work like a unit. In one example, the server computer may be a database server coupled to a network server. The server computer may be coupled to a database and may include any hardware, software, other logic, or combination of the foregoing for servicing requests from one or more client computers. The server computer may include one or more computing devices and may use any of a variety of computing structures, arrangements, and compilations to service requests from one or more client devices.

[0016] A key may refer to a piece of information used in a cryptographic algorithm to transform input data into another representation. A cryptographic algorithm may be an encryption algorithm that transforms original data into an alternative representation, or a decryption algorithm that transforms encrypted information back into original data. Examples of cryptographic algorithms may include Advanced Encryption Standard (AES), Data Encryption Standard (DES), Triple Data Encryption Standard / Algorithm (TDES / TDEA), or other suitable algorithms. The key used in the cryptographic algorithm may have any suitable length (e.g., 56 bits, 128 bits, 169 bits, 192 bits, 256 bits, etc.). In some embodiments, longer key lengths may provide more secure encryption that is less susceptible to hacker attacks.

[0017] Trusted memory may refer to a memory with tamper-resistant features to protect the memory and / or its contents. The memory may be implemented using DRAM, SRAM, flash or other suitable memory or storage medium. The tamper-resistant features may be implemented using hardware, software or a combination of both. Examples of tamper-resistant features may include sensors for sensing abnormal temperature, voltage, current, etc., on-chip encryption / decryption functions, on-chip authentication, backup power supplies (e.g., batteries), tamper-triggered automatic erase features, internal data reorganization features, etc. The memory portion of the trusted memory may be implemented using DRAM, SRAM, flash or other suitable memory or storage medium. Compared to standard memory, trusted memory may generally have lower density and higher cost.

[0018] Unprotected memory may refer to a device capable of storing data. Unprotected memory may lack any special features to protect the memory or its contents. Unprotected memory may be implemented using DRAM, SRAM, flash, or other suitable memory or storage media. Unprotected memory may typically have higher density and lower cost than trusted memory.

[0019] Oblivious random access memory (ORAM) may refer to a memory system that is capable of masking memory access patterns. For example, ORAM may continuously reshuffle and re-encrypt data contents as they are being accessed to obscure which data elements are being accessed at the time. ORAM may employ symmetric encryption or homomorphic encryption.

[0020] Figure 1 A block diagram of a data screening system 100 according to some embodiments is shown. The system 100 may include a data source device 110, a data consumption device 120, and a data screening device 150. The data source device 110 may include a data storage or processing device that provides a data stream to a data consumer. The data source device 110 may be implemented using a server computer or other type of computing device that can service data requests from data consumers. Examples of the data source device 110 may include a database system that stores and provides database entries in response to queries, a content server that provides streaming media content, a network-enabled device that provides network packets, and the like. When the data source device 110 receives a data request from the data consumption device 120, the data source device may respond by obtaining the requested data from the data storage area and providing a data stream containing the requested data to the data consumption device 120.

[0021] The data consumption device 120 may be a device that consumes or uses data obtained from a data source. The data consumption device 120 may be implemented as a client computing device, or other type of computing or communication device that provides data or information to an end user. Examples of data consumption devices 120 may include a user computer, a portable communication device such as a smart phone, an Internet of Things (IoT) device, a media device such as a television, a display, and / or a speaker, etc. The data consumption device 120 may send a request for data to the data source device 110, receive a data stream corresponding to the requested data, and present or otherwise provide the requested data to the end user. In some embodiments, the data consumption device 120 may also process the data received from the data source and provide the processed data to another device for consumption.

[0022] Data screening device 150 can be a device that receives an input data stream and provides an output data stream derived from the input data stream. In some embodiments, data screening device 150 can be an independent device that is coupled between data source device 110 and data consumption device 120 in a communication manner. For example, data screening device 150 can be a network device, such as a router coupled between a networked data source and a client computer. As another example, data screening device 150 can be a set-top box coupled between a media source and a display device. In some embodiments, data screening device 150 can be integrated as a part of data source device 110 or data consumption device 120. For example, data screening device 150 can be implemented as a content filter at a content server, or implemented as a firewall at a client device.

[0023] The data screening device 150 may include a screening engine 152, a memory system 154, an input interface 156, and an output interface 158. In an embodiment where the data screening device 150 is a stand-alone device, the input interface 156 and the output interface 158 may be physical interfaces that provide wired or wireless connections to external devices. Examples of such physical interfaces may include data I / O ports, electrical contacts, antennas, and / or transceivers, etc. In an embodiment where the data screening device 150 is integrated into a data source or a data consumption device, one or both of the input interface 156 and the output interface 158 may be an application programming interface (API) that connects to a software component such as an application, an operating system, etc.

[0024] The screening engine 152 is coupled between the input interface 156 and the output interface 158 in a communication manner. The screening engine 152 can be implemented using software, hardware, or a combination of the two, and can perform screening functions on the input data stream received at the input interface 156, and provide an output data stream on the output interface 158. For example, the input data stream may include a series of data input elements (e.g., data frames, network packets, database entries, etc.), and the screening engine 152 can process each data input element to determine whether the data input element meets the screening condition. If the data input element meets the screening condition, the screening engine 152 can output the data input element on the output data stream. In some embodiments, the screening engine 152 can parse the data input element, and / or perform logic or calculation operations on the data input element to determine whether the data input element meets the screening condition. If the data input element meets the screening condition, the screening engine 152 can provide the unchanged data input element to the output data stream, or the screening engine 152 can modify the data input element before providing the data input element to the output data stream.

[0025] In some cases, it is possible that each data input element of the input data stream satisfies the filter condition. In other cases, the output data stream may contain only a subset of the input data elements because some of the data input elements do not satisfy the filter condition. Examples of filter conditions may include determining whether a network packet is intended for a particular destination, whether a database entry satisfies a database query, whether one or more parameters of the input data meet a particular threshold, whether an integrity check of the input data element verifies that the input data element has no errors, whether a user is authorized to access the data input element, etc. In some embodiments, the filter condition may be programmable or configurable. In some embodiments, the filter condition may be hard-coded.

[0026] The screening engine 152 can utilize the memory subsystem 154 when performing the screening function. For example, the memory system 154 can be used as a buffer to temporarily store processed data input elements before providing the data input elements to the output data stream. The memory subsystem 154 can be implemented using DRAM, SRAM, flash or other suitable memory or a combination thereof, and can include trusted memory (also referred to as secure memory), unprotected memory, or a combination of the two. Trusted memory can include tamper-proof features to protect the contents of the memory, while unprotected memory can lack such tamper-proof features. Therefore, unprotected memory may be vulnerable to attacks that can scan or read the contents of the memory. However, compared to trusted memory, unprotected memory may be less expensive and can be used at much higher densities. Various memory eviction techniques for how to read data from memory can be implemented to achieve inadvertent screening, thereby preventing attackers from inferring which data input elements pass the screening conditions. These techniques will be further discussed below.

[0027] To evaluate the effectiveness of oblivious filtering, consider a threat model where the database application is compromised and the attacker has fine-grained observation of its memory (both content and access patterns), and assumes that the enclave memory cannot be observed. In this threat model, the attacker has fine-grained observation of the memory area containing the output table to infer which records passed the filter. Some schemes for oblivious filtering may include using oblivious sorting, using enclaves to perform multiple passes on the input database while outputting a constant number of matching records in each pass, etc. However, these schemes may not be suitable for streaming databases, or for large databases that cannot all fit in memory, and may also suffer from performance overhead (e.g., due to oblivious sorting or multiple passes).

[0028] In today's extremely data-driven world, in order to provide the required privacy and security guarantees, it may be necessary to perform computations on encrypted databases while hiding access patterns and associated side channels from eavesdropping adversaries. While many solutions exist in offline settings (requiring access to the entire database), it is not entirely clear whether these results can be directly extended to online settings where lookahead is not allowed while retaining similar guarantees. The techniques disclosed in this article can be used in online settings, where only a single pass over the entire database or incoming data stream is allowed. For a memory of size b, a lower bound on the information leakage of an inadvertent filtering operation can be established as Ω(logb / b) bits of information per input element. While some solutions assume that the memory used by the system runs within a (expensive) trusted execution environment, the use of ORAM (inadvertent random access memory) can still help reduce the size of trusted memory by carefully accessing unprotected memory.

[0029] According to the technology disclosed in this article, when the size τ of the input element is larger than the logarithm of the size M of the unprotected memory, only a size of at most If w-window order preservation is allowed (order preservation across windows but not necessarily within a window), then only a size of at most ) of trusted memory. This means that for a data stream containing a million positive elements (each 1Mb in size), the trusted memory can be reduced from 1Tb (e.g., to store all elements in trusted memory) to about 3Mb, which is about 6 orders of magnitude smaller. Combined with the fact that every unit reduction in buffer size reduces latency by one unit, this can provide an almost 10-fold faster 5 times the output.

[0030] Overview of Casual Screening

[0031] This section introduces the concept of oblivious filtering. Figure 2 A system 200 is shown containing a trusted module that performs screening and unprotected memory that stores input and output data streams, which may be limited or unlimited depending on the application. The trusted module may implement screening operations and act as a component of an encrypted database engine, which may have similar components for other operators. The system 200 may be configured to use parameterized length trusted memory buffers to reduce timing leaks. In this embodiment, both the code and the memory buffers may be hosted within an enclave.

[0032] Consistent with the enclave threat model, it can be assumed that an attacker cannot observe the contents or access patterns of trusted memory and cannot tamper with the trusted code that implements the screening algorithm. While this assumption may apply to some platforms, such as the Sanctum enclave, in SGX-based library implementations, oblivious memory techniques can be used to hide access patterns to trusted memory. On the other hand, an attacker is able to fully observe the access patterns and contents of unprotected memory; in other words, the attacker may be able to observe the timing of each input consumption and output production at a fine-grained clock granularity.

[0033] The input and output streams may be encrypted under a symmetric key known to the client and provided by the client to the trusted module (e.g., sent via an authenticated channel established using remote attestation). The purpose of oblivious filtering is to hide records in the input stream that match the filter, which leaks some information about the plaintext input. In a vulnerable implementation, the trusted module iterates over the entire input stream, taking one record at a time, and produces each matching record in the output stream. To prevent obvious attacks, the trusted module may re-encrypt the matching records before writing them to unprotected memory. However, this alone is not sufficient, as the attacker only needs to verify that the output was produced within the duration between consecutive inputs.

[0034] Oblivious filtering can be used to prevent a (polynomial-time) adversary from creating correlations between inputs and outputs, since the attacker may only learn the size of the output stream and nothing more. The filter can temporarily store matching records in a trusted memory buffer and evict entries based on techniques that minimize leakage while taking into account constraints such as buffer size and order preservation.

[0035] Data Flow and Information Leakage

[0036] The input to the system may be a data stream, where each element comprises some (encrypted) content and a timestamp associated with it. In some embodiments, the timestamp of the data input element need not be considered part of the content. Instead, the timestamp may be assumed to be a publicly readable property that is immutable (or cannot be modified). The encrypted content is hidden from an adversary. A collision-resistant secure encryption scheme may be employed so that: (1) it is impossible to decrypt the content without knowledge of the key; and (2) it is impossible to produce two different contents with the same encryption. It may be assumed that the key used for this encryption (and decryption) is not available to an adversary and is known only to the system. It should be noted that in some embodiments, the encryption scheme may not necessarily support applying a filter to the encrypted content to infer the output that would be obtained if the filter were applied to the plaintext content.

[0037] In some embodiments, data input elements may arrive in sequence and in discrete time units. A clock may be modeled as advancing as these inputs arrive, so that the system can only produce output when the inputs arrive. This may be referred to as an event-driven model for streaming media. This model may be used to ensure that an adversary does not control the input arrival times, and that the adversary cannot pause the input stream or delay these elements to observe the output of the current element in the buffer. The techniques disclosed herein may be used for infinite streams, or for finite stream settings, in which case the event-driven nature may be enforced by appending an unalterable stop symbol to the last input element and ensuring that the system eventually outputs the element in the buffer once this symbol is encountered.

[0038] Filter tasks can handle queries that verify whether a given predicate holds true given the content of an element. Such queries can use the values ​​of content attributes to compute some predicate, whose output determines whether a given element satisfies the query. The solution described in this article can support arbitrarily complex predicate queries as long as they are not a function of the timestamp of the incoming source element and do not require the computation of other elements of the stream (i.e., the output of the query about the element can be computed entirely using the content of this element). An input element is said to match the filter (or is a positive element) if the query deployed by the filter returns true when run against the element's content. All other elements are said to be negative.

[0039] The system can produce an output stream containing the elements re-encrypted after processing by the system and with timestamps associated with them. It should be noted that since the system has access to the keys used for the encryption / decryption algorithm, queries can be computed on the decrypted content, and the resulting output can be re-encrypted with new randomness before being released to the output stream. Computing the output of the query about the element determines whether the element will appear in the output stream - the system can discard elements that do not match. It can be assumed that this computation is instantaneous compared to the time between the arrival of successive inputs, so that the output (if all) can be produced immediately after the input is seen (if the algorithm specifies so).

[0040] In some embodiments, the screening algorithm may have one or more of the following properties for the output stream:

[0041] Liveness — Every input element that matches the query is eventually output by the system. This ensures a steady stream of output and prevents the system from retaining any input in its buffer indefinitely. This property also helps characterize the latency of the filtering algorithm.

[0042] Order Preservation — The output stream preserves the relative order of the matching elements in the input stream. This preserves the integrity of the output stream, especially for applications that are sensitive to the order in which elements arrive.

[0043] Type consistency - The system can only output valid elements that match the filter. This property ensures that no additional (uneasy) layers of filtering are applied to the output produced by the system.

[0044] For the purposes of the analysis described in this paper, unlike the input stream, it is allowed to produce multiple (or no) elements as output at a given time. This can be analogous to models that produce at most one output at each time step, in that an output stream containing multiple outputs can be flattened by producing one output at a time using a buffer of at most twice the size. This may not affect the asymptotic expected information leakage to the adversary.

[0045] For the execution of the filter, the system has access to a trusted memory module, which will also be referred to as a buffer, which can store up to b elements at any point in time. The system can use this buffer to store elements before outputting them (at which point the elements will be removed from the buffer). The size of this buffer is specified by the deployed algorithm and is assumed to be known to the adversary. However, the contents of this buffer and its access patterns can always be hidden from the adversary. In other words, it can be assumed that the trusted memory module is side-channel secure, so that the adversary cannot use any information other than knowledge of the algorithm, the buffer size, and the timestamps on the input and observed output streams to obtain information about the contents or access patterns of the buffer. Since the scalability of such trusted memory modules can be expensive, it may be desirable to keep the memory buffer size as small as possible. The technology disclosed herein provides a system that can provide similar privacy guarantees by utilizing unprotected memory in conjunction with smaller trusted modules.

[0046] The adversary can be assumed to be a computationally constrained adversary whose goal is to obtain the output stream generated by the system in real time and to Figure 3 Determine which input elements match the query with full knowledge of the query, the algorithm, and the buffer sizes deployed as shown. It can be assumed that the adversary may not be able to observe the contents and access patterns of the trusted module / memory (buffer) at any time, nor control the algorithm's private coin toss (if any). Furthermore, it can be assumed that the adversary cannot control the input stream (or the inter-arrival times of the positive elements) and cannot inject any elements that were not in the original stream. It can also be assumed that the adversary may not be able to decrypt the contents of any element of both the input and output streams. Furthermore, the system takes care to re-encrypt the positive elements before producing them to the output stream, so that the adversary may not make any obvious correlation between the two streams unless the timestamps reveal a correlation. Unless otherwise specified, it can be assumed that the adversary's knowledge of the query allows for some prior joint distribution σ over the inputs, such that σ twill provide a priori likelihood that the element input at time t is positive, as believed by the adversary. In some embodiments, it is not assumed that the input stream is actually sampled from the distribution σ. In this case, the adversary's belief is independent of the actual source of the input stream.

[0047] refer to Figure 3 ,After observing the output stream for a certain number of time steps and knowing the algorithm deployed by the system, the adversary can try to infer which inputs are likely to match a given data filter. In some embodiments, the system aims to make the true input stream as indistinguishable as possible within the set of all such possibilities.

[0048] Inadvertent Flow - Filtering

[0049] set up is the set of all possible input elements. Algorithm (or solution) The input stream can be said to be screened with respect to query Q and with a buffer size of b if it satisfies the three properties specified previously - liveness, consistency and order preservation. In the following, subscripts Q and b can be removed from the equations when the context is clear.

[0050] If the input at time t matches the filter, the scheme can eventually output this element at some time t′≥t in such a way that the relative order of the output element is the same as the relative order of the corresponding input elements. For the time difference t′-t, the input element resides in a memory buffer used by the scheme, and the element is removed from the memory buffer once it is produced in the output stream. The discussion in this article will focus on how the scheme uses buffers for different elements to produce the output stream, rather than the entire operation of first decrypting the element, checking whether it matches the filter, and re-encrypting it if it can produce output. Therefore, the scheme can play the role of modifying the time interval between positive elements in the input stream.

[0051] Streaming solution leaks

[0052] The leakage of a screening scheme quantitatively measures the amount of information an adversary can infer about the input stream after observing the system's output for some time. It can be assumed that the adversary knows the scheme deployed by the system and knows the length of the input that the system has consumed when it starts making inferences.

[0053] In some embodiments, when an adversary observes a sequence of outputs, the number of possible input sequences that could produce the outputs is large given the scheme and buffer size, so the adversary has negligible probability of making a correct guess. This information-theoretic security focuses on prefixes of the input stream, not individual elements. Representation scheme , given the buffer size b, the prior joint distribution σ of the adversary’s beliefs about the input streams, and the length n of the input streams consumed by the system. For clarity, some parameters can be removed from this notation when the context is clear.

[0054] set up is the set of all input sequences of length n that the system can receive as input. The adversary’s prior belief is that is sampled from distribution σ, so the information it currently has is given by Given this measure The entropy of . Let O be the solution The output produced. Given and b's knowledge, conditional entropy measures the amount of information the adversary now has about the input streams that the system actually encounters. Then, given an output O, it will leak is calculated as the difference between these two entropies, which may correspond to the bits of information leaked to the adversary:

[0055]

[0056] Even for simple schemes, exact computations (or even closed-form bounds) are not trivial to compute. Therefore, one goal could be to optimize the leakage of the expectation, where the expectation occupies the space of possible output sequences that the system could produce:

[0057]

[0058] The next section describes the characteristics of oblivious streaming algorithms that allow for small values ​​of expected leakage.

[0059] Screening algorithm analysis

[0060] This section introduces four different schemes and analyzes their expected (asymptotic) leakage. The schemes are organized in an incremental way so that later schemes have lower leakage than earlier ones (whenever possible). Unless otherwise specified, all inputs are assumed to arrive in an iid (independent and identically distributed) fashion, where each element has probability p = 1 / 2 of matching the filter (this model is denoted by σ 1 / 2 denoted by ). The analysis will show a lower bound on the leakage of any algorithm that inadvertently filters the data stream.

[0061] A. Match-time eviction solution (EUM)

[0062] First, we discuss the simple scheme where the system continuously adds matching input elements to the buffer until it is full, and then each time an input element matches, the system evicts the oldest element from the buffer and then adds the current matching element to the buffer. This scheme is called the evict-on-match scheme and is denoted as EUMb , where b is the buffer size used by the system. The following pseudocode presents this scheme.

[0063] Eviction on match:

[0064] For t=1,2…:

[0065] 1. If the buffer is full, the oldest element in the buffer is evicted to the output stream.

[0066] 2. Input events are added to the buffer only if they match the filter.

[0067] It can be shown that this scheme may have high leakage to the adversary, because every time an eviction occurs, the adversary learns that a matching element may have arrived, allowing the adversary to learn complete information about all elements that arrived when the buffer is full.

[0068] Assume t 1 is the first time the adversary observes the output, and ignores (temporarily) all t>t 1 The opponent at t = t 1 The following can be learned at time t. Starting from the first step, output is only produced when the buffer is full and another input element matches the filter. Therefore, the adversary immediately knows: (1) At time t 1 The arriving element may have matched the filter—this is the maximum information an adversary can learn about any element in this model; (2) at time t 1 -1 Exactly b elements before it may match the filter—the adversary believes that t 1} -1b choices are equally likely; and (3) for time t 1 For all elements that arrive thereafter, matching elements cause eviction, while non-matching elements do not cause eviction (there is a one-to-one correspondence between whether an input element matches and when eviction occurs).

[0069] For b<n, EUM b The leakage is per element information bits, n is high probability. To calculate σ 1 / Download EUM b Stream leak:

[0070]

[0071] In the last step above, you can use logn n / 2 =n-Ω(logn). 1 / 2 In the model, t 1 represents a random variable that represents the length of the input sequence that ends with a positive element and contains a total of b+1 positive elements. n is high probability, which means This means that with high probability,

[0072] Observe that b = n minimizes leakage, in which case we have When b = o(n), then This means that when a system deploys EUM, any buffer size that is any (constant) fraction smaller than the total length of the exposed output sequence will (strictly) leak at least a constant fraction of bits of information to the adversary.

[0073] B. Eviction at Full Time (EWF):

[0074] Next, we discuss the case where not every element that arrives after the first eviction is fully exposed to the adversary (regarding whether it matches the filter or not). This can be performed by flushing the buffer whenever an eviction occurs. This scheme is called the eviction-when-full scheme, or EWF for short. b , where b is the size of the deployed buffer. The different steps in this scheme are described below.

[0075] Full time expulsion:

[0076] For t=1,2…:

[0077] 1. If the buffer is full, output all events in the buffer and flush the buffer.

[0078] 2. Input events are added to the buffer only if they match the filter.

[0079] We analyze the scheme in the case 1b<n (the case b≥n is irrelevant since the adversary does not observe the output in this case). Let t 1 ,…,t k is the time the adversary observes the output. Since the output is only produced when the buffer is full and only positive events are stored in the buffer, each output contains b events. However, this immediately leaks t = t to the adversary. 1 ,…,t k The input may be positive at times (or no output will be produced at these times), but not necessarily between these times (unlike EUM). For ease of notation, let t 0 = 0. In σ 1 / In the model, the following results hold true when n is high probability. Can be used as Abbreviation of .

[0080] For b<n, EWF b The leakage is per element To prove this, observe that the leakage function can be written as This is because it is observed that any b-1 elements may match in the time interval between successive evictions. Using n≠0, we get The fact that the following results are obtained:

[0081]

[0082] The following lemma can be used to further simplify the above expression.

[0083] Lemma 1: For every n,n′,k≥0, if n≥n′, then n k ≥n′ k . Assume that for some integer r ≥ 0, n = n′ + r. Use n k =n-1 k +n-1 k-1 , it can be proved that Because the combined word is always non-negative.

[0084] Lemma 2: For every a,b,c,d ≥ 0, a b c d ≤a+c b +d. This proof follows directly from the Chu-Vandermonde identity, from which we can conclude Separating the terms with k=b and exploiting the fact that the remaining terms are all non-negative, we can get the desired result.

[0085] Lemma 3: For any b<n, we have According to Lemma 2:

[0086]

[0087] Now, according to Lemma 1, since t k ≤n, we get Lemma then follows from n kb ≤n n / 2 and log 2 n n / 2 =n-Ω(logn) starts.

[0088] Using Theorem 3, we observe that EWF b The leakage can be defined as follows:

[0089]

[0090] where k′ is the number of cycles in which at least one negative element reaches the input stream. 1 / 2 In the model, the high probability is Indicates high probability.

[0091] Corollary 1: For any choice of buffer size, the full-time eviction scheme leaks bits of information to the adversary. Specifically, when using a constant-sized buffer, the leakage is Θ(1) per element. The buffer size that balances the leakage is This comes from the observation that (here, The sign hides the logarithmic factor. )

[0092] The above analysis shows that large buffers provide low leakage. The largest buffer possible is b = n, in which case, This works for any b = Ω(n / logn). However, this is probably too large a buffer to be deployed in practice. On the other hand, a constant-size buffer leaks Ω(n) bits of information, which is consistent with EUM o(n) The leak sequence is the same.

[0093] Therefore, for the same buffer size, EWF can provide asymptotically lower leakage than the previously discussed EUM. Intuitively, this is because instead of always keeping the buffer full after encountering enough positive elements, the EWF scheme empties the buffer each time it reaches saturation, thus essentially starting over with the elements arriving now, reducing the correlation (if any) with previously seen elements. However, the reduced leakage is still susceptible to smaller buffer sizes, in which case the leakage is asymptotically no better than the previous scheme.

[0094] C. Collection and Eviction Scheme (CE)

[0095] Building on the concepts discussed above, we will discuss a scheme similar to EWF, except that it adds all tuples to the buffer instead of just those that match the filter. This change in the scheme will be called the collect and evict scheme, or CE for short for buffer sizes b ≤ n. b , which helps to further reduce the asymptotic worst-case leakage. The following pseudo-code summarizes the main steps in this scheme.

[0096] Collection and eviction

[0097] For t=1,2…:

[0098] 1. If the buffer is full, output all matching events in the buffer and flush the buffer.

[0099] 2. Add the input event to the buffer (regardless of whether it matched or not).

[0100] CE observedb An output is generated every b time steps - this uniformity in the output generation helps reduce leakage, as differences in eviction times (as shown above) can be a source of high leakage. Let the output generated at time t = kb (where k ≥ 1) be denoted by c k For simplicity, assume that n is a multiple of b. This means that when the adversary guesses the matching distribution in the input stream attached to the system at time n, a total of n / b outputs are known to the adversary. We will use As Assume t 0 =0.

[0101] For b≤n, CE b The leakage is per element information bits, and the probability is high at time n. i Observe that c i output, so CE b The flow leakage can be calculated as After simplification, we get:

[0102] Finally, using the actual (where n≥3), we get

[0103] It should be noted that this establishes a tight bound on leakage, rather than a lower bound as in the previous scheme. This means that the collection and eviction scheme may always leak for any choice of buffer size. Similar to EWF, the inference is as follows.

[0104] Corollary 2: For sufficiently large n, Scheme CE b (in ) leaks each element information bits, n is a high probability. get

[0105] Observe that when b < n, the leakage rate of the collection and eviction scheme is asymptotically lower than that of the full-time eviction scheme. This follows directly from the fact that when b = O(n), However, as b / n increases, the leakage of both schemes becomes almost the same, while when b / n is low, the leakage of the collection and eviction scheme is much lower. As discussed in more detail below, it is demonstrated that CE b All schemes for screening flows in the model according to embodiments of the present invention have asymptotically optimal leakage.

[0106] D. Randomized Collection and Eviction (URCE)

[0107] So far, we have discussed fully deterministic schemes in the sense that the system makes no random choices when processing its input stream. Schemes that randomize the timing of evictions and the number of elements to be evicted are analyzed to determine whether it is possible to achieve a performance better than CE. b This scheme is called the randomized collection and eviction scheme and is used with URCE when deploying a buffer of size b. b It is usually assumed that the adversary does not pay attention to the random coin flips done by the algorithm, but here it can be assumed that the adversary knows the probability of an eviction occurring and the randomization process that chooses to evict the stated number of elements.

[0108] The following pseudo code represents the steps according to some embodiments of this scheme:

[0109] Random Collection and Eviction

[0110] For t=1,2…:

[0111] 1. Use probability p e , selects a (uniformly) random number of the earliest elements from the buffer and evicts them to the output stream.

[0112] 2. If the input element matches the filter, it is added to the buffer. If the buffer is full before this addition, a (uniformly) random number of the oldest elements are first evicted from the buffer, and then the input element is added to the buffer.

[0113] Next, we discuss some properties of the random processes generated by this scheme, which are used to infer the reasons for information leakage to the adversary. For a given discrete random variable X over positive integers, let U(X) denote a uniform random variable in {1,…,X}.

[0114] Lemma 4: For a positive integer i, let p i ≥0 represents the probability of X taking value i, then, we get the following:

[0115]

[0116] First, the expected buffer size at any given time can be determined. This information could be a source of leakage to an adversary.

[0117] Fix the buffer size b, and let p be the probability that an input element matches the filter (same and independent of other elements). For t ≥ 1, let X t (p) = 1, the probability is p, otherwise it is 0. Let B i represents the buffer size at time i. Let B 0 = 0. Then, for t ≥ 0,

[0118] B t+1=B t +X t (p)-Y t

[0119] Among them, if B t <b, then Y t =X t (p e )·U(B t ), otherwise it is U(B t ). B t When <b, we get the following result (using linear expectation and Lemma 4):

[0120]

[0121] Use B 0 = 0 Solving this recursion, we get

[0122] Observe that if The expected buffer size will not exceed b / 2. Therefore, from now on, you can set This may require b > 2. This may result in some enforced evictions (with negligible probability), but the buffer can be augmented with arbitrarily large unprotected memory so that a reasonable p can be set without increasing leakage. e value. Figure 4 For b = 100, where n = 10 7 The plot of the time steps shows that in practice, p e This choice does not enforce a high probability of eviction.

[0123] Figure 4 Shown for URCE b Graph 410 shows the buffer size and expected (amortized) element-level leakage for solutions with b=100, p=0.5, and p e =10 / (b+2) for URCE b The buffer size indicates that under high probability, the forced eviction will not occur within the eviction probability p. e Graph 420 shows the effect of b = 100 on URCE when the adversary observes a single output produced at the end of n time steps. b The expected leakage (per element).

[0124] Let M t A random variable representing the number of outputs observed by the adversary at time step t. Assume that B t Stay below the maximum buffer size b until time t, then M t =X i(p e )·U(B i ), where p e =10 / (b+2). Then we can get the following result: This gives the following expression for the expected number of outputs:

[0125]

[0126] After computing the expected buffer size and the expected number of observed outputs, the leakage can be computed when the adversary observes the first output. Consider a positive element input to the system at time t. To compute the probability of output at time t′=t+k (where k≥1) (assuming the buffer is not full during this time), consider the size of the buffer at time t′.

[0127] Note that this element is evicted at time t′=t+1 with probability (Because at least B t Suppose the system produces only k ≥ 1 outputs at time t = n ≥ 2.

[0128] Set t Indicates that the system receives the first input element at time t, Out k,t represents the elements of k outputs generated at time t. Then, Using Bayes' rule, we can get the following results:

[0129]

[0130] In addition, let ε t is the element that the adversary believes the input at time t matches the filter. Obviously, Pr(ε t )=p. For Prε t |Out 1,n ), you can use (where C = In t ). Therefore, the following results are obtained:

[0131]

[0132] The opponent at time t 2 =n, the flow leakage when an output is observed can be calculated as follows:

[0133]

[0134] in, In addition, Pr(ε t |Out1,n ) is at most This means the following:

[0135]

[0136] For p = 1 / 2 and p e =10 / (b+2),

[0137]

[0138] and

[0139]

[0140] Figure 4 The graph in the figure depicts the 20 and per-element bounds for a maximum buffer size of 100. Observe that both the upper and lower bounds are asymptotically similar to the green graph (logn / n). This suggests that the expected leakage per element when observing only one output is of the same order of logn / n. Any further observations will only leak more information (if any), suggesting that URCE b There is no way around the tight bounds on leakage, as observed with collect and evict. This means that randomized algorithms cannot be more effective than fully deterministic schemes when it comes to the amount of information inferred by the adversary. Therefore, collect and evict can be the best oblivious screening algorithm according to some embodiments.

[0141] Leaks caused by exposing output streams

[0142] This section discusses the lower bound of leakage for any oblivious screening algorithm given a buffer size b ≤ n. It can be seen that in the threat model, any oblivious screening algorithm may leak Ω(logb / b) bits of information per element. Throughout the discussion, it is assumed that the input distribution is σ 1 / 2 The general approach is to show that for a fixed buffer size, the leakage of any algorithm is at least as much as the leakage of an algorithm that produces output only at the end of every b time steps, similar to CE b .

[0143] For any plan When all outputs are produced at the end of n time steps, the leakage to the adversary can be minimized. Here we can assume that the buffer size is n. Let At time step t 1 <…t k ≤n generates output c 1 ,…,c k , so that And each c i ≥1. For the convenience of notation, let t 0 = 0. cannot produce more output than the total number of positive tuples seen so far, so for each i, it can be c i ≤t i situation.

[0144] When observing these outputs, assume that the adversary infers that the input stream contains i-1 +1 to t=t i m i positive tuples, where for all j, ∑ j c j ≤∑ j m j (The equation for j=k holds true.) Then, the leakage to the adversary is Using Lemma 1 and Lemma 2 and n k By maximizing at k = n / 2, this expression can be simplified as follows:

[0145]

[0146] When b < n, no more than b elements can be stored in the buffer at any time. Therefore, the output can be generated at most b time steps apart. In these sequences of b time steps, the above analysis shows that the leakage is minimal when all the outputs for these time steps are generated simultaneously at the end of b time steps. The leakage here is Ω(logb). Finally, since all inputs arrive independently (and identically) of each other, the leakage over n time steps is the sum of the leakage over each of these time slots of b time steps, which is equal to Ω((n / b)logb). Averaging over all elements, the minimum leakage per element is an amortized Ω(logb / b) bits of information.

[0147] Delay control during screening

[0148] This section analyzes the delays introduced by different filtering schemes. For latency-critical applications, such as video streaming, it may be preferable that inadvertent filtering operations do not introduce large delays that render the service unavailable. There is a tradeoff between how often an adversary observes the output and how much of the input stream cannot be inferred. In this regard, it is assumed that the maximum delay of any positive tuple in the input stream is at most Δ ≥ 0. The implications of buffer size and leakage are discussed below. Furthermore, if the buffer size is fixed, the minimum leakage for a given b and Δ will be determined. It can be assumed that each input independently matches the filter with a (constant) probability p > 0.

[0149] We first discuss the eviction-on-match scheme. Recall that this algorithm can produce an output each time the buffer is full, and only evict matching elements from the buffer. The expected time before the buffer is full for the first time is b / p, and thereafter, a positive element is expected to arrive every 1 / p time steps. Therefore, any element will experience an expected delay of b / p-1 time steps, because it will be evicted when it is the oldest element in the buffer. To tolerate a given Δ, this means using b=p(Δ+1). Independent of the choice of b and Δ, the leakage of this scheme is Ω(1 information bits per element.

[0150] For the full-time eviction scheme, recall that eviction occurs when the buffer is full, at which time the buffer contents are cleared. Therefore, for an element arriving at time t, we write t = Qb + R, where and At this point, the buffer contains the expected Therefore, the buffer can hold at most elements, which means that the delay introduced by this element is time steps. This delay is maximum when t is a multiple of b (= b / p) and minimum when t = (b / p) - 1. To allow for Δ in the worst case, b = pΔ. This results in a leakage of Ω(logn / n) + Θ(1 / Δ) bits of information per element.

[0151] For the collect and evict scheme, an output is produced every b time steps, and elements are added to the buffer regardless of whether they match the filter. Thus, for an element arriving at time t = Qb + R (where and ), the buffer will time steps later, this is therefore the delay experienced by this element. Note that this delay is independent of p. The maximum value of this delay is b, which means that to allow for Δ, we can have b = Δ. This means that each element leaks information bits.

[0152] Finally, for the randomized collection and eviction scheme, recall that the eviction process is probabilistic, where at each time step, with probability p e Evict a random number of elements from the buffer. At time t, set p appropriately. e so that the buffer does not overflow (with high probability), the expected buffer size is Since the buffer contains only matching elements, the delay introduced by an element arriving at time t is the same as that of at least The expected number of outputs produced at time t is at most p. Therefore, it takes at least Furthermore, since each eviction causes at least one element to be evicted, the maximum delay of an element at time t is set up The maximum expected delay is 3p(b+2) 2 / 100. Therefore, to allow for Δ, we need From the lower bound, we know that the leakage in this case is per element information bits.

[0153] ORAM-assisted use of unprotected memory

[0154] The lower bounds discussed above provide the best possible leakage for the collection and eviction scheme in the threat model given the buffer size. However, according to Corollary 2, there is a tradeoff between the buffer size and the length of the input stream processed, i.e., low leakage comes at the expense of potentially very large buffer sizes. The algorithms discussed so far have assumed that the buffers are trusted memory modules, which makes it impractical to provide low leakage in reality, as side-channel-free trusted memory can be expensive and overkill if used only for screening purposes.

[0155] This section discusses optimizations to the collection and eviction scheme. Oblivious RAM (ORAM) is used to replace larger trusted memory with unprotected memory. According to some embodiments, the construction makes the size of the trusted memory vary only with the size of the elements in the input stream, and has nothing to do with how many input elements need to be processed. Since unprotected memory is readily available and inexpensive, the leakage can be arbitrarily close to the lower limit, which depends mainly on the limit on system latency. The construction details are as follows.

[0156] Let τ be the size of the encrypted elements from the input stream, and B, M be the sizes (in bits) of the trusted and unprotected memory (ORAM), respectively, and T be the number of time steps after which eviction occurs. It can be assumed that the ORAM is organized into T+1 blocks, each of length E(τ+1) bits, where E(x) is the size of the encrypted x-bit string. Block 0 is used to store dummy elements, and the remaining blocks are used to store (encrypted) elements from the input stream. For each block, the encrypted (τ+1)-bit string consists of 1 indicator bit, and the remaining τ bits correspond to the content of the input string. Without loss of generality, it can be assumed that no 0 bits occur in the input stream. τ+1 string, so ORAM is initialized by Enc(0||0 in each block (new encryption for each block) τ )composition.

[0157] Figure 5A system 500 for an ORAM-assisted oblivious filtering algorithm according to some embodiments is shown. The algorithm may proceed as follows: For each input element, a write is issued to the ORAM. If the input element matches the filter, the input element is written to the ORAM, otherwise a dummy write is issued. This hides information about which type of element (positive or negative) is written to the ORAM. The dummy element is always written to position 0 of the ORAM, and the i-th positive element is written to position i. The indicator bit on these elements is set to 1 (for a positive element e, the string written to the ORAM is Enc(1||e)). This will last for T time steps, after which the eviction process will be triggered.

[0158] During eviction, the system can read one element at a time from the ORAM starting at position 1 and continue doing so until it reads an element whose indicator bit is 0. For each element read, the system also reads Enc(0||0 τ ) is written to the ORAM at the corresponding location. Once the element is in trusted memory, it is re-encrypted and produced in the output stream. This way, when the eviction process is over, the ORAM is set back to its initial state (except for a new encryption of all zeros).

[0159] Although the above algorithm may look similar to the collect and evict scheme, it should be noted that eviction does not occur after B time steps, but after T time steps. Therefore, the leakage of this algorithm is The expected delay for each element is O(T) time steps (the delay introduced by ORAM can be made o(T) by using recursive path ORAM). The trusted buffer can process input elements and outputs in parallel, of course, it needs to store counters to store the next element in unprotected memory. Therefore, if recursive path ORAM is used, this algorithm can use only B=2max{τ+1),E(τ+1)}+log(T / τ)=Θ(log(T / τ)) bits of protected memory, and M=(T+1)E(τ+1)=Θ(Tτ) bits. This means that the size of the trusted memory is independent of the length of the stream, but only depends on the size of a single element in the stream and the logarithm of the size of the unprotected memory. For applications that are not sensitive to latency, T can be chosen based on the affordable M and B. Otherwise, based on the maximum affordable latency, T can be set to this latency.

[0160] If w-window order preservation is allowed (order preservation across windows but not necessarily within a window), then only a maximum size of This means that for a stream containing a million positive elements, each of size 1Mb, the trusted memory has been reduced from 1Tb (as used to store all elements in the trusted memory) to about 3Mb, which is about 6 orders of magnitude smaller. Combined with the fact that every unit reduction in buffer size reduces latency by one unit, this can provide a throughput that is almost 10 times faster. 5 times the output.

[0161] Replace leakage with bandwidth efficiency

[0162] In the algorithms discussed above, the output stream contains only those elements that match the filters in the system, thus maintaining full bandwidth efficiency (relative to the number of additional elements in the output stream). The fact that the lower bound is not easy to determine is partly due to this limitation. However, adding a few dummy elements to the output stream while maintaining the same buffer size may not reduce leakage. Even adding one dummy element to the output stream may increase leakage to the adversary in the threat model. It should be noted that adding dummy elements may require another filter at the end of the application where the output stream is consumed, however, this filter is much simpler than the inadvertent filtering task considered here. The simplest way to perform this filter is to decrypt the elements: the dummy elements do not decrypt to any meaningful elements, but rather to the elements corresponding to the positive elements from the input stream.

[0163] Assuming that the adversary cannot distinguish encrypted matching elements from dummy elements, if any given output stream O contains at most εm dummy elements, then the algorithm It is (ε,m)-bandwidth efficient for ε∈[0,1], where m is the number of unmatched elements in the input stream that produces O.

[0164] For example, consider the basic scheme of adding a dummy element for each unmatched element in the input stream. Note that this scheme achieves zero leakage, since the adversary observes the output at each time step but cannot distinguish between the real and dummy outputs. Furthermore, in this case, the trusted buffer only needs to be able to hold a single element (for decryption and checking whether the element matches the filter), so this algorithm can use only b = Θ(τ) bits. However, for this scheme, ε = 1, which makes it very bandwidth inefficient.

[0165] To understand the effect of adding dummy elements to the output, consider a variation of the collect and evict scheme where each time an eviction occurs, if the buffer contains m unmatched elements, the number of outputs produced is b-m+εm=bm(1-ε). Therefore, assuming the adversary knows ε<1, then when k outputs are observed, the adversary's guess for m is (bk) / (1-ε), making the leakage b-log 2b(bk) / (1-ε)>0. Through a more rigorous analysis (when p=1 / 2, taking the average of the possible values ​​of m), it can be seen that the expected leakage of this scheme is Therefore, any ε < 1 will induce non-zero leakage to the adversary. Depending on the application, if ε needs to be bounded (e.g., networking applications), the leakage will depend on this bandwidth efficiency. On the other hand, this result also shows that using virtual packets with fractions of ε has (asymptotically) similar leakage to using a (larger) buffer of size b(2ε) / (2ε-εb), in which case no bandwidth amplification occurs.

[0166] Multiplexed Stream Filtering

[0167] The techniques described herein can also multiplex multiple input streams and perform filtering in a way that the output elements do not reveal what input stream they came from. In the multiplexed input, the elements arriving at each time step are m-tuples, where m is the number of different (independent, and potentially with different prior distributions) streams being multiplexed. Assuming all these streams are generated honestly and there are no adversarial influences on them, the system essentially treats this input as m different inputs arriving at (finer) time steps separated by 1 / m units.

[0168] For applications that do not rely on the relative order of these m inputs, the system can simply shuffle the outputs corresponding to these inputs and produce the outputs in the output stream in a scheme-specified manner. The shuffling operation occurs within a trusted module, so an adversary cannot associate the produced outputs with their membership to some given input stream (better than random guessing). The technical caveat here is to ensure that finer time steps in the input stream do not produce outputs immediately, but rather the system processes the inputs at stages, and at each stage, all m inputs are processed simultaneously, and the outputs are first shuffled and then stored in a buffer (if any). In this way, the leakage of the system on the m multiplexed input streams is no more than m times the leakage of a single stream (because m times as many inputs are now processed).

[0169] For applications that may need the output streams to respect the relative position of the multiplexed streams, one solution is to insert dummy elements for streams that do not contain the required number of positive elements. This will help reduce the leakage of the system at the expense of bandwidth efficiency - usually in privacy-preserving settings there is a strict trade-off between redundancy and leakage.

[0170] Exemplary Methods

[0171] Figure 6A process 600 of inadvertent filtering according to some embodiments is shown. For example, the process 600 may be performed by a screening device, such as a computing device that performs a data filtering function. The screening device may include a memory subsystem that serves as a buffer between an input data stream and an output data stream, and the memory system may include both a trusted memory and an unprotected memory. The unprotected memory may have a larger memory size than the trusted memory, and the size of the trusted memory may depend on the length of the input element and the logarithm of the size of the unprotected memory.

[0172] In block 602, an input data stream having a plurality of input elements is received. For example, an input element may be a network packet, a database entry, a media content, etc. In some embodiments, the input elements may be received at periodic time intervals, wherein one input element is received at each time interval. In some embodiments, the input data stream is multiplexed as a data stream, wherein input elements for different intended consumers are interleaved with each other.

[0173] In block 604, for each of the received input elements, determine whether the input element satisfies the filter condition. In some embodiments, the data input element can be parsed, and a logical or computational operation can be performed on the data input element to determine whether the data input element satisfies the filter condition. Examples of filter conditions may include determining whether a network packet is intended for a particular destination, whether a database entry satisfies a database query, whether one or more parameters of the input data meet a particular threshold, whether an integrity check of the input data element verifies that the input data element has no errors, whether a user is authorized to access the data input element, etc. In some embodiments, the filter condition may be programmable or configurable. In some embodiments, the filter condition may be hard-coded.

[0174] In block 606, for each of the received input elements that satisfy the filter condition, a write operation is performed to store the input element in the memory subsystem. The write operation may include concatenating a positive indicator bit (e.g., a binary 1 bit) with the input element that satisfies the filter condition, encrypting the result of the concatenation, and storing the encrypted result in the memory subsystem. The input elements that satisfy the filter condition may be written to a trusted memory or to an unprotected memory. In some embodiments, if the order of the input elements needs to be maintained at the output, the location or memory address in the memory subsystem may be written in order so that the trusted memory is filled first, then the unprotected memory is filled, or the unprotected memory is filled first, then the trusted memory is filled.

[0175] In box 608, for those received input elements that do not meet the screening condition, a virtual write operation is performed on the memory subsystem. In some embodiments, a virtual write operation can be performed on each input element that does not meet the screening condition. In such embodiments, the total number of writes including both the actual write operation and the virtual write operation will be equal to the total number of received input elements. The virtual write operation may include cascading a negative indicator bit (e.g., a binary 0 bit) with a data string (e.g., a zero string) having the same length as the input element, encrypting the result of the cascade, and storing the encrypted result in the memory subsystem. The virtual write operation can all be written to a specified location in the memory subsystem, such as address 0 of the memory subsystem. Depending on how the system maps addresses to the memory subsystem, address 0 may correspond to a trusted memory or an unprotected memory.

[0176] In block 610, when the memory subsystem is full, the memory subsystem is ejected to provide the content to the output data stream. For example, except for the specified address of the virtual write, each address of the memory subsystem can be read out and provided to the output data stream. In some embodiments where the virtual write is written elsewhere, only the data elements with positive indicator bits read from the memory subsystem are provided to the output data stream. In addition, the indicator bit can be deleted from the data element before the data element is output to the output data stream. In some embodiments, before the data element is sent to the data consumer, other modifications (e.g., decryption, re-encryption, header insertion, compression, formatting or format conversion, etc.) can be performed on the data element.

[0177] Various entities or components described herein may be associated with one or more computer devices or operate the computer devices to facilitate the functions described herein. Some entities or components including any server or database may use any suitable number of subsystems to facilitate the functions. Examples of such subsystems or components may be interconnected via a system bus. Additional subsystems may include printers, keyboards, fixed disks (or other memories including computer-readable media), monitors that may be coupled to display adapters, and other devices. Peripheral devices and I / O devices coupled to an input / output (I / O) controller (which may be a processor or other suitable controller) may be connected to a computer system, such as a serial port. For example, a serial port or external interface may be used to connect a computer device to a wide area network such as the Internet, a mouse input device, or a scanner. Interconnection through a system bus allows a central processor to communicate with each subsystem and controls the execution of instructions from a system memory or fixed disk and information exchange between subsystems. System memory and / or fixed disk may embody computer-readable media.

[0178] It should be understood that the technology of the present invention as described above can be implemented using computer software (stored in a tangible physical medium) in a modular or integrated manner in the form of control logic. Based on the disclosure and teachings provided herein, those of ordinary skill in the art will know and understand other ways and / or methods of implementing the technology of the present invention using hardware and a combination of hardware and software.

[0179] Any software components or functions described in this application may be implemented as software code executed by a processor using any suitable computer language such as Java, C, C++, C#, Objective-C, Swift, or a scripting language such as Perl or Python using, for example, conventional or object-oriented techniques. The software code may be stored as a series of instructions or commands on a computer-readable medium for storage and / or transmission, suitable media including random access memory (RAM), read-only memory (ROM), magnetic media such as a hard drive or floppy disk, or optical media such as a compact disk (CD) or digital versatile disk (DVD), flash memory, etc. The computer-readable medium may be any combination of such storage or transmission devices.

[0180] Such programs can also be encoded and transmitted using carrier signals adapted to be transmitted by wired, optical and / or wireless networks that meet multiple protocols including the Internet. Therefore, the computer-readable medium according to an embodiment of the present invention can be created using data signals encoded with such programs. The computer-readable medium encoded with program code can be encapsulated with compatible devices or provided separately from other devices (e.g., downloaded via the Internet). Any such computer-readable medium can reside on or within a single computer product (e.g., hard drive, CD or entire computer system), and can be present on or within different computer products in a system or network. The computer system can include a monitor, a printer, or other suitable displays for providing any result mentioned herein to the user.

[0181] The above description is illustrative and non-restrictive. After those skilled in the art have read this disclosure, many variations of the present invention will become apparent to them. Therefore, the scope of the present invention should not be determined with reference to the above description, but should be determined with reference to the pending claims and their full scope or equivalents.

[0182] One or more features of any embodiment may be combined with one or more features of any other embodiment without departing from the scope of the present invention.

[0183] As used herein, the use of "a," "an," or "the" is intended to mean "at least one" unless clearly indicated to the contrary.

Claims

1. A method for screening a data stream, the method comprises: receiving, by a computing device, an input data stream having a plurality of input elements; for each of the received input elements, determining whether the input element satisfies a screening condition; for each of the received input elements that satisfy the screening condition, performing a write operation to store the input element in a memory subsystem; for those received input elements that do not satisfy the screening condition, performing at least a virtual write operation on the memory subsystem; and when the memory subsystem is full, evicting the memory subsystem to an output data stream.

2. The method according to claim 1, wherein the write operation comprises concatenating a positive indicator bit with the input element.

3. The method according to claim 2, wherein the write operation further comprises encrypting the result of the concatenation and storing the encrypted result in the memory subsystem.

4. The method according to claim 1, wherein the virtual write operation is performed for each of the received input elements that do not satisfy the screening condition.

5. The method according to claim 1, wherein the virtual write operation comprises concatenating a negative indicator bit with a data string having the same length as the input element.

6. The method according to claim 5, wherein the virtual write operation comprises encrypting the result of the concatenation and storing the encrypted result in the memory subsystem.

7. The method according to claim 5, wherein the data string is a zero string.

8. The method according to claim 1, wherein the memory subsystem comprises a trusted memory and an unprotected memory.

9. The method according to claim 8, wherein at least one of the input elements that satisfy the screening condition is stored in the trusted memory.

10. The method according to claim 8, wherein at least one of the input elements that satisfy the screening condition is stored in the unprotected memory.

11. The method according to claim 8, wherein the virtual write operation writes to a specified location in the memory subsystem.

12. A computing system, comprises: a memory subsystem; a processor; and a non-transitory computer-readable storage medium storing code that, when executed by the processor, causes the processor to perform operations including the following: receiving an input data stream having a plurality of input elements; for each of the received input elements, determining whether the input element satisfies a screening condition; for each of the received input elements that satisfy the screening condition, performing a write operation to store the input element in the memory subsystem; for those received input elements that do not satisfy the screening condition, performing at least a virtual write operation on the memory subsystem; and when the memory subsystem is full, evicting the memory subsystem to an output data stream.

13. The computing system according to claim 12, wherein the write operation includes concatenating a positive indicator bit with the input element, encrypting the concatenated result, and storing the encrypted result in the memory subsystem.

14. The computing system according to claim 12, wherein the virtual write operation is performed for each of the received input elements that do not meet the screening criteria.

15. The computing system according to claim 12, wherein the virtual write operation includes concatenating a negative indicator bit with a data string having the same length as the input element, encrypting the concatenated result, and storing the encrypted result in the memory subsystem.

16. The computing system according to claim 15, wherein the data string is a zero string.

17. The computing system according to claim 12, wherein the memory subsystem includes a trusted memory and an unprotected memory.

18. The computing system according to claim 17, wherein at least one of the input elements that meet the screening criteria is stored in the trusted memory.

19. The computing system according to claim 17, wherein at least one of the input elements that meet the screening criteria is stored in the unprotected memory.

20. The computing system according to claim 17, wherein the virtual write operation writes to a specified location in the memory subsystem.

21. The computing system according to claim 17, wherein the size of the trusted memory depends on the length of the input element and the logarithm of the size of the unprotected memory.

22. A method for screening a data stream, the method comprising: receiving, by a computing device, a first input element of an input data stream having a plurality of input elements; determining that the first input element meets a screening criteria; performing a write operation to store the first input element in a memory subsystem; receiving a second input element of the input data stream; determining that the second input element does not meet the screening criteria; performing a virtual write operation to a specified location in the unprotected memory of the memory subsystem; and when the memory subsystem is full, evicting the memory subsystem to an output data stream.

23. The method according to claim 22, wherein the first input element is stored in the trusted memory of the memory subsystem.

24. The method according to claim 22, wherein the second input element is stored in the unprotected memory of the memory subsystem.

25. The method according to claim 22, wherein the memory subsystem further includes a trusted memory, and the size of the trusted memory depends on the length of the input elements of the input data stream and the logarithm of the size of the unprotected memory.

Citation Information

Patent Citations

  • Secure query processing over encrypted data

    US20140281512A1