PUF-based obfuscation scheme for in-memory architecture
By introducing a PUF array and FeFET to generate a secure key in the IMC circuit, the weights of the neural network are scrambled and decrypted, thus solving the security vulnerability of the IMC architecture and achieving confidentiality protection on edge devices.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ROBERT BOSCH GMBH
- Filing Date
- 2025-10-21
- Publication Date
- 2026-04-21
AI Technical Summary
Existing in-memory computing (IMC) architectures have security vulnerabilities on edge devices, making them susceptible to software and hardware attacks. This can lead to the leakage of intellectual property and data related to neural network architectures, affecting the confidentiality of security-critical applications.
By introducing a physically unclonable function (PUF) array and key mechanism into the IMC circuit, the weights of the neural network architecture are pre-scrambled and decrypted and encrypted at runtime. The randomness of FeFET is used to generate a secure key, thereby achieving obfuscation and protection of the neural network architecture.
It effectively prevents attackers from extracting structural information and datasets from the neural network architecture through the output, ensuring the confidentiality of the algorithm while maintaining low overhead in terms of area and power consumption.
Smart Images

Figure CN121902884A_ABST
Abstract
Description
Background Technology
[0001] Following recent advancements in neural networks (NNs) for complex tasks such as image recognition and natural language processing, there is a growing interest in developing efficient hardware platforms to support their further progress. The increased complexity and data volume associated with existing NN architectures pose significant challenges to IoT devices operating at the edge, where computational resources and storage capabilities are limited. This has led to on-chip memory becoming a key constraint on power consumption and latency for modern digital embedded AI accelerators. These limitations have spurred research efforts to develop new computing paradigms that can overcome the memory bottleneck. Non-von Neumann architectures, such as in-memory computation (IMC), enable logical and arithmetic operations to be executed directly in memory, thus significantly reducing the amount of data transferred to the execution unit.
[0002] While IMC-based accelerators offer numerous advantages, it is essential to address the security vulnerabilities inherent in these architectures to ensure the confidentiality of running AI algorithms can be maintained. Potential leaks could allow attackers to steal the intellectual property of neural network architectures without needing datasets that are typically proprietary due to their high commercial value. Furthermore, this information could be used to craft more sophisticated attacks targeting the integrity of the algorithms. Violations of this nature can have severe consequences, potentially leading to hazards in security-critical applications. Previous research has demonstrated the feasibility of reverse engineering critical information about neural network architectures through attacks at both the software and hardware levels. At the software level, it is possible to generate adversarial queries by observing predicted outputs, which reveal structural information about the neural network architecture. This vulnerability could allow attackers to reveal internal properties of a model, such as those described in “PRADA: Protecting against DNN Model Stealing Attacks” (arXiv:1805.02628) by M. Juuti et al., March 2019, or even extract information about the dataset used during the training phase of the algorithm, including potentially sensitive data, such as those described in “Adapting Membership Inference Attacks to GNN for Graph Classification: Approaches and Implications” (arXiv:2110.08760) by B. Wu et al., October 2021. Adversarial queries can be crafted to construct model extraction attacks that generate alternative neural network architectures with similar behavior. In B. Wu et al.'s "Model Extraction Attacks on GraphNeural Networks: Taxonomy and Realization" (arXiv:2010.12751) published in November 2021, and N. Papernot et al.'s "Practical Black-Box Attacks against Machine Learning" (ACM) published in April 2017, pp. 506-51, it is demonstrated that this can also be achieved by treating the target NN as a black box and using the input / output to train the behavior of a second model, rather than using the raw data. Conversely, hardware attacks are closely related to inherent vulnerabilities in the specific architectural design of embedded AI accelerators.The non-volatile nature of the cells constituting the crossbars of many existing IMC accelerators makes these devices inevitably vulnerable to probing attacks, as they remain programmed even after the devices are powered off, as described, for example, in P. Roberts' "MIT discarded hard drives yield private info," published in Computerworld in January 2003. These attacks involve direct access to or manipulation of memory to extract sensitive information, as described, for example, in S. Chhabra et al.'s "i-NVMM: a secure non-volatile main memory system with incremental encryption," published in the proceedings of the 38th International Symposium on Computer Architecture at ACM in San Jose, California, June 2011, pp. 177-188. This can be achieved by physically tampering with the memory chip or by using specialized tools to read the contents of specific cell locations. While these attacks may seem impractical because they require decapsulation of the device, similar results have been demonstrated through non-intrusive attacks. Similar to DRAM, resistive random access memory (RRAM) technology has been shown to be vulnerable to rowhammer attacks due to thermal crosstalk, as described, for example, in "NeuroHammer: inducing bit-flips in memristive crossbarmemories" by F. Staudigl et al., published in IEEE Design, Automation and Test Europe Conference & Exhibition 2022 (DATE), pp. 1181-1184. This injection of faults into the values mapped in the cells leads to a potential leakage of information from the operating model, as described, for example, in "SNIFF: reverse engineering of neural networks with fault attacks" by J. Breier et al., published in IEEE Reliability Journal, 2011.
[0003] Furthermore, in "Side-Channel Attack Analysis on In-Memory Computing Architectures," published by Z. Wang et al. in the IEEE Communications on Emerging Topics in Computational Intelligence (pp. 1-13) in 2023, the feasibility of side-channel attacks was demonstrated by extracting valuable information about network structure by exploiting device power consumption and execution time. Instead, in "A Method for Reverse Engineering Neural Network Parameters from Compute-in-Memory Accelerators," published by J. Read et al. in July 2022 at the IEEE Annual VLSI Workshop (ISVLSI) in Nicosia, Cyprus (pp. 302-307), the values of weights mapped inside the memory were reverse-engineered using photon emissions from peripheral circuitry around the crossbar switch. This attack was proven successful in retrieving over 99% of the weights of a 128×128 matrix with as few as 350 trajectories. These attacks highlight the importance of research into the implementation of security countermeasures specific to modern IMC architectures, which will help mitigate potential threats and ensure the confidentiality of embedded NN algorithms. Summary of the Invention
[0004] An in-memory computing (IMC) circuit, an AI accelerator based on in-memory computing, and a method for in-memory computing are disclosed, which aim to reduce the defects of the above-mentioned devices and methods, and in particular allow for increased security of in-memory computing, while requiring less overhead in terms of area and power consumption.
[0005] As used below, the terms “have,” “contain,” or “include,” or any of their grammatical variations, are used in a non-exclusive manner. Thus, these terms can all refer to a situation where, apart from the features introduced by these terms, there are no further features in the entity described in this context, and also to a situation where one or more further features exist. For example, the phrases “A has B,” “A contains B,” and “A includes B” can all refer to a situation where, apart from B, there are no other elements in A (i.e., A is constituted solely and exclusively by B), and also to a situation where, apart from B, entity A has one or more further elements, such as element C, elements C and D, or even further elements.
[0006] Furthermore, it should be noted that the terms "at least one," "one or more," or similar wording indicating that a feature or element may typically appear once or more will only be used once when describing the corresponding feature or element. In the following text, in most cases, the terms "at least one" or "one or more" will not be repeated when referring to the corresponding feature or element, even though the corresponding feature or element may appear only once or more.
[0007] Furthermore, as used herein, the terms “preferably,” “more preferably,” “particularly,” “more particularly,” “specifically,” “more specifically,” or similar terms are used in combination with optional features without limiting the possibility of substitution. Therefore, the features introduced by these terms are optional features and are not intended to limit the scope of the claims in any way. As those skilled in the art will understand, the invention is carried out by using alternative features. Similarly, features introduced in the phrase “in embodiments of the invention” or similar wording are intended as optional features without any limitation on alternative embodiments of the invention, without any limitation on the scope of the invention, and without any limitation on the possibility of combining features introduced in this way with other optional or non-optional features of the invention.
[0008] In-memory computing (IMC) circuitry includes: - An array comprising a matrix of memory cells having multiple rows and columns, wherein memory cells in the same column are connected by a common bit line and memory cells in the same row are connected by a common word line, wherein the memory cells are configured to store weights of a trained neural network architecture, wherein the order of the weights is pre-shuffled; - At least one decoder is configured to use a key to output the number of shift operations to be performed by each of the multiple shift registers; - Multiple shift registers are configured to shift the output of the array based on the output of the decoder.
[0009] Neural network architectures can be implemented using memory cell matrices and matrix operations. After training the neural network architecture, the weights to be used in the trained architecture are determined, a process also known as "pre-training." Pre-shuffling can include reordering the pre-trained weights of the neural network architecture, particularly reordering the importance of bits before deployment to the IMC cross-connector. The order of the weights can be pre-shuffled as follows: Weights can be programmed into cells such that the column order is pre-shuffled. The order of the weights can also be pre-shuffled using a key. The key must be known before the weights are programmed into the array. The key can be used to shuffle the weights so that they can be corrected later using a shift register (also known as a shifter). The key can be specific to each hardware, for example, because it originates from a physically non-cloning function array. The key can also be deliberately chosen and used for multiple hardware instances.
[0010] Weights programmed within an array can be intentionally placed in an incorrect order (e.g., in the wrong column), particularly by using and / or according to a key. This shuffling can be corrected using the key via a shift register.
[0011] For example, the trained neural network is a convolutional neural network.
[0012] Shift registers can use multiple input values generated through matrix operations on an array, and can shift these input values based on information contained in the key. For example, the multiple shift registers can be configured to shift column indices based on the decoder's output.
[0013] The array may include multiple columns that are used as dummy columns. Dummy columns may not be used to store the weights of the trained neural network architecture. Dummy columns may not be used by the software, particularly during training and / or deployment. The deployment process may intentionally omit certain columns. In other words, in terms of hardware, these columns are treated the same as every other column, and dummy columns are identical to other columns. However, the results of dummy columns may not contribute to the neural network computation. In addition to the columns used to store the weights of the trained neural network architecture, dummy columns can also be shuffled. The presence of shuffled columns can be determined by the training and / or deployment process.
[0014] The key can be stored in a secure portion of the computational circuitry within memory. For example, the key can be a random number stored in a secure portion of the IMMMC chip (i.e., tamper-proof). The key can be a pre-known key, such as in the case of operation in a trusted environment. However, other options for the key are possible. For example, as described in detail below, the key can include an applied challenge concatenated with bits of the PUF array's response.
[0015] For example, the array can be an analog array. For example, the array can be a digital array. For example, the array can have a hybrid architecture.
[0016] For example, the in-memory computing circuitry further includes: - Challenge input, configured to apply at least one challenge to a physically non-cloning function (PUF) array; The array includes a physically non-clonable function array, which is configured to generate the output. The decoder is configured to use a key to output the number of shift operations to be performed by each of the multiple shift registers, wherein the key includes a challenge and the output of a physically unclonable array of functions; - Multiple shift registers are configured to shift the output of the array based on the output of the decoder.
[0017] For example, the PUF can receive a challenge number. The challenge selects the output generated by the PUF. The challenge and output are called the "key," which can then be used as input to the decoder. The decoder can then generate the number of shift cycles to be performed by the shift register.
[0018] For example, each memory cell includes a ferroelectric field-effect transistor (FeFET), wherein the physically non-cloning function array further includes multiple current-sensing amplifiers. The current-sensing amplifiers can be connected to the corresponding ferroelectric field-effect transistor in the last row and can be configured to measure the response of the memory cell matrix. However, other sources of the key are possible.
[0019] Each current sensing amplifier can be configured to measure the current difference between the two columns separately.
[0020] The in-memory computing circuitry may include at least one analog-to-digital converter (ADC) configured to convert the output of the current-sensing amplifier into a digital form. The in-memory computing circuitry may include an analog FeFET cross switch, a mixed-signal block (ADC), and digital blocks such as adders.
[0021] For example, a decoder may include a series of multiplexers configured to output corresponding shift values based on a key used as a selector. For instance, a decoder may include a simple combination of lookup tables or Boolean functions.
[0022] In-memory computing circuitry can be configured to perform runtime encryption and / or decryption mechanisms. It has been found that the proposed encryption and / or decryption mechanisms are effectively protected against both software and hardware attacks, as described in more detail below. The encryption and / or decryption mechanisms can be used to completely obfuscate the entire neural network architecture, or only obfuscate specific layers crucial for inference computation.
[0023] This invention proposes pre-shuffling the column order during the deployment phase of algorithm parameters in the non-volatile memory of an IMC chip, for example, by randomly shuffling the columns of a crossbar switch (also known as a matrix). This allows the importance of bits to be masked. Therefore, if the value of each programming unit is revealed, an attacker must guess the number of shift operations required to obtain the correct result of a multiply-accumulate (MAC) operation. When values are provided serially, this shuffling operation can be performed not only on the output columns but also on the bits of the input signals, as in the architecture, such as in the case of a crossbar switch as described in "FELIX: A Ferroelectric FET Based Low Power Mixed-Signal In-Memory Architecture for DNN Acceleration" by Soliman et al.
[0024] A physically non-cloning function array can be configured to receive at least one selected challenge as input, and the key is generated using the response of a matrix. The decoder can be configured to use the key to output the number of shift operations to be performed by each shift register. This can be performed in the factory-prepared state and / or in a programmed state. For example, this can be performed in a metastable state (i.e., a state that is neither 1 nor 0), where the programming algorithm does not select programming parameters that make such a state more likely to occur. It may not verify the actual state after programming as in the normal case. This can allow for random states to occur within the cells.
[0025] In subsequent stages utilizing in-memory computation circuitry, the memory cell matrix can be configured to receive activation inputs to be processed for neural network architecture inference and to perform multiply-accumulate computations. Each shift register can be configured to perform a shift operation corresponding to the decoder's output. The decoder can use a key to output the number of shift operations to be performed by each of the multiple shift registers.
[0026] The in-memory computing circuitry may further include at least one adder and / or counter configured to accumulate the output of the array. The in-memory computing circuitry may also be configured to transmit the accumulated output of at least one further (e.g., digital, logic) array for further processing.
[0027] The proposed in-memory computation circuitry can be configured for decryption at runtime. To perform decryption at runtime, the key can be decoded to determine the correct number of shift operations to be performed after transforming the output from each column in the digital domain. This operation is performed during inference execution.
[0028] The key can be a response to a PUF mechanism, such as those implemented in "Leveraging Ferroelectric Stochasticity and In-Memory Computing for DNN IP Obfuscation" by Mankali et al. or "Hardware Security Primitives using Passive RRAMCrossbarArray: Novel TRNG and PUF Designs" by Singh et al., or it can be a random number stored in a secure (i.e., tamper-proof) part of the chip. Other options for the key are also possible. For example, the key can be pre-selected during programming or alternatively generated. For instance, the key can be a random number stored in a secure part of the chip.
[0029] In-memory computational circuitry may include at least one framework configured to detect anomalies in a trained neural network architecture. For example, external tools, such as FACER (see "FACER: Universal Framework for Detecting Anomalous Operation of Deep Neural Networks" by Schorn et al.), can be combined with the running neural network architecture to perform inference computation. Whenever an anomalous input that might correspond to an adversarial example is fed to the neural network, a flag signal can be emitted to the decoder. This can cause all output values of the cross-switch to be shifted out, thereby disrupting the predictive computation in the neural network. This can prevent attackers from extracting valuable information about the algorithm's architecture by looking at the output.
[0030] The present invention further proposes an in-memory computing-based artificial intelligence accelerator, which includes a pipeline of multiple in-memory computing circuits according to the present invention. For definitions and embodiments of the in-memory computing-based artificial intelligence accelerator, refer to the description of the in-memory computing circuits described above or below in more detail.
[0031] An AI accelerator may include at least one control unit configured to control the pipeline of in-memory computing circuitry. The control unit may be configured to initiate the pipeline and / or generate different control signals, such as shift and reset signals, for different system blocks. The control unit may be embodied, for example, as described by T. Soliman et al. in “FELIX: A Ferroelectric FETBased Low Power Mixed-Signal In-Memory Architecture for DNN Acceleration”, published in ACM Embedded Computing Systems, Vol. 21, No. 6, pp. 1-25, November 2022.
[0032] The present invention further proposes a method for in-memory computation.
[0033] In this method, an in-memory computing circuit according to the present invention is used. For the definition and embodiments of the method, refer to the description of the in-memory computing circuit described above or below in more detail. The method includes the following method steps, which, as examples, may be performed in a given order. However, different orders are also possible. Furthermore, two or more method steps may be performed simultaneously or in a time-overlapping manner. Additionally, it is possible to repeat one, more, or even all of the method steps.
[0034] The method includes the following steps: i. Retrieve the security key; ii. The decoder uses the key to output the number of shift operations to be performed by each shift register; iii. Transfer the decoder output to the shift register.
[0035] The importance of bits in the trained neural network architecture is reordered before deployment on the IMC cross-switch. This method can be used to deobfuscate the pre-trained weights of the neural network architecture.
[0036] Step i. may include: retrieving at least one challenge using a challenge input, and applying the challenge to a physically non-clonable function array of in-memory computing circuitry. The method may further include measuring the output of the physically non-clonable function array. Step ii. may include having a decoder use the challenge and the output of the physically non-clonable function array as a key to output the number of shift operations to be performed by each shift register.
[0037] The method may further include shifting the column using a shift register based on the decoder's output.
[0038] The method may further include pre-shuffling the weights of the trained neural network architecture according to a key, and storing the pre-shuffled weights in an array of in-memory computational circuitry, particularly during the deployment phase of algorithm parameters in the non-volatile memory of the IMC chip. Therefore, the pre-shuffling can be performed before step i. of the method. The key may be known in advance before shuffling.
[0039] This method can be used to completely obfuscate the neural network architecture, or to obfuscate only the specific layers that are crucial for inference computation.
[0040] This invention proposes a novel security countermeasure and defense mechanism that deobfuscates values mapped to a physically non-clonable function array (also known as an IMC cross switch) at runtime, directing them to predefined locations. In most IMC architectures, only a portion of the bits of each cross switch column are represented by a multiply-accumulate (MAC) operation. This feature can be used for obfuscation design and to mask the bit importance of each value by scrambling the column indices. This invention allows authorized users to obtain logical inference computations even after the device is powered off, while maintaining the confidentiality of the algorithm against the aforementioned attacks. This can be achieved by performing an appropriate number of shift operations on each MAC result before propagating each MAC result to subsequent neural network layers. This number can be represented by a specific bit sequence that constitutes the key, which is forwarded to a register at the end of each cross switch column. Since the key length scales with the depth of the neural network, the size of the tamper-proof secure region required to store the key incurs significant overhead. Transmitting keys to edge devices, such as through cloud communication, may be impractical and potentially risky due to eavesdropping, as described by F. Rottenberg et al. in “CSI-Based Versus RSS-Based Secret-Key Generation Under Correlated Eavesdropping,” published in IEEE Trans. Commun. Vol. 69, No. 3, pp. 1868-1881, March 2021. Therefore, in one embodiment, the present invention proposes integrating a physically unclonable function (PUF) based on a ferroelectric field-effect transistor (FeFET) and using the extracted response as a digital fingerprint. The suitability of FeFETs for PUFs is illustrated, for example, in “Exploiting FeFET Switching Stochasticity for Low-Power Reconfigurable Physical Unclonable Function,” published by X. Guo et al. in IEEE ESSCIRC 2021-IEEE 47th European Solid State Circuits Conference (ESSCIRC) in Grenoble, France, September 2021, pp. 119-122. FeFETs have been found to possess suitable properties and inherent randomness, and FeFET PUFs have been proven to generate reliable and unpredictable responses to predefined challenges. This invention utilizes combinations of these values to generate a secure key that can be used to reveal the correct shift amount of the output for each column during runtime. In particular, this invention proposes a novel FeFET-based PUF design to generate secure and reliable keys with minimal overhead.By applying a custom shift operation to the result of each MAC operation, an efficient runtime deobfuscation mechanism for secure computation in IMC-based AI accelerators can be implemented. A design space probing method will be described, which can be used to identify the minimum overhead required for PUF design based on the level of security offered against different attack strategies.
[0041] The challenge input can be any input for a physically non-clonable function array. The challenge can be optional, for example, determined by a control unit.
[0042] Convolutional neural networks, such as those described in "FELIX: A Ferroelectric FETBased Low Power Mixed-Signal In-Memory Architecture for DNN Acceleration" by T. Soliman et al., published in ACM Embedded Computing Systems, Vol. 21, No. 6, pp. 1-25, November 2022, can be a class of deep neural networks (DNNs). Convolutional neural networks can include three types of layers: convolutional layers, pooling layers, and fully connected layers. Convolutional layers can use multiple sets of kernels or weights to process the input feature map and produce an output feature map. Convolutional layers can be the core element of these architectures. Each element out(x,y,z) of the output feature map can be computed using the following equation: (1) Among them, the variable act in and w z This corresponds to the kernel input and weights at depth z. The indices k and Cin represent the kernel size and the depth of the input feature map, respectively. The core operation described by this equation is commonly referred to as MAC, which handles most of the computational workload.
[0043] Existing artificial intelligence (AI) hardware accelerators, such as those described in "DianNao: ASmall-Footprint High-Throughput Accelerator for Ubiquitous Machine-Learning" by T. Chen et al., are typically designed to distribute tasks across multiple execution units (called processing elements (PEs)), each computes a subset of the workload in parallel within an independent cluster. IMC-based embedded AI accelerators replace the digital logic required for arithmetic computations encapsulated in each PE with matrices (also known as cross switches). The terms "in-memory computing circuitry" and "PE" are used interchangeably herein.
[0044] Each crossbar switch comprises a matrix of memory cells that store partial bit representations of the weights of a (pre-)trained neural network architecture. Cells within the same column and row of the matrix are connected by shared vertical output lines (also called word lines) and horizontal output lines (also called bit lines). As input activations are rerouted to each row, the corresponding signals can be propagated to each memory cell, which multiplies the two values. The results can then be accumulated over each shared column and processed by the peripheral circuitry surrounding the crossbar switch. The computed MAC calculation can be performed in different domains, depending on the characteristics of the cells and the nature of the peripheral circuitry.
[0045] Emerging technologies for non-volatile cells, such as FeFETs, can enable more efficient parallel MAC computations performed in the analog / mixed domain. However, mapping the full value of each weight to the resistance state of a single cell with high accuracy is challenging due to the immaturity of the manufacturing process and analog fluctuations in the device. Furthermore, the process of converting the accumulated analog signal to the digital domain requires a very high-precision analog-to-digital converter (ADC), which often constitutes a major performance bottleneck in these architectures, as described in A. Shafiee et al.'s paper, "ISAAC: A Convolutional Neural Network Accelerator with In-Situ Analog Arithmetic in Crossbars," published in IEEE, June 2016, pp. 14-26. This overhead can be mitigated at the hardware level by allocating more resources and distributing the computational complexity of MAC operations. This can be performed, for example, by dividing arithmetic multiplication into smaller operands at the bit level, as described in, for example, "Analog / Mixed-Signal HardwareError Modeling for DeepLearning Inference" by AS Rekhi et al., pp. 1–6, of the proceedings of the 56th Annual Conference on Design Automation, ACM, Las Vegas, Nevada, USA, June 2019, and "Increasing Throughput of In-Memory DNNAccelerators by Flexible Layerwise DNN Approximation" by CD l. Parra et al., pp. 17–24, of IEEE Micro, Vol. 42, No. 6, November 2022. For example, it is possible to perform bit decomposition on the MAC operation in equation (1) in the following manner: (2) (3) In equation (2), and These are the p bits activated by the input at loop c and the r bits of the weight of the i-th kernel, respectively. k is the number of parallel MAC operations in each loop, and quant is the quantization used for the weight parameters. It is a partial p-bit representation of the output feature map computed at loop c. In equation (3), P is the quantization used to represent the input activation, and D is the total number of MAC operations required to compute out, representing the complete output feature map. As shown, for example, in “FELIX: A Ferroelectric FET Based Low Power Mixed-SignalIn-Memory Architecture for DNNAcceleration” published by T. Soliman et al. in November 2022, Vol. 21, No. 6, pp. 1-25, and in “15.3 A 351TOPS / Wand 372.4GOPS Compute-in-Memory SRAM Macro in 7nm FinFET CMOS for Machine-Learning Applications” published by Q. Dong et al. in February 2020, at the IEEE International Solid-State Circuits Conference (ISSCC) in San Francisco, California, USA, pp. 242-244, these architectures can utilize the bit decomposition of MAC operations to optimize ADC overhead and overcome technical limitations. They select a binary representation for each cell, map each column of the crossbar to a particular position's importance representation, and assign digital adders and shifters as peripheral circuits that perform the operations described in equations (2) and (3).
[0046] As outlined above, each memory cell may include a ferroelectric field-effect transistor (FeFET). FeFETs can be fabricated by depositing a ferroelectric layer (FE layer) on top of a gate insulating layer (typically silicon dioxide (SiO2)) of a ferroelectric material (such as hafnium dioxide (HfO2)) in a conventional metal-oxide-semiconductor field-effect transistor (MOSFET). The non-centrosymmetric structure of the ferroelectric material can lead to spontaneous polarization within its cell, creating a net dipole moment that can be redirected by an external electric field. This redirectable dipole moment in the FE layer can achieve the threshold voltage (Vth) of the underlying transistor. TH The FeFET exhibits two stable polarization states, namely low V. TH (LVT) and high V TH(HVT), which depends on the magnitude and polarity of the programming pulse applied to the gate terminal. Both polarization states can be maintained even after the gate pulse is removed, enabling its use as a non-volatile memory device. This maintenance allows the FeFET to store information in the form of polarization states determined by the applied write or programming pulse. The FE layer can include multiple domains, each contributing to the overall polarization. The switching behavior of these domains can be influenced by the magnitude of the programming pulse. As the magnitude of the applied programming voltage increases, more domains can align their polarization with the electric field direction. This multi-domain switching mechanism can create an intermediate V between the low threshold voltage (LVT) state and the high threshold voltage (HVT) state. TH The state depends on the strength of the applied programming voltage.
[0047] Physically unclonable function arrays (PUFs) can be hardware security primitives configured to generate device-specific challenge-response pairs (CRPs) based on process variations and non-ideals of the underlying technology. PUFs can be embodied as described in “Hardware Security Primitives using Passive RRAM Crossbar Array: Novel TRNG and PUF Designs”, eprint:2211.03526, by S. Singh et al., 2022. Traditional CMOS-based PUFs utilize inherent device-to-device variations generated during manufacturing. However, once manufactured, the randomness of these device characteristics remains fixed and immutable, which can compromise the long-term reliability and effectiveness of PUFs as they cannot adapt to evolving security requirements. Non-volatile memories such as RRAM, spin-transfer torque magnetic RAM (STT-MRAM), and FeFETs offer promising approaches to developing reconfigurable PUFs. These technologies exhibit inherent cycle-to-cycle (C2C) variations within the device, providing a cost-effective means for dynamic reconfiguration, thereby enhancing customizable security solutions. For example, FeFET can be an HfO2-based FeFET. HfO2-based FeFETs are compatible with CMOS technology and feature a high ION / IOFF ratio, efficient power utilization, and field-dependent programmability.
[0048] The switching dynamics of FeFETs can exhibit inherently stochastic behavior, meaning that the polarization state of the ferroelectric (FE) layer can differ between devices even when subjected to the same programming pulse. This stochasticity in polarization switching results in varying leakage currents (IDS) across different devices for the same programming pulse. Additionally, FeFETs can exhibit significant device-to-device variability, contributing a significant standard deviation to their VTH distribution. Combined with the stochastic nature of polarization switching, this variability can be leveraged to generate challenge-response pairs (CRPs) in hardware security applications.
[0049] The PUF registration scheme can be implemented as follows. The FeFET-based CRP generation array can utilize the registration phase to register specific V using differential IDS sensing and reprogramming techniques. TH A state is assigned to each FeFET. This process can begin with the application of an intermediate write pulse, which partially switches the polarization of all FeFETs, thus generating V due to the inherent random switching. TH And slight variations in IDS. A differential amplifier can be used to measure the IDS difference between pairs of FeFETs. This process can be repeated across all FeFET pairs in the array, thus generating a unique V for each registration stage. TH Figure. The highly reconfigurable nature of this scheme allows for the generation of entirely new V in each subsequent registration stage. TH The distribution is characterized by its cyclical randomness.
[0050] The IMC circuit can be configured for CRP evaluation. After the registration phase, the IMC circuit can be prepared for CRP generation. Due to significant device-to-device variations, FeFETs exhibit unique IDS values even when programmed to the same state during the registration phase. To generate different CRPs, the differential current between the bit lines (BLs) can be calculated across all possible combinations within the cross-switch array. This can be achieved using sense amplifiers, where each sense amplifier evaluates the current difference as I... diff,xy =I DS,x - I DS,y , where I DS,x and I DS,y This represents the accumulated current flowing through the corresponding bit lines BLx and Bly of the array. In the case of an 8×8 crossbar switch array, this would generate a total of 28 different current difference combinations across the 8 BL lines. For clarity, we... Figure 3 The diagram illustrates the current difference specific to BL1. Figure 3 Seven sense amplifiers are employed. The inherent variability of FeFETs ensures that this configuration produces a unique Idiff,xy response for each possible challenging input, thus facilitating the generation of robust and unique CRPs.
[0051] After receiving a response from the PUF array, a key can be generated by concatenating the bits of the CRP. This bit vector can then be forwarded to a decoder that receives the key as input and outputs a combination corresponding to the number of shift operations to be performed on each register. When the key consists of the expected CRP, the decoder outputs the correct combination necessary for computing logical reasoning. The decoder's internal structure can vary. For example, it can implement an AES encryption scheme, such as the "Advanced Encryption Standard (AES)" technical report published by MJ Dworkin in NIST FIPS 197-upd1 at the National Institute of Standards and Technology (NIST) in Gaithersburg, Maryland, in May 2023, which generates a combination as plaintext when the key is applied to a previously stored bit string. For example, the decoder can be implemented as a device-specific programmable logic array, such as described in "Logic Design of Programmable Logic Arrays" by Kambayashi, published in IEEE Transactions on Computers, Vol. C-28, No. 9, pp. 609-617, September 1979, or as a small SRAM-IMC crossbar switch, as described in "New security challenges on machine learning inference engine: Chip cloning and model reverse engineering" by S. Huang et al., published in arXiv preprint arXiv:2003.09739, 2020. The optimal configuration of the decoder can be chosen according to requirements.
[0052] The decoder's output can then be connected to a shift register. For example, the shift register can be implemented as a barrel shifter located at the end of each column of the crossbar switch, such as described on page 436 of "Designalternatives for barrel shifters" by MR Pillmeier et al., published in Seattle, Washington, December 2002, and edited by FT Luk. These components can be implemented to perform combined shifts of input values to compute the number of times specified as additional input parameters. Non-secure IMC architectures with binary units, such as “FSPA: An FeFET-based Sparse Matrix-Dense Vector Multiplication Accelerator” published by X. Zhang et al. in July 2023 at the 60th ACM / IEEE Design Automation Conference (DAC) in San Francisco, California, USA, pp. 1-6; “AFerroelectricFET-Based Processing-in-Memory Architecture for DNNAcceleration” published by Y. Long et al. in December 2019 in IEEE Solid State Computing Devices & Circuits Exploration, Vol. 5, No. 2, pp. 113-122; and “FELIX: A Ferroelectric FET Based Low Power Mixed-Signal In-Memory Architecture for DNNAcceleration” published by T. Soliman et al. in November 2022 in ACM Transactions on Embedded Computing Systems, Vol. 21, No. 6, pp. 1-25, pp. 1-25, pp. 1-25, 202 ...3, 2023, 2023, IEEE Transactions on Embedded Computing Systems, 2023, 2023, 2023, 2023, 2023, 2023, 2023, 2023, 2023, 2023, 2023, 2023, 2023, 2023, 2023, 2023, 2023
[0053] The IMC architecture can operate in the analog / mixed-signal domain and can perform MAC calculations using plaintext values of parameters. Therefore, to ensure that the weights of the neural network architecture are still stored in a fully encrypted format, frequent loops of decryption and encryption are required during runtime. This process not only negatively impacts cell aging but also introduces additional peripheral circuitry, resulting in significant area and energy overhead. In "Enabling Securein-Memory Neural Network Computing by Sparse FastGradient Encryption," published by Y. Cai et al. in ICCAD (pages 1-8) in 2019, a sparse fast gradient encryption method for weights crucial to the final prediction calculation was proposed. The weights are decrypted during runtime to perform inference calculations using plaintext values, and then re-encrypted after use. This design ensures security against direct cell readout attacks even after power-off. However, if the chip is suddenly powered off during execution, certain parts of the device may still be vulnerable. Furthermore, the problem of repeatedly rewriting encrypted and decrypted values in memory remains unresolved, which will negatively affect the aging process of FeFET cells.To avoid the need for frequent rewriting, several works have proposed cell-level XOR encryption mechanisms for the following: SRAM, such as "XOR-CIM: compute-in-memory SRAM architecture with embedded XOR encryption" published by S. Huang et al. in ACM, November 2020, pp. 1-6; FeFET, such as "IMCE: An In-Memory Computing and Encrypting Hardware Architecture for Robust Edge Security" by H. Shao et al.; FinFET, such as "Novel Ferroelectric Tunnel FinFET based Encryptionembedded Computing-in-Memory for Secure AI with High Area-and Energy-Efficiency" published by J. Luo et al. in IEEE, pp. 36.5.1-36.5.4, December 2022; and RRAM technology, such as "Secure-RRAM: A 40nm 16kb Compute-in-Memory" published by W. Li et al. in IEEE, April 2021, pp. 1-2. "Macro with Reconfigurability, Sparsity Control, and Embedded Security" is used to map neural network architectures on memory in cryptographic form. While this countermeasure fully protects the design from cell readouts even during runtime execution, simulation results are computed in plaintext. In this case, side-channel information of the peripheral circuitry is related to the detailed outputs of the crossbar switch columns. Therefore, secret weight parameters can be reverse engineered via side-channel attacks, as described, for example, in "Side-Channel Attack Analysis on In-Memory Computing Architectures" by Z. Wang et al., published in IEEE Emerging Topics in Computing Intelligence, pp. 1–13, 2023. Another solution to overcome the cost of performing runtime-secure computation is to obfuscate the values mapped in the design.For example, in "A Low Cost Weight Obfuscation Scheme for Security Enhancement of ReRAM Based Neural Network Accelerators" published by Y. Wang et al. in ACM, January 2021, pp. 499-504, and "Security enhancement for reRAM computing system through obfuscating crossbarrow connections" published by M. Zou et al. in IEEE Design, Automation and Test Europe 2020 (DATE), pp. 466-471, 2020, this is accomplished by obfuscating the connections between rows or columns of different IMC crossbar switches, respectively. At runtime, the true location of the value is revealed by using a key as a selector for a large multiplexer, and then the value is forwarded to the corresponding connection. Alternatively, J. Zhang et al., in their July 2022 paper "WESCO: Weight-encoded Reliability and Security Codesign for In-memory Computing Systems" published in IEEE pp. 296-301, proposed an obfuscation mechanism involving bit swapping of mapped weights. The same encoding process can be applied to each cross-switch row and its corresponding input to maintain consistent matrix multiplication. The drawbacks of these implementations relate to overhead in terms of area, latency, and power consumption, resulting from the number of additional circuits or redundant loops required to perform the same operation for different encodings. Furthermore, if the number of potential combinations is constrained, the design remains vulnerable to potential side-channel attacks targeting restricted areas of the chip. Conversely, if the design is overly complex, new significant architectural constraints will necessitate complex signal rerouting. However, none of the aforementioned countermeasures are effective against software threats such as adversarial attacks and model-stealing attacks. In their 2020 paper "New security challenges on machine learning inference engine: Chip cloning and model reverse engineering" published on the arXiv preprint arXiv:2003.09739, S. Huang et al. proposed a strategy based on model-driven neural network fine-tuning to cover device-specific non-ideals via a PUF approach. More specifically, the models are retrained to adapt to the offset variations of the implemented ADC, making them unique and cloning-incompatible.If the model is stolen and run under different or no changes, accuracy drops between 20% and 75%, depending on the type and technique of the ADC. However, such countermeasures require costly ad-hoc retraining each time the model is deployed on a different chip. In "Leveraging Ferroelectric Stochasticity and In-Memory Computing for DNN IP Obfuscation," published by L. Mankali et al. in IEEE Solid State Computing Devices & Circuits Exploration, Vol. 8, No. 2, pp. 102-110 in December 2022, responses obtained from a FeFET-based PUF architecture are used to decode the values of pre-mapped weights without further preservation. This mechanism can be implemented as a defensive feature that destroys the value upon detection of an adversarial attack, but it makes the design vulnerable to hardware attacks on peripheral circuitry because MAC calculations are performed using plaintext values.
[0054] The pipelined in-memory computing circuitry (i.e., each cluster of the pipeline) can be implemented in the same way as each other (particularly from an architectural perspective) without any security countermeasures. Each PE can include a matrix connected to the same input signal, and each output column is converted to a digital domain and then shifted n positions according to the importance of the bits it represents. The order of the importance of the bits represented by each column can be shuffled before deployment on the device. The shuffled configuration adopted can then be known to the service provider but not to the user. The column-specific bit importance can only be revealed at runtime using the correct key provided by the device-specific CRP through the PUF. An unauthorized user maliciously attempting to steal architectural information would need to guess the correct combination of shift operations to be performed at the end of each column in order to perform meaningful reasoning. In the worst-case scenario, an attacker could identify the column belonging to each PE by observing the peripheral circuitry and performing power-side channel analysis on each crossbar switch, for example, as described in “Side-Channel Attack Analysis on In-Memory Computing Architectures” by Z. Wang et al., published in IEEE Emerging Topics in Computing Intelligence, pp. 1–13, 2023. This would require the attacker to crack the number of combinations for each cluster, which can be calculated using the following formula.
[0055] (4) in n quant It is used for the quantization of the parameters of the deployed neural network, and N PEsIt represents the number of PEs present on each cluster.
[0056] Physically non-clonable function arrays can include multiple dummy columns that are not used to store weights of trained (e.g., convolutional) neural network architectures. In addition to the columns used to store weights of trained (e.g., convolutional) neural network architectures, the dummy columns can also be shuffled. Since the parameters of equation (4) are highly dependent on the device use case, the number of possible combinations may not be high enough to ensure a satisfactory level of protection. To address this, an additional “dummy column” is proposed to be activated in each PE during runtime inference computation to increase the number of combinations that need to be guessed. These redundant columns can be mapped to random values or to unused weights from other layers to avoid introducing further area overhead. At the end of each loop, the key is propagated to a shift register that maps the computed “dummy” MAC value to a number of shift operations greater than or equal to the number of quantization operations used for each result. This can lead to register overflow, which sets values to zero that do not contribute to the computed inference. In this case, the number of combinations that need to be guessed can be calculated using the following equation: (5) Where n dummy It is the number of "virtual columns" introduced.
[0057] The methods and apparatus according to the invention, particularly for deploying and utilizing the aforementioned security countermeasures, can be used in both trusted and untrusted environments. An environment can be considered trusted when it is reasonably assumed that an attacker cannot maliciously manipulate or observe the behavior of the device. Conversely, an environment can be considered untrusted when this assumption is not true, such as when the device is not actually possessed.
[0058] For example, in a trusted environment, to obfuscate the bit importance of each block of weights in a neural network architecture deployed on the same PE cross-switch, the indices of each column can be randomly shuffled. While monitoring this process, the final configuration can be mapped, and the new combinations of shift operations to be applied to the MAC results of each cross-switch column can be stored as an integer vector. This sequence constitutes the key required by the authorized user to compute a logical output prediction. The PUF design can be programmed according to the above specifications. Different CRPs can then be selected and registered to calibrate the decoder.
[0059] For example, in an untrusted environment, whenever the in-memory computation circuitry is in its initial state, a selected challenge can be applied to the PUF array and concatenated with the response bits to generate a key. The key can then be processed by a decoder that outputs the correct combination of shift operations and transfers the value to the corresponding shift register. Subsequent stages of execution can continue in the same manner as in an unprotected design. Specifically, the activation input to be processed for NN inference can be passed to each obfuscated PE, which performs MAC computation. Each register then performs a shift operation corresponding to the sequence of values initially passed by the decoder. The correct output can then be accumulated and propagated to the remaining digital logic for further processing.
[0060] The IMC according to the invention was tested to withstand attacks. The assumptions about the nature of the attack model were made with the intention of covering as many attacks as possible as those described in the preceding sections. More specifically, the IMC according to the invention was tested to protect the confidentiality of a neural network (NN) algorithm in a white-box scenario. It can be assumed that an attacker has the following resources and capabilities: knows the values of the cells of the IMC cross-switch of the design and the corresponding positions of these cells. These values correspond to the parameters of the (pre-)trained NN performing the inference computation. Has actual possession of the device, allowing them to apply any input query to the device and observe the final output from the device. Knows the entire hardware architecture of the chip. Conversely, the attacker has no access to the responses of our PUF architecture, which constitute the key. This assumption can be based on previous research on such FeFET-based PUF designs, which has demonstrated the robustness of their inherent properties, including randomness, uniqueness, reliability, and reconfigurability, as illustrated in, for example, "Exploiting FeFET Switching Stochasticity for Low-Power Reconfigurable Physical Unclonable Function" by X. Guo et al., published in Grenoble, France, September 2021, at IEEE ESSCIRC 2021 – the 47th European Conference on Solid State Circuits (ESSCIRC), pages 119–122, and their security against various machine learning modeling attacks, as described in, for example, "IMCE: An In-Memory Computing and Encrypting Hardware Architecture for Robust Edge Security" by H. Shao et al. This assumption does not apply to challenges that can instead be considered publicly available. Exemplary results of tested attacks are shown in detail in the figures below.
[0061] This invention proposes a novel lightweight security countermeasure based on an obfuscation method targeting IMC-based embedded AI accelerators. The proposed method and apparatus allow for low-overhead FeFET-based PUF designs that generate reliable CRPs for use as keys. These pairs can be used to decode appropriate sequences of shift operations necessary for deobfuscating the bit importance of the deployed NN at runtime. Experimental results demonstrate that the proposed implementation is robust against all tested attack strategies with less than 3% area overhead.
[0062] The proposed obfuscation mechanism eliminates the need for decoding before each MAC computation. It is suitable for IMC use cases because it avoids multiple rewritings of values on the crossbar switches, which could potentially compromise the durability of non-volatile cells. Compared to decryption schemes, such as Huang et al.'s "XOR-CIM: Compute-In-Memory SRAM Architecture with Embedded XOR Encryption," the proposed implementation requires less overhead in terms of area and power consumption. This can be achieved by utilizing shifters. These shifters can be components already used in several unprotected architectures. Additionally, the decoder implemented in this context may require less logic because the range of possible values to be routed to the shift inputs is much smaller compared to Wang et al.'s "A Low Cost Weight Obfuscation Scheme for Security Enhancement of ReRAM Based Neural Network Accelerators" and Zou et al.'s "Security Enhancement for RRAM Computing System through Obfuscating Crossbar RowConnections." Regarding Zhang et al.'s "WESCO: Weight-encoded Reliability and Security Co-design for In-memory Computing Systems," the proposed implementation is more flexible because it does not introduce any architectural constraints or latency overhead due to computing corresponding neurons with the same type of encoding. Compared to Mankali et al.'s "Leveraging Ferroelectric Stochasticity and In-Memory Computing for DNN IP Obfuscation," this invention extends protection from adversarial attacks to hardware attacks, such as side-channel attacks or probing attacks, to keep the IP of the running NN architecture from being exposed. If an attacker obtains information about the programmed values of a cell or the MAC calculation results, they will be unable to compute the correct amount of shift operations to obtain consistent values.
[0063] It should be understood that the preferred embodiments of the present invention may also be any combination of the dependent claims or the above embodiments and the corresponding independent claims. Attached Figure Description
[0064] Further optional features and embodiments will be described in more detail in the following description of the embodiments, preferably in conjunction with the dependent claims. Here, the respective optional features may be implemented in an isolated manner and in any feasible combination, as will be understood by those skilled in the art. The scope of the invention is not limited to the preferred embodiments. Embodiments are depicted schematically in the figures. Here, the same reference numerals in these figures refer to the same or functionally comparable elements.
[0065] In the attached diagram: Figure 1 a) to Figure 1 d) shows (a) the structure of the FeFET based on 22nm FDSOI, (b) the polarization switching dynamics, and (c) I DS -V GS Characteristics and (d) V in FeFET devices TH Variability; Figure 2 An embodiment of an in-memory computing circuit and registration process for a FeFET-based PUF array is shown; Figure 3 An embodiment of a FeFET-based PUF for CRP generation is shown; Figure 4 An implementation of the obfuscated n-PE cluster architecture is shown; Figure 5 , 6 a) and 6b) illustrate exemplary mapping methods from parameters to IMC cross switches; Figure 7 The simulation I obtained through Monte Carlo simulation is shown. diff ; Figure 8 The accuracy results of tests against attacks on Resnet20, Resnet32, Neural Networks (NiN), and Lenet5 architectures are shown; and Figure 9 The time required for different key lengths is shown for Resnet20, Resnet32, Neural Networks (NiN), and Lenet5 architectures. Detailed Implementation
[0066] Figure 1 a) An exemplary ferroelectric field-effect transistor (FeFET) 110 is shown, which is suitable for use as a memory cell 112 of a physically non-cloning function array 114 (also referred to as an IMC cross switch). Specifically, Figure 1a) illustrates the structure of a FeFET based on 22nm FDSOI. The FeFET 110 can be fabricated by depositing a ferroelectric layer (FE layer) on top of a gate insulating layer (typically silicon dioxide (SiO2)) of a ferroelectric material (such as hafnium dioxide (HfO2)) in a conventional metal-oxide-semiconductor field-effect transistor (MOSFET). The non-centrosymmetric structure of the ferroelectric material leads to spontaneous polarization within its unit cell, creating a net dipole moment that can be redirected by an external electric field. This redirectable dipole moment in the FE layer achieves the threshold voltage (Vth) of the underlying transistor. TH The FeFET 110 exhibits two stable polarization states, namely low V. TH (LVT) and high V TH (HVT), which depends on the magnitude and polarity of the programming pulse applied to the gate terminal. Even after the gate pulse is removed, both polarization states are retained, enabling its use as a non-volatile memory device. This retention allows the FeFET to store information in the form of polarization states determined by the applied write or programming pulse. The FE layer comprises multiple domains, each contributing to the overall polarization. The switching behavior of these domains is influenced by the magnitude of the programming pulse. As the magnitude of the applied programming voltage increases, more domains can align their polarization with the electric field direction; see [reference needed]. Figure 1 (b). Figure 1 (b) illustrates the polarization switching dynamics. This multi-domain switching mechanism can create an intermediate V between the low threshold voltage (LVT) state and the high threshold voltage (HVT) state. TH The state depends on the strength of the applied programming voltage; see [link / reference]. Figure 1 (c). Figure 1 (c) shows the I of a 500nm × 500nm n-type FeFET. DS -V GS This characteristic reveals the variation of V after programming with pulses of different magnitudes. TH This behavior is caused by the field-dependent partial polarization of the ferroelectric domains within the device. FeFET 110 exhibits significant device-to-device variability, contributing significantly to the standard deviation of its VTH distribution, such as... Figure 1 As shown in (d). Figure 1 (d) shows the V in the FeFET device obtained from a Monte Carlo SPICE simulation of 1000 samples. TH Variability.
[0067] Figure 2An embodiment of an in-memory computing circuitry 116 and a registration process for a FeFET-based PUF array 114 is illustrated. The in-memory computing circuitry 116 may be an element of a pipeline 118 of multiple in-memory computing circuitry 112. The physically non-cloning function array 114 comprises a matrix of memory cells 112 having multiple rows and columns. Memory cells 112 in the same column are connected via a common bit line, and memory cells 112 in the same row are connected via a common word line. The memory cells 112 are configured to store weights of a trained (e.g., convolutional) neural network architecture. The column order is pre-shuffled. Each memory cell 112 may include a ferroelectric field-effect transistor (FeFET) 110. The physically non-cloning function array 114 further includes multiple current-sensing amplifiers 120. The current-sensing amplifiers 120 are connected to the corresponding ferroelectric field-effect transistor 110 in the last row and are configured to measure the response of the matrix of memory cells 112.
[0068] For example, Figure 2 The registration scheme shown can be implemented as follows. FeFET-based PUF array 114 (e.g.) Figure 2 (As shown) The registration phase can be utilized to identify specific V using differential IDS sensing and reprogramming techniques. TH State assignment is performed on each FeFET. This process begins with the application of an intermediate write pulse, which partially switches the polarization of all FeFET 110s, thereby generating V due to the inherent random switching. TH and I DS Slight changes. Differential amplifier 120 is used to measure paired FeFETs ( Figure 2 M in 11 and M 12 The difference in IDS between ( ). For example, if I DS,1 >I DS,2 Then FeFET M 11 It was reprogrammed to a strong low threshold voltage (LVT) state, and M 12 It is reprogrammed to a strong high threshold voltage (HVT) state, and if I DS,1 DS,2 Conversely, if the condition is true, then the condition is false. This process is repeated across all FeFET 110 pairs in the array, thus generating a unique V for each registration stage. TH Figure. The highly reconfigurable nature of this scheme allows for the generation of entirely new V in each subsequent registration stage. TH The distribution is characterized by its cyclical randomness.
[0069] Figure 2 Challenge input 122 is further shown, which is configured to apply at least one challenge to the physically unclonable function (PUF) array 114. Additionally, control unit 124 is depicted.
[0070] Figure 3 An example of CRP generation is illustrated. Following the registration phase, the system is prepared for CRP generation. Due to significant device-to-device variations, the FeFET 110 exhibits unique IDS values even when programmed to the same state during the registration phase. To generate different CRPs, we calculate the differential current between the bit lines (BLs) within the crossbar switch across all possible combinations. This is achieved using sense amplifiers, where each sense amplifier evaluates the current difference as I... diff,xy =I DS,x - I DS,y Here, I DS,x and I DS,y This represents the accumulated current flowing through the corresponding bit lines BLx and Bly of the array. In the case of an 8×8 crossbar switch array, this would generate a total of 28 different current difference combinations across the 8 BL lines. For clarity, in Figure 3 The diagram illustrates the current difference specific to BL1. Figure 3 Seven sense amplifiers are employed. The inherent variability of FeFETs ensures that this configuration produces a unique Ig for each possible challenging input. diff,xy This facilitates the generation of robust and unique CRPs.
[0071] After receiving a response from PUF array 114, a key may be generated by bit concatenating the CRP. This bit vector is then forwarded to decoder 126. Decoder 126 is configured to use the key to output the data to be shifted by multiple shift registers 128 (e.g., such as...). Figure 4 The number of shift operations performed by each shift register in the diagram. The key includes the response, or another key.
[0072] Figure 4 An architectural implementation of a confused n-PE cluster based on an in-memory computation-based AI accelerator 130 is illustrated. Decoder 126 receives a key as input and outputs a combination corresponding to the number of shift operations to be performed on each register. When the key consists of the expected CRP, decoder 126 outputs the correct combination required to compute logical reasoning. The internal structure of decoder 126 can vary.
[0073] like Figure 4As further shown, the physically non-clonable function array 114 may include multiple virtual columns that are not used to store the weights of the trained convolutional neural network architecture. In addition to the columns used to store the weights of the trained convolutional neural network architecture, the virtual columns may also be shuffled. Since the parameters of equation (4) are highly dependent on the device's use case, the number of possible combinations may not be high enough to ensure a satisfactory level of protection. To address this, it is proposed to activate additional “virtual columns” in each PE during runtime inference computation to increase the number of combinations that need to be guessed. These redundant columns may be mapped to random values or to unused weights from other layers to avoid introducing further area overhead. At the end of each loop, the key is propagated to a shift register that maps the computed “virtual” MAC value to a number of shift operations greater than or equal to the number of quantization operations used for each result. This can lead to register overflow, which sets values to zero that do not contribute to the computed inference. In this case, the number of combinations that need to be guessed can be calculated using the following equation: (5) Where n dummy It is the number of "virtual columns" introduced.
[0074] Figure 5 , Figure 6 a) and Figure 6 (b) illustrates an exemplary mapping method from parameters to the IMC cross switch 114. The aim is to implement a security countermeasure that will deter any potential attacker from attempting a brute-force attack intended to reverse engineer the NN architecture. Assuming an attacker can identify and extract raw data from each PE, the initial naive attempt from their side could include directly guessing the custom shift operation performed by each register. However, this strategy is extremely inefficient given the vast number of possible scenarios that can be computed using Equations 4 and 5. For example, in the context of an IMC device, such as as described in “FELIX: A Ferroelectric FET Based Low Power Mixed-Signal In-Memory Architecture for DNNAcceleration” by T. Soliman et al., published in ACM Embedded Computing Systems, Vol. 21, No. 6, pp. 1-25, November 2022, in n quant = 4 and N PEs In the case of 4.8 = 32, the number of combinations required to expose only one cluster will be (4!). 32 >10 44If four virtual columns are activated for each PE, this number can even increase to [(4 + 4)! / (4!)]. 32 >10 103 However, since the overhead would be excessive otherwise, it is assumed that each CRP generated by the PUF design decodes multiple shift operations simultaneously. This provides an attacker with an alternative, more efficient strategy: targeting the response bit before it is forwarded to decoder 126. By logically propagating each value through decoder 126, the attacker can obtain the corresponding combination of shift operations, thus gradually revealing parts of the hidden NN architecture with less guesswork. To evaluate the effectiveness of this gradual reveal in terms of inference accuracy, it is necessary to consider the actual mapping of parameters to the IMC cross switches. However, given a huge number of potential scenarios depending on the specific choices made by the hardware designer, in Figure 5 and Figure 6 The text presents two mapping-agnostic strategies that attackers can follow for their brute-force tactics.
[0075] Figure 5 The architectural approach is illustrated. Light-colored circles represent exposed neurons, while shaded circles represent unexposed neurons. The first approach consists of a naive strategy that targets a register storing the bits of the response generated from the PUF. In this case, the attacker randomly tries to guess each bit sequentially, and after running a full inference with test batches of each combination, the attacker observes the predicted output to detect meaningful changes in accuracy. As the attacker progressively guesses the correct parts of the key, he also exposes multiple neurons in the NN architecture. However, the distribution of these neurons is not necessarily uniform, but rather randomly distributed across all layers of the network, such as... Figure 5 As shown.
[0076] Figure 6 a) and Figure 6 b) illustrates the algorithmic approach. Light-colored circles represent exposed neurons, while shaded circles represent unexposed neurons. Figure 6 In section a), the output-to-input method is shown, while... Figure 6In b), the input-to-output method is shown. The algorithmic approach is more compatible with the algorithmic structure of the NN architecture. In this case, it is assumed that the attacker can identify the layer position of the parameters mapped in the columns of different cross switches (e.g., by analyzing the side-channel information of the chip, as described in "Side-Channel Attack Analysis on In-Memory Computing Architectures" published by Z. Wang et al. in IEEE Communications on Emerging Topics in Computing Intelligence, pp. 1-13, 2023). In this case, the examination of the hardware architecture allows the attacker to identify the specific part of the key that is related to the region of interest of the chip. Based on this assumption, two potential strategies against the attacker are derived, such as... Figure 6 a and Figure 6 As shown in b. First, the attacker targets the output layer of the NN architecture, as this is where the most meaningful part of the inference is computed. Once the output layer has been successfully exposed, the attacker can proceed from the output layer to the input layer, or from the input layer to the output layer. Similar to the architectural approach, the attacker runs a full batch of test cases to detect at which guessed bit combinations the algorithm outputs meaningful predictions.
[0077] Since the key length is proportional to the overhead introduced by the PUF logic, our goal is to investigate design space probing resulting from the trade-off between redundant logic and security levels introduced by our implementation. To evaluate this, each inference phase of the neural network was run using a full batch of tests, and the accuracy obtained was observed. Multiple scenarios were simulated, with the percentage of the exposed architecture progressively increased, and points where accuracy exceeded a certain threshold were identified. This threshold was defined as the level at which the observed output became distinct from random noise otherwise obtained when the design was fully obfuscated. This indicates to an attacker that a significant portion of the key has been correctly guessed, and for that attacker, it can represent the trigger point for starting to fine-tune their exposed strategy. Once the number of undisclosed weights required to achieve the selected threshold accuracy is identified, it can be assumed that it reflects the percentage of key bits that should be correctly guessed in the same manner. Equation 6 below can be used to calculate the time required for the attacker to achieve the selected threshold accuracy: Time thr = Time inf * 2 Keybits∗Percentagethr (6) Where Time inf and Time thrThese represent the time required to compute the complete test batch inference and reach the predefined accuracy threshold, respectively; keybits is the total number of secret bits to be guessed; and Percentagethr is the desired level of deobfuscation. To find the minimum number of columns required to generate Keybits secret bits and compute the corresponding overhead, equations 7 and 8 can be solved: (7) (8) Where N columns This represents the number of columns in the PUF design, Nchall represents the number of selected challenges, and Area represents... col It is the area required to implement each column, Area per It is the area of the peripheral circuitry (including the sense amplifier and decoder) associated with the PUF, and the overhead. PUF It is the total area of the PUF.
[0078] The experimental setup is as follows. The FeFET-based PUF was implemented using GlobalFoundries' 22nm technology. An 8×8 FeFET cross-switch array (see...) Figure 3 The simulation was performed in SPICE using a compact model of a ferroelectric capacitor developed with the GlobalFoundries 22nm process development kit. During the registration phase, random programming pulses were applied across all WLs in the array, causing the FeFET to be randomly programmed into intermediate states due to the random switching behavior of the ferroelectric domain. The differential current (IL) between adjacent bits BL was sensed. diff Based on its magnitude, the corresponding FeFET is reprogrammed to a strong HVT or LVT state according to the registration scheme described above. After registration, a challenge input is applied to the WL, where a sense amplifier connected to the BL detects the current difference across 28 possible combinations of the 8 BLs. Monte Carlo simulations are performed on 1000 samples to evaluate I... diff The variation takes into account the threshold voltage variation with a standard deviation of 57mV. The inherent variability of FeFETs ensures that all 256 possible challenging input combinations of the 8×8 FeFET cross-switch array produce a unique response (I0). diff For simplicity, only data from... Figure 7 I obtained from Monte Carlo simulation diff,12 The average value. This clearly shows that for all 256 possible combinations of challenge inputs, PUF generates a unique I. diff,12 .
[0079] The behavior of the PUF cross-switching was included in multiple simulations using ProxSim (C. De laParra et al., “ProxSim: GPU-based simulation framework for cross-layer approximate DNN optimization”, published at the IEEE 2020 European Conference and Exhibition on Design, Automation and Test (DATE), pp. 1193-1198), a GPU-based framework for hardware-aware retraining and evaluation of DNNs. During the algorithm's runtime inference phase, this framework was modified to include our deobfuscation scheme. When the expected response is returned, the lookup table outputs the correct sequence of shift operations to be applied to the values computed from each PE. The inference phases of multiple NN architectures were run using the attack strategy described above, with a different percentage of the public key for each scenario. More specifically, ResNet20 (see K. He et al., “Deep Residual Learning for Image Recognition”, eprint: 1512.03385), ResNet32 (see K. He et al., “Deep Residual Learning for Image Recognition”, 2015), and ResNet32 (see K. He et al.) were tested. The architectures used in He et al.’s 2015 publication, “Deep Residual Learning for Image Recognition” (eprint: 1512.03385), Neural Networks (NiN) (see M. Lin et al.’s 2014 publication, “Network In Network” (eprint: 1312.4400), and Lenet5 (see Y. Lecun et al.’s “Gradient-based learning applied to document recognition”, published in Proc. IEEE Vol. 86 No. 11, November 1998, pp. 2278-2324), were trained using the CIFAR10 (see A. Krizhevsky’s “Convolutional Deep Belief Networks on CIFAR-10”) and MNIST (see L. Deng’s 2012 publication, IEEE Journal of Signal Processing Vol. 29 No. 6, pp. 141-142) datasets.
[0080] To demonstrate the robustness of our security measures to highly compressed neural networks, we utilize quantization of 4-bit integers for the weights and 8-bit integers for the activation inputs. For each simulation, we save the accuracy obtained for different percentages of the disclosed architecture and collect the time required to perform a complete inference on Proxsim. This timeframe is used as a reference because it represents the worst-case scenario, where an attacker gains access to the GPU (which can achieve approximately 10 times the throughput of existing IMC chips, as described in "FELIX: A Ferroelectric FET Based Low Power Mixed-Signal In-Memory Architecture for DNN Acceleration" by T. Soliman et al., November 2022, ACM Journal of Embedded Computing Systems, Vol. 21, No. 6, pp. 1-25, or "A Ferroelectric FET-Based Processing-in-Memory Architecture for DNN Acceleration" by Y. Long et al., December 2019, IEEE Solid State Computing Devices & Circuits, Vol. 5, No. 2, pp. 113-122), and similar simulation frameworks optimized for performing low-latency inference computations.
[0081] Figure 8 and Figure 9 The results presented show the accuracy and time performance achieved for different scenarios. By observing the accuracy graphs for different simulations, we experimentally identified that after the 20% mark, all accuracy curves begin to increase rapidly and distinguish themselves from noise. This percentage of accuracy is set as the threshold level that an attacker aims to achieve, as it provides an indication of a partially successful brute-force attack. The results demonstrate that a targeted strategy of starting with the output layer and gradually exposing the rest of the neural network is extremely inefficient for all algorithms. However, this does not apply to the other two strategies, which show different benefits depending on the type of architecture. Using the ResNet architecture, an attacker can achieve higher accuracy much faster by gradually exposing layers from input to output. For example, in the case of ResNet32, exposing 75% of the architecture is sufficient to retrieve more than 70% of the original accuracy. In contrast to architectures such as NiN and Lenet5, gradually exposing the random portion of the algorithm is slightly more convenient. Using this technique, for example, Lenet5 achieves more than 50% of the original accuracy with 87% of the key exposed. Once the optimal brute-force strategy has been identified for each architecture, Equation 6 is used to plot the time required to reach the threshold accuracy for different bit lengths on a logarithmic scale, such as... Figure 9 As shown in the results, even for lightweight architectures like Lenet5, it would take an attacker over a billion years to test accuracy using sufficient combinations of a 64-bit key. This also applies to other architectures using different strategies. For ResNet32, for example, using a key with the same number of bits, testing all combinations to reach the threshold accuracy would take more than 100 million years.
[0082] Defense mechanisms could be as follows. As mentioned above, there are other possible attacks on IMC accelerators that threaten the information protection of running algorithms. By simply examining the calculated predicted output, attackers could potentially extract valuable information about the NN architecture, thus circumventing the need for side-channel or probing attacks on the cross-connect circuitry and subsequent brute-force attacks. This can be accomplished through carefully crafted adversarial queries (such as those described in “PRADA: Protecting against DNN Model Stealing Attacks” arXiv:1805.02628 by M. Juuti et al., March 2019, and “Model Extraction Attacks on GraphNeural Networks: Taxonomy and Realization” arXiv:2010.12751 by B. Wu et al., November 2021), or by injecting faults into the logic (such as those described in “NeuroHammer: inducing bit-flips in memristive crossbar memories” by F. Staudigl et al., 2022, at IEEE Design, Automation and Test Europe Conference & Exhibition (DATE), pp. 1181-1184). If tools such as FACER (e.g., as described in "FACER: A Universal Framework for Detecting Anomalous Operation of Deep Neural Networks" by C. Schorn et al., published in September 2020 at the 23rd International Conference on Intelligent Transportation Systems (ITSC) in Rhodes, Greece, pp. 1-6) are used to detect such attacks, it is also possible to exploit this security countermeasure as a defense mechanism similar to that described in L. Mankali et al., published in December 2022 in "Leveraging Ferroelectric Stochasticity and In-Memory Computing for DNN IP Obfuscation" in IEEE Solid State Computing Devices and Circuits, Vol. 8, No. 2, pp. 102-110. This can be achieved by setting a flag in the decoder, which then outputs random shift operations and transmits them to all PE registers. Therefore, the computation performed by the IMC accelerator will output noise that becomes useless against potential model-stealing attacks.
[0083] Regarding overhead, Table I presents a comparative analysis of the overhead introduced by our security countermeasures applied to the two architectures with existing technologies. These two architectures are “FELIX: A Ferroelectric FET Based Low Power Mixed-SignalIn-Memory Architecture for DNN Acceleration” published by T. Soliman et al. in ACM Embedded Computing Systems, Vol. 21, No. 6, pp. 1-25 in November 2022, and “A Ferroelectric FET-Based Processing-in-Memory Architecture for DNN Acceleration” published by Y. Long et al. in IEEE Solid State Computing Devices & Circuits, Vol. 5, No. 2, pp. 113-122 in December 2019.
[0084] Table I: .
[0085] To calculate the area introduced by the PUF, we start from... Figure 7A subset of the most stable and diverse combinations of CRPs was selected, and Equations 7 and 8 were solved. As a reference for the decoder, the area introduced by the necessary number of 128-bit AES encryption components was used, for example, as described in the first edition of "Top-Down Digital VLSI Design: From Architectures to Gate-Level Circuits and FPGAs" by H. Kaeslin, published in 2014 by Morgan Kaufmann in San Francisco, California, USA; these components were synthesized at 22nm. The results show that even by selecting only 2% of all potential CRPs, it is possible to generate a 512-bit key with an area overhead of less than 3% for two chips. For comparison, values for other prior art solutions have been reported and adjusted to reflect the settings according to the invention, where the entire NN architecture is obfuscated and a single binary state is represented by each 1T1R unit. In contrast to "A Low Cost Weight Obfuscation Scheme for Security Enhancement of ReRAM Based Neural Network Accelerators" published by Y. Wang et al. in ACM, January 2021, pp. 499-504, and "Security enhancement for reRAM computing system through obfuscating crossbar row connections" published by M. Zou et al. in IEEE Design, Automation and Test Europe 2020 (DATE), pp. 466-471, 2020, our solution introduces overhead independent of the specific type of NN architecture, but only proportional to the bit length of the key. Additionally, in contrast to XOR encryption—which requires a crossbar row array twice the size of a 1T1R FeFET cell to store the same number of parameters (e.g., as described by M. Zou et al. in “Security enhancement for rram computing system through obfuscating crossbar row connections”, published in IEEE 2020 European Conference and Exhibition on Design, Automation and Test (DATE), pp. 466-471)—the proposed implementation introduces consistent overhead regardless of the cell structure or technology employed.Compared to other solutions, our proposed design offers more comprehensive protection against a variety of attack types while maintaining a comparable (if not reduced) area overhead.
Claims
1. An in-memory computing (IMC) circuit (116), comprising: - An array comprising a matrix of memory cells (112) having multiple rows and columns, wherein memory cells (112) in the same column are connected by a common bit line and memory cells in the same row are connected by a common word line, wherein the memory cells (112) are configured to store weights of a trained neural network architecture, wherein the order of the weights is pre-shuffled; - At least one decoder (126) is configured to use a key to output the number of shift operations to be performed by each of the multiple shift registers (128); - The plurality of shift registers (128) are configured to shift the output of the array according to the output of the decoder (126).
2. The in-memory computing circuit (116) according to the preceding claim, wherein the array (144) includes a plurality of columns used as virtual columns, which are not used to store the weights of the trained neural network architecture, wherein the virtual columns are also shuffled except for the columns used to store the weights of the trained neural network architecture.
3. The in-memory computing circuit (116) according to any one of the preceding claims, wherein the trained neural network is a convolutional neural network.
4. The in-memory computing circuit (116) according to any one of the preceding claims, wherein the key is a key stored in a secure portion of the in-memory computing circuit (116).
5. The in-memory computing circuit (116) according to any one of the preceding claims, wherein the in-memory computing (IMC) circuit (116) further comprises: - Challenge input (122), configured to apply at least one challenge to the Physically Unclonable Function (PUF) array (114), The array mentioned above includes the physically non-clonable function array (114), which is configured to generate output. The decoder (126) is configured to use the key to output the number of shift operations to be performed by each of the plurality of shift registers (128), wherein the key includes the challenge and the output of the array of physically unclonable functions.
6. The in-memory computing circuit (116) according to the preceding claim, wherein each memory cell (112) includes a ferroelectric field-effect transistor (FeFET) (110), wherein the physically unclonable function array (114) further includes a plurality of current-sensing amplifiers (120), wherein the current-sensing amplifiers (120) are connected to the corresponding ferroelectric field-effect transistor (110) in the last row and are configured to measure the response of the matrix of the memory cells (112), wherein each of the current-sensing amplifiers (120) is configured to measure the current difference between two columns respectively.
7. The in-memory computing circuit (116) according to any one of the preceding two claims, wherein the in-memory computing circuit (116) includes at least one analog-to-digital converter (ADC) configured to convert the output of the current sensing amplifier (120) into a digital form.
8. The in-memory computing circuit (116) according to any one of the preceding three claims, wherein the array of physically unclonable functions (114) is configured to receive at least one selected challenge as input, and the key is generated by using the response of the matrix, wherein the decoder (126) is configured to use the key to output the number of shift operations to be performed by each shift register (126).
9. The in-memory computing circuit (116) according to the preceding claim, wherein in a subsequent stage of using the in-memory computing circuit, the memory cell matrix is configured to receive activation inputs to be processed for neural network architecture inference and to perform multiplication and accumulation calculations, wherein each shift register (128) is configured to perform a shift operation corresponding to the key.
10. The in-memory computing circuit (116) according to any one of the preceding claims, wherein the in-memory computing circuit (116) includes at least one frame configured to detect anomalies in the trained neural network architecture.
11. An artificial intelligence accelerator (130) based on in-memory computing, comprising a pipeline (118) of a plurality of in-memory computing circuits (116) according to any one of the preceding claims.
12. The AI accelerator (130) based on in-memory computing according to the preceding claim, wherein the AI accelerator (130) includes at least one control unit (124) configured to control the pipeline (118) of the in-memory computing circuit (116).
13. A method for in-memory computing, wherein the in-memory computing circuit (116) according to any of the preceding claims relating to in-memory computing circuitry is used, wherein the method comprises the following steps: i. Retrieve the security key; ii. The decoder (126) uses the key to output the number of shift operations to be performed by each shift register (128); iii. The output of the decoder (126) is transmitted to the shift register.
14. The method according to the preceding claim, wherein step i. comprises retrieving at least one challenge by using a challenge input and applying the challenge to a physically non-clonable function array (114) of the in-memory computing circuitry (116), wherein the method further comprises measuring the output of the physically non-clonable function array (114), wherein step ii. comprises the decoder (126) using the challenge and the output of the physically non-clonable function array as a key to output the number of shift operations to be performed by each shift register (128).
15. The method according to any one of the preceding two claims, wherein the method comprises pre-shuffling the weights of a trained neural network architecture according to a key, and storing the pre-shuffled weights in an array of in-memory computing circuitry (116).