Dual-core TEE security system constructed based on open source E902 processor
By designing a dual-core TEE security system based on open source E902 processor, the difficulties of IoT device security and integrity protection are solved, and multi-level security guarantee and efficient secure startup process are achieved.
Patent Information
- Application Number
- CN202510197044.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is difficult to effectively protect the security and integrity of IoT devices and prevent devices from being attacked, data breaches and physical devices from being invaded.
A dual-core TEE security system based on open source E902 processor is designed, including secure boot, physical address access firewall IOPMP, main modules such as Mailbox, SM2, SHA256&SM3, AES&SM4, and chip-side channel defense technology.
It realizes independent execution space, secure startup and legality verification, physical address access firewall, modular design, chip-side channel defense and full random verification, enhancing the security and reliability of the system.
Smart Images

Figure CN120074789A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of security systems, and specifically refers to a dual-core TEE security system built based on an open-source E902 processor. Background Art
[0002] With the continuous development of the Internet of Things, the number of devices is increasing and the functions are constantly expanding, and the security threats it faces are also increasing, such as device attacks, data leakage, physical device intrusion, etc. The trusted execution environment can protect the security and integrity of physical devices in the Internet of Things. For example, it can prevent devices from being taken over by unauthorized parties or being maliciously attacked, protect sensitive data in the devices from being stolen, and reduce the impact area and losses of intrusion events. Therefore, the trusted execution environment is crucial for ensuring the security and reliability of the Internet of Things. Summary of the Invention
[0003] The technical problem to be solved by the present invention is to provide a dual-core TEE security system built based on an open-source E902 processor in view of the deficiencies mentioned in the above background art.
[0004] To solve the above technical problem, the technical solution provided by the present invention is: a dual-core TEE security system built based on an open-source E902 processor, which includes the following:
[0005] One: Secure Boot: including image signature, image signature verification, and secure boot process;
[0006] Two: Physical Address Access Firewall IOPMP;
[0007] Three: Main Modules: including Mailbox module, SM2 module, SHA256&SM3 module, and AES&SM4 module;
[0008] Four: Chip Side-Channel Defense Technology: including an attack-resistant SM2 algorithm architecture and an attack-resistant SM4 algorithm architecture;
[0009] Five: Full Random Verification.
[0010] Furthermore, the image signature is to extract the digest of the SPL program, TEE program, and REE program images using SHA256, and use the SM2 algorithm to perform the signature process on the hash values of each image through the private key of the image publisher, and store the generated digest value, signature value, and image in the Flash.
[0011] The mirror signature verification mentioned above means that for the program loaded from Flash, the public key in Bootrom and the transported signature value are first used to verify the hash value of the mirror. If the signature verification passes, the integrity of the mirror can be further verified. Otherwise, the startup fails. If the signature verification passes, the hash value of the mirror is recalculated and compared with the hash value hash(image) stored in Flash. If they are equal, it indicates that the integrity verification of the mirror passes. Otherwise, the startup fails.
[0012] Furthermore, the secure startup process includes the following steps:
[0013] A. Before secure startup, the following steps need to be taken:
[0014] 1. First, call the national cryptography algorithm library of Python, generate the signature value and digest value using the bin file generated by the REE program code as ciphertext, and the public key PublicKey required by the SM2 algorithm, and solidify it in BootRom.
[0015] 2. In Flash, the TEE, REE program code, and the hash value (digest value) and signature value generated through the mirror signature process are stored. The key used for Python signature and the public key stored in BootRom are a pair.
[0016] B. The secure startup process is roughly divided into the following steps:
[0017] 1. When the secure core is powered on and reset, the chip starts running from the starting address 0x0, that is, runs the code in BootRom, names BootRom as ICACHE0. At this time, ICACHE0 will access Flash through the secure SPI serial port to read the signature, digest, and SPL program for the mirror signature verification process, complete the software mirror legality check. If the legality verification passes, the SPL program is transported to the ICACHE1 address in the TEE ICACHE, and the TEE CPU jumps to this address to execute the SPL program segment.
[0018] 2. In the SPL program, there is also a code program for transporting data. Similarly, the second-stage Bootloader will also transport the code program. Here, the TEE initialize and Run program, as well as the signature and digest, are transported. After passing the authentication and data integrity verification, they are transported to ICACHE2, and the running address also jumps to ICACHE2 to start the secure configuration of the TEE system, such as the configuration of IOPMP.
[0019] 3. After the secure core completes the security configuration, it will continue to transfer the code program. Here, the REE software program, as well as the signature and digest, are transferred. After passing the authentication and data integrity verification, the REE program digest and digest signature value are transferred to ICACHE2. The TEE CPU verifies the legality of the REE program and configures the REE reset signal. The REE program code is transferred to ISRAM. When the TEE CPU raises the REE reset signal, the REE system starts to initialize and run. If the TEE CPU does not complete the configuration of the REE CPU reset signal, the REE will wait continuously, ensuring that the non-secure core is started after the secure core completes the security configuration. In summary.
[0020] 4. The dual-core TEE security system based on the open-source E902 processor according to claim 1, wherein the physical address access firewall IOPMP specifically implements the following functions:
[0021] A. Only allow the TEE to configure the IOPMP;
[0022] B. When it is determined that the SID is the TEE, it must be allowed. We default that the TEE has the permission to access all peripherals;
[0023] C. When it is determined that the SID is the REE, check whether the address is within the configured range. If not, it must be rejected. If so, then check whether it has the corresponding permission. If it has, access is allowed. If not, access is rejected;
[0024] D. When it is determined that the SID is neither the TEE nor the REE, access is rejected.
[0025] Furthermore, the TEE CPU and the REE CPU of the Mailbox module use their own independent shared memories;
[0026] When the TEE CPU transfers data to the REE CPU through Mailbox (TEE2REE), it first writes the data type, data length, and data into the FIFO of Mailbox (TEE2REE). After the writing is completed, Mailbox (TEE2REE) generates a tee2ree_intr request signal to notify the REE CPU to receive the data from Mailbox (TEE2REE);
[0027] When the REE CPU finishes receiving data, Mailbox (TEE2REE) generates a tee2ree_rsp reply signal to notify the TEE CPU that the data has been received by the REE CPU and the communication ends. Similarly, when the REE CPU transmits data to the TEE CPU, there are corresponding ree2tee_intr request signals and ree2tee_rsp reply signals to complete the communication;
[0028] In order to quickly respond to inter-core communication, this design uses the method of interrupts. Therefore, in order to implement dual-core communication, two interrupt response functions of Mailbox need to be added to each of the two processor cores respectively;
[0029] Furthermore, the SM2 module includes the following architecture:
[0030] A. Hardware architecture of the digital signature system;
[0031] B. Hardware architecture of the cryptographic protocol layer;
[0032] C. Hardware architecture of the elliptic curve operation layer;
[0033] D. Hardware of the binary field modular arithmetic layer;
[0034] Furthermore, the SHA256&SM3 module designs reconfigurable RTL code. Specifically, when the SM3&SHA256 module is called during the secure boot process, the TEE CPU first inputs the selected encryption mode to the control register to control the selection of the SM3 or SHA256 algorithm;
[0035] This module can encrypt any plaintext length. When the input plaintext length is greater than or equal to 512bit, it is padded to a multiple of 512, and then the padded data is block-processed in 512bit chunks. Each time 512bit is input for encryption, the status register is used to monitor whether the encryption is completed. After a group of encryption is completed, a new 512bit plaintext is input for encryption until all groups are encrypted and then the hash value is output.
[0036] Furthermore, the AES&SM4 module can arbitrarily select four modes (AES128, AES192, AES256, and SM4). The reconfigurable design reduces resource usage, and a method based on the composite field is used to implement the S-box;
[0037] Anti-attack means are realized by adding random masks. Among them, the linear operations use XOR masks, and the non-linear part uses a reconfigurable masked S-box. The protection based on random masks will randomize the key before its use. After the final result is obtained through calculation, the mask is removed through the mask correction term to restore the expected output and achieve the anti-attack effect.
[0038] Furthermore, the anti - attack SM2 algorithm architecture is specifically as follows:
[0039] Introduce a random quantity to convert the input point of the elliptic curve point multiplication into a random projective coordinate point (xz 2 , yz 2 , z), so that the operation power consumption of the same point is different each time, and balance the operation amounts of point addition and point doubling; introduce a random number r during scalar operation, calculate k=(k + r)-r, when doing point multiplication, kP=(k + r)P - rP = kP, and the final calculation result is the same but the intermediate operation power consumption is disrupted. Introduce a random shift register RSR to generate random variables, perform corresponding pseudo - operations according to the random variables, and finally verify the Q coordinate of the result obtained after the point multiplication operation to check whether the result falls on the elliptic curve. If an error occurs, immediately terminate the operation and report an error;
[0040] The anti - attack SM4 algorithm architecture is specifically as follows:
[0041] First, introduce random number masks m1 and m2 to perform XOR masking on the input plaintext and key respectively, and perform demasking restoration operations before the encryption and decryption results are output and after the key expansion is completed, so that the entire calculation process is covered by the mask, ensuring that the values calculated each time are different;
[0042] Second, introduce a random number mask m3 to perform XOR masking on the input of the S - box and then perform the pre - affine transformation. After the post - affine transformation is completed, perform demasking to obtain the output of the S - box, thus ensuring the randomization of the S - box calculation;
[0043] Finally, since it is difficult to accurately inject faults at the same point of the same operation simultaneously, redundant hardware circuit modules are added to detect whether errors occur in the linear transformation. If the results of the calculation path and the redundant detection path are different, immediately terminate the operation and report an error.
[0044] Furthermore, the full - random verification is specifically to write a python model through the Jupyter Lab tool to perform full - random verification on the encryption algorithm, and automatically complete the encryption and decryption detection of a large amount of data.
[0045] After adopting the above content, the present invention has the following advantages: 1. Independent execution space: TEE and REE have their own execution spaces, ensuring the safe operation of the program, preventing the leakage of sensitive data and unauthorized access;
[0046] 2. Secure startup and legality verification: The system realizes a secure startup process, including image signature and image signature verification, ensuring the legality and integrity of the software image;
[0047] 3. Physical Address Access Firewall IOPMP: Restricts the access of non-secure cores to secure world resources such as memory and peripherals through IOPMP, enhancing the security of the system;
[0048] 4. Modular Design: Includes Mailbox module, SM2 module, SHA256&SM3 module, and AES&SM4 module, achieving modular design for easy maintenance and upgrade;
[0049] 5. Chip Side-Channel Defense Technology: Adopts attack-resistant SM2 and SM4 algorithm architectures, improving the system's defense against side-channel attacks;
[0050] 6. Reconfigurable Cryptographic Algorithm IP: Achieves reconfigurable design of AES and SM4 algorithms, reducing resource usage and improving the flexibility and efficiency of the algorithms;
[0051] 7. Full Random Verification: Writes a python model through the Jupyter Lab tool to perform full random verification on encryption algorithms, automatically completing encryption and decryption detection of a large amount of data, improving the accuracy and reliability of algorithm verification;
[0052] 8. Wide Application Scenarios and Optimized Resource Usage: Applicable to various scenarios such as cloud computing, Internet of Things, and mobile devices that require high trust and security. Through reconfigurable design and optimized hardware implementation, it reduces hardware resource consumption and improves performance;
[0053] 9. Improved Security: Enhances the security of the system through mechanisms such as secure boot, IOPMP, and Mailbox communication, protecting devices and data from malicious attacks;
[0054] 10. Practicality and Innovation: This application is not only innovative in theory but also highly practical in actual applications, providing strong technical support for Internet of Things security. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 It is a schematic diagram of the overall autonomous design architecture of a dual-core TEE security system built based on the open-source E902 processor.
[0056] Figure 2 It is a schematic diagram of the mirror signature process of a dual-core TEE security system built based on the open-source E902 processor.
[0057] Figure 3 It is a schematic diagram of the authentication & data integrity verification process of a dual-core TEE security system built based on the open-source E902 processor.
[0058] Figure 4It is a schematic diagram of the secure boot process of a dual-core TEE security system built based on the open-source E902 processor.
[0059] Figure 5 It is a schematic diagram of the IOPMP function process of a dual-core TEE security system built based on the open-source E902 processor.
[0060] Figure 6 It is a schematic diagram of the PMP configuration register form of a dual-core TEE security system built based on the open-source E902 processor.
[0061] Figure 7 It is a schematic diagram of the Mailbox communication of a dual-core TEE security system built based on the open-source E902 processor.
[0062] Figure 8 It is a schematic diagram of the signature generation algorithm of a dual-core TEE security system built based on the open-source E902 processor.
[0063] Figure 9 It is a schematic diagram of the verification algorithm process of a dual-core TEE security system built based on the open-source E902 processor.
[0064] Figure 10 It is a schematic diagram of the SM2 hardware architecture of a dual-core TEE security system built based on the open-source E902 processor.
[0065] Figure 11 It is a schematic diagram of the hardware architecture of the cryptographic protocol layer of a dual-core TEE security system built based on the open-source E902 processor.
[0066] Figure 12 It is a schematic diagram of the dot product operation architecture of a dual-core TEE security system built based on the open-source E902 processor.
[0067] Figure 13 It is a schematic diagram of the hardware architecture of the lazy reducer of a dual-core TEE security system built based on the open-source E902 processor.
[0068] Figure 14 It is a schematic diagram of the Montgomery algorithm process of a dual-core TEE security system built based on the open-source E902 processor.
[0069] Figure 15 It is a schematic diagram of the algorithm process of a dual-core TEE security system built based on the open-source E902 processor.
[0070] Figure 16 It is a schematic diagram of the hardware circuit of the SHA256&SM3 module of a dual-core TEE security system built based on the open-source E902 processor.
[0071] Figure 17It is a schematic diagram of the SM3 & SHA256 encryption process of a dual-core TEE security system built based on the open-source E902 processor.
[0072] Figure 18 It is a summary schematic diagram of the SM4 algorithm of a dual-core TEE security system built based on the open-source E902 processor.
[0073] Figure 19 It is a schematic diagram of the F and F' functions of a dual-core TEE security system built based on the open-source E902 processor.
[0074] Figure 20 It is a schematic diagram of the key expansion process and g function processing of a dual-core TEE security system built based on the open-source E902 processor.
[0075] Figure 21 It is a schematic diagram of the AES encryption and decryption process of a dual-core TEE security system built based on the open-source E902 processor.
[0076] Figure 22 It is a schematic diagram of the AES & SM4 hardware component modules of a dual-core TEE security system built based on the open-source E902 processor.
[0077] Figure 23 It is a schematic diagram of the overall architecture of the AES_SM4 system of a dual-core TEE security system built based on the open-source E902 processor.
[0078] Figure 24 It is a schematic diagram of the anti-attack SM2 dot multiplication algorithm architecture of a dual-core TEE security system built based on the open-source E902 processor.
[0079] Figure 25 It is a schematic diagram of the anti-attack SM4 algorithm architecture of a dual-core TEE security system built based on the open-source E902 processor. Specific implementation manners
[0080] Based on the single-core wujian100, an additional core is added as a non-secure core, and the original E902 serves as the secure core. The CPU running in the secure world is called the TEE CPU, and the CPU running in the non-secure world is called the REE CPU. Each E902 core can run its own program separately. The non-secure core runs the non-secure world program, and the secure core runs the secure world program. Among them, the secure world program includes calling cryptographic IPs such as SM2 and SHA256 for encryption and decryption; implementing secure boot; configuring IOPMP, etc.
[0081] Each E902 core has its own independent storage area, ensuring that the dual-core can run independent software programs separately. The secure core adopts a tightly coupled system, and an ICACHE is added inside the CPU. Since the ICACHE is directly connected to the instruction bus and does not need to go through the bus matrix, the tightly coupled system can accelerate the instruction fetching, thereby improving the running efficiency of the TEE. The non-secure core uses the ISRAM and DSRAM under the bus to store programs. Since the TEE adopts a tightly coupled system, it ensures that the non-secure CPU cannot access the secure world program. The ICACHE stores the initial running bootloader program. The ISRAM and DSRAM store the REE execution program.
[0082] The access permissions of the non-secure core to the serial port and the Crypto encryption and decryption module are restricted through IOPMP. Since the TEE adopts a tightly coupled system, there is no need to restrict the REE's access to the memory under the bus through IOPMP, achieving the purpose of saving resources. This way of restricting access permissions can effectively prevent malware from attacking the system, thereby protecting the privacy and data security of users.
[0083] One SPI interface is hung under the AHB2APB1 bridge. This interface is used for data interaction with the external Flash, enabling data transfer and program loading; and the IOPMP is used to restrict the access permissions of the non-secure core to the peripherals mounted under the APB1 bridge.
[0084] Crypto encryption and decryption modules are mounted on the AHB bus, namely the AES / SM4 module, the SM2 module, and the SM3 / SHA256 module, which are used to encrypt and decrypt programs, sign and verify signatures, and verify the integrity of data, realizing the function of verifying the legality of software images.
[0085] Communication between the secure core and the non-secure core is carried out through the Mailbox. The TEE CPU uses the TEE2REE Mailbox to transmit data to the REE CPU, and the REE CPU uses the REE2TEE Mailbox to transmit data to the TEE CPU.
[0086] The security solution is as follows:
[0087] The dual-core SOC platform we designed is comprehensively constructed with a variety of security solutions. For example, a physical address firewall is used to restrict the access space of the non-secure core, and a Crypto encryption and decryption module is used to complete the secure boot process to ensure the credibility of the system. As can be seen from the previous section, the dual-core SoC of this design adopts a two-layer hierarchical bus design. The CPU and Icache located in the isolated system TEE form a tightly coupled system. Therefore, the isolated system is a completely isolated environment, and the internal data cannot be directly accessed by external master devices. This can ensure that the program executed by the TEE CPU cannot be tampered with by the REE CPU, guaranteeing the isolation of the TEE CPU.
[0088] The second-layer bus layer, the Bus Matrix, is used to build the TEE system. The TEE CPU can access all slave devices mounted on this bus through this bus. At the same time, in order to make the whole system more secure, the TEE CPU controls the access rights of the REE CPU to some RAM under this bus by configuring the IOPMP access firewall.
[0089] The third-layer bus layer is the sub-bus AHB LS BUS of the Bus Matrix. The TEE CPU can access all slave devices under this bus, while the REE CPU restricts the access rights to the Crypto encryption and decryption module through the IOPMP, as well as restricts the access rights to the SPI peripherals of the AHB2APB1 bridge. Therefore, the REE CPU cannot tamper with the data in the Flash at will, protecting the data in the Flash. The following will introduce each part of the security solution in detail.
[0090] I. Secure Boot
[0091] In the secure boot solution, the secure core TEE needs to perform a legality check on the software image and verify the loading program step by step. Therefore, the secure boot will start from the secure core first. Only when the secure core starts, completes the legality check of the image and completes the configuration, will the non-secure core be started.
[0092] 1. Image Signing and Image Signature Verification
[0093] Extract the digest of the SPL program, TEE program, and REE program images using SHA256, and perform the signature process on the hash values of each image using the SM2 algorithm through the private key of the image publisher. Store the generated digest value, signature value, and image in the Flash. The basic process of image signing is as Figure 2 shown;
[0094] The program loaded from Flash first needs to use the public key in the Bootrom and the transported signature value to verify the hash value of the image. If the verification passes, the integrity of the image can be further verified. Otherwise, the startup fails. If the verification passes, the hash value of the image is recalculated and compared with the hash value hash(image) stored in Flash. If they are equal, it indicates that the integrity verification of the image has passed; otherwise, the startup fails. Authentication and data integrity are as Figure 3 shown below:
[0095] 2. Secure Boot Process
[0096] Before secure boot, the following steps need to be taken:
[0097] 1). First, call the national cryptography algorithm library of Python to generate a signature value and a digest value using the bin file generated by the REE program code as ciphertext. The public key PublicKey required for the SM2 algorithm is solidified in the BootRom. In addition, a small piece of code is solidified in the BootRom for power-on self-start. Its main function is to transport the SPL (Second Program Loader, the second-stage program loader) in Flash, and perform authentication and data integrity verification. After decryption is completed, secondary transportation will be carried out;
[0098] 2). In Flash, the TEE, REE program codes, and the hash value (digest value) and signature value generated through the image signature process are stored. The key used for Python signature and the public key stored in the BootRom are a pair.
[0099] The secure boot process is roughly divided into the following steps:
[0100] 1). The secure core is powered on and reset. The chip will start running from the starting address 0x0, that is, run the code in the BootRom. Our BootRom is named ICACHE0. At this time, ICACHE0 will access Flash through the secure SPI serial port to read the signature, digest, and SPL program for the image verification signature process, complete the software image legality check. If the legality verification passes, the SPL program will be transported to the ICACHE1 address in the TEE ICACHE, and the TEE CPU will jump to this address to execute the SPL program segment;
[0101] 2) In the SPL program, there is also a code program for data transfer. Similarly, the Bootloader in the second stage will also transfer the code program. Here, the program transferred is the TEE initialize and Run program, as well as the signature and digest. After passing the authentication and data integrity verification, they are transferred to ICACHE2, and the running address also jumps to ICACHE2 to start the security configuration of the TEE system, such as the configuration of IOPMP;
[0102] 3) After the security core completes the security configuration, it will continue to transfer the code program. Here, the program transferred is the REE software program, as well as the signature and digest. After passing the authentication and data integrity verification, the REE program digest and digest signature value are transferred to ICACHE2. The TEE CPU verifies the legality of the REE program and configures the REE reset signal. The REE program code is transferred to ISRAM. When the TEE CPU raises the REE reset signal, the REE system starts to initialize and run. If the TEE CPU does not complete the configuration of the REE CPU reset signal, the REE will wait all the time, which ensures that the non-secure core starts after the security core completes the security configuration;
[0103] In summary, the entire secure boot process is as Figure 4 shown.
[0104] II. IOPMP of Physical Address Access Firewall
[0105] IOPMP (Input / Output Physical Memory Protection) is used to protect the physical memory in the system from unauthorized access by malware or devices. These malicious behaviors can lead to problems such as data leakage, system crashes, and security vulnerabilities. By implementing the IOPMP mechanism, a computer system can regulate the memory access behaviors of peripherals with DMA functions or Bus Master ports, identify and prevent potential malicious access, thereby enhancing security;
[0106] IOPMP is usually implemented by hardware and configured and managed by the CPU. Its implementation principle is that the CPU configures the physical memory through the IOPMP interface, divides the physical memory into multiple regions, and assigns an access permission to each region. These regions are usually divided based on the concept of Memory Domain. Then, it sets the access permissions of each memory domain through the IOPMP interface, including reading, writing, executing, and protecting, etc.;
[0107] In terms of specific implementation, first, only the access permissions between the TEE and the REE are controlled, and access from other master devices is uniformly denied; second, since the bus only has read or write operations, only the read and write permissions are controlled; third, the allow signal output by the IOPMP is used to control whether to perform read or write operations on the peripheral or memory. To sum up, the IOPMP functional flow is as Figure 5 shown below. The functions we implemented are as follows:
[0108] 1. Only allow the TEE to configure the IOPMP;
[0109] 2. When it is determined that the SID is the TEE, access is definitely allowed. We default that the TEE has the permission to access all peripherals;
[0110] 3. When it is determined that the SID is the REE, check whether the address is within the configured range. If not, access is definitely denied; if it is, then check whether it has the corresponding permission. If it does, access is granted; if not, access is denied;
[0111] 4. When it is determined that the SID is neither the TEE nor the REE, access is denied.
[0112] Since the SoCs we designed all use peripherals or memory through read and write operations on the bus, theoretically, access from the master device to the slave device can be controlled by controlling the bus. Therefore, in addition to protecting physical memory, the IOPMP can also be used to protect other system resources, such as peripherals.
[0113] The IOPMP realizes its functions by configuring two types of registers: the configuration register and the address register.
[0114] As Figure 6 shown below, where R, W, and X respectively correspond to read, write, and execute permissions. When the value is 1, the permission exists; when the value is 0, the permission does not exist. For L, when L is 0, both the physical memory protection setting register and the address register of this table entry can be modified; when L is 1, all contents of this table entry, including the L bit, cannot be modified before the CPU resets.
[0115] The function of the A field needs to be explained in conjunction with the address register. The values of the A field are shown in Table 1. The specific explanations are as follows:
[0116] When A = OFF, this PMP entry is in the disabled state and does not match any address.
[0117] When A = TOR, the address range controlled by this PMP entry is determined by the address register of the previous PMP entry (with the value pmpaddr i-1 ) and the address register of this PMP entry (with the value pmpaddri ) Co - determination. It matches any address y that satisfies the following conditions:
[0118] pmpaddr i-1 <y < pmpaddr i
[0119] When A = NA4, if the value of pmpaddr is yyyy...yyyy, then the address space controlled by this PMP entry is 4 bytes starting from yyyy...yyyy00, namely the four addresses yyyy...yyyy00, yyyy...yyyy01, yyyy...yyyy10, yyyy...yyyy11, and each address stores one - byte data.
[0120] When A = NAPOT, as shown in Table 2, first calculate the number of consecutive 1s starting from the low - order bits of pmpaddr, and then deduce the address - space size. If the value of pmpaddr is yyyy...yyy0, that is, the number of consecutive 1s is 0, then the address space controlled by this PMP entry is 8 bytes starting from yyyy...yyy000. If the value of pmpaddr is yyyy...yy01, that is, the number of consecutive 1s is 1, then the address space controlled by this PMP entry is 16 bytes starting from yyyy...yy0000, and so on. If the value of pmpaddr is y...y01...1, and the number of consecutive 1s is n, then the address space controlled by this PMP entry is 2 n+3 bytes.
[0121] Table 1 Encoding of the A field in the IOPMP configuration register
[0122]
[0123] Table 2 Encoding ranges represented by the IOPMP address and configuration registers
[0124]
[0125] III. Introduction of Main Modules
[0126] 1. Mailbox Module
[0127] In a dual - core SoC architecture, inter - processor communication (IPC) is crucial for achieving data communication, event control, and resource sharing. As Figure 7As shown in the figure, in this design, Mailbox communication based on the combination of shared memory and inter-core interrupt is used, and the shared memory is implemented using FIFO. The dual-core process can transfer data through the mailbox, that is, perform read and write communication according to the interaction protocol agreed upon by both parties. To solve the read-write consistency problem of the shared memory and protect the data security of the secure core in the shared area, the TEE CPU and REE CPU use their own independent shared memories.
[0128] As Figure 7 It can be seen that when the TEE CPU transfers data to the REE CPU through Mailbox (TEE2REE), it first writes the data type, data length, and data into the FIFO of Mailbox (TEE2REE). After the writing is completed, Mailbox (TEE2REE) generates a tee2ree_intr request signal to notify the REE CPU to receive the data from Mailbox (TEE2REE). When the REE CPU finishes receiving the data, Mailbox (TEE2REE) generates a tee2ree_rsp reply signal to notify the TEE CPU that the data has been received by the REE CPU, and the communication ends. Similarly, when the REE CPU transfers data to the TEE CPU, there are corresponding ree2tee_intr request signals and ree2tee_rsp reply signals to complete the communication. To be able to respond quickly to inter-core communication, this design uses the method of interrupt. Therefore, to implement dual-core communication, two interrupt response functions for Mailbox need to be added to each of the two processor cores respectively.
[0129] 2. SM2 Module
[0130] The function of the design of the SM2 module is to verify the legality of the image during the secure boot process, that is, to verify the identity of the image publisher through the SM2 module.
[0131] As an asymmetric cryptographic algorithm, the SM2 algorithm uses different encryption keys and decryption keys to sign and verify the signature of data during application. The SM2 algorithm is also known as the digital signature algorithm. The digital signature algorithm generates a digital signature for data by a signer and verifies the reliability of the signature by a verifier. Each signer has a public key and a private key, where the private key is used to generate the signature, and the verifier uses the public key of the signer to verify the signature. Digital signature verification is divided into two parties, one is the signing party that sends the information, and the other is the verifying party that receives the information. Before starting the digital signature, the signing party needs to send its public key to the verifying party for subsequent verification by the verifying party.
[0132] The main functions of the digital signature algorithm are to verify data integrity and identify identities. The algorithm includes two technologies: a one-way hash function and an elliptic curve. Among them, the one-way hash function is also called the hash algorithm. It can compress plaintext information of any length through continuous iteration and output information of a fixed required length. It has the following properties:
[0133] When inputting plaintext information, the process of calculating the hash value is simple, but its reverse process, that is, calculating the plaintext from the hash value, is very difficult, which is why it is called a one-way function.
[0134] It is very difficult to find two different plaintext messages that can generate the same hash value.
[0135] No information related to the plaintext message can be seen from the hash value.
[0136] Therefore, according to these properties of the hash function above, the hash value generated by the hash function can correspond one-to-one to the message itself, and the hash value also has the advantage of short bit width. Therefore, in digital signature, we do not need to sign the plaintext message itself, but only need to sign the hash value of the plaintext, which can greatly speed up the signature speed. The SM2 digital signature generation algorithm process is as Figure 8 shown.
[0137] After the verification party receives the information M′ and the signature information (r′, s′), it also first needs to compress the information through the hash function. The specific verification algorithm process is as Figure 9 shown.
[0138] The specific design idea is as follows:
[0139] 2.1 Hardware Architecture Design of the Digital Signature System
[0140] The entire hardware architecture can be divided into four parts. The first part is the AHB bus layer, which is mainly composed of four registers. Among them, the control register is used to configure parameters for the entire digital signature system. The input / output register is used to input and output data. The status register is used to display the working status of the digital signature system. Design an AHB bus interface at this layer to mount the entire system IP onto the bus. The CPU transmits input data and control data to the system IP through the AHB bus by instructions, controls the IP to perform protocol operations such as key generation, digital signature, or digital signature verification. After the operation is completed, the CPU obtains the operation result of the IP through the AHB bus again.
[0141] The second part is the cryptographic protocol layer, which mainly controls the operation processes of key generation, digital signature, and digital signature verification for the entire system.
[0142] The third part is the elliptic curve operation layer, which includes a control module for point multiplication operation and a point operation control module. Since the operation methods of the two curves are different, the point operation control module is designed in two parts.
[0143] The fourth part is the modular operation module, which includes modular addition, modular multiplication, modular squaring modules, and a lazy reducer for Koblitz curves.
[0144] 2.2 Hardware Architecture Design of the Cryptographic Protocol Layer
[0145] The main controller of the cryptographic protocol layer has two functions. The first is to connect to the AHB bus layer and obtain parameter configuration information from the bus layer. The second is to schedule key generation, digital signature, and digital signature verification operations. A main control state machine controls the key generation state machine, digital signature state machine, and digital signature verification state machine. Then, by calling the point multiplication operation module, point addition operation module, and modular operation unit, the operations of the three protocols are completed.
[0146] 2.3 Hardware Architecture Design of the Elliptic Curve Operation Layer
[0147] The point multiplication operation is the core module of the entire system and has two functions. One is to receive data information sent by the protocol layer controller, and the other is to call the point operation controller to complete the point multiplication operation. Its hardware architecture diagram is as Figure 12 shown. The entire module completes the following control process: 1. Perform parameter configuration to determine which curve to use; 2. For general random curves, start the point multiplication operation after scanning the scalar k; for Koblitz curves, start lazy reduction, perform τNAF generation operation, and then start the point multiplication operation; 3. Start the point multiplication operation; 4. Control the modular inverse module and the basic operation module to perform coordinate inverse transformation.
[0148] 2.4 Hardware Design of the Binary Field Modular Operation Layer
[0149] Modular addition: The addition operation in the binary field corresponds to the exclusive OR of polynomial coefficients, and subtraction is equivalent to addition. Therefore, a single exclusive OR gate can be used to complete modular addition and subtraction operations.
[0150] Lazy reducer: Design a modular reducer based on a subtractor, an adder, two registers, and a shift register. The hardware architecture diagram is as Figure 13 shown.
[0151] Modular squaring: The property that addition in the binary field is equal to exclusive OR results in only even terms remaining after squaring, with all odd terms being zero. Therefore, the squaring operation is implemented using the zero-insertion method. After obtaining the squaring result, it is substituted into the modular reduction module to obtain the modular squaring result.
[0152] Modular multiplication: Modular multiplication is the operation with the highest complexity and also the key operation that determines the system performance. Here, the Montgomery modular multiplication algorithm is adopted, and the result of the modular multiplication operation is obtained through shifting. The algorithm is as Figure 14 shown.
[0153] 3. SHA256 & SM3 Module
[0154] 3.1 SM3 Module
[0155] The specific operation process of the SM3 algorithm can be divided into several stages: padding, iterative compression, and generating a hash value. The following will briefly describe these stages. The algorithm flow is as Figure 15 shown.
[0156] The input data is a message m with a length of l (l < 2^64) bits. The SM3 hash algorithm generates a hash value after padding and iterative compression. The length of the hash value is 256 bits.
[0157] Padding: Assume the length of the message m is l bits. First, add the bit "1" to the end of the message, and then add k "0"s. k is the smallest non - negative integer that satisfies l + 1 + k ≡ 448 mod 512. Then add a 64 - bit bit string, which is the binary representation of the length l. The bit length of the padded message m' is a multiple of 512. For example: for the message 0110000101100010 01100011, its length l = 24, and after padding, the bit string obtained is:
[0158]
[0159] Group expansion: The group expansion process first performs a grouping operation on the padded message in blocks of 512 bits, and then expands the 512 - bit message blocks after grouping to generate 132 words for subsequent compression function processing. The specific process of group expansion is as follows: Group the padded data m' in groups of 512 bits each. The number of groups after grouping depends on the message length l, that is, n = (l + k + 65) / 512. Then expand the above - mentioned message blocks according to the message expansion rule to generate 132 words such as W0, W1, …, W67, W0', W1', …, W63'. The specific rules are as follows:
[0160] Table 1 Group Expansion Rules
[0161]
[0162] Compression function: The input of the compression function CF is a 512 - bit message block after grouping. Each group of messages needs to go through 64 rounds of cyclic iteration, and this process is also called state update. The compression process of the message is as follows:
[0163] FOR i = 0 TO n - 1
[0164] V (i+1) = (V (i) , B (i) )
[0165] ENDFOR.
[0166] The specific process of the SHA256 algorithm has some similarities with SM3. The general process is divided into: preprocessing the message, including padding and adding length information; grouping processing, where each group goes through multiple rounds of compression functions and various bit operations to obtain a 256-bit hash value; and concatenating the hash values of all groups to obtain the final 256-bit hash value.
[0167] 3.3. Hardware Circuit Implementation of SHA256 & SM3
[0168] The hardware circuit of the SHA256 & SM3 module is as Figure 16 shown. The core of the control unit is the counter circuit. The control unit determines the number of rounds of iteration after block expansion according to the external control signal. Since the number of iteration rounds required by SM3 and SHA256 is inconsistent, in addition, the control unit also selects the operation module. The data path is mainly composed of modules such as message padding, message expansion, iterative compression, output transformation, ROM, and compression function registers. The input data is sent to the iterative compression module for operation after message padding and message expansion. The calculated data is sent to the compression function register through output transformation, and it is judged whether the specified number of rounds is reached according to the number of iteration rounds determined by the control unit. If the specified number of rounds is reached, the message is output. The ROM module is used to store the special constant Kt and the initial hash value required for hash operation.
[0169] We designed the reconfigurable RTL code for SM3 & SHA256. The overall process of SM3 & SHA256 is as Figure 17 shown. When the SM3 & SHA256 module is called during the secure boot process, the TEE CPU first inputs the selected encryption mode to the control register to control the selection of the SM3 or SHA256 algorithm. This module can encrypt any plaintext length. When the input plaintext length is greater than or equal to 512 bits, it becomes a multiple of 512 after padding, and then the padded data is block-processed by 512 bits. Each time 512 bits are input for encryption, the status register is used to monitor whether the encryption is completed. After a group of encryption is completed, a new 512-bit plaintext is input for encryption until all blocks are encrypted and the hash value is output.
[0170] 4. AES & SM4 Modules
[0171] 4.1. SM4 Module
[0172] The SM4 algorithm is a block symmetric cipher algorithm. The block length of this algorithm is 128 bits (i.e., 16 bytes, 4 words), and the key length is also 128 bits. Its encryption and decryption processes adopt a 32-round iteration mechanism. Each round requires a round key, and the generation of the round key is obtained through iteration from the initial key. Both encryption and decryption need to go through 32 rounds of iteration and an endian transformation.
[0173] The encryption process of the algorithm is as follows. For details, see Figure 18 :
[0174] Input 4 words of plaintext (X 0 , X 1 , X 2 , X 3 ) (where X i represents a 32-bit word), and 4 words of key (MK 0 , MK 1 , MK 2 , MK 3 );
[0175] After 32 rounds of iteration, each round of iteration requires a 1-word round key, denoted as (rk 0 , rk 1 , …, rk 31 ). The iteration process is to continuously use the round function F to calculate the next word backward. The round function F is F(X i , X i+1 , X i+2 , X i+3 , rk i );
[0176] After 32 iterations, we can get a total of 36 words (X 0 , X 1 , …, X 35 ). Finally, an endian transformation is performed to reverse the last four words (X 32 , X 33 , X 34 , X 35 ) obtained from the iteration to get the final ciphertext (Y 0 , Y 1 , Y 2 , Y 3 ) = (X 35 , X 34 , X 33 , X 32 ).
[0177] The decryption process of the algorithm is exactly the same as the encryption process, and it also includes 32 rounds of iteration and an endian transformation. Only when iterating, the round keys need to be used in reverse order.
[0178] The round function F is composed of the composite permutation T, and T is composed of the non-linear transformation t and the linear transformation L. After one non-linear transformation and one linear transformation, one iteration is completed and the next word is calculated. The construction of the F function is as Figure 19 shown.
[0179] Similarly, the key expansion module also obtains 32 round keys through a similar iterative process. First, the original key (Mk 0 , Mk 1 , Mk 2 , MK 3 ) is XORed with the system parameter FK to obtain (K 0 , K 1 , K 2 , K 3 ). Then, through 32 rounds of iteration, each round of iteration requires a fixed parameter round constant of 1 word, denoted as (CK 0 , CK 1 , …, CK 31 ). The iterative process is to continuously use the key round function F’ to calculate the next word backward. The key round function F’ is F’(K i , K i+1 , K i+2 , K i+3 , CK i ), where K i+4 = rk i ; Similarly, the key round function F’ is also composed of the composite permutation T’, and T’ is composed of the non-linear transformation and the linear transformation. The construction of F’ is as Figure 19 shown.
[0180] 4.2, AES Module
[0181] AES is also a symmetric block iterative encryption and decryption algorithm, and the key is the foundation for the AES algorithm to implement encryption and decryption. Currently, the AES key grouping supports 128 bits, 192 bits, and 256 bits, and the length of the plaintext to be encrypted or the ciphertext to be decrypted is fixed at 128 bits. Different key lengths correspond to 10 rounds, 12 rounds, and 14 rounds of iteration respectively. The AES round operation has a substitution-permutation structure, which is composed of round key addition, byte substitution, row shift, and column mixing. The encryption process of AES is as Figure 21 shown on the left.
[0182] In addition, key expansion is also required, and the key expansion process is as Figure 20As shown on the left, taking the 128-bit key expansion as an example, a total of 11 sub-round keys are generated. The g function in the key expansion algorithm serves two purposes. One is to increase the non-linearity in the key scheduling, and the other is to eliminate the symmetry in AES. Both of these properties are necessary to resist certain block cipher attacks. The processing flow of the g function is as Figure 20 shown on the right.
[0183] The AES decryption process is the reverse deduction operation of the encryption process. As Figure 21 shown on the right, the decryption process also requires iterative operations with round keys. The decryption operations are respectively composed of inverse row shift, inverse byte substitution, add round key, and inverse column mixing. The ciphertext is first XORed with the sub-round key obtained in the last round, and then undergoes 9 rounds of the same decryption operations. The inverse column mixing is not required in the last round, thus obtaining the plaintext.
[0184] 4.3, Reconfigurable Design of SM4 & AES
[0185] The hardware circuit structures of AES and SM4 are introduced above. Here, a strategy for realizing reconfigurable design is proposed - to achieve this purpose by integrating similar or identical modules in the two algorithms. The hardware component modules of the AES algorithm are as Figure 22 shown on the left. In the top-level structure, the bus is used to obtain the data and control signals sent by the upper-layer software. The encryption / decryption controller and the key expansion controller are implemented by state machines and are respectively used to control the iterations of encryption / decryption and key expansion, located in the second-layer control layer. The sub-operation layer includes operations such as the S-box, row shift, column mixing, and add round key, which all play important roles in the encryption / decryption round transformation. Among them, the S-box is an important operation in the key expansion. The hardware component modules of SM4 are very similar to those of AES. Therefore, they have similar interfaces and control layer architectures, and both include sub-operation layer structures such as the S-box and add round key. The hardware component modules of SM4 are as Figure 22 shown on the right.
[0186] The role of the interface layer is mainly responsible for the input and output of bus data, including the data of the control register, the data of the status register, and the result data. Since AES supports three key bit widths and algorithm selection, here a key register with the highest bit width of 256 bits is configured, and a control register for algorithm selection and key bit width selection is added. The role of the control layer is mainly for algorithm selection, and it is controlled by a state machine to select different algorithms, and jumps to different states according to the data in the interface layer to select different algorithms for encryption / decryption. The role of the sub-operation layer is mainly responsible for the operation implementation of the algorithm, and the most area-consuming S-box is implemented by calculation to reduce the circuit area, Figure 23 which is the overall architecture diagram of the encryption / decryption system for the AES_SM4 algorithm.
[0187] Based on the idea of software-hardware co - operation, the CPU is responsible for providing data and control signals. The CPU uses the bus addressing method to read and write the corresponding hardware registers through the AHB bus, controlling the selection of AES128, 192, 256 and SM4 algorithms, encryption / decryption operations, sending plaintext / ciphertext and keys, reading the hardware operation status, and reading the encryption / decryption results.
[0188] IV. Chip Side - Channel Defense Technology
[0189] As a non - invasive attack method, side - channel attack will not damage the chip and is an attack method with low cost and high efficiency. Its basic idea is to obtain the voltage or electromagnetic signal changes of the chip in various working states, thereby reflecting the power consumption changes of the chip during different operations. Combining the chip input and output data, using statistical methods to process the power consumption curve, and then extracting useful key information.
[0190] 1. Anti - attack SM2 Algorithm Architecture
[0191] To counter side - channel attacks, the anti - attack techniques adopted in the design of SM2 are as follows: introducing a random quantity to convert the input point of the elliptic curve point multiplication into a random projective coordinate point (xz 2 , yz 3 , z), so that the power consumption of the same point in each operation is different; balancing the operation amounts of point addition and point doubling; introducing a random number r during scalar operation, calculating k=(k + r)-r, and during point multiplication, kP=(k + r)P - rP = kP, the final calculation results are the same but the intermediate operation power consumption is disrupted; introducing a random shift register RSR to generate random variables and perform corresponding pseudo - operations according to the random variables; finally, verifying the result Q coordinates obtained after the point multiplication operation to check whether the result falls on the elliptic curve. If an error occurs, the operation is immediately terminated and an error is reported.
[0192] According to the anti - attack scheme designed above, the anti - attack SM2 point multiplication algorithm architecture is as Figure 24 shown.
[0193] 2. Anti - attack SM4 Algorithm Architecture
[0194] This scheme mainly adopts full masking and error detection for anti - attack design. Figure 25It is an architecture diagram of the SM4 algorithm resistant to attacks. First, random number masks m1 and m2 are introduced to perform XOR masking on the input plaintext and key respectively, and the masking is restored before the encryption / decryption results are output and after the key expansion is completed, so that the entire calculation process is covered by the mask, ensuring that the values calculated each time are different; Second, a random number mask m3 is introduced to perform XOR masking on the input of the S-box and then perform the forward affine transformation. After the backward affine transformation is completed, the mask is removed to obtain the output of the S-box, thus ensuring the randomization of the S-box calculation; Finally, since it is difficult to accurately inject faults at the same point of the same operation simultaneously, redundant hardware circuit modules are added to detect whether errors occur in the linear transformation. If the results of the calculation path and the redundant detection path are different, the operation is immediately terminated and an error is reported.
[0195] V. Full random verification
[0196] When verifying the rationality of the cryptographic algorithm, we used several groups of data for verification. This detection method lacks accuracy. We wrote a python model through the JupyterLab tool to perform full random verification on the encryption algorithm, automatically completing the encryption / decryption detection of hundreds of groups of data to verify the reliability of the algorithm.
[0197] The present invention proposes a dual-core TEE security system based on the open-source E902 processor. This architecture includes a secure core (TEE) and a non-secure core (REE), which execute programs independently, and the IOPMP restricts the non-secure core's access to secure resources. This is the core of the system security. The specific advantages are as follows:
[0198] 1. Secure boot and legality verification:
[0199] The system realizes a secure boot process based on the trust chain, combines the SHA256 and SM2 algorithms to extract the digest and verify the signature of the software image, ensuring the secure boot of the system and the legality of the software.
[0200] 2. Physical address access firewall IOPMP:
[0201] The use of IOPMP effectively restricts the non-secure core's access to secure resources and enhances the physical security protection of the system.
[0202] 3. Reconfigurable cryptographic algorithm IP:
[0203] The present invention realizes the reconfigurable design of the AES and SM4 algorithms, reduces resource consumption, and improves the flexibility and efficiency of the algorithms. This design allows the algorithms to be dynamically adjusted under different security requirements.
[0204] 4. Chip side-channel defense technology:
[0205] For side-channel attacks, the present invention adopts anti-attack SM2 and SM4 algorithm architectures, and improves the system's defense ability against side-channel attacks through random masking and error detection mechanisms.
[0206] 5. Mailbox communication mechanism:
[0207] The communication between the secure core and the non-secure core implemented through the Mailbox module ensures the security of data exchange and the stability of the system.
[0208] 6. Full random verification:
[0209] The present invention uses the Jupyter Lab tool for full random verification, automatically completes the encryption and decryption detection of a large amount of data, and improves the accuracy and reliability of algorithm verification.
[0210] The above advantages are specifically realized in the following ways:
[0211] 1. The secure core TEE adopts a tightly coupled structure, and the storage space of the TEE core is an independent ICACHE, without passing through the system bus matrix, to achieve a completely independent TEE core environment.
[0212] 2. Realize the dual-core operation of independent programs and be able to print output through the serial port, and the TEE CPU controls the REECPU to start.
[0213] 3. This dual-core SOC integrates AES, SM4, SM2, SM3, SHA256 cryptographic algorithm IPs for encrypting, decrypting, signing, and verifying signatures of programs.
[0214] 4. Based on the trust chain, combined with SHA256 to extract the program digest and SM2 signature verification to implement the secure boot process.
[0215] 5. Realize communication between the two cores based on the shared memory and inter-core interrupt scheme.
[0216] 6. Through IOPMP, the non-secure core REE cannot control the Crypto module and some peripherals such as the SPI serial port.
[0217] 7. Realize the reconfigurability of the AES and SM4 algorithms, and four modes (AES128, AES192, AES256, and SM4) can be arbitrarily selected. The reconfigurable design reduces resource usage, and the S-box is implemented using a method based on the composite domain.
[0218] 8. AES and SM4 implement anti - attack measures by adding random masks. Among them, the linear operations use XOR masks, and the non - linear part uses a reconfigurable mask S - box. The protection based on random masks randomizes the key before its use. After obtaining the final result through calculation, the mask is removed by the mask correction term to restore the expected output, achieving the anti - attack effect.
[0219] 9. Implement the reconfiguration of the SM3 and SHA256 algorithms. The reconfigurable design reduces resource usage. This module is used during the self - start process to verify the integrity of the self - start.
[0220] 10. When verifying the rationality of the cryptographic algorithm, we used several groups of data for verification. This detection method lacks accuracy. By writing a Python model using the JupyterLab tool to perform full - random verification on the encryption algorithm, automatically complete the encryption and decryption detection of hundreds of groups of data to verify the reliability of the algorithm.
[0221] The above describes the present invention and its implementation manners. This description is not restrictive, and the actual structure is not limited thereto. Generally speaking, if those of ordinary skill in the art are inspired by it and, without departing from the purpose of the present invention, creatively design a structural manner and embodiments similar to this technical solution, they shall fall within the protection scope of the present invention.
Claims
1. A dual-core TEE security system built on the open source E902 processor, featuring: It includes the following:
1. Secure boot: including image signature, image signature verification and secure boot process; 2. Physical address access firewall IOPMP; 3. Main modules: including Mailbox module, SM2 module, SHA256&SM3 module and AES&SM4 module; 4. Chip side channel defense technology: including the anti-attack SM2 algorithm architecture and the anti-attack SM4 algorithm architecture; Five: Fully random verification.
2. The dual-core TEE security system based on the open source E902 processor according to claim 1 is characterized in that: The image signature is to use SHA256 to extract the summary of the SPL program, TEE program, and REE program images, and use the SM2 algorithm to execute the signature process on the hash value of each image through the private key of the image publisher, and store the generated summary value, signature value, and image in Flash; The image signature verification is that the program loaded from Flash first needs the public key in Bootrom and the transferred signature value to verify the hash value of the image. If the signature verification passes, the image integrity can be further verified, otherwise, the startup fails. If the signature verification passes, the hash value of the image is recalculated and compared with the hash value hash(image) stored in Flash. If they are equal, it indicates that the integrity verification of the image has passed, otherwise the startup fails.
3. The dual-core TEE security system based on the open source E902 processor according to claim 1 is characterized in that: The secure boot process includes the following steps: A. The following steps need to be performed before safe boot:
1. First, call the Python national secret algorithm library, use the bin file generated by the REE program code as ciphertext to generate the signature value and digest value, and use the public key PublicKey required by the SM2 algorithm, and solidify it in BootRom; 2. The TEE and REE program codes and the hash value (digest value) and signature value generated by the image signing process are stored in Flash. The key used for Python signature is a pair with the public key stored in BootRom. B. The secure boot process is roughly divided into the following steps:
1. After the security core is powered on and reset, the chip will start running from the starting address 0x0, that is, running the code in BootRom. BootRom is named ICACHE0. At this time, ICACHE0 will access the Flash through the secure SPI serial port to read the signature, summary, and SPL program to perform the image verification process and complete the software image legitimacy verification. If the legitimacy verification passes, the SPL program will be moved to the ICACHE1 address in the TEE ICACHE, and the TEE CPU will jump to this address to execute the SPL program segment; 2. In the SPL program, there is also a code program for moving data. Similarly, the second-stage Bootloader will also move code programs. Here, the TEE initialize and run program, as well as the signature and digest, are moved to ICACHE2 after passing the identity authentication and data integrity verification. The running address also jumps to ICACHE2 to start the security configuration of the TEE system, such as the configuration of IOPMP.
3. After the security core completes the security configuration, it will continue to move the code program. Here, the REE software program, signature, and summary are moved. After passing the identity authentication and data integrity verification, the REE program summary and summary signature value are moved to ICACHE2. The TEE CPU verifies the legitimacy of the REE program and configures the REE reset signal. The REE program code is moved to ISRAM. When the TEE CPU pulls up the REE reset signal, the REE system starts to initialize and run. If the TEE CPU has not completed the configuration of the REE CPU reset signal, the REE will wait. This ensures that the security core completes the security configuration before starting the non-security core.
4. The dual-core TEE security system based on the open source E902 processor according to claim 1 is characterized in that: The physical address access firewall IOPMP specifically implements the following functions: A. Only TEE is allowed to configure IOPMP; B. When the SID is determined to be TEE, it must be allowed. We assume that TEE has access to all peripherals. C. When the SID is determined to be REE, determine whether the address is within the configuration range. If not, it must be rejected. If it is, determine whether it has the corresponding permission. If so, access is granted. If not, access is denied. D. When it is determined that the SID is neither TEE nor REE, access is denied.
5. The dual-core TEE security system based on the open source E902 processor according to claim 1 is characterized in that: The Mailbox module TEE CPU and REE CPU use their own independent shared memory; When the TEE CPU transmits data to the REE CPU through the Mailbox (TEE2REE), it first writes the data type, data length, and data into the FIFO of the Mailbox (TEE2REE). After writing, the Mailbox (TEE2REE) generates a tee2ree_intr request signal to notify the REE CPU to receive data from the Mailbox (TEE2REE). When the REE CPU completes data reception, Mailbox (TEE2REE) generates a tee2ree_rsp reply signal to notify the TEE CPU that the data has been received by the REE CPU and the communication is over. Similarly, when the REE CPU transmits data to the TEE CPU, there are corresponding ree2tee_intr request signals and ree2tee_rsp reply signals to complete the communication. In order to quickly respond to inter-core communication, this design adopts the interrupt method. Therefore, in order to achieve dual-core communication, both processor cores need to add two Mailbox interrupt response functions respectively.
6. The dual-core TEE security system based on the open source E902 processor according to claim 1 is characterized in that: The SM2 module includes the following structure: A. Digital signature system hardware architecture; B. Hardware architecture of cryptographic protocol layer; C. Elliptic curve operation layer hardware architecture; D. Binary domain modular operation layer hardware.
7. The dual-core TEE security system based on the open source E902 processor according to claim 1 is characterized in that: The SHA256&SM3 module is designed with reconfigurable RTL code. Specifically, when the SM3&SHA256 module is called during the secure boot process, the TEE CPU first inputs the encryption mode selection to the control register to control whether to select the SM3 or SHA256 algorithm. This module can encrypt plaintext of any length. When the input plaintext length is greater than or equal to 512 bits, it is padded to become a multiple of 512, and then the padded data is divided into 512-bit blocks. Each time 512 bits are input for encryption, the status register is used to monitor whether the encryption is completed. After a group of encryption is completed, a new 512-bit plaintext is input for encryption until all groups are encrypted and the hash value is output.
8. The dual-core TEE security system based on the open source E902 processor according to claim 1 is characterized in that: The AES&SM4 module arbitrarily selects four modes (AES128, AES192, AES256 and SM4), the reconfigurable design reduces resource usage, and adopts a composite domain-based approach to implement the S-box; Anti-attack measures are implemented by adding random masks, in which linear operations use XOR masks, and the nonlinear part uses reconfigurable masked S-boxes. The protection based on random masks will randomize the key before it is used. After the final result is calculated, the mask is removed through the mask correction term to restore the expected output and achieve anti-attack effect.
9. The dual-core TEE security system based on the open source E902 processor according to claim 1 is characterized in that: The anti-attack SM2 algorithm architecture is as follows: Introduce a random quantity to convert the elliptic curve point multiplication input point into a random projection coordinate point (xz 2 , yz 2 , z), so that the power consumption of the same point is different each time, and the amount of calculation of the point addition and the point doubling is balanced; the random number r is introduced in the scalar operation, k = (k + r) - r is calculated, and kP = (k + r)P - rP = kP is used for point multiplication. The final calculation result is the same, but the intermediate calculation power consumption is disrupted. The random shift register RSR is introduced to generate random variables, and the corresponding pseudo operation is performed according to the random variables. Finally, the result Q coordinate obtained after the point multiplication operation is completed is verified to verify whether the result falls on the elliptic curve. If an error occurs, the operation is terminated immediately and an error is reported; The architecture of the SM4 algorithm that is resistant to attack is as follows: First, random number masks m1 and m2 are introduced to perform XOR masking on the input plaintext and key respectively. Before the encryption and decryption results are output and after the key expansion is completed, the masking and restoration operations are performed, so that the entire calculation process is covered by the mask, ensuring that the value of each calculation is different; Secondly, a random number mask m3 is introduced to XOR mask the input of the S-box and then perform a pre-affine transformation. After the post-affine transformation is completed, the mask is removed to obtain the S-box output, thereby ensuring the randomization of the S-box calculation; Finally, since it is difficult to accurately inject faults into the same point of the same operation at the same time, redundant hardware circuit modules are added to detect whether errors occur in linear transformations. If the results of the calculation path are different from those of the redundant detection path, the operation is terminated immediately and an error is reported.
10. The dual-core TEE security system based on the open source E902 processor according to claim 1 is characterized in that: The fully random verification is specifically to write a python model through the Jupyter Lab tool to perform fully random verification on the encryption algorithm, and automatically complete the encryption and decryption detection of large amounts of data.
Citation Information
Cited By
Trusted execution environment control circuit and artificial intelligence chip
CN121069856A