Method for performing software scrambling by binary rewriting and electronic equipment for performing software scrambling by binary rewriting
By rewriting and coroutine management of binary software, the problems of defense against bypass attacks and anti-tampering are solved, and random scrambling and anti-piracy of cryptographic operations are realized, which is suitable for ordinary microcontrollers.
Patent Information
- Application Number
- CN202311845262.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2025-07-01
AI Technical Summary
The existing technology is difficult to effectively defend against bypass attacks, especially energy analysis attacks, and there are insufficient anti-tampering and anti-piracy measures for cryptography operations and MCU firmware.
By rewriting the binary software, selecting some instructions to replace them with a jump instruction sequence using the deformation number and original code rewriting information, jumping to the additional code unit to complete the logical function and then returning, combining coroutine management and random scrambling mechanism, random scrambling and tamper-proof monitoring of cryptographic operations can be realized.
Without increasing hardware costs, it is effectively defended against energy analysis attacks, suitable for ordinary microcontrollers, realizes anti-pirated version and tamper-proof of one machine and one code, and adapts to various cryptographic operations.
Smart Images

Figure CN120234782A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of computer software technology, cryptography, microcontrollers, and integrated circuits. Background Art
[0002] Currently, cryptography is widely used in the field of computer technology and is everywhere in life. Examples include symmetric ciphers such as AES and SM4, asymmetric ciphers such as RSA and ECC, and hash algorithms such as sha256 and sm3. Attacks against cryptography are a very important research area.
[0003] Side channel attack, also known as power analysis attack, is to utilize the data specificity shown by confidential information in aspects such as time, power consumption, and electromagnetic signal dissipation during cryptographic operations, and through statistical analysis, to obtain various confidential information during the operations. It includes SPA, DPA (Differential Power Analysis), SEMA, DEMA, Timing attatck, etc.
[0004] For specific knowledge, reference can be made to
Energy Analysis Attacks, Translated Series of Mathematical Classics, Translated by Feng Dengguo, Zhou Yongbin, Liu Jiye, etc., Science Press, ISBN number: 9787030281357
Advanced DPA Theory and Practice, ISBN number: 9787514913125
[0005] The anti-tampering and anti-piracy of MCU firmware are also long-standing requirements. Summary of the Invention
[0006] The purpose of this application is to provide a method for rewriting binary software and an electronic device for rewriting binary software.
[0007] To achieve this goal, the basic idea is to rewrite information based on the deformation number and the original code, select a part of the instructions in the original binary code and replace them with a jump instruction sequence. The jump instruction sequence jumps to the additional code. After the additional code completes the same logical function as the replaced instructions, it then jumps back to the position after the replaced instructions and continues to execute. The additional code unit can also run additional calculations, external coroutines, and other additional logical functions.
[0008] According to the basic idea, this application adopts the following technical solutions: A software method for rewriting binary software, which creates an original code unit, an original code rewriting information unit, a deformation number unit, an original code rewriting execution unit, an image code unit, and an additional code unit; The original code unit, the image code unit, and the additional code unit are all collections of binary instructions and data; The image code unit and the additional code unit form a rewritten code unit, and the rewritten code unit performs the same logical function as the original code unit; The original code rewriting information unit is generated by binary instruction analysis of the original code unit, and stores information such as which code elements can be rewritten and the allowed rewriting types; The deformation number unit is data that is different each time binary software rewriting is performed, or randomly generated data; The original code rewriting execution unit, based on the original code rewriting information unit and the deformation number unit, selects a part of the instructions of the original code unit to be replaced with a jump instruction sequence to generate an image code unit; The jump instruction sequence jumps to the additional code unit. After the additional code unit completes the same logical function as the replaced instruction, it then jumps back to the position after the replaced instruction and continues to execute.
[0009] Furthermore, a context save unit, an additional calculation unit, and a context restore unit are also created; in addition to performing the same logical function as the replaced instruction, the additional code unit also runs the additional calculation unit to execute logical functions not available in the original code unit. The context save unit is run before running the additional calculation unit, and the context restore unit is run afterwards to prevent the existence of the additional calculation unit from affecting the logical function of the original code unit.
[0010] Furthermore, the rewritten code unit serves as the main coroutine unit, and a coroutine switching unit is also created in the additional calculation unit. Additionally, several external coroutine units and a coroutine management unit are created; The external coroutine unit also includes a context save unit, a coroutine switching unit, and a context restore unit, and also includes an external stack unit and an external calculation unit. The external stack unit enables the operation of the external coroutine unit not to affect the main stack used by the main coroutine unit, and the external calculation unit completes the calculation tasks of the external coroutine unit; The coroutine management unit includes a coroutine information storage unit and a coroutine selection unit corresponding to the main coroutine unit and the external coroutine unit; The coroutine switching units of the main coroutine unit and the external coroutine unit both jump to the coroutine management unit. The coroutine management unit saves the information of the jumped coroutine unit to the corresponding coroutine information storage unit, and then the coroutine selection unit selects a coroutine information storage unit, loads the information of the coroutine unit and jumps to the corresponding main coroutine unit or external coroutine unit.
[0011] Furthermore, the deformation number unit is a random number generator, and the binary codes of the main coroutine unit and / or the external coroutine unit are the same rewritten code unit for cryptographic operations; One of the coroutine units serves as a cryptographic operation coroutine unit to perform cryptographic operations on the real key and data and generate the required output results; The other coroutine units serve as random scrambling coroutine units to perform the same cryptographic operations on random input data and discard the generated output results; The coroutine management unit randomly selects the next coroutine unit to run each time a coroutine switch occurs based on the random number generated by the random number generator.
[0012] Furthermore, the deformation number unit generates a deformation number from the unique data of the processor or electronic device running the rewritten code unit.
[0013] An electronic device for performing binary software rewriting includes an original code unit, an original code rewriting information unit, a deformation number unit, an original code rewriting execution unit, an image code unit, and an additional code unit; The original code unit, the image code unit, and the additional code unit are all collections of binary instructions and data; The image code unit and the additional code unit constitute a rewritten code unit, and the rewritten code unit completes the same logical functions as the original code unit; The original code rewriting information unit is generated by analyzing the binary instructions of the original code unit and stores information such as which code elements can be rewritten and the allowed rewriting types; The deformation number unit is data that is different each time binary software rewriting is performed, or randomly generated data; The original code rewriting execution unit selects a part of the instructions of the original binary code unit to be replaced with a jump instruction sequence based on the original code rewriting information unit and the deformation number unit, and generates an image code unit; The jump instruction sequence jumps to the additional code unit. After the additional code unit completes the same logical functions as the replaced instructions, it jumps back to the position after the replaced instructions and continues to execute.
[0014] Furthermore, it also includes a live saving unit, an additional calculation unit, and a live restoring unit. In addition to completing the same logical functions as the replaced instructions, the additional code unit also runs the additional calculation unit to execute logical functions not available in the original code unit. The live saving unit is run before running the additional calculation unit, and the live restoring unit is run after that to avoid the existence of the additional calculation unit affecting the logical functions of the original code unit.
[0015] Furthermore, the rewritten code unit serves as the main coroutine unit, and a coroutine switching unit is also created in the additional calculation unit. In addition, it includes several external coroutine units and a coroutine management unit; The external coroutine unit also includes a context save unit, a coroutine switching unit, and a context restoration unit, and further includes an external stack unit and an external computing unit. The external stack unit enables the operation of the external coroutine unit not to affect the main stack used by the main coroutine unit, and the external computing unit completes the computing tasks of the external coroutine unit; The coroutine management unit includes a coroutine information storage unit and a coroutine selection unit corresponding to the main coroutine unit and the external coroutine unit; The coroutine switching units of the main coroutine unit and the external coroutine unit both jump to the coroutine management unit. The coroutine management unit saves the information of the jumped coroutine unit to the corresponding coroutine information storage unit, and then the coroutine selection unit selects a coroutine information storage unit, loads the information of the coroutine unit and jumps to the corresponding main coroutine unit or external coroutine unit.
[0016] Further, the deformation number unit is a random number generator, and the binary codes of the main coroutine unit and / or the external coroutine unit are the same rewritten code unit for performing cryptographic operations; One of the coroutine units serves as a cryptographic operation coroutine unit to perform cryptographic operations on the real key and data and generate the required output result; The other coroutine units serve as random scrambling coroutine units to perform the same cryptographic operations on random input data and discard the generated output results; The coroutine management unit randomly selects the next running coroutine unit each time a coroutine switch occurs according to the random number generated by the random number generator.
[0017] Further, the deformation number unit generates a deformation number from the unique data of the processor or electronic device running the rewritten code unit.
[0018] Compared with the prior art, the embodiments of the present application have the following beneficial effects: Through binary rewriting, random scrambling of cryptographic operations is achieved, which is applicable to ordinary general microcontrollers, does not increase hardware costs, and has a good effect of preventing energy analysis attacks. At the same time, it is adaptable to various cryptographic operations and does not require studying the principles of cryptographic operations.
[0019] By combining the unique data of the microcontroller chip or electronic device, binary rewriting of the software running therein is performed to achieve binary code obfuscation, which can achieve one code for one machine, prevent piracy from running on other chips, and anti-tampering monitoring can be performed in the additional code. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0021] Figure 1 A flow chart of a method for performing binary rewriting in Embodiment 1 of the present application; Figure 2 This is a flow chart of a method for completing additional calculations using binary rewriting in Example 2 of the present application; Figure 3 This is a flow chart of a method for rewriting a plug-in coroutine using binary in Example 3 of the present application; Figure 4 This is a flow chart of a method for performing random scrambling using multiple coroutines in Example 4 of the present application; Figure 5 This is a flow chart of a method for fully masking confidential information using multiple coroutines in Example 5 of the present application; Figure 6 This is a flow chart of the method for forced salt addition in Example 6 of the present application. Implementation
[0023] In order to make the purpose, features, and advantages of the invention of this application more obvious and easy to understand, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the embodiments described below are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0024] The technical solution of the present application is further explained below with reference to the accompanying drawings and through specific implementation methods. Embodiment 1
[0025] This application takes the ARM Cortex-M0 core as an example to illustrate a method for binary rewriting.
[0026] For the description of the Cortex-M0 core, please refer to the ARMv6-M Architecture Reference Manual, https: / / developer.arm.com / documentation / ddi0419 / latest / . The CPU has registers such as R0~R12, SP, LR, PC, APSR, etc. that are related to computational task codes. The instruction set is mainly 16-bit wide instructions, and there are also a small number of 32-bit wide instructions (MSR MRS UDF BL DSB DMB ISB, etc. 7 instructions. According to the 16-bit encoding rule, the 16-bit value starting from 0xe800 and the next 16 bits together constitute a 32-bit instruction).
[0027] The original code unit 102 is any binary code of Cortex-M0, which is generally compiled by C language and assembly language, and generates obj file after being compiled by compiler and assembler, and then linked by linker to generate binary code in the format of hex, s19, bin, etc. After being imported into MCU's FLASH or SRAM, it can be run to complete the predetermined logic function.
[0028] The original code unit can be an entire binary software or a part of the binary software, such as a function, a function library, etc., corresponding to the part of the binary software.
[0029] Performing binary instruction analysis on the original code unit 102 can generate the original code rewrite information unit 103. The purpose is to analyze the properties of each instruction storage unit (16 bits wide for the M0 core, which can be called a code element), so as to clarify which code elements can be binary rewritten and the binary rewriting method, so as to simplify the work of the original code rewrite execution unit 101.
[0030] For the 16-bit basic storage unit of the M0 core binary code (which can be called a code element, for the 8051 core, the code element is 8 bits. There are also MCUs with pure 32-bit instructions, and the code element is 32 bits), the two most important attributes are: instruction or data, 16 bits or 32 bits. There are four situations: data, 16-bit instruction, the first 16 bits of a 32-bit instruction, and the last 16 bits of a 32-bit instruction. We can call the attributes of the code element the code class.
[0031] Analyzing the code class attributes of code elements is time-consuming and difficult (ARM RISV and others are relatively simple, CISC instruction sets such as 8086 system are more complex, and the simplest are equal-width instruction sets such as PIC16). It needs to be analyzed in advance and saved in the original code rewrite information unit 103.
[0032] Binary instruction analysis can utilize the information in files such as lst, obj, and map generated by compilers, assemblers, and linkers. For example, write a "code class attribute extraction software" to automatically extract the code class attributes of each code element from these files.
[0033] The author of this article plans to develop this code class attribute extraction software for two compilers, IAR EWARM and Keil MDK.
[0034] Currently, the mainstream compilers are produced by foreign manufacturers, mainly Keil MDK, IAR EWARM, and GCC.
[0035] With the development of domestic MCU chips, our country has also introduced its own compiler environment (mostly based on GCC). In the future, a code class list file for binary code can also be generated by the compiler, clearly listing the code class attributes of each code element.
[0036] Take a small C function as an example: int TestTask(int para0) { volatile int k = 0; for (int i = 0; i < 6; i++) { AsmTaskYield(); k = k ^ 0x5a69abcd; k++; } TaskEnd(); return 0; } The corresponding lst file generated after compilation is as follows: \ In section.text, align 2, keep-with-next 26 int TestTask(int para0) 27 { \ TestTask: (+1) \ 0x0 0xB538 PUSH {R3 - R5,LR} \ 0x2 0x0004 MOVS R4,R0 28 volatile int k = 0; \ 0x4 0x2500 MOVS R5,#+0 \ 0x6 0x9500 STR R5,[SP, #+0] 29 for (int i = 0; i < 6; i++) \ ??TestTask_0: (+1) \ 0x8 0x2D06 CMP R5, #+6 \ 0xA 0xDA0A BGE ??TestTask_1 30 { 31 AsmTaskYield(); \ 0xC 0x....'.... BL AsmTaskYield 32 k = k ^ 0x5a69abcd; \ 0x10 0x9800 LDR R0, [SP, #+0] \ 0x12 0x.... LDR R1,??DataTable8_2 ;;0x5a69abcd \ 0x14 0x4041 EORS R1, R1, R0 \ 0x16 0x9100 STR R1, [SP, #+0] 33 k++; \ 0x18 0x9800 LDR R0, [SP, #+0] \ 0x1A 0x1C40 ADDS R0, R0, #+1 \ 0x1C 0x9000 STR R0, [SP, #+0] 34} \ 0x1E 0x1C6D ADDS R5, R5, #+1 \ 0x20 0xE7F2 B ??TestTask_0 35 TaskEnd(); \ ??TestTask_1: (+1) \ 0x22 0x....'.... BL TaskEnd 36 return 0; \ 0x26 0x2000 MOVS R0, #+0 \ 0x28 0xBD32 POP {R1, R4, R5, PC} ;; return 37} The corresponding information in the map file after linking is as follows: TestTask 0x800'0081 0x2a Code Gb main.o [1] Image in MCU memory, 42 bytes (0x2a) starting from address 0x800'0080, corresponding to 21 16-bit code elements: 0x38,0xB5,0x04,0x00,0x00,0x25,0x00,0x95,0x06,0x2D,0x0A,0xDA,0x00,0xF0,0x2A,0xF9, 0x00,0x98,0x57,0x49,0x41,0x40,0x00,0x91,0x00,0x98,0x40,0x1C,0x00,0x90,0x6D,0x1C, 0xF2,0xE7,0xFF,0xF7,0xE2,0xFF,0x00,0x20,0x32,0xBD, Actually, the immediate value 0x5a69'abcd is also placed at address 0x800'01f0 Information generated by the disassembler: TestTask: 0x800'0080: 0xb538 PUSH {R3-R5, LR} 0x800'0082: 0x0004 MOVS R4, R0 volatile int k=0; 0x800'0084: 0x2500 MOVS R5, #0 0x800'0086: 0x9500 STR R5, [SP] for(int i=0;i<6;i++) 0x800'0088: 0x2d06 CMP R5, #6 0x800'008a: 0xda0a BGE.N 0x800'00a2 AsmTaskYield(); 0x800'008c: 0xf000 0xf92a BL AsmTaskYield ; 0x800'02e4 k=k^0x5a69abcd; 0x800'0090: 0x9800 LDR R0, [SP] 0x800'0092: 0x4957 LDR.N R1, [PC, #0x15c] ; 0x5a69'abcd 0x800'0094: 0x4041 EORS R1, R1, R0 0x800'0096: 0x9100 STR R1, [SP] k++; 0x800'0098: 0x9800 LDR R0, [SP] 0x800'009a: 0x1c40 ADDS R0, R0, #1 0x800'009c: 0x9000 STR R0, [SP] for(int i=0;i<6;i++) 0x800'009e: 0x1c6d ADDS R5, R5, #1 0x800'00a0: 0xe7f2 B.N 0x800'0088 TaskEnd(); 0x800'00a2: 0xf7ff 0xffe2 BL TaskEnd ; 0x800'006a return 0; 0x800'00a6: 0x2000 MOVS R0, #0 0x800'00a8: 0xbd32 POP {R1, R4, R5, PC} This immediate data is placed at another discontinuous address: 0x800'01f0: 0x5a69'abcd DC32 0x5a69'abcd As can be seen from the disassembly text, among the 21 bytes from 0x800'0080 to 0x800'00a8, most of the bytes are 16-bit instructions (17 bytes), and there are only two 32-bit instructions (4 bytes), namely BL AsmTaskYield and BL TaskEnd. And the only immediate number is stored at another discontinuous address 0x800'01f0.
[0037] The most basic form of the original code rewriting information unit containing the code class information of each symbol can be an array containing N 2-bit elements (with values ranging from 0 to 3), corresponding to the N symbols of the original code unit; 0 represents data or an unknown symbol, 1 represents a 16-bit instruction, 2 represents the high 16 bits of a 32-bit instruction, and 3 represents the low 16 bits of a 32-bit instruction.
[0038] There is the original code rewriting information unit 103 containing the code class information of each symbol. The original code rewriting execution unit 101 can already, based on the data provided by the transformation number unit 104, select a part of the instructions in the original code unit 102, and according to the encoding rules of the instruction set (taking the Cortex-M0 kernel as an example, which is not too complex), determine whether replacement is possible, and select the jump instruction sequence for replacement, replace it with the jump instruction sequence, and generate the image code unit 106. The most basic rules are: 1. Data is not allowed to be rewritten; 2. A 32-bit instruction must be entirely replaced by a jump instruction sequence, and it is not allowed to rewrite only half; 3. There must be a replacement instruction sequence to achieve the same logical function as the replaced instruction sequence.
[0039] The jump instruction sequence jumps to the additional code unit 110. After the additional code unit 110 completes the same logical function as the replaced instruction, it then jumps back to the position after the replaced instruction and continues to execute.
[0040] The additional code unit 110 and the image code unit 106 together form the rewritten code unit 105, which realizes the same logical function as the original code unit.
[0041] The transformation number unit 104 can be a random number generator, or it can be based on the unique data of the MCU chip or electronic device, or a pseudo-random number sequence generated based on the unique data.
[0042] The original code rewriting execution unit 101 can run in the software deployment device of the MCU chip, such as a programmer. After binary rewriting, the rewritten code unit 105 is burned into the target MCU, which is generally used for anti-piracy of MCU firmware; It can also run on the same MCU chip as the rewritten code unit 105. In fact, the original code rewriting execution unit 101 can also be a part of the original code unit 102 and can be the object of binary rewriting. This method is generally used for random scrambling in cryptographic operations, but can also be used for anti-piracy of MCU firmware.
[0043] More analysis can be performed on the symbols of the original code unit to generate more information and store it in the original code rewriting information unit to achieve a faster and better instruction replacement effect.
[0044] For example, instructions can be classified according to the following attributes: Whether it is related to the PC? Whether it is a NOP or an equivalent NOP instruction? An arithmetic or a data transfer instruction? Instructions related to the PC can be further divided into branch instructions, LDR to fetch the literal pool instructions, etc.
[0045] Branch instructions can be divided into conditional branch instructions, unconditional branch instructions, subroutine call instructions, long-distance branch instructions, short-distance branch instructions, etc.
[0046] Among them, whether the execution logic of the instruction is related to the PC (Program Counter) is very important. If the execution logic of the instruction is not related to the PC (Program Counter) (also known as ROPI), then after jumping to the additional code unit, the instruction can be directly copied into the additional code unit to achieve the same logic as the instruction replaced in the original code unit. Such binary rewriting is very simple, and we can call it copy-based binary rewriting, which is the type of binary rewriting we mainly use. In fact, we can observe that most instructions are not related to the PC and are suitable for copy-based binary rewriting.
[0047] In the example code, the following several instructions are related to the PC and cannot be simply copied to the additional code unit for execution, while other instructions can be rewritten using copy-based binary rewriting: 0x800'008a: 0xda0a BGE.N 0x800'00a2 0x800'008c: 0xf000 0xf92a BL AsmTaskYield ; 0x800'02e4 0x800'0092: 0x4957 LDR.N R1, [PC, #0x15c] ; 0x5a69'abcd 0x800'00a0: 0xe7f2 B.N 0x800'0088 0x800'00a2: 0xf7ff 0xffe2 BL TaskEnd ; 0x800'006a
[0048] In the case of only the binary data of the original code unit, binary instruction analysis can be carried out by manual disassembly one by one, by virtual running after the virtual machine of the Cortex-M0 kernel loads the original code unit, or by the simulation debugger loading the original code unit into the actual Cortex-M0 kernel MCU chip for single-step execution, or a combination of multiple means can also be used.
[0049] In fact, sometimes when we only have binary code, it may not be possible to fully determine the code class of each code element. We can express it as "unknown state" in the corresponding information or classify it as data to avoid binary rewriting of this code element (generally, binary rewriting of data is not allowed).
[0050] For CPUs with other instruction sets, the code element width may be different, and the types of code classes may also be different.
[0051] The jump instruction sequence used to replace the original code unit needs to be carefully selected. There are two aspects: on the one hand, the jump instruction sequence is preferably a single 16-bit code element that can replace any instruction; on the other hand, it is hoped that the jump distance is as far as possible to facilitate placing the attached code unit at a relatively far address.
[0052] In actual applications, it is difficult to balance these two requirements.
[0053] The M0 core has a 16-bit instruction B.N xxxx short-distance unconditional jump instruction sequence, but it can only jump to addresses within the range of +-2KB. This requires that the attached code unit must be able to be placed at a relatively close address.
[0054] The more complex M3 core than M0 has a 32-bit long-distance unconditional jump instruction sequence B.W xxxxxx, which can jump to addresses within the range of +-16MB. However, since it is a 32-bit double code element, when replacing a single 16-bit instruction, it must be replaced together with the next replaceable instruction.
[0055] The M0 core does not have a long-distance unconditional jump instruction sequence, but some alternative methods can be thought of. For example, BL.W xxxxxx is a long-distance subroutine call instruction that can jump to the +-16MB address range, but it will destroy the content of the LR return address register. Therefore, two instructions, PUSH {Rm-Rn,LR} + BL.W xxxxxx, can be used as a long-distance replacement instruction sequence occupying 3 code elements. In the attached program unit, the content of Rm-Rn LR saved in the stack is used to restore the Rm-Rn LR registers. Since the replacement instruction sequence occupies 3 code elements, it must be 3 consecutive replaceable 16-bit instructions, or 2 consecutive 16-bit or 32-bit instructions, to replace and rewrite together. Rm-Rn can be 0 to 7 registers.
[0056] If the BL.W xxxxxx instruction is used to replace the original BL.W xxxxxx instruction, it can achieve a perfect replacement because the content of the LR register must have been saved when the original BL.W xxxxxx instruction was called.
[0057] Jump instruction sequences that are further than +-16MB require the use of instruction sequences such as PUSH {R0}+LDR.N R0,Address+BXR0+Address (literal pool) to implement, which account for 5 to 6 code elements (the literal pool is 32-bit address-aligned and requires an additional code element space). In the additional program unit, the content of R0 saved in the stack is used to restore the R0 register.
[0058] It is also possible to use the SWI software interrupt instruction to achieve replacement, but this method will affect the entire CPU interrupt system and has a greater risk of affecting the logical functions of the original code unit.
[0059] When jumping back from the additional code unit to the back of the replaced instruction, the available jump instructions are more flexible. First pushing the return address onto the stack and then POP {PC} is the best choice.
[0060] In addition to the code class attributes of the code elements, the execution frequency of the code elements is also important information for binary rewriting. Code elements with a low execution frequency have fewer execution opportunities for the additional code units they jump to after replacement. Code elements with a very high execution frequency have more execution times for the additional code units they jump to after replacement, which will have a greater impact on the running speed of the code.
[0061] If the algorithm structure of the original code unit is understood, such as the AES encryption and decryption algorithm, which has 11 rounds and each round is divided into several steps, the corresponding instruction positions suitable for rewriting can be found and stored in the original code rewriting information unit, and targeted binary rewriting can be performed based on the logical functions.
[0062] The most extreme approach is that the original code rewriting information unit directly stores the positions of which code elements can be rewritten in what forms, and does not need to store information about code elements that cannot be binary rewritten. For example, storing a table T0~Tn, with code element addresses D0~Dn, and rewrite forms R0~Rn, and storing D0 and R0 at table T0. In this way, the work of the original code rewriting execution unit 101 is very simple.
[0063] The original code rewriting information unit can be stored in any memory, such as internal FLASH, external FLASH, or obtained through a network connection before performing binary rewriting. Since the original code rewriting information unit is data, it cannot be the object of binary rewriting.
[0064] In addition to having the same logical functions, binary code rewriting will inevitably still produce various differences from the original binary code before rewriting, which need to be noted: A: The additional code unit 110 will necessarily occupy extra FLASH or SRAM space. When rewriting, it must be clear which address spaces are free spaces not used by the original code unit. The free spaces should be used to place the FLASH or SRAM used by the additional code unit. Subsequent embodiments will add various logic functions to the additional code unit and also use additional SRAM space to store additional variables, which should also be placed in the free SRAM space. In most cases, the original code unit belongs to the party that performs binary rewriting for some positive purposes (such as anti-piracy, random scrambling, etc.), so it is very clear which storage spaces are free. Or specifically reserve some FLASH and SRAM spaces for binary rewriting. Only when the binary rewriting is carried out by an adversary for some reverse purposes will it be unclear which storage spaces are free. A safe and feasible method is to select an MCU model with a larger storage space than the MCU on which the original code unit runs. For example, if it was originally 64KB FLASH and 8KB SRAM, an MCU model with 128KB FLASH and 16KB SRAM can be selected. For a binary original code unit, if it is to be prevented from being binary rewritten by an adversary, an MCU model with the highest resource amount can be deliberately used and all the storage space of the MCU chip can be used up.
[0065] B: After binary code rewriting, generally, the time consumed to complete the logic function will increase. Unless it is possible to clearly know the logic function of the original code and implement the same logic function in a faster way, binary rewriting only serves as a patch.
[0066] C: The long-distance replacement instruction sequence of M0 uses the stack space to save and restore the LR register, posing a risk of stack overflow. If the original code has taken some stack monitoring measures, it may be affected.
[0067] D: The image code unit and the original code unit may cause exceptions if some code integrity monitoring measures are carried out due to some instructions being replaced.
[0068] Embodiment 2 Based on Embodiment 1, Embodiment 2 adds additional calculations to complete additional logic functions that the original code unit does not have.
[0069] As attached Figure 2 It also includes the original code unit 202, the original code rewriting information unit 203, the deformation number unit 204, the original code execution unit 201, the image code unit 206, and the rewritten code unit 205.
[0070] Based on Embodiment 1, a on-site save unit 211, an additional calculation unit 212, and a on-site recovery unit 213 are added to the additional code unit 210.
[0071] The function of the on-site save unit 211 is to save registers such as R0 to R12, LR, and APSR. The function of the on-site recovery unit 213 is to restore registers such as R0 to R12, LR, and APSR. This avoids the execution of the additional calculation unit from damaging these registers.
[0072] Without these three units, the additional code unit can only have the effect of adding some random delays.
[0073] After on-site protection, the additional calculation unit 212 can arbitrarily use the CPU for various calculations and logical functions.
[0074] For example, for copyright detection, anti-tampering, anti-piracy, and so on.
[0075] For copyright detection, the copyright information string stored at a certain address can be checked in the additional calculation unit, such as "Made in China, 2023", etc.
[0076] For anti-tampering, the cumulative checksum of a certain section or the entire MCU FLASH space can be calculated in the additional calculation unit. If it is not the expected value, it is considered to have been tampered with intentionally or unintentionally.
[0077] For anti-piracy, on the one hand, the unique data of the MCU (such as the 96-bit UID that many MCUs have now, some calibration words, etc.) or the unique data of the peripheral devices of the MCU is used as the variant number of the variant number unit to participate in binary rewriting; on the other hand, anti-tampering detection is performed on the rewritten code unit after binary rewriting in the additional calculation unit.
[0078] Embodiment 3 Based on Embodiment 2, an external coroutine is added in Embodiment 3.
[0079] As shown in the appendix Figure 3 It also includes an original code unit 302, an original code rewriting information unit 303, a variant number unit 304, an original code execution unit 301, an image code unit 306, and a rewritten code unit 305. An additional code unit 310, a on-site save unit 311, an additional calculation unit 312, and a on-site recovery unit 313. Based on Embodiment 2, the additional calculation unit in the additional code unit is further expanded, and the mechanism of an external coroutine is introduced.
[0080] Coroutine is a concept in some high-level languages that emerged in recent years. It, together with processes and threads, belongs to the concept of implementing multitasking through CPU time-sharing, but is more lightweight and resource-consuming than processes and threads. Specific knowledge can be searched on the Internet for "coroutine, thread, process". In this application, the coroutine is binary-hooked to the code of the original code unit, and the coroutine context switch is performed through the corresponding local stack of the coroutine, and they run alternately.
[0081] In the additional computing unit 312 of the additional code unit 310, it is added or replaced with a coroutine switching unit 314, and jumps to the coroutine management unit 330.
[0082] The coroutine management unit 330 includes a coroutine selection unit 331, a coroutine information storage unit - main 332, and a coroutine information storage unit - external (there can be multiple groups).
[0083] In addition, an external coroutine unit 320 is created. Similar to the additional code unit 310, it also includes a context save unit 321, an external computing unit 322, a context restore unit 323, a coroutine switching unit 324, and also includes an external stack unit 325. There can be multiple groups of external coroutine units 320, corresponding to multiple external coroutines.
[0084] The external stack unit 325 is used to store the task switching context of the external coroutine unit 320, local variables allocated by the stack, etc. It can be a globally allocated SRAM area, a temporary SRAM area allocated by the main stack, a SRAM area dynamically allocated on the heap such as malloc, etc.
[0085] In distinction from the external coroutine unit 320, the additional code unit 310 can be called the main coroutine unit and uses the original main stack.
[0086] After jumping from the coroutine switching unit 314 in the additional code unit 310 or the coroutine switching unit 324 in the external coroutine unit 320 to the coroutine management unit 330, first save the coroutine information to the corresponding coroutine information storage unit - main 332 or coroutine information storage unit - external 333; then the coroutine selection unit 331 selects a group of coroutine information storage units, retrieves the coroutine information inside, and thus jumps back to the coroutine switching unit of the corresponding coroutine to continue running.
[0087] Coroutine selection is generally an integer variable (which can be denoted as int CoroutineIndex). 0 represents the main coroutine, 1 represents the first external coroutine, and so on.
[0088] Coroutine information is generally an array that stores the stack pointers of each coroutine (which can be denoted as Sp_Val[N], where N is the number of external coroutines + 1, and Sp_Val[0] is generally the stack pointer of the main coroutine). The main information of each coroutine (such as CPU registers, stack variables, etc.) is stored in its respective stack.
[0089] For example, when jumping from the main coroutine to the coroutine management unit 330 with CoroutineIndex being 0, the coroutine information of the main coroutine, that is, the value of the SP register when the main coroutine comes, is saved in the first position Sp_Val[0] of the array. Then, according to certain rules, such as taking a value from the transformation number unit 304 and after calculation, the coroutine selection unit 331 decides to run the external coroutine 1 next. It then modifies CoroutineIndex to 1, retrieves the value of Sp_Val[1], restores it to the SP register, points to the stack of the external coroutine 1, and then jumps to the coroutine switching unit of the external coroutine 1 to run the external coroutine 1.
[0090] The external calculation unit in the external coroutine unit can use the same binary code entity as the main coroutine, but only the data information in the external stack unit 324 is different from that of the main coroutine. This is very suitable for the random scrambling in cryptographic operations.
[0091] The code of the external coroutine and the main coroutine can call coroutine suspension functions such as TaskYield to return to the coroutine management unit and provide the opportunity for other coroutines to run. It can call coroutine end functions such as TaskEnd to notify the coroutine management unit that the external coroutine has completed all calculation tasks. Unless the corresponding coroutine is created again, do not switch to this external coroutine anymore.
[0092] Functions such as TaskYield and TaskEnd can be explicitly called in the original code or can be hooked into the original code in the way of binary rewriting.
[0093] Creating an external coroutine can be called the TaskCreate function, which can be done in the original code or can be hooked into the original code in the way of binary rewriting.
[0094] External coroutines can be used for random scrambling in cryptographic operations and can also be used for anti-piracy monitoring of MCU code, etc.
[0095] Example 4 In Example 4, the external coroutine in Example 3 is used for cryptographic operations for random scrambling to counter energy analysis attacks. Energy analysis attacks use the specificity of confidential information on the energy traces of the MCU (including the energy of the power supply and the electromagnetic energy emitted outward) to guess a part of the confidential information and thus crack the confidential information.
[0096] To counter energy analysis attacks, it is necessary to reduce the specificity of the confidential information in the energy traces of the MCU. On the one hand, it is to reduce the dissipated energy, which can be called shielding; on the other hand, it is to increase the noise component, which can be called masking.
[0097] A simple and inaccurate statement is to reduce the signal-to-noise ratio of the confidential information.
[0098] As shown in the appendix Figure 4 , create an encryption operation coroutine unit 411 to implement the cryptographic operations we need, such as AES encryption, AES decryption, SM3 hash algorithm, SM4 block encryption algorithm, SM2 asymmetric encryption algorithm, etc.
[0099] The encryption operation coroutine unit 411 performs cryptographic operations on the input parameter 410 to generate an output result 412.
[0100] The random scrambling coroutine unit 421 is an operation code unit identical to the encryption operation coroutine unit 411, but it performs cryptographic operations on the random input parameter 420 to generate a random output result 422 that is not used.
[0101] The random selection coroutine management unit 401 randomly selects the encryption operation coroutine unit 411 and the random scrambling coroutine unit 421 according to the random numbers generated by the random number generator unit 402 for the next stage of operation. After the encryption operation coroutine unit 411 and the random scrambling coroutine unit 421 run some instructions, they jump back to the random selection coroutine management unit 401 to continue the next random operation.
[0102] The random number generator unit 402 generates random numbers as the random input parameter 420 on the one hand, and generates random numbers as the input of the random task switching unit 401 on the other hand. Of course, the random numbers required by these two units can also use different random number generators.
[0103] The encryption operation coroutine unit 411 generates side-channel signals P0~Pn, and the random scrambling coroutine unit 421 generates side-channel signals R0~Rn, which can both be observed by the attacker. However, R0~Rn are random data, and at the same time, R0~Rn are randomly mixed with P0~Pn in time, making it difficult for the attacker to distinguish them and difficult to obtain confidential information such as keys from the side-channel signals. It is equivalent to R0~Rn masking P0~Pn.
[0104] The number of random scrambling coroutine units can be increased to reduce the signal-to-noise ratio of the confidential information and achieve a stronger masking effect.
[0105] In addition to the random scrambling coroutines using the same operation code units, random scrambling coroutines with different operation codes can also be added, which is equivalent to using Embodiment 3 and Embodiment 4 in combination.
[0106] For asymmetric cryptographic operations, the code volume is relatively large (from more than 10 KB to dozens of KB), the operations are complex, and it is relatively difficult to determine where to perform binary rewriting, resulting in the operation of a certain coroutine for a long time. A timer unit 403 can be added to introduce a time value into the random selection coroutine management unit 401. After running a certain coroutine for more than a certain time, in the way of generating an interruption by the timer, the coroutine is randomly forced to switch, and the time for the next interruption generated by the timer is randomly configured.
[0107] Embodiment 5 On the basis of Embodiment 4, Embodiment 5 introduces the concept of full masking of confidential information.
[0108] Taking AES encryption as an example, it consists of various 8-bit operations, such as 8-bit exclusive OR, 8-bit look-up table, 8-bit shift, 8-bit substitution, etc. The input of each operation has 8-bit confidential information and other non-confidential information, and generally generates an 8-bit result.
[0109] The confidential information includes both the AES key and the intermediate results that can deduce part of the AES key.
[0110] We first decompose the AES encryption process into one 8-bit operation after another.
[0111] See the appendix Figure 5 , for each 8-bit operation, the 8-bit confidential information 502 generates 255 other values out of 256 values as the masking information 512. The confidential information 502 and the non-confidential information 504 are sent into the operation coroutine 501 to generate the operation result 503.
[0112] The masking information 512 is also sent into 255 masking coroutines 511 together with the non-confidential information 504 to perform the same operation, generating discarded operation results 513. The operation codes of the 256 coroutines can be the same, but the running order is random.
[0113] After all the coroutines have finished running, only the operation result 503 is taken.
[0114] Then continue with the next 8-bit operation until the entire AES operation is completed.
[0115] Such an approach can be called the full masking of confidential information, that is, all possible values of the confidential information are sequentially calculated with non-confidential information in a random order. If the running order of 256 coroutines is truly random, statistically speaking, no matter how many energy traces of operations are collected in a power analysis attack, the data specificity brought by the confidential information during the operation cannot be found, and the confidential information cannot be deduced.
[0116] Of course, it should be noted that when loading the confidential information 502 and the masking information 512, and saving the operation result 503 and the discarded operation result 513, data comparison, data transfer, and memory reading and writing are used, which will inevitably generate some data specificity. However, these two processes are standardized operations, which are convenient for taking targeted masking measures.
[0117] On the one hand, masking coroutines can be used to perform masking data comparison, data transfer, and memory reading and writing. However, these operations do not have as much statistical specificity as 8-bit operations, and do not require as much masking information as 256 coroutines for full masking. On the other hand, the MCU generally also has modules such as DMA for data transfer, which can perform data transfer simultaneously with the CPU for masking.
[0118] We can design three dedicated software or hardware processes for full masking: A: Help generate masking values in a random order, such as generating 256 values from 0 to 255, but their order is randomly distributed. At the same time, during the generation process, the dissipated energy has the smallest data specificity of 256 values.
[0119] B: Load the confidential information 502 into the masking values in a random order according to the confidential information. At the same time, during the loading process, the dissipated energy has the smallest data specificity of the confidential information.
[0120] C: Save the operation result 503 from 256 operation results in a random order according to the confidential information. At the same time, during the saving process, the dissipated energy has the smallest data specificity of the confidential information.
[0121] Although 256 coroutines can achieve the full masking of confidential information, the operation speed is theoretically slowed down by 256 times. In fact, considering the coroutine switching time, it may be slowed down by thousands of times. It is only suitable for cryptographic operations at critical points, such as key derivation and other cryptographic operations.
[0122] On the other hand, the coroutine stacks of 256 coroutines will also occupy a large amount of SRAM space. Assuming that each coroutine stack is 256 bytes, 64KB is required.
[0123] Full masking of 8-bit confidential information does not require 256 randomly running coroutines for implementation. It is also entirely possible to use a single coroutine to calculate 256 possible values in a completely random order. It is also possible to use several coroutines, with each coroutine assigned several possible values for calculation. This only slows down the calculation speed of the entire cryptographic operation without excessive consumption of storage space.
[0124] On the other hand, all operations can be implemented using addition and shifting, and addition and shifting of multi-bit data can be implemented using addition and shifting of fewer bits. For example, 32-bit operations can be implemented using 8-bit operations. For single-bit, there is only addition, and single-bit shifting is actually data transfer.
[0125] In this way, we can decompose all operations into single-bit or 2-bit or 4-bit operations, and use the method of random multi-coroutines for full masking, with only 2 or 4 or 16 coroutines.
[0126] It should be noted that the multi-bit four arithmetic operations composed of full-masked single-bit addition and shifting operations may also generate data specificities related to confidential information and need to be carefully designed.
[0127] For example, for table lookup operations, the SBOX of AES, which combines single-bit table lookups into 8-bit table lookups, it is impossible to eliminate the data specificities of the table lookup operations of 8-bit confidential information by decomposing them into fewer bits.
[0128] If non-confidential information is masked, that is, the same input operation coroutine and masking coroutine as the confidential information are used, and the non-confidential information is input into the operation coroutine, while the other 255 values of the non-confidential information are input into the masking coroutine, there should also be some masking effects.
[0129] Example 5 can also be used in combination with Examples 1, 2, 3, and 4 for more random scrambling.
[0130] Example 6 Based on Example 5, Example 6 performs key derivation in a full masking manner and sets a forced salting mechanism based on one-way variable data comparison to ensure that the keys used in each decryption operation are different, combined with binary rewriting for random scrambling in each decryption operation. Thus, it not only has a strong ability to resist power analysis attacks but also eliminates the disadvantage of extremely slow speed caused by full masking.
[0131] As shown in the appendix Figure 6, The unidirectional change data 601 stores data such as 64 bits. The ciphertext 610 sent from the encryption end is appended with the same 64-bit salt 602. The salt judgment unit 603 compares the salt 602 with the unidirectional change data 601 to check whether it conforms to the unidirectional change rule. If it does not conform, the decryption operation is refused.
[0132] If the salt 602 conforms to the unidirectional change rule, on the one hand, the salt 602 is saved to the unidirectional change data 601, and on the other hand, the salt 602 is sent to the key derivation unit 605. The key derivation unit 605 uses the salt 602 and the confidential information 604 (RootKey) with the key derivation algorithm to generate a one-time working key 606 (WorkKey). The key derivation process adopts a fully masked method to effectively deal with high-order DPA attacks.
[0133] The decryption operation unit 611 decrypts the ciphertext 610 with the working key 606 to generate the plaintext 612. The decryption operation can adopt a certain weak masking to be able to resist SPA, that is, simple power analysis.
[0134] The unidirectional change rule of the salt can be simple mathematical increment or decrement, that is, each time the encryption end encrypts, a working times counter is incremented or decremented by 1 and used as the salt value.
[0135] Since the current time is a unidirectional change quantity, the current time can be used as the salt value, and the encryption end does not need to save a working times counter. If the 32-bit value is in seconds, it can represent 135 years. Using a 64-bit value to represent time, even in microseconds, it can represent tens of thousands of years.
[0136] The solution of Embodiment 6 does not provide mandatory protection for the encryption end because the encryption end itself holds the plaintext and the confidential information RootKey, and generally will be in a secure environment and is not easily attacked by opponents.
[0137] The symmetric encryption key, that is, the confidential information RootKey, is stored at both the encryption end and the decryption end. The current MCU storage space is relatively large and can store much more information than the AES key. For example, it is 16 times the AES128 key, which is only 256 bytes. The key derivation algorithm derives a 16-byte AES key from the 256-byte confidential information and the salt value, which can greatly increase the difficulty of recovering all 256 bytes of confidential information from the key derivation link and the 256-byte confidential information reading link, and can also simplify the operation of the key derivation algorithm.
[0138] In fact, we can design a special key derivation algorithm to derive a relatively short working key of 32 bytes from a large sample data of 256 bytes according to the salt value. Each time of derivation, not all the information of each bit of the 256 bytes is used.
[0139] Example Seven Based on Example One, in Example Seven, binary rewriting indication information (referred to as the rewriting indication code element block) is added to the binary code in the original code unit according to certain rules to simplify the working complexity of the code class attribute extraction software, improve the pertinence of binary rewriting, and enhance the working effect.
[0140] This approach is very suitable for binary rewriting for positive purposes, such as MCU firmware anti-piracy and random scrambling of cryptographic operations.
[0141] The first method is to use a specific code element sequence (instruction sequence) that cannot appear in normal code, leaving a code element space that does not execute actual logical functions for placing additional code units.
[0142] For example, in the Cortex_M0 kernel, the following instruction series will not appear in practice: b.N .+2+6+16 ; Skip the subsequent abnormal instruction sequence b.N .+2 b.N .-2 ; Two instructions form an infinite loop Nop ; Alignment DC32 0xabcdee11 ; 16-byte UUID identification code DC32 0xabcdee22 DC32 0xabcdee33 DC32 0xabcdee44 DC32 StartAddressOfSramForRewriting / / Reserved SRAM variable space (for the stack of the external co-routine, etc.) DC32 RewritingTypeOfHere / / Various other data to indicate the operation of the original code rewriting execution unit The first short jump is used to skip the subsequent abnormal instruction sequence. The subsequent [b.N .+2; b.N.-2] instructions form an infinite loop, which cannot appear in normal code. However, such a code element sequence may appear in the constant area of the C language.
[0143] We can use another method. In the constant definition of the DC32 part, include a UUID identification code (Universally Unique Identifier, similar to the GUID in Windows). For example, for the string "Rewrite Indicator Code Element Block 2023", use the sha1 hash algorithm to generate a 128-bit hash value (31AA90D2F205D4E66346A66602C9374374C5DE24) as the UUID identification code. The probability of this hash value appearing in the actual original code is extremely low (practically can be considered as not occurring). The code class attribute extraction software can search for this UUID from the binary code and then locate the rewrite indicator code element block. Place a B.N unconditional jump instruction in front of the UUID identification code of the rewrite indicator code element block to the instruction after the rewrite indicator code element block, so that the rewrite indicator code element block does not affect the normal execution of the original code unit. The Universally Unique Identifier can be longer, such as 256 bits, 512 bits. Or carefully select appropriate shorter values such as 32 bits, 48 bits, 56 bits, 64 bits, 80 bits, etc. as the Universally Unique Identifier. The shorter the UUID, the higher the false detection probability, and the longer it is, the lower the false detection probability.
[0144] The rewrite indicator code element block can be embedded in a C language function in the form of inline assembly, so that it can be integrated with the binary code of the C function, remain within the short-distance jump range, and facilitate binary rewriting using short-distance jump instructions. It can also be defined directly in the C code in the form of a constant array.
[0145] In addition to reserving the code element space, the rewrite indicator code element block can also add additional indication information to indicate the reserved SRAM variable space (for the stack of external coroutines, etc.) and various other data to indicate the operation of the original code rewrite execution unit and meet the purposes that users need to achieve through binary rewriting.
[0146] For example, define a __root unsigned char CoroutineVal
[1024] to reserve 1024 bytes of variable space, and place the address and array size in the rewrite indicator code element block for binary rewriting.
[0147] It is also possible to define a function that is specifically for binary rewriting and does not contain actual logical functions (only contains the rewrite indicator code element block) void RewritingFunction(void). Then, any BL instruction that calls this RewritingFunction function indicates that the binary rewriting specified in the rewrite indicator code element block needs to be performed.
[0148] 68 __root void RewritingFunction(void) 69 { 70 asm(" b.N .+2+6+16"); \ RewritingFunction: (+1) \ 0x0 0xE00A b.N .+2+6+16 71 asm(" b.N .+2"); \ 0x2 0xE7FF b.N .+2 72 asm(" b.N .-2"); \ 0x4 0xE7FD b.N .-2 73 asm(" nop"); \ 0x6 0xBF00 nop 74 asm(" DC32 0xabcdee11"); ^ \ 0x8 0xABCD'EE11 DC32 0xabcdee11 75 asm(" DC32 0xabcdee22"); ^ \ 0xC 0xABCD'EE22 DC32 0xabcdee22 76 asm(" DC32 0xabcdee33"); \ 0x10 0xABCD'EE33 DC32 0xabcdee33 77 asm(" DC32 0xabcdee44"); \ 0x14 0xABCD'EE44 DC32 0xabcdee44 78} \ 0x18 0x4770 BX LR ;; return By formulating appropriate rules and embedding rewrite instruction code element blocks into the original code units, the code class attribute extraction software can be made much simpler, and the original code rewrite execution unit can also more flexibly and accurately implement a variety of functions.
[0149] The rewrite instruction code element blocks can be located by searching for UUIDs, and the original code rewrite information units can be directly extracted from each rewrite instruction code element block without performing code class attribute analysis on each code element one by one. Even the original code rewrite information units do not need to be stored separately, but are constructed by temporarily searching for UUIDs during binary rewriting.
[0150] The method of searching for UUID is also a kind of symbol analysis for binary instructions, and it also needs to be carried out symbol by symbol.
[0151] Taking the AES algorithm as an example, several types of empty functions are defined with the rewritten instruction symbol block. At the position of each round, the corresponding empty functions are called respectively to perform binary rewriting such as adding delay, adding random operations, adding external coroutines, and adding external fully masked coroutines (corresponding to Embodiment One, Two, Three, Four, and Five respectively), so as to achieve random scrambling and can well resist power analysis attacks.
[0152] Another example is a certain MCU binary code, which reserves 128 additional code unit spaces internally. 128 empty functions are defined, and one of the 128 empty functions is called at 512 places. According to the 96-bit UID unique serial number of the MCU chip, 128 different functions are filled into the 128 additional code units, and the chip unique serial number CHIP_UID detection, the integrity detection of the entire rewritten code unit, anti-tampering, etc. are carried out in different ways. The code obfuscation of one code per machine is realized, effectively preventing the MCU code from being copied to another MCU for running. For example, it is equivalent to burying 128 landmines in the MCU binary code. To copy the MCU code to another chip for running, it is necessary to remove these 128 landmines first, which consumes a lot of workload and time.
[0153] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A software method for binary software rewriting, characterized in that, Create an original code unit, an original code rewriting information unit, a variant number unit, an original code rewriting execution unit, an image code unit, and an additional code unit; The original code unit, the image code unit, and the additional code unit are all collections of binary instructions and data; The image code unit and the additional code unit constitute a rewritten code unit, and the rewritten code unit completes the same logical function as the original code unit; The original code rewriting information unit is generated by analyzing binary instructions in the original code unit and stores information such as which code elements can be rewritten and the allowed types of rewriting; The variant number unit is data that is different each time binary software rewriting is performed, or randomly generated data; The original code rewriting execution unit selects a part of the instructions in the original code unit and replaces them with a jump instruction sequence based on the original code rewriting information unit and the variant number unit to generate an image code unit; The jump instruction sequence jumps to the additional code unit. After the additional code unit completes the same logical function as the replaced instruction, it then jumps back to the position after the replaced instruction and continues execution.
2. The software method for binary software rewriting according to claim 1, characterized in that Also create a context save unit, an additional calculation unit, and a context restoration unit; in addition to completing the same logical function as the replaced instruction, the additional code unit also runs the additional calculation unit to execute logical functions not available in the original code unit. The context save unit is run before running the additional calculation unit, and the context restoration unit is run afterwards to prevent the existence of the additional calculation unit from affecting the logical function of the original code unit.
3. The software method for binary software rewriting according to claim 2, wherein The rewritten code unit serves as the main coroutine unit. A coroutine switching unit is also created in the additional calculation unit, and several external coroutine units and a coroutine management unit are additionally created; Each external coroutine unit also includes a context save unit, a coroutine switching unit, and a context restoration unit, and also includes an external stack unit and an external calculation unit. The external stack unit enables the operation of the external coroutine unit not to affect the main stack used by the main coroutine unit, and the external calculation unit completes the calculation tasks of the external coroutine unit; The coroutine management unit includes a coroutine information storage unit and a coroutine selection unit corresponding to the main coroutine unit and the external coroutine units; The coroutine switching units of the main coroutine unit and the external coroutine units both jump to the coroutine management unit. The coroutine management unit saves the information of the jumped coroutine unit to the corresponding coroutine information storage unit, and then the coroutine selection unit selects a coroutine information storage unit, loads the information of the coroutine unit, and jumps to the corresponding main coroutine unit or external coroutine unit.
4. The software method for binary software rewriting according to claim 3, characterized in that, The variant number unit is a random number generator, and the binary codes of the main coroutine unit and / or the external coroutine units are the same rewritten code unit for performing cryptographic operations; One of the coroutine units serves as a cryptographic operation coroutine unit to perform cryptographic operations on the real key and data and generate the required output result; The other coroutine units serve as random scrambling coroutine units to perform the same cryptographic operations on random input data and discard the generated output results; Based on the random number generated by the random number generator, the coroutine management unit randomly selects the next coroutine unit to run each time a coroutine switch occurs.
5. The software method for binary software rewriting according to claim 1, characterized in that, The deformation number unit generates a deformation number from the unique data of the processor or electronic device that runs the rewritten code unit.
6. An electronic device for performing binary software rewriting, characterized in that, It includes an original code unit, an original code rewriting information unit, a deformation number unit, an original code rewriting execution unit, an image code unit, and an additional code unit; The original code unit, the image code unit, and the additional code unit are all collections of binary instructions and data; The image code unit and the additional code unit constitute the rewritten code unit, and the rewritten code unit completes the same logical function as the original code unit; The original code rewriting information unit is generated by analyzing the binary instructions of the original code unit and stores information such as which code elements can be rewritten and the allowed rewriting types; The deformation number unit is data that is different each time binary software rewriting is performed, or randomly generated data; The original code rewriting execution unit selects a part of the instructions in the original binary code unit and replaces them with a jump instruction sequence according to the original code rewriting information unit and the deformation number unit, generating an image code unit; The jump instruction sequence jumps to the additional code unit. After the additional code unit completes the same logical function as the replaced instruction, it jumps back to the position after the replaced instruction and continues to execute.
7. An electronic device for performing binary software rewriting according to claim 6, characterized in that, It also includes a context save unit, an additional calculation unit, and a context restore unit. In addition to completing the same logical function as the replaced instruction, the additional code unit also runs the additional calculation unit to execute the logical function not available in the original code unit. The context save unit is run before running the additional calculation unit, and the context restore unit is run after that to prevent the existence of the additional calculation unit from affecting the logical function of the original code unit.
8. An electronic device for binary software rewriting according to claim 7, characterized in that, The rewritten code unit serves as the main coroutine unit. A coroutine switching unit is also created in the additional calculation unit, and it also includes several external coroutine units and a coroutine management unit; The external coroutine unit also includes a context save unit, a coroutine switching unit, and a context restore unit, and also includes an external stack unit and an external calculation unit. The external stack unit enables the operation of the external coroutine unit not to affect the main stack used by the main coroutine unit, and the external calculation unit completes the calculation tasks of the external coroutine unit; The coroutine management unit includes a coroutine information storage unit and a coroutine selection unit corresponding to the main coroutine unit and the external coroutine units; The coroutine switching units of the main coroutine unit and the external coroutine units both jump to the coroutine management unit. The coroutine management unit saves the information of the jumped coroutine unit to the corresponding coroutine information storage unit, and then the coroutine selection unit selects a coroutine information storage unit, loads the information of the coroutine unit, and jumps to the corresponding main coroutine unit or external coroutine unit.
9. An electronic device for performing binary software rewriting according to claim 8, characterized in that, The deformation number unit is a random number generator, and the binary codes of the main coroutine unit and / or the external coroutine units are the same rewritten code unit for cryptographic operations; One of the coroutine units serves as a cryptographic operation coroutine unit to perform cryptographic operations on the real key and data and generate the required output result; Other coroutine units perform the same cryptographic operations on random input data as the random scrambling coroutine unit and discard the generated output results; The coroutine management unit randomly selects the next coroutine unit to run each time a coroutine switch occurs, based on the random number generated by the random number generator.
10. An electronic device for binary software rewriting according to claim 6, characterized in that, The deformation number unit generates a deformation number from the unique data of the processor or electronic device that runs the rewritten code unit.
Citation Information
Cited By
Method for realizing confidential burning by using SRAM (Static Random Access Memory) power-on data
CN120687106A