Model training and task execution method and device, storage medium and equipment
By performing multi-level mask processing and model training on ARM machine code, the problem of large disassembly errors in the prior art is solved, and higher accuracy and security are achieved.
Patent Information
- Application Number
- CN202411996736.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-31
AI Technical Summary
The existing disassembly technology is difficult to accurately restore the assembly code corresponding to ARM machine code, resulting in large errors in program code analysis, affecting security and reliability.
The training samples are masked through various masking methods, and the model is trained based on the masked data. The disassembly model is used to mask the ARM machine code, determine the characteristic representation of the mask data, and predict the position and content of the mask data by decoding the model, and optimize the disassembly model to reduce errors.
It improves the disassembly accuracy of ARM machine code, reduces errors, and enhances the security and reliability of program code.
Smart Images

Figure CN119938062A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a method, apparatus, storage medium and device for model training and task execution. Background Art
[0002] In recent years, the Advanced RISC Machine (ARM) program architecture has gradually become the mainstream standard for mobile and embedded systems, playing an increasingly important role in the fields of artificial intelligence, the Internet of Things, and smart driving. In order to ensure the security and reliability of the business and provide more comprehensive protection for private data, it is usually necessary to disassemble the machine code of the ARM architecture (ARM machine code) to obtain assembly code that developers can understand. By analyzing the assembly code, possible vulnerabilities and security risks in the code can be discovered and handled in a timely manner.
[0003] However, due to the complexity of the ARM architecture itself, the richness of the ARM instruction set and the unique execution mechanism of the ARM architecture, existing disassembly technology is difficult to accurately restore the actual assembly code corresponding to the ARM machine code. The resulting assembly code has large errors, making it difficult for developers to accurately analyze the program code based on the assembly code obtained by existing methods.
[0004] Therefore, how to accurately disassemble ARM machine code to ensure the security and reliability of program code is an urgent problem to be solved. Summary of the invention
[0005] This specification provides a model training and task execution method, device, storage medium and equipment. The training samples are masked in a variety of masking methods and the model training is performed based on the masked data to fully improve the model performance.
[0006] This manual adopts the following technical solutions:
[0007] A model training method, comprising:
[0008] Acquire an advanced reduced instruction set ARM machine code, wherein the ARM machine code includes a plurality of instruction data, and the plurality of instruction data are executed according to a preset logical execution order;
[0009] Inputting the ARM machine code into a preset disassembly model, so as to perform mask processing on at least part of the instruction data contained in the ARM machine code through the disassembly model, so as to obtain first mask data according to the masked instruction data and the unmasked instruction data, and determine the feature representation corresponding to the first mask data as the first feature, and, for each instruction data, perform mask processing on at least part of the data units contained in the instruction data through the disassembly model, so as to obtain second mask data corresponding to the instruction data, and determine the feature representation of the second mask data corresponding to the instruction data as the second feature corresponding to the instruction data;
[0010] Inputting the first feature into a preset decoding model, so that the decoding model determines the position information corresponding to the masked instruction data in the ARM machine code according to the first feature as the predicted position information, and inputting the second feature corresponding to each instruction data into the decoding model, so that the decoding model predicts the masked data unit in each instruction data according to the second feature corresponding to each instruction data to obtain a predicted value;
[0011] Based on the deviation between the predicted value and the actual value corresponding to the masked data unit in each instruction data, and the deviation between the predicted position information and the actual position information corresponding to the masked instruction data in the ARM machine code, the target loss function is determined, and the disassembly model is at least trained with minimizing the target loss function as the optimization goal.
[0012] Optionally, determining the feature representation corresponding to the first mask data specifically includes:
[0013] A feature representation corresponding to the first mask data is determined through a first feature extraction network set in the disassembly model.
[0014] Optionally, determining a feature representation of the second mask data corresponding to the instruction data as a second feature corresponding to the instruction data specifically includes:
[0015] Through the second feature extraction network set in the disassembly model, the feature representation of the second mask data corresponding to the instruction data is used as the second feature corresponding to the instruction data.
[0016] Optionally, for each instruction data, mask processing is performed on at least part of the data units contained in the instruction data through the disassembly model to obtain second mask data corresponding to the instruction data, specifically including:
[0017] For each instruction data, mask processing is performed on at least part of the data units contained in the instruction data according to a preset mask ratio through the disassembly model to obtain second mask data corresponding to the instruction data.
[0018] Optionally, for each instruction data, masking is performed on at least part of the data units contained in the instruction data by using the disassembly model to obtain second mask data corresponding to the instruction data, specifically including:
[0019] For each instruction data, determine the mask strategy corresponding to the instruction data according to the data type corresponding to the instruction data through the disassembly model;
[0020] According to the mask strategy corresponding to the instruction data, determine the field type corresponding to the data unit to be masked in the instruction data as the target field type;
[0021] The data units of the instruction field in the instruction data that belong to the target field type are masked to obtain second mask data corresponding to the instruction data.
[0022] Optionally, for each instruction data, masking is performed on at least part of the data units contained in the instruction data by using the disassembly model to obtain second mask data corresponding to the instruction data, specifically including:
[0023] For each instruction data, mask processing is performed on at least part of the data units contained in the instruction data according to a preset mask ratio through the disassembly model to obtain third mask data corresponding to the instruction data, and, for each instruction data, a mask strategy corresponding to the instruction data is determined according to the data type corresponding to the instruction data; according to the mask strategy corresponding to the instruction data, a field type corresponding to the data unit to be masked in the instruction data is determined as a target field type; and data units of the instruction field in the instruction data belonging to the target field type are masked to obtain fourth mask data corresponding to the instruction data;
[0024] The second mask data corresponding to the instruction data is determined according to the third mask data and the fourth mask data corresponding to the instruction data.
[0025] Optionally, determining a feature representation corresponding to the first mask data as a first feature specifically includes:
[0026] For each instruction data included in the first mask data, determining position information corresponding to the instruction data in the first mask data, the position information including a position code corresponding to the instruction data in the first mask data and boundary information, the boundary information being used to indicate a starting character and an ending character of the instruction data in the first mask data;
[0027] The first feature is determined according to the position information corresponding to each instruction data.
[0028] Optionally, a feature representation of the second mask data corresponding to the instruction data is determined as a second feature corresponding to the instruction data:
[0029] For each data unit included in the second mask data corresponding to the instruction data, determine a one-hot encoding corresponding to the data unit and a position encoding corresponding to the data unit in the instruction data to which the data unit belongs;
[0030] Determine an initial feature corresponding to the data unit according to the one-hot encoding and the position encoding;
[0031] Determine a feature representation corresponding to the data unit according to a weight between the data unit and each other data unit included in the second mask data corresponding to the instruction data, and an initial feature corresponding to each data unit included in the second mask data corresponding to the instruction data;
[0032] According to the feature representation corresponding to each data unit in the instruction data, a second feature corresponding to the data unit is determined.
[0033] Optionally, the method further comprises:
[0034] Determine a target feature corresponding to the ARM machine code according to the first feature and the second feature;
[0035] According to the target feature, determining the data type corresponding to each instruction data as the predicted data type;
[0036] Determining a classification loss value according to a deviation between the predicted data type and an actual data type corresponding to each instruction data;
[0037] The disassembly model is retrained with minimizing the classification loss value as an optimization goal.
[0038] Optionally, the method further comprises:
[0039] Determine a target feature corresponding to the ARM machine code according to the first feature and the second feature;
[0040] Determine, according to the target feature, function boundary information corresponding to the ARM machine code as predicted function boundary information;
[0041] Determine a boundary loss value according to a deviation between the predicted function boundary information and the actual function boundary information corresponding to the ARM machine code;
[0042] The disassembly model is retrained with minimizing the boundary loss value as an optimization goal.
[0043] Optionally, the method further comprises:
[0044] Determine a target feature corresponding to the ARM machine code according to the first feature and the second feature;
[0045] According to the target feature, determining the data type corresponding to each instruction data as the predicted data type, and determining the function boundary information corresponding to the ARM machine code as the predicted function boundary information;
[0046] Determine a comprehensive loss value according to a deviation between the predicted function boundary information and the actual function boundary information corresponding to the ARM machine code, and a deviation between the predicted function boundary information and the actual function boundary information corresponding to the ARM machine code;
[0047] The disassembly model is retrained with minimizing the comprehensive loss value as an optimization goal.
[0048] This specification provides a task execution method, including:
[0049] Get the ARM machine code that needs to be disassembled as the target machine code;
[0050] Inputting the target machine code into a pre-trained disassembly model to determine a feature representation corresponding to the target machine code through the disassembly model, wherein the disassembly model is trained by the above-mentioned model training method;
[0051] According to the feature representation corresponding to the target machine code, the target machine code is disassembled to obtain the assembly code corresponding to the target machine code, and the task is executed according to the assembly code.
[0052] Optionally, determining the feature representation corresponding to the target machine code through the disassembly model specifically includes:
[0053] Determine a first feature corresponding to the target machine code through a first feature extraction network set in the disassembly model, and determine a second feature corresponding to the target machine code through a second feature extraction network set in the disassembly model;
[0054] A feature representation corresponding to the target machine code is determined according to the first feature and the second feature.
[0055] Optionally, according to the feature representation corresponding to the target machine code, disassembling the target machine code to obtain the assembly code corresponding to the target machine code specifically includes:
[0056] Determine, according to the characteristic representation corresponding to the target machine code, function boundary information corresponding to the target machine code and a data type corresponding to each instruction data contained in the target machine code;
[0057] According to the function boundary information corresponding to the target machine code and the data type corresponding to each instruction data contained in the target machine code, the target machine code is disassembled to obtain the assembly code.
[0058] Optionally, disassembling the target machine code according to the function boundary information corresponding to the target machine code and the data type corresponding to each instruction data contained in the target machine code to obtain the assembly code specifically includes:
[0059] For each instruction data, determining whether the data type corresponding to the instruction data is instruction data, wherein the instruction data includes: ARM instructions and Thumb instructions;
[0060] If so, the instruction data is disassembled.
[0061] Optionally, the method further comprises:
[0062] For each instruction data included in the target machine code, if the data type corresponding to the instruction data is inline data, the instruction data is not disassembled.
[0063] This specification provides a model training device, including:
[0064] An acquisition module is used to acquire an advanced reduced instruction set ARM machine code, wherein the ARM machine code includes a plurality of instruction data, and the plurality of instruction data are executed according to a preset logical execution order;
[0065] An input module, used for inputting the ARM machine code into a preset disassembly model, so as to perform mask processing on at least part of the instruction data contained in the ARM machine code through the disassembly model, so as to obtain first mask data according to the masked instruction data and the unmasked instruction data, and determine the feature representation corresponding to the first mask data as the first feature, and, for each instruction data, perform mask processing on at least part of the data units contained in the instruction data through the disassembly model, so as to obtain second mask data corresponding to the instruction data, and determine the feature representation of the second mask data corresponding to the instruction data as the second feature corresponding to the instruction data;
[0066] A prediction module, used for inputting the first feature into a preset decoding model, so that the decoding model determines the position information corresponding to the masked instruction data in the ARM machine code according to the first feature as the predicted position information, and inputting the second feature corresponding to each instruction data into the decoding model, so that the decoding model predicts the masked data unit in each instruction data according to the second feature corresponding to each instruction data, and obtains a predicted value;
[0067] A training module is used to determine a target loss function based on the deviation between the predicted value and the actual value corresponding to the masked data unit in each instruction data, and the deviation between the predicted position information and the actual position information corresponding to the masked instruction data in the ARM machine code, and to train at least the disassembly model with minimizing the target loss function as the optimization goal.
[0068] This specification provides a task execution device, including:
[0069] The acquisition module is used to obtain the ARM machine code that needs to be disassembled as the target machine code;
[0070] A determination module, used for inputting the target machine code into a pre-trained disassembly model to determine a feature representation corresponding to the target machine code through the disassembly model, wherein the disassembly model is trained by the above-mentioned model training method;
[0071] The execution module is used to disassemble the target machine code according to the feature representation corresponding to the target machine code, obtain the assembly code corresponding to the target machine code, and execute the task according to the assembly code.
[0072] This specification provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned model training and task execution method is implemented.
[0073] This specification provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned model training and task execution methods when executing the program.
[0074] At least one of the above technical solutions adopted in this specification can achieve the following beneficial effects:
[0075] In the model training method provided in the present specification, the server obtains an advanced reduced instruction set ARM machine code and inputs a disassembly model, performs mask processing on part of the instruction data contained in the ARM machine code to obtain first mask data, and determines a first feature corresponding to the first mask data, and performs mask processing on at least part of the data units contained in the instruction data to obtain second mask data, and determines a second feature of the second mask data; based on the first feature, the predicted position information is determined, and based on the second feature corresponding to each instruction data, the masked data unit in each instruction data is predicted to obtain a predicted value; the disassembly model is trained based on the deviation between the predicted value and the actual value corresponding to the masked data unit, and the deviation between the predicted position information and the actual position information corresponding to the masked instruction data.
[0076] It can be seen from the above method that in the process of model training, this scheme performs mask processing on the data units in each instruction data on the one hand, and on the other hand, performs mask processing on the entire instruction data, and predicts the location information of the masked data unit and the instruction data respectively according to the data features determined by the two mask processing methods, and then performs self-supervisory training on the disassembly model. Through this training method, the disassembly model can fully learn the semantic information between the internal unit structure of the instruction data and the logical relationship between different instructions, so as to have the ability to extract deeper semantic features, and based on the extracted semantic features, it can perform more accurate disassembly tasks on ARM machine code. BRIEF DESCRIPTION OF THE DRAWINGS
[0077] The drawings described herein are used to provide a further understanding of this specification and constitute a part of this specification. The illustrative embodiments and descriptions of this specification are used to explain this specification and do not constitute an improper limitation on this specification. In the drawings:
[0078] Figure 1 A flowchart of a model training method provided in this specification;
[0079] Figure 2 A schematic diagram of a program code of an ARM architecture provided in this specification;
[0080] Figure 3 A schematic diagram of a field structure of a machine code provided in this specification;
[0081] Figure 4 A schematic diagram of second mask data after field-level mask processing provided in this specification;
[0082] Figure 5 A schematic diagram of masked data provided in this specification;
[0083] Figure 6A schematic diagram of the overall process of pre-training KMLM provided in this specification;
[0084] Figure 7 A flowchart of a task execution method provided in this specification;
[0085] Figure 8 A schematic diagram of a model training device provided in this specification;
[0086] Fig. 9 A schematic diagram of a task execution device provided in this specification;
[0087] Fig.10 A method provided in this specification corresponding to Figure 1 or Figure 7 Schematic diagram of electronic equipment. DETAILED DESCRIPTION
[0088] In order to make the purpose, technical solutions and advantages of this specification more clear, the technical solutions of this specification will be clearly and completely described below in combination with the specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of this specification, not all of them. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this specification.
[0089] There are two main traditional disassembly strategies: Linear Sweep and Recursive Traversal. Among them, Linear Sweep processes all bytes in the code segment linearly. However, since it does not understand control flow information, linear scanning is prone to errors when processing inline code and instruction set switching, which is common in ARM binary files.
[0090] Recursive Traversal starts disassembling from the program's entry point and follows the branches. Recursive traversal performs better when processing inline data, but it may miss indirect branches determined at runtime and may suffer from low code coverage.
[0091] The technical solutions provided by the embodiments of this specification are described in detail below in conjunction with the accompanying drawings.
[0092] Figure 1 A flow chart of a model training method provided in this specification includes the following steps:
[0093] S100: Acquire an advanced reduced instruction set ARM machine code, wherein the ARM machine code includes a plurality of instruction data, and the plurality of instruction data are executed according to a preset logical execution order.
[0094] Since its birth, the ARM architecture has gone through many generations of development, from ARMv4 to the latest ARMv9, and the instruction set functions have been continuously expanded. Each CPU architecture version introduces new features and instructions. Developers can select the target CPU architecture of their program through the compilation option (-march).
[0095] The instruction set design of the ARM architecture has evolved several times since its inception. The original ARMv3 introduced the 32-bit Arm instruction set (A32) in the early 1990s. To improve code density, ARMv4T introduced the 16-bit Thumb (T16) instruction set as a subset of A32 in the late 1990s. Thumb was gradually extended to 32 bits with the introduction of ARMv6T2, forming the T32 instruction set. More recently, ARMv8 introduced the 64-bit Arm instruction set (A64) for the new AArch64 execution state in 2011, while maintaining backward compatibility with A32 and T32 to support 32-bit code.
[0096] Among them, the ARM instruction set is known for its fixed-length instruction encoding and rich general-purpose registers. This design enables the ARM processor to perform data processing and calculations efficiently, but it also brings unique challenges to disassembly. The ARM architecture also includes multiple modes and exception handling mechanisms, which need to be correctly identified and handled during disassembly. For example, the ARM processor supports multiple operating modes, including user mode, system mode, FIQ (fast interrupt) mode, etc., each of which has its specific registers and instruction sets. ARM's branch and jump instructions play an important role in the program execution process, and they can change the execution flow of the program. However, these instructions also make disassembly more complicated because the disassembler needs to accurately identify these instructions and track the execution path of the program.
[0097] Currently, the ARM architecture has become the mainstream standard for mobile and embedded systems, and these platforms include smartphones, tablets, IoT devices, robotic vehicles, etc. However, the firmware software that supports these platforms often remains in a closed source state. Due to the closed nature of the firmware software, researchers cannot directly access the source code, which makes it more difficult to understand the internal logic and structure of the binary file. Therefore, it is particularly important to disassemble the ARM machine code to obtain the assembly code that developers can understand. Among them, disassembly mainly refers to the process of recovering lost high-level program constructs from binary code, such as assembly instructions, function boundaries, function signatures, and control flow graphs.
[0098] Due to the characteristics of the ARM architecture, the disassembly of ARM machine code faces many difficulties, such as:
[0099] Runtime instruction set switching: The 32-bit ARM architecture provides two instruction sets, ARM and Thumb / Thumb-2. These two instruction sets are usually intertwined in real binary files. Instruction set switching can occur explicitly or implicitly at runtime, which poses a huge challenge to static analysis-based disassemblers.
[0100] Inline data: In AArch32 / AArch64, inline data is more common than in x86 / x64. Inline data refers to data that is directly embedded in the code segment, rather than stored in a separate data segment. Due to the existence of inline data, a static analysis-based disassembler may mistakenly identify a piece of inline data as a valid instruction set, resulting in misleading disassembly of subsequent machine code.
[0101] Implicit function call: Unlike the call instruction in x86 / x64, the Thumb instruction set lacks explicit function call instructions. For example, the BL (Branch with Link) instruction can represent both a direct jump and a function call. This implicit function call mechanism increases the difficulty of disassembly because the disassembler needs to determine the specific meaning of the BL instruction based on the context.
[0102] Based on this, this specification provides a model training method that embeds ARM domain knowledge into model pre-training during the training of the disassembly model. The model training process is achieved by designing a multi-level mask strategy and adding another pre-training target (span boundary object, SBO) for ARM characteristics.
[0103] ARM binary files (ARM machine code) are randomly masked by multi-level masks, including unit level, field level and span level. Then, the knowledge is unsupervisedly pre-trained on the masked data to obtain a knowledge-aware masked language model (KMLM). KMLM is composed of the basic mask model (MLM basic ) and span mask model (MLM span ) are jointly optimized in the pre-training stage. basic Able to understand the semantic relationship between the fields within the instruction, while MLM span Aims to capture the logical relationship between consecutive instructions.
[0104] In this specification, the execution subject used to implement a model training and task execution method can be a designated device such as a server. For the sake of ease of description, the following will only take the server as the execution subject as an example to illustrate a model training and task execution method provided in this specification.
[0105] During the training of the disassembly model, the server can obtain the advanced reduced instruction set ARM machine code, which can be an instruction data sequence containing several binary instruction data. In actual applications, each instruction data in the instruction data sequence is executed according to a preset logical execution order.
[0106] For ease of understanding, this specification provides a schematic diagram of a program code of an ARM architecture, such as Figure 2 shown.
[0107] Figure 2 This is a schematic diagram of a program code of an ARM architecture provided in this specification.
[0108] In practical applications, Figure 2 Each instruction data in is executed in sequence according to its corresponding execution order (as shown in the first column), and the second column indicates the data type corresponding to the instruction data, where $a, $t, and $d are mapping symbols corresponding to each data type, $a indicates ARM instruction, $t indicates Thumb instruction, and $d indicates inline data. Inline data refers to data values directly embedded in the ARM program code, rather than accessing the data by referencing the memory address of the data. In the disassembly process, there is no need to disassemble the inline data.
[0109] The third column in the figure represents the storage address, and the fourth column represents the ARM machine code, which is obtained by compiling the ARM assembly code. For each ARM instruction data, the instruction data is encoded by a 32-bit fixed length, in which each character corresponds to a 4-bit data unit.
[0110] by Figure 2 Take the ARM instruction "e08f3003" on line 10 as an example. The prefix "0x" indicates that its value is in hexadecimal form. The ARM has only 32 bits in total, each of which corresponds to a four-bit binary code, representing a data unit.
[0111] The fifth column in the figure shows the assembly code before compilation. Taking "e08f3003" as an example, its corresponding assembly code is "add, r3, pc, r3", which means adding the value of the program counter (pc) and register r3, and storing the result back in r3.
[0112] It should be noted that all the above data (including data types, machine codes, and assembly codes) can be obtained during the model training process, but only ARM machine codes can be obtained during the actual execution of the disassembly task.
[0113] Furthermore, for machine codes, different data units correspond to different instruction fields, and the instruction fields in each instruction data may include:
[0114] 32-bit ARM instructions:
[0115] (Cond, Special, [Opcode], [Oprnd dest ],[Oprnd src1 ],[Oprnd src2 ]);
[0116] (Cond, Special, Reg dest , Reg src1 , Reg src2 );
[0117] (Cond, Special, Reg dest , Reg src , Imm);
[0118] (Cond, Special, Reg src1 );
[0119] (Cond, Special, Imm).
[0120] 16-bit Thumb instructions:
[0121] (Special, [Cond], [Opcode], [Oprnd src1 ],[Oprnd src2 ],[Oprnd dest ]).
[0122] 32-bit Thumb instructions:
[0123] (Special1, [Opcode], [Oprnd src1 ]), (Special2, [Oprnd src2 ],[Oprnd dest ]).
[0124] The Cond field is a conditional code that specifies the execution condition of the instruction. For example, 'eq' stands for "equal" (execute when equal), 'ne' stands for "not equal" (execute when not equal), etc. If omitted, the instruction is considered to be executed unconditionally.
[0125] The Special field refers to some special instruction flags or prefixes, which are used to mark some non-standard or special-purpose instructions, such as the specific prefix of the coprocessor instruction.
[0126] The Opcode field is the operation code: This is the core part of the instruction, defining the basic operation to be performed. The operation code determines the specific type of operation that the processor will perform. dest Indicates the destination operand, that is, the register or memory location where the result is stored after the instruction is executed. src Represents the source operand, that is, the source of data provided to the instruction for operation.
[0127] [Oprnd_dest] (destination operand): This is the register or memory location where the result of the instruction is stored after execution. In ARM assembly, this is usually a register (such as R0-R15), indicating where the result of the calculation is stored.
[0128] Reg dest , Reg src1 , Reg src2 Field is a more specific description of the operand, where Reg dest Specifies the destination register, and Reg src1 and Reg src2 The first and second source registers are specified respectively. These are the registers that are directly operated on when the instruction is executed.
[0129] The Imm field is an immediate value. An immediate value refers to a numerical value directly encoded in an instruction. It is directly embedded in a machine instruction as part of an operation.
[0130] For ARM instructions, each instruction starts with a 4-bit condition field {Cond}, which indicates which condition will be checked. For example, 0000 (mnemonic EQ) means that the zero flag must be set. The subsequent 5-8-bit {Opcode} field encodes the operation category, including data processing, branching, and other operations. The rest of the instruction is {Operand}, which specifies the operands of the instruction, including a flexible combination of immediate values, registers, and addressing modes.
[0131] The Thumb instruction set (T16) uses a more compact 16-bit encoding. It omits the conditional execution field and reduces {Opcode} to 6 bits, leaving 10 bits for {Operand}, which limits the available operations and addressing modes. To expand functionality, Thumb-2 (T32) supplements the 16-bit Thumb encoding by adding a new 32-bit format while retaining the 16-bit Thumb to support more complex behaviors. For ease of understanding, this specification provides a field structure diagram of a machine code, such as Figure 3 shown.
[0132] Figure 3 A schematic diagram of a field structure of a machine code provided in this specification.
[0133] Here also Figure 2 Take the ARM instruction data "e08f3003" in as an example, where the data unit "e" corresponds to the four-bit instruction field {Cond}, encoded as "1110", indicating unconditional execution, the data unit "08" corresponds to the eight-bit instruction field {Opcode}, encoded as "00000100", and the data units "f" and "3" are register operands, where {Rn}=r3, {Rd}=r15 (indicating pc), and {Rm}=r3.
[0134] In addition, Figure 2 Take the machine code "ea030a97" in line 8 as an example. The condition code of the first 4 bits of the {Cond} field is 1110, indicating that the instruction is executed unconditionally. The encoding of the 5-8 bits of the {Special} field is 1010, including the reserved bit (101) and the function bit L is 0, indicating that this is a branch (b) instruction without saving the return address. The remaining 24 bits are the immediate value imm32, which is used to calculate the offset of the current program counter (pc).
[0135] S102: Input the ARM machine code into a preset disassembly model to perform mask processing on at least part of the instruction data contained in the ARM machine code through the disassembly model, so as to obtain first mask data based on the masked instruction data and the unmasked instruction data, and determine the feature representation corresponding to the first mask data as the first feature, and, for each instruction data, perform mask processing on at least part of the data units contained in the instruction data through the disassembly model to obtain second mask data corresponding to the instruction data, and determine the feature representation of the second mask data corresponding to the instruction data as the second feature corresponding to the instruction data.
[0136] In this specification, a knowledge-aware masked language model (KMLM) may be provided in the disassembly model. The knowledge-aware masked language model is composed of two parts, including a first feature extraction network and a second feature extraction network. The first feature extraction network may be composed of an MLM. span (Span Mask Language Model), used to capture the logical relationship between continuous instruction data, the second feature extraction network can be composed of an MLM basic (Basic Mask Language Model) is used to understand the relationship between fields within instruction data.
[0137] During the model training process, a mask network may be provided in the disassembly model, and the mask network is used to perform mask processing on the ARM machine code. Since mask processing on the target machine code is not required in practical applications, the mask network may be deleted after completing the pre-training of the KMLM.
[0138] Specifically, the mask network in the disassembly model can mask at least part of the instruction data contained in the ARM machine code to determine the first mask data based on the masked instruction data and the unmasked instruction data. For the sake of ease of description, the masking method is uniformly described as "cross-instruction masking" below.
[0139] In binary analysis tasks, it is often necessary to reason about the relationship between instructions. For example, in the Instruction Set Architecture (ISA) and inline data identification tasks, the ldreq r0, [r5] instruction appears before the bx r0 instruction, and this order can inform the mapping symbol type $a of the ldreq instruction. For self-supervised learning, modeling such dependencies under cell-level masks is more challenging because the original masked language model (MLM) only captures the relationship at the cell level.
[0140] Specifically, during the cross-instruction masking process, each 32-bit instruction data is randomly masked. Instead of directly masking 32 bits into 8 [MASK] and expecting the model to infer their values, the server introduces a custom Span Boundary Object (SBO) to explicitly train the pre-trained model to predict the complete instruction span given the context instruction. SBO describes the boundaries of the instruction, forcing the model to learn the overall multi-unit pattern. This makes the modeling of the relationship between instructions richer.
[0141] In the SBO task, the model needs to predict the boundaries of the masked instruction span rather than the specific tokens (or instruction data) within the span. This helps the model understand the relationship between instructions at a higher level, so as to better perform tasks that require cross-instruction reasoning. SBO depicts instruction boundaries, forcing the model to learn holistic multi-unit patterns. This enables richer modeling of dependencies between instructions compared to individual unit predictions.
[0142] In the AArch32 architecture, a 32-bit binary sequence can represent various elements, including an ARM instruction, two 16-bit Thumb instructions, a 32-bit Thumb-2 instruction (for ARMv6 and later ARM binaries), or an inline data. All cells in the cross-instruction mask are replaced with [MASK].
[0143] For an ARM machine code including several instruction data, randomly selected 32-bit (8 data units) aligned instruction data are masked to obtain the first mask data. The process can be expressed as:
[0144]
[0145] Where i∈N, index i is the input instruction data sequence x m_span Mid span (x {4i} , ..., x _{4i+7} ) location. During this process, the instruction data can be masked at a rate of 15%, but this replacement is done at the instruction span level, not at the per-unit level.
[0146] With MLM basic Different from focusing on a single data unit, MLM span The focus is on the instruction data sequence within the span, that is, a continuous instruction block (x s , ..., x e ), these blocks are processed by the overall mask. The model is based on the external boundary cell x s-1 and x e+1 The span embedding stored at is used to predict the contents of the masked instruction span. Here, the span embedding contains contextual information spanning multiple instructions, enabling the model to learn and infer high-level semantics spanning multiple instructions, such as data transfers between registers or control flow conversions, even in the presence of instruction set switching. MLM span is designed to capture and exploit implicit relationships across instructions, which is necessary to understand complex code sequences.
[0147] During this process, the first feature extraction network can determine the position information corresponding to each instruction data contained in the first mask data, the position information including the position code corresponding to the instruction data in the first mask data and boundary information (the boundary information is used to indicate the starting character and the ending character of the instruction data in the first mask data), and then determine the first feature according to the position information corresponding to each instruction data.
[0148] Specifically, when a masked instruction data sequence MASK(x s , ..., x e ), where (s, e) represents the start and end positions of the sequence, and the first feature extraction network can use the output encoding of the outer boundary units of the span, as well as the target span P i-s+1 The position encoding is used to represent each instruction data x within the span i The server can determine the feature representation y corresponding to the position information through a two-layer feed-forward network with GeLU activation function and layer normalization (j) , as the first feature. The first feature can be expressed as:
[0149] y (j) =f(x s-1 , x e+1 , P i-s+1 )
[0150] Among them, the position codes P1, P2, ... mark the position relative to the left boundary x s-1 The relative position of the masked instruction data. In this specification, the first feature extraction network can first construct a vector h0 containing external boundary information and position encoding, where:
[0151] h0=[x s-1 ;x e+1 ;P i-s+1 ]
[0152] Next, we apply layer normalization and GeLU activation function to update h0 to get h1, where:
[0153] h1=LayerNorm(GeLU(w1h0)
[0154] Finally, layer normalization and GeLU activation function are applied again to get y (j) ,in:
[0155] y (j) =LayerNorm(GeLU(w2h1))
[0156] At the same time, for each instruction data contained in the ARM machine code, the mask network can perform mask processing on at least part of the data units contained in the instruction data, thereby determining the second mask data corresponding to the instruction data based on the masked data units and unmasked data units in the instruction data.
[0157] Specifically, the mask network can perform unit-level masking on each instruction data, and regard each instruction data as a unit sequence of size n, represented as x=(x1, ..., x n ), for any x i (i∈N), which represents a 4-bit data unit. Each data unit x i ∈x is mapped to a 16-bit one-hot encoding. For example, an input unit b is encoded as "0000100000000000". In other words, the server can define a total of 18 encodings as vocabularies, which include [MASK] and [DATA] representing masked instruction sets (including ARM instructions and Thumb instructions) and masked inline data, in addition to representing hexadecimal numbers through one-hot encoding of 0-15 bits.
[0158] In the process of masking the instruction data, the mask network can randomly mask at least part of the data units in the instruction data according to a preset mask ratio, so as to obtain the second mask data corresponding to the instruction data. For the convenience of description, the masking method is collectively referred to as "unit-level masking" below. The masking process can be expressed as:
[0159] D = {x m_unit |(x i , …, MASK unit(xi) , …, x n ), i∈P unit}
[0160] Among them, x m_unit Indicates the second mask data, MASK(x i ) function indicates masking the i-th data unit in the given instruction data. unit is the set of masked cell positions in x. Given an ARM machine code containing several instruction data, where |D|=N, the second feature extraction network is trained using the second mask data in D.
[0161] In this specification, the preset mask ratio may be set to 15%, that is, in the ARM machine code, a total of 15% of the data units are masked.
[0162] Among the masked cells, 80% of the cells can be replaced with "[MASK]", 10% of the cells can be replaced with random cells, and the remaining 10% can remain as they are. Of course, the server can also replace all masked instruction cells with "[MASK]" and all masked inline data with "[DATA]".
[0163] In addition, the mask network can also use another mask strategy to mask the data units in each instruction data, that is, for each instruction data, through the mask network, according to the data type corresponding to the instruction data, determine the mask strategy corresponding to the instruction data, and according to the mask strategy corresponding to the instruction data, determine the field type corresponding to the data unit to be masked in the instruction data as the target field type, and mask the data unit of the instruction field in the instruction data belonging to the target field type to obtain the second mask data corresponding to the instruction data. For the convenience of description, the masking method is collectively referred to as "field-level masking" below.
[0164] The data types corresponding to the instruction data include: ARM instructions and Thumb instructions, and the field types may include: {Cond} field, {Opcode} field, {Operand} field, {Special} field, etc.
[0165] Specifically, for ARM instructions, the corresponding mask strategies may include:
[0166] For each 32-bit instruction, the data unit corresponding to the 4-bit condition code field {Cond} is always masked;
[0167] For data processing instructions, the data unit corresponding to the mask {Opcode} field;
[0168] For instruction data containing three register operands, the data unit corresponding to each 4-bit {Operand} field can be randomly masked;
[0169] For instruction data with two register operands and one immediate operand, the data units corresponding to the two register fields can be masked, but the data units corresponding to the immediate field are not masked;
[0170] For a branch instruction with an immediate value as the target, the data units corresponding to the condition code {Cond} field and the {Special} field may be masked, while the offset remains unchanged.
[0171] Since Thumb (T16) instruction encoding is more compact than A32. For example, most Opcode fields are 2 bits (4 bits in A32), and all register fields are encoded as 3 bits. Unlike A32 where the condition code Cond is the beginning of each instruction, only conditional branch instructions contain {Cond}. In order to reuse the 4-bit mask [MASK], the server can add padding to the field for alignment.
[0172] Therefore, for Thumb instructions, the corresponding mask strategies may include:
[0173] All data units corresponding to the {Special} field can be masked;
[0174] For register {Oprndsrc1}, the server can add a padding number to its least significant bit for masking;
[0175] The data unit corresponding to the immediate value (i.e., the offset of the operand displacement or as a branch target) is never masked;
[0176] The data unit corresponding to the {Opcode} field in the mask conditional branch;
[0177] For push and pop operations, the register list is masked as {Opndsrc1}.
[0178] It should be noted that the field types corresponding to all instruction fields that can be masked in the above mask strategy can be considered as target field types.
[0179] For ease of understanding, this specification provides a schematic diagram of the second mask data after field-level mask processing, such as Figure 4 shown.
[0180] Figure 4 This is a schematic diagram of second mask data after field-level mask processing provided in this specification.
[0181] Among them, for the instruction data "e08f3003", the instructions contained therein, add, r3, pc, r3, have five instruction fields: {cond} (condition code), {op1} (usually part of Opcode or a special opcode identifier), {Rn} (the second register operand), {Rm} (the second register operand, which may be an immediate value for some instructions) and {Rd} (the destination register). According to the above masking strategy, the data unit corresponding to its condition code {cond} is masked, and the register operand is randomly selected for masking.
[0182] For the instruction data "ea030a97", the instruction is a branch instruction, in which the data unit corresponding to the {cond} field is always masked, and the data unit corresponding to the {special} field is randomly masked.
[0183] In this specification, the mask network may perform unit-level masking on the ARM machine code only to obtain the second mask data, or may perform field-level masking on the ARM machine code only to obtain the second mask data.
[0184] Of course, the disassembly model can perform unit-level masking and field-level masking on the ARM machine code at the same time, so as to obtain the second mask data. Among them, for each instruction data, through the second mask network, according to the preset mask ratio, at least part of the data units contained in the instruction data are masked to obtain the third mask data corresponding to the instruction data, and, for each instruction data, according to the data type corresponding to the instruction data, the mask strategy corresponding to the instruction data is determined; according to the mask strategy corresponding to the instruction data, the field type corresponding to the data unit to be masked in the instruction data is determined as the target field type; the data unit of the instruction field in the instruction data belonging to the target field type is masked to obtain the fourth mask data corresponding to the instruction data, and then the third mask data and the fourth mask data are merged to obtain the second mask data.
[0185] In addition, the mask network can first mask the ARM machine code through unit-level masking to obtain unit-level mask data, and then further mask the unit-level mask data through field-level masking to obtain second mask data, or first mask the ARM machine code through field-level masking to obtain field-level mask data, and then further mask the field-level mask data through unit-level masking to obtain second mask data.
[0186] It should be pointed out that in the process of masking by unit-level masking and field-level masking, some instruction fields will contain two or more data units. Therefore, each data unit in each instruction data can be masked by one or more [MASK]s, that is, the data unit is masked once by unit-level masking and once by field-level masking.
[0187] After unit-level masking and field-level masking, the internal structure and grammatical knowledge of each instruction are integrated into the pre-trained data.
[0188] In this way, the model can learn facts and knowledge about the ARM instruction set architecture (ISA) and real binary files during the pre-training phase, including the meaning of the various fields of the instruction, the relationship between them, and how they are combined into valid instructions. This helps the model better understand, analyze, and process ARM machine code in subsequent tasks.
[0189] For ease of understanding, the manual provides a schematic diagram of masked data, such as Figure 5 shown.
[0190] Figure 5 This is a schematic diagram of masked data provided in this specification.
[0191] Each light-colored square represents a 4-bit data unit, while the dark-colored square is the masked data unit. The first row is the first masked data after cross-instruction-level masking, and the second to fourth rows are the second masked data after field-level masking.
[0192] Four code snippets are used here to demonstrate the masking effect:
[0193] $a: ldreq r0, [r5] (0x00009505): This is an ARM instruction that means "if r5 is equal to 0, load r0 with the contents of the memory address pointed to by r5". Here, ARMED may selectively mask certain fields in the instruction, such as the operation code (opcode) or register identifiers.
[0194] $a:bx r0(0x10ff2fe1): This is also an ARM instruction, meaning "jump to the address pointed to by the r0 register". Again, ARMED can selectively mask certain parts of the instruction to train the model.
[0195] $d: 0x1b6898f0: This is an inline data value, not an instruction. In this case, ARMED may selectively mask certain bits or bytes in the data so that the model can learn the structure and pattern of the data.
[0196] $t: push{r3, lr} (0x0865) and $t: movs{r0, #8} (0x0820): These two refer to Thumb instructions (a subset of the ARM instruction set, used for 16-bit encoded instructions), which respectively mean "push the contents of the r3 and lr registers onto the stack" and "move the immediate value 8 (which may be a signed number) to the r0 register". Again, ARMED selectively masks certain parts of these instructions to train the model.
[0197] The server may further determine a feature representation corresponding to the second mask data as the second feature through a second feature extraction network.
[0198] In this process, for each data unit contained in the second mask data, the second feature extraction network can determine the one-hot encoding corresponding to the data unit and the position encoding corresponding to the data unit in the instruction data to which the data unit belongs; determine the initial feature corresponding to the data unit based on the one-hot encoding and the position encoding; perform weighted summation based on the initial features corresponding to each data unit in the second mask data to which the data unit belongs, and the weight between each data unit in the second mask data to which the data unit belongs and the data unit, to determine the feature representation corresponding to the data unit, and then determine the second feature based on the feature representation corresponding to each data unit.
[0199] Among them, the position code is used to represent the position information of the data unit in the instruction data to which it belongs, and the weight between each data unit and the data unit is used to represent the correlation between other data units and the data unit. The greater the correlation, the greater the weight, and vice versa.
[0200] Specifically, each input data unit is converted into a one-hot vector and combined with its position embedding to generate the initial features. This means that each unit is represented as a 16-bit vector, in which only one element is 1 and the rest are 0 to represent a specific 4-bit binary value or special mark (such as [MASK] or [DATA]). Position encoding is to provide the position of each unit in the instruction data in the sequence.
[0201] The initial features take into account the information of other units in the binary sequence to form a context-aware embedding. This step helps the model understand the dependencies between units.
[0202] Next, the Transformer unit in the second feature extraction network dynamically updates the initial features from the second layer to the last layer through a multi-head self-attention mechanism. The self-attention mechanism allows the model to capture correlations between embeddings at different positions, enhances the model's understanding of long-distance dependencies in the sequence, and outputs a deeper second feature by the last layer.
[0203] In this specification, the position encoding is learned, and we then stack a feed-forward network f(x) with one hidden layer to incorporate the position embedding E pos and one-hot encoding E unit .
[0204] The masking process can be expressed as:
[0205] E 1,i =MLP(concat(Eunit(xi) , E pos(i) ))
[0206] Among them, E1, i Refers to the embedded representation of xi in the second layer of attention mechanism, E unit(xi) represents the embedding applied to the one-hot encoded unit xi, and E pos(i) Represents applying the learned positional encoding to x i The corresponding position i.
[0207] It should be noted that the disassembly model in this specification may also only have one feature extraction network, and through this feature extraction network, the first mask data is obtained and the first feature corresponding to the first mask data is determined, and the second mask data is obtained and the second feature corresponding to the second mask data is determined.
[0208] In addition, unit-level mask, field-level mask and cross-instruction mask can be masked using different mask networks or the same mask network. Of course, when training the disassembly model, it is also possible not to add a mask network to the disassembly model, but to pre-process the ARM machine code for mask operations through a preset script or an external mask model, and then input the pre-processed mask data into the disassembly model for feature extraction and prediction.
[0209] S106: Input the first feature into a preset decoding model so that the decoding model determines the position information corresponding to the masked instruction data in the ARM machine code according to the first feature as predicted position information, and input the second feature corresponding to each instruction data into the decoding model so that the decoding model predicts the masked data unit in each instruction data according to the second feature corresponding to each instruction data to obtain a predicted value.
[0210] S108: Determine a target loss function based on the deviation between the predicted value and the actual value corresponding to the masked data unit in each instruction data, and the deviation between the predicted position information and the actual position information corresponding to the masked instruction data in the ARM machine code, and train at least the disassembly model with minimizing the target loss function as the optimization goal.
[0211] During the model training process, an additional decoding model can be set. The decoding model can be set inside the disassembly model or inside the disassembly model. When set inside the disassembly model, the decoding model can be a Softmax network, whose input dimension is the same as the output dimension.
[0212] The first feature extracted by the first feature extraction network can be input into the decoding model so that the decoding model can be based on the first feature y (j) Predict the position information of the masked instruction data in the instruction data sequence included in the entire ARM machine code as the predicted position information.
[0213] Of course, the first feature extraction network can also be obtained through the first feature y (j) To predict the masked data unit, the predicted value corresponding to the masked data unit can be expressed as:
[0214]
[0215] in, is the predicted value of each masked data unit in the span, MLM span Predict the masked data units in each span, which are located in the position set P span middle.
[0216] The second feature output by the second feature extraction model can also be input into the above decoding model, so that the decoding model predicts the original value of the data unit masked by the [MASK] mark according to the second feature as the predicted value. For each masked data unit, the predicted value corresponding to the data unit can be expressed as:
[0217]
[0218] in, is the masked data unit x i The corresponding predicted value.
[0219] It should be noted that since the masked data units do not need to be restored during the disassembly process, the above decoding model may not be used in the disassembly process after the training is completed.
[0220] Furthermore, the server can determine the loss function based on the deviation between the predicted value corresponding to the masked data unit and the actual value corresponding to the masked data unit, and the deviation between the predicted position information corresponding to the masked instruction data and the actual position information corresponding to the masked instruction data, and then train the disassembly model with minimizing the loss function as the optimization goal.
[0221] Specifically, the server may train the first feature extraction network and the second feature extraction network respectively.
[0222] For the first feature extraction network, the corresponding loss function can be expressed as:
[0223]
[0224] Where P(y i |y i ) is the feature representation y based on the span boundary i Under this condition, the prediction result is the actual position information y i Based on the loss function, the server can minimize the deviation between the predicted position information corresponding to the masked instruction data and the actual position information corresponding to the masked instruction data.
[0225] For the second feature extraction network, its loss function can be expressed as:
[0226]
[0227] in, Denotes the loss function corresponding to the first mask network, D unit Indicates the ARM machine code containing several instruction data, P unit is the set of masked cell positions in x, is the masked data unit x i The server may minimize the deviation between the predicted value corresponding to the masked data unit and the actual value corresponding to the masked data unit based on the above loss function.
[0228] Of course, in order to optimize the first feature extraction network and the second feature extraction network at the same time, the server can also add the above two losses to jointly train the above two models. Specifically, for any data unit x i , which uses the same input embedding in the first feature extraction network and the second feature extraction network. The objective loss function can be expressed as:
[0229] L pre(xi) =L MLMbasic(xi) +L MLMspan(xi)
[0230] Among them, L pre(xi) represents the target loss of the knowledge-aware masked language model KMLM, which is optimized by gradient descent. The two parts here are: L MLMbasic(xi) Represents data unit x i Its own prediction loss L MLMspan(xi) Represents the loss predicted based on location information. This joint optimization strategy ensures that the model can not only understand the context of a single instruction field, but also capture high-level semantics across multiple instructions, thereby achieving better performance in ARM binary disassembly tasks. For ease of understanding, this manual provides a schematic diagram of the overall process of pre-training KMLM, as shown in Figure 6 shown.
[0231] Figure 6 The figure is a schematic diagram of the overall process of pre-training KMLM provided in this specification.
[0232] After the server inputs the target machine code into the disassembly model, KMLM can pre-process the target machine code, that is, perform unit-level masking, field-level masking and cross-instruction masking respectively, and then input the second mask data after the unit-level masking and field-level masking into the second feature extraction network (L MLMbasic ), the first mask data after the cross-instruction mask processing is input into the first feature extraction network (L MLMspan ), and then pre-train the KMLM based on the first feature output by the first feature extraction network and the second feature output by the second feature extraction network.
[0233] Furthermore, after pre-training the KMLM, the server can deploy it to the disassembly model and fine-tune the disassembly model.
[0234] In the fine-tuning stage, for the ARM binary disassembly task, the focus is on two basic tasks: instruction identification and function boundary recovery. These two tasks are the basis of disassembly. Other more complex operations such as the generation of control flow graph (CFG) and call graph rely on accurate instruction and function recovery results. Therefore, the ARMED framework designs two fine-tuning tasks to assist the disassembly of ARM binaries: Mapping Symbol Classification and Function Boundary Detection.
[0235] For the mapping conformance classification task, the server can determine the target feature corresponding to the ARM machine code based on the first feature and the second feature, and then determine the data type corresponding to each instruction data based on the target feature as the predicted data type, and then determine the classification loss value based on the deviation between the predicted data type and the actual data type corresponding to each instruction data, and retrain the disassembly model with minimizing the classification loss value as the optimization goal.
[0236] Specifically, the server can first use the pre-trained KMLM and input a piece of ARM machine code. KMLM will generate a series of embedding vectors E l (j) =(E l,1 , E l,2 , ..., E l,n ), where each E l,i A high-level representation of a portion of the original code.
[0237] For each fine-tuning task, there is a corresponding ground truth label sequence y, whose length is consistent with the embedding vector sequence, that is, y = {y (j) |y (j) ∈C y}. C y Represents the set of all possible categories, each embedding vector E l , i After processing, it will be mapped to a label y (j) superior.
[0238] Among them, the mapping conformance classification task aims to identify specific mapping symbols (such as $a, $t, and $d) in ARM machine code, and the function boundary detection task focuses on identifying the start and end positions of the function, which is crucial for understanding the program structure and logic.
[0239] Due to the characteristics of the Reduced Instruction Set Computer (RISC) architecture, it is easy to determine the boundaries of each ARM instruction, but it is more difficult to distinguish between code and data at certain addresses. To solve this problem, ARM uses mapping symbols. Mapping symbols are an architecture-specific extension of the ARM ELF file generated by the compiler, which are used to identify the conversion between instructions (A32, T16, and T32) and inline data. For example, the mapping symbol $t 0x0001043c indicates that the offset 0x0001043c contains Thumb code. There are three main types of mapping symbols: (1) $a: The starting address of the ARM code area. (2) $t: The starting address of the Thumb code area. (3) $d: The starting address of the data area.
[0240] Inline data is ubiquitous in ARM binaries. To distinguish inline data (include data) from data segments ($d.realdata), inline data detection and instruction set identification can be combined by merging their labels. In particular, we add a new label D($d) to the label set |y|, which represents inline data. Therefore, for each 32-bit aligned binary code snippet, the set of possible class labels |y| = {A, T, D} represents ARM (A) instructions, Thumb (T) instructions, or inline data (D), respectively.
[0241] Specifically, through f task represents the model to be fine-tuned (i.e., disassembly), which receives each embedding E l , i As input, the label y (j) , where y i =f task (E l,i ). Let f taskDefined by the model parameter Θtask: f task (E l,i ; Θtask), therefore, training and fine-tuning the model f task The goal is to find Θtask that minimizes the cross entropy between the predicted labels and their actual labels for all inputs. In other words, the goal is to optimize the parameter Θtask so that the model l , i The label predictions made are as close as possible to the true label values. This is achieved by minimizing the cross entropy loss function, which measures the difference between the predicted distribution and the true distribution. The loss function can be expressed as:
[0242]
[0243] in, represents f task The predicted value of the output is
[0244]
[0245] For the function boundary detection task, the disassembly model can determine the target feature corresponding to the ARM machine code based on the first feature and the second feature; determine the function boundary information corresponding to the ARM machine code based on the target feature as the predicted function boundary information; then determine the boundary loss value based on the deviation between the predicted function boundary information and the actual function boundary information corresponding to the ARM machine code, and retrain the disassembly model with minimizing the boundary loss value as the optimization goal.
[0246] Specifically, the server can define an end-to-end function boundary detection task, where all units are divided into three categories: |Cy| = {S, E, N}, representing the start of a function (S), the end of a function (E), and neither the start of a function nor the end of a function (N). The function boundary is calculated based on the probability distribution of these three possible labels, and the predicted label is the one with the highest probability.
[0247] After the preprocessing and pretraining phases, the deep semantics between units, fields, and instruction sequences have been implicitly embedded in the weight parameters of the underlying neural network. In order to extract the learned semantics for various downstream tasks, the embedding vector output by the last layer of the pretrained language model (KMLM) can be used to stack a multi-layer perceptron (MLP) to predict different disassembly subtasks. The specific operation is as follows: For each embedding E in the last layer of the pretrained model, l,i , which is calculated as E l,i =KMLM (xi) .
[0248] Next, we define f task The function uses these embedded vectors to make classification predictions, and the prediction results can be expressed as:
[0249] f task (E l,i ; θ1,θ2)=Softmax(tanh(E l,i ·θ1)·θ2)
[0250] Among them, θ1 is a shape d emb ×d emb , which is used to transform the embedding vector; θ2 is a matrix with a shape of d emb ×|C y ∣ is a matrix for classification subtasks, where d emb is the dimension of the embedding vector, |C y ∣ is the number of output categories, which may vary depending on the fine-tuning subtask.
[0251] Therefore, the final output shape of the MLP is n×|C y ∣, where n is the length of the input instruction data sequence. In this way, the model can generate probability distributions corresponding to different categories for each position in each input sequence, which is suitable for processing downstream tasks such as sequence labeling or classification.
[0252] Of course, the server can also fine-tune the disassembly model through the above two subtasks at the same time, determine the comprehensive loss value according to the above classification loss value and boundary loss value, and retrain the disassembly model with minimizing the comprehensive loss value as the optimization goal.
[0253] Through the above two fine-tuning tasks, the capabilities of the pre-trained model can be further refined to make it more adaptable to the specific needs of ARM binary disassembly and improve the accuracy and efficiency of disassembly. During the fine-tuning process, the model parameters will be adjusted according to the specific task objectives to optimize its performance in instruction recognition and function boundary detection.
[0254] After the training and fine-tuning of the disassembly model are completed, the server can deploy the disassembly model, so as to execute the disassembly task for the ARM machine code through the disassembly model. For ease of understanding, this specification provides a flowchart of a task execution method applied to the above disassembly model, such as Figure 7 shown.
[0255] Figure 7 A flowchart of a task execution method provided in this specification includes the following steps:
[0256] S700: Obtain the ARM machine code to be disassembled as the target machine code.
[0257] In the process of executing the disassembly task, the server usually intelligently obtains the binary machine code of the ARM architecture, and disassembles the ARM machine code to obtain its corresponding assembly code and other information such as assembly instructions, function boundaries, function signatures, and control flow graphs.
[0258] S702: Input the target machine code into a pre-trained disassembly model to determine a feature representation corresponding to the target machine code through the disassembly model, wherein the disassembly model is trained by the above-mentioned model training method.
[0259] The server can input the target machine code into a pre-trained disassembly model, determine the first feature corresponding to the target machine code through a first feature extraction network set in the disassembly model, and determine the second feature corresponding to the target machine code through a second feature extraction network set in the disassembly model, and then determine the feature representation corresponding to the target machine code based on the first feature and the second feature.
[0260] For example, the server can determine the weight corresponding to the first feature and the weight corresponding to the second feature by weighted summation, and then perform weighted summation of the first feature and the second feature according to their respective corresponding weights to obtain the final feature representation. Of course, the server can also directly fuse the first feature and the second feature to obtain the final feature representation.
[0261] S704: Disassemble the target machine code according to the feature representation corresponding to the target machine code to obtain an assembly code corresponding to the target machine code, and execute the task according to the assembly code.
[0262] In the process of disassembling the target machine code, the function boundary information corresponding to the target machine code and the data type corresponding to each instruction data contained in the target machine code are determined according to the feature representation corresponding to the target machine code. Then, according to the function boundary information corresponding to the target machine code and the data type corresponding to each instruction data contained in the target machine code, the target machine code is disassembled to obtain the assembly code.
[0263] For each instruction data included in the target machine code, if the data type corresponding to the instruction data is inline data, the instruction data is not disassembled.
[0264] After obtaining the ARM assembly code, the server can perform tasks based on the assembly code. For example, the server can analyze the assembly code to determine the security vulnerabilities therein and then handle the security vulnerabilities. For another example, the server can analyze the overall logic of the assembly code and optimize the program code based on the analysis results.
[0265] From the above method, it can be seen that this scheme designs a multi-level masking strategy for ARM machine code so that the masked area can represent the common and possible changes between instruction fields in real ARM instructions. The input ARM binary file is randomly masked by multi-level masks, including unit level, field level and span level. Then, the knowledge is unsupervisedly pre-trained on the masked data to obtain a knowledge-aware masked language model (KMLM). KMLM is composed of the basic MLM (MLM basic ) and Span MLM (MLM span ) are jointly optimized in the pre-training stage. basic Able to understand the changes between the fields within the instruction, and MLM span It aims to capture the implicit relationship between consecutive instructions. The disassembly model trained in the above way can achieve accurate disassembly of ARM machine code.
[0266] The above are one or more methods for implementing model training or task execution in this specification. Based on the same idea, this specification also provides corresponding model training or task execution devices, such as Figure 8 or Fig. 9 shown.
[0267] Figure 8 A schematic diagram of a model training device provided in this specification, comprising:
[0268] The acquisition module 800 is used to acquire an advanced reduced instruction set ARM machine code, wherein the ARM machine code includes a plurality of instruction data, and the plurality of instruction data are executed according to a preset logical execution order;
[0269] An input module 802 is used to input the ARM machine code into a preset disassembly model, so as to perform mask processing on at least part of the instruction data contained in the ARM machine code through the disassembly model, so as to obtain first mask data according to the masked instruction data and the unmasked instruction data, and determine the feature representation corresponding to the first mask data as the first feature, and, for each instruction data, perform mask processing on at least part of the data units contained in the instruction data through the disassembly model to obtain second mask data corresponding to the instruction data, and determine the feature representation of the second mask data corresponding to the instruction data as the second feature corresponding to the instruction data;
[0270] The prediction module 804 is used to input the first feature into a preset decoding model so that the decoding model determines the position information corresponding to the masked instruction data in the ARM machine code according to the first feature as the predicted position information, and input the second feature corresponding to each instruction data into the decoding model so that the decoding model predicts the masked data unit in each instruction data according to the second feature corresponding to each instruction data to obtain a predicted value;
[0271] The training module 806 is used to determine the target loss function based on the deviation between the predicted value and the actual value corresponding to the masked data unit in each instruction data, and the deviation between the predicted position information and the actual position information corresponding to the masked instruction data in the ARM machine code, and to train at least the disassembly model with minimizing the target loss function as the optimization goal.
[0272] Optionally, the input module 802 is specifically configured to determine a feature representation corresponding to the first mask data through a first feature extraction network provided in the disassembly model.
[0273] Optionally, the input module 802 is specifically configured to use, through a second feature extraction network provided in the disassembly model, a feature representation of the second mask data corresponding to the instruction data as a second feature corresponding to the instruction data.
[0274] Optionally, the input module 802 is specifically configured to, for each instruction data, perform mask processing on at least part of the data units contained in the instruction data according to a preset mask ratio through the disassembly model to obtain second mask data corresponding to the instruction data.
[0275] Optionally, the input module 802 is specifically used to, for each instruction data, determine, through the disassembly model, a mask strategy corresponding to the instruction data according to the data type corresponding to the instruction data; determine, according to the mask strategy corresponding to the instruction data, a field type corresponding to the data unit to be masked in the instruction data as a target field type; and mask the data units of the instruction field in the instruction data that belong to the target field type to obtain second mask data corresponding to the instruction data.
[0276] Optionally, the input module 802 is specifically used to, for each instruction data, mask at least part of the data units contained in the instruction data according to a preset mask ratio through the disassembly model to obtain third mask data corresponding to the instruction data, and, for each instruction data, determine the mask strategy corresponding to the instruction data according to the data type corresponding to the instruction data; determine the field type corresponding to the data unit to be masked in the instruction data as the target field type according to the mask strategy corresponding to the instruction data; mask the data units of the instruction field in the instruction data that belong to the target field type to obtain fourth mask data corresponding to the instruction data; and determine the second mask data corresponding to the instruction data according to the third mask data and the fourth mask data corresponding to the instruction data.
[0277] Optionally, the input module 802 is specifically used to determine, for each instruction data contained in the first mask data, position information corresponding to the instruction data in the first mask data, the position information including the position code and boundary information corresponding to the instruction data in the first mask data, and the boundary information is used to indicate the starting character and the ending character of the instruction data in the first mask data; determine the first feature according to the position information corresponding to each instruction data.
[0278] Optionally, the input module 802 is specifically used to determine, for each data unit contained in the second mask data corresponding to the instruction data, the one-hot encoding corresponding to the data unit and the position encoding corresponding to the data unit in the instruction data to which the data unit belongs; determine the initial feature corresponding to the data unit based on the one-hot encoding and the position encoding; determine the feature representation corresponding to the data unit based on the weight between the data unit and each other data unit contained in the second mask data corresponding to the instruction data, and the initial feature corresponding to each data unit contained in the second mask data corresponding to the instruction data; determine the second feature corresponding to the data unit based on the feature representation corresponding to each data unit in the instruction data.
[0279] Optionally, the training module 806 is also used to determine the target feature corresponding to the ARM machine code based on the first feature and the second feature; determine the data type corresponding to each instruction data as the predicted data type based on the target feature; determine the classification loss value based on the deviation between the predicted data type and the actual data type corresponding to each instruction data; and retrain the disassembly model with minimizing the classification loss value as the optimization goal.
[0280] Optionally, the training module 806 is also used to determine the target feature corresponding to the ARM machine code based on the first feature and the second feature; determine the function boundary information corresponding to the ARM machine code based on the target feature as predicted function boundary information; determine the boundary loss value based on the deviation between the predicted function boundary information and the actual function boundary information corresponding to the ARM machine code; and retrain the disassembly model with minimizing the boundary loss value as the optimization goal.
[0281] Optionally, the training module 806 is also used to determine the target feature corresponding to the ARM machine code based on the first feature and the second feature; determine the data type corresponding to each instruction data based on the target feature as the predicted data type, and determine the function boundary information corresponding to the ARM machine code as the predicted function boundary information; determine the comprehensive loss value based on the deviation between the predicted function boundary information and the actual function boundary information corresponding to the ARM machine code, and the deviation between the predicted function boundary information and the actual function boundary information corresponding to the ARM machine code; and retrain the disassembly model with minimizing the comprehensive loss value as the optimization goal.
[0282] Fig. 9 A schematic diagram of a model training device provided in this specification, comprising:
[0283] The acquisition module 900 is used to acquire the ARM machine code to be disassembled as the target machine code;
[0284] A determination module 902 is used to input the target machine code into a pre-trained disassembly model to determine a feature representation corresponding to the target machine code through the disassembly model;
[0285] The execution module 904 is used to disassemble the target machine code according to the feature representation corresponding to the target machine code, obtain the assembly code corresponding to the target machine code, and execute the task according to the assembly code.
[0286] Optionally, the determination module 902 is specifically used to determine a first feature corresponding to the target machine code through a first feature extraction network set in the disassembly model, and to determine a second feature corresponding to the target machine code through a second feature extraction network set in the disassembly model; and to determine a feature representation corresponding to the target machine code based on the first feature and the second feature.
[0287] Optionally, the execution module 904 is specifically used to determine, according to the feature representation corresponding to the target machine code, function boundary information corresponding to the target machine code and the data type corresponding to each instruction data contained in the target machine code; according to the function boundary information corresponding to the target machine code and the data type corresponding to each instruction data contained in the target machine code, disassemble the target machine code to obtain the assembly code.
[0288] Optionally, the execution module 904 is specifically used to determine, for each instruction data, whether the data type corresponding to the instruction data is instruction data, the instruction data including: ARM instructions and Thumb instructions; if so, disassemble the instruction data.
[0289] Optionally, the execution module 904 is further configured to, for each instruction data included in the target machine code, not disassemble the instruction data if the data type corresponding to the instruction data is inline data.
[0290] This specification also provides a computer-readable storage medium, which stores a computer program, which can be used to execute the above Figure 1 or Figure 7 A model training or task execution method provided.
[0291] This manual also provides Fig.10 The one shown corresponds to Figure 1 or Figure 7 A schematic diagram of the electronic device. Figure 4 As mentioned above, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory, and may also include other hardware required for the business. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to achieve the above Figure 1 or Figure 7 The model training or task execution method described. Of course, in addition to the software implementation, this specification does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc., that is, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0292] In the 1990s, improvements to a technology could be clearly distinguished as hardware improvements (for example, improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the method flow). However, with the development of technology, many improvements to the method flow today can be regarded as direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement in a method flow cannot be implemented using a hardware entity module. For example, a programmable logic device (PLD) (such as a field programmable gate array (FPGA)) is such an integrated circuit whose logical function is determined by the user's programming of the device. Designers can "integrate" a digital system on a PLD by programming it themselves, without having to ask a chip manufacturer to design and produce a dedicated integrated circuit chip. Moreover, nowadays, instead of manually making integrated circuit chips, this kind of programming is mostly implemented by "logic compiler" software, which is similar to the software compiler used when developing and writing programs, and the original code before compilation must also be written in a specific programming language, which is called hardware description language (HDL). There is not only one HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also know that it is only necessary to program the method flow slightly in the above-mentioned hardware description languages and program it into the integrated circuit, and then it is easy to obtain the hardware circuit that implements the logic method flow.
[0293] The controller can be implemented in any appropriate manner, for example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing a computer-readable program code (such as software or firmware) that can be executed by the (micro)processor, a logic gate, a switch, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in a purely computer-readable program code manner, the controller can be implemented in the form of a logic gate, a switch, an application-specific integrated circuit, a programmable logic controller, and an embedded microcontroller by logically programming the method steps. Therefore, this controller can be considered as a hardware component, and the devices included therein for implementing various functions can also be regarded as structures within the hardware component. Or even, the devices for implementing various functions can be regarded as both software modules for implementing the method and structures within the hardware component.
[0294] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0295] For the convenience of description, the above device is described in various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0296] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0297] This specification is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0298] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0299] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0300] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0301] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0302] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0303] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0304] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0305] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0306] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0307] The above description is only an embodiment of the present specification and is not intended to limit the present specification. For those skilled in the art, the present specification may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present specification shall be included in the scope of the claims of the present specification.
Claims
1. A model training method, comprising: Acquire an advanced reduced instruction set ARM machine code, wherein the ARM machine code includes a plurality of instruction data, and the plurality of instruction data are executed according to a preset logical execution order; Inputting the ARM machine code into a preset disassembly model, so as to perform mask processing on at least part of the instruction data contained in the ARM machine code through the disassembly model, so as to obtain first mask data according to the masked instruction data and the unmasked instruction data, and determine the feature representation corresponding to the first mask data as the first feature, and, for each instruction data, perform mask processing on at least part of the data units contained in the instruction data through the disassembly model, so as to obtain second mask data corresponding to the instruction data, and determine the feature representation of the second mask data corresponding to the instruction data as the second feature corresponding to the instruction data; Inputting the first feature into a preset decoding model, so that the decoding model determines the position information corresponding to the masked instruction data in the ARM machine code according to the first feature as the predicted position information, and inputting the second feature corresponding to each instruction data into the decoding model, so that the decoding model predicts the masked data unit in each instruction data according to the second feature corresponding to each instruction data to obtain a predicted value; Based on the deviation between the predicted value and the actual value corresponding to the masked data unit in each instruction data, and the deviation between the predicted position information and the actual position information corresponding to the masked instruction data in the ARM machine code, the target loss function is determined, and the disassembly model is at least trained with minimizing the target loss function as the optimization goal.
2. The method according to claim 1, determining the feature representation corresponding to the first mask data, specifically comprising: A feature representation corresponding to the first mask data is determined through a first feature extraction network set in the disassembly model.
3. The method according to claim 1, determining the feature representation of the second mask data corresponding to the instruction data as the second feature corresponding to the instruction data, specifically comprising: Through the second feature extraction network set in the disassembly model, the feature representation of the second mask data corresponding to the instruction data is used as the second feature corresponding to the instruction data.
4. The method according to claim 1, for each instruction data, performing mask processing on at least part of the data units contained in the instruction data through the disassembly model to obtain second mask data corresponding to the instruction data, specifically comprising: For each instruction data, mask processing is performed on at least part of the data units contained in the instruction data according to a preset mask ratio through the disassembly model to obtain second mask data corresponding to the instruction data.
5. The method according to claim 1, for each instruction data, performing mask processing on at least part of the data units contained in the instruction data through the disassembly model to obtain second mask data corresponding to the instruction data, specifically comprising: For each instruction data, determine the mask strategy corresponding to the instruction data according to the data type corresponding to the instruction data through the disassembly model; According to the mask strategy corresponding to the instruction data, determine the field type corresponding to the data unit to be masked in the instruction data as the target field type; The data units of the instruction field in the instruction data that belong to the target field type are masked to obtain second mask data corresponding to the instruction data.
6. The method according to claim 1, for each instruction data, performing mask processing on at least part of the data units contained in the instruction data through the disassembly model to obtain second mask data corresponding to the instruction data, specifically comprising: For each instruction data, mask processing is performed on at least part of the data units contained in the instruction data according to a preset mask ratio through the disassembly model to obtain third mask data corresponding to the instruction data, and, for each instruction data, a mask strategy corresponding to the instruction data is determined according to the data type corresponding to the instruction data; according to the mask strategy corresponding to the instruction data, a field type corresponding to the data unit to be masked in the instruction data is determined as a target field type; and data units of the instruction field in the instruction data belonging to the target field type are masked to obtain fourth mask data corresponding to the instruction data; The second mask data corresponding to the instruction data is determined according to the third mask data and the fourth mask data corresponding to the instruction data.
7. The method according to claim 1, determining the feature representation corresponding to the first mask data as the first feature, specifically comprising: For each instruction data included in the first mask data, determining position information corresponding to the instruction data in the first mask data, the position information including a position code corresponding to the instruction data in the first mask data and boundary information, the boundary information being used to indicate a starting character and an ending character of the instruction data in the first mask data; The first feature is determined according to the position information corresponding to each instruction data.
8. The method according to claim 1, determining a feature representation of the second mask data corresponding to the instruction data as the second feature corresponding to the instruction data: For each data unit included in the second mask data corresponding to the instruction data, determine a one-hot encoding corresponding to the data unit and a position encoding corresponding to the data unit in the instruction data to which the data unit belongs; Determine an initial feature corresponding to the data unit according to the one-hot encoding and the position encoding; Determine a feature representation corresponding to the data unit according to a weight between the data unit and each other data unit included in the second mask data corresponding to the instruction data, and an initial feature corresponding to each data unit included in the second mask data corresponding to the instruction data; According to the feature representation corresponding to each data unit in the instruction data, a second feature corresponding to the data unit is determined.
9. The method of claim 1, further comprising: Determine a target feature corresponding to the ARM machine code according to the first feature and the second feature; According to the target feature, determining the data type corresponding to each instruction data as the predicted data type; Determining a classification loss value according to a deviation between the predicted data type and an actual data type corresponding to each instruction data; The disassembly model is retrained with minimizing the classification loss value as an optimization goal.
10. The method of claim 1, further comprising: Determine a target feature corresponding to the ARM machine code according to the first feature and the second feature; Determine, according to the target feature, function boundary information corresponding to the ARM machine code as predicted function boundary information; Determine a boundary loss value according to a deviation between the predicted function boundary information and the actual function boundary information corresponding to the ARM machine code; The disassembly model is retrained with minimizing the boundary loss value as an optimization goal.
11. The method of claim 1, further comprising: Determine a target feature corresponding to the ARM machine code according to the first feature and the second feature; According to the target feature, determining the data type corresponding to each instruction data as the predicted data type, and determining the function boundary information corresponding to the ARM machine code as the predicted function boundary information; Determine a comprehensive loss value according to a deviation between the predicted function boundary information and the actual function boundary information corresponding to the ARM machine code, and a deviation between the predicted function boundary information and the actual function boundary information corresponding to the ARM machine code; The disassembly model is retrained with minimizing the comprehensive loss value as an optimization goal.
12. A task execution method, comprising: Get the ARM machine code that needs to be disassembled as the target machine code; Inputting the target machine code into a pre-trained disassembly model to determine the feature representation corresponding to the target machine code through the disassembly model, wherein the disassembly model is trained by the method according to any one of claims 1 to 11 above; According to the feature representation corresponding to the target machine code, the target machine code is disassembled to obtain the assembly code corresponding to the target machine code, and the task is executed according to the assembly code.
13. The method according to claim 12, wherein determining the feature representation corresponding to the target machine code through the disassembly model comprises: Determine a first feature corresponding to the target machine code through a first feature extraction network set in the disassembly model, and determine a second feature corresponding to the target machine code through a second feature extraction network set in the disassembly model; A feature representation corresponding to the target machine code is determined according to the first feature and the second feature.
14. The method according to claim 12, wherein the target machine code is disassembled according to the feature representation corresponding to the target machine code to obtain the assembly code corresponding to the target machine code, specifically comprising: Determine, according to the characteristic representation corresponding to the target machine code, function boundary information corresponding to the target machine code and a data type corresponding to each instruction data contained in the target machine code; According to the function boundary information corresponding to the target machine code and the data type corresponding to each instruction data contained in the target machine code, the target machine code is disassembled to obtain the assembly code.
15. The method according to claim 14, disassembling the target machine code according to the function boundary information corresponding to the target machine code and the data type corresponding to each instruction data contained in the target machine code to obtain the assembly code, specifically comprising: For each instruction data, determining whether the data type corresponding to the instruction data is instruction data, wherein the instruction data includes: ARM instructions and Thumb instructions; If so, the instruction data is disassembled.
16. The method of claim 15, further comprising: For each instruction data included in the target machine code, if the data type corresponding to the instruction data is inline data, the instruction data is not disassembled.
17. A model training device, comprising: An acquisition module is used to acquire an advanced reduced instruction set ARM machine code, wherein the ARM machine code includes a plurality of instruction data, and the plurality of instruction data are executed according to a preset logical execution order; An input module, used for inputting the ARM machine code into a preset disassembly model, so as to perform mask processing on at least part of the instruction data contained in the ARM machine code through the disassembly model, so as to obtain first mask data according to the masked instruction data and the unmasked instruction data, and determine the feature representation corresponding to the first mask data as the first feature, and, for each instruction data, perform mask processing on at least part of the data units contained in the instruction data through the disassembly model, so as to obtain second mask data corresponding to the instruction data, and determine the feature representation of the second mask data corresponding to the instruction data as the second feature corresponding to the instruction data; A prediction module, used for inputting the first feature into a preset decoding model, so that the decoding model determines the position information corresponding to the masked instruction data in the ARM machine code according to the first feature as the predicted position information, and inputting the second feature corresponding to each instruction data into the decoding model, so that the decoding model predicts the masked data unit in each instruction data according to the second feature corresponding to each instruction data, and obtains a predicted value; A training module is used to determine a target loss function based on the deviation between the predicted value and the actual value corresponding to the masked data unit in each instruction data, and the deviation between the predicted position information and the actual position information corresponding to the masked instruction data in the ARM machine code, and to train at least the disassembly model with minimizing the target loss function as the optimization goal.
18. A task execution device, comprising: The acquisition module is used to obtain the ARM machine code that needs to be disassembled as the target machine code; A determination module, used for inputting the target machine code into a pre-trained disassembly model to determine a feature representation corresponding to the target machine code through the disassembly model, wherein the disassembly model is trained by the method according to any one of claims 1 to 11 above; The execution module is used to disassemble the target machine code according to the feature representation corresponding to the target machine code, obtain the assembly code corresponding to the target machine code, and execute the task according to the assembly code.
19. A computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the method according to any one of claims 1 to 15.
20. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 15 when executing the program.
Citation Information
Patent Citations
Cross-instruction architecture binary code similarity detection method based on semantics
CN112596736A
Semantic representation model pre-training method and device, electronic equipment and storage medium
CN113239705A
Training method and device for pre-training language model
CN114579699A
Model training method and device of pre-training model, storage medium and electronic equipment
CN117476113A
Performance prediction model training method and device and performance prediction method and device
CN118690824A
Cited By
Communication fault event argument extraction method based on large model event focusing and term enhancement
CN120805902A
A Communication Failure Event Argument Extraction Method Based on Large Model Event Focusing and Terminology Enhancement
CN120805902B