Disassembling method and device

CN120153350APending Publication Date: 2025-06-13HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380078027.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-07-06
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Existing disassembly technology easily recognizes convergent data as instructions when processing object code, resulting in analysis troubles, and recursive disassembly cannot completely cover object code, consumes a lot of computing power and memory, and has low accuracy.

Method used

Byte offset fallback technology is used to roll back one or more bytes from the target address when garbled code is encountered, and redisassemble one or more bytes are generated, multiple disassembly codes are filtered out according to the number of instructions and the target address of the jump instruction. disassembly code.

Benefits of technology

Improves code coverage, reduces computing power and memory consumption, and ensures the accuracy and efficiency of selected disassembled code.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120153350A_ABST
    Figure CN120153350A_ABST
Patent Text Reader

Abstract

The disassembling method comprises the following steps: disassembling machine codes / binary files one by one to obtain first disassembling codes; under the condition that messy codes exist in the first disassembling codes, one or more different bytes are returned from a first target address in the machine code / binary file for disassembling again, and at least one second disassembling code is obtained, wherein the first target address is an address corresponding to the last byte in the messy code or a corresponding position of a current disassembling address in a machine code / binary file; a target disassembly code is selected from the at least one second disassembly code. According to the method, one or more disassembling codes are obtained by adopting a byte offset rollback technology during disassembling, and a correct disassembling code is selected from the multiple disassembling codes according to the number of instructions in the disassembling codes and the target address of the jump, so that the purpose of realizing a relatively large code coverage rate with less computing power and memory is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Disassembly method and device Technical Field

[0001] The present invention relates to the field of computer technology, and more particularly to a method and apparatus for disassembling Background Art

[0002] Disassembly is the process of converting target code into assembly code, or rather, converting machine language code into assembly language code. During disassembly, since machine instructions are identical to ordinary binary values, it's easy to mistake values ​​that aren't instructions for machine instructions. To minimize these errors, disassembly algorithms often employ two modes: linear disassembly and recursive disassembly. While linear disassembly can cover the target code as closely as possible, it can easily misidentify inline data within the target code as instructions, complicating subsequent analysis. Recursive disassembly, on the other hand, lacks complete coverage of the target code.

[0003] Summary of the Invention

[0004] The present application provides a disassembly method, apparatus, computer-readable storage medium, and computer program product, which adopt byte offset back-off technology during disassembly to obtain one or more disassembly codes, and select the correct disassembly code from multiple disassembly codes based on the number of instructions in the disassembly code and the target address of the unconditional jump instruction, thereby achieving a higher code coverage rate with less computing power and memory.

[0005] In a first aspect, the present application provides a disassembly method, comprising: disassembling the obtained machine code / binary file line by line to obtain a first disassembly code; when there is garbled code in the first disassembly code, backing off one byte or multiple different bytes from the first target address in the machine code / binary file and re-disassembling line by line to obtain at least one second disassembly code, wherein the first target address is the address corresponding to the last byte in the garbled code or the address corresponding to the current address of disassembly in the machine code / binary file; filtering out the target disassembly code from the at least one second disassembly code; and outputting the target disassembly code.

[0006] In this solution, when garbled characters appear in the generated disassembled code during the disassembly process of machine code / binary files, a byte offset rollback technique is used. Starting from the first target address in the machine code / binary file, the machine code / binary file is reassembled one or more bytes later, resulting in one or more versions of disassembled code. The correct disassembled code is then selected from these multiple versions, achieving higher code coverage with less computing power and memory.

[0007] In one possible implementation, filtering out the target disassembly code from the at least one second disassembly code includes: filtering out the target disassembly code from the at least one second disassembly code based on the number of instructions contained in each code in the at least one second disassembly code and / or the second target address of the jump instruction jump.

[0008] That is to say, selecting the target disassembly code based on the number of instructions in the second disassembly code can reduce the difficulty of selecting the target disassembly code. Selecting the target disassembly code based on the second target address of the jump in the second disassembly code can ensure that the selected target disassembly code is correct. Selecting the target disassembly code based on the number of instructions in the second disassembly code and the second target address of the jump can reduce the difficulty of selecting the target disassembly code while ensuring the accuracy of the selected target disassembly code.

[0009] In one possible implementation, based on the number of instructions contained in each code in at least one second disassembly code, a target disassembly code is screened out from at least one second disassembly code, including: selecting a second disassembly code with the largest number of instructions from at least one second disassembly code as the target disassembly code.

[0010] That is to say, when selecting a target disassembly code from at least one second disassembly code, by comparing the number of instructions in each version of the second disassembly code in the at least one second disassembly code, the disassembly code corresponding to the version with the largest number of instructions is selected as the target disassembly code. While taking into account the correctness of the selected target disassembly code, the difficulty of selecting the target disassembly code from the at least one second disassembly code is reduced.

[0011] In one possible implementation, based on the target address of the jump instruction jump contained in each code in at least one second disassembly code, a target disassembly code is filtered out from at least one second disassembly code, including: selecting a second disassembly code with a correct second target address of the jump instruction from at least one second disassembly code as the target disassembly code.

[0012] That is, since the second target address of the jump instruction in the correct version of the target disassembly code is necessarily correct, by determining the second target address of the jump instruction contained in each version of the disassembly code in at least one second disassembly code, and selecting the disassembly code corresponding to the version with the correct second target address of the jump instruction as the target disassembly code, it can be ensured that the selected target disassembly code is necessarily correct.

[0013] In one possible implementation, based on the number of instructions contained in each code in at least one second disassembly code and the target address of the jump instruction jump, a target disassembly code is screened out from at least one second disassembly code, including: selecting n second disassembly codes with the largest number of instructions from at least one second disassembly code, where n is a natural number greater than or equal to 2; and selecting a second disassembly code with a correct second target address of the jump instruction as the target disassembly code from the n second disassembly codes.

[0014] That is to say, by judging the number of instructions contained in each code in the second disassembly code to select the target disassembly code, the difficulty of selecting the target disassembly code can be reduced, but the accuracy may not be so high. By judging whether the target address of the jump in each code in the second disassembly code is correct to select the target disassembly code, the correctness of the selected target disassembly code can be guaranteed, but the computing power required to be consumed is relatively large. Therefore, by judging the number of instructions contained in each code in the second disassembly code, at least one second disassembly code is screened for the first time. Then, by judging whether the jump target address in the screened disassembly code is correct, the target disassembly code is selected, while ensuring the correctness of the target disassembly code obtained, computing power is saved.

[0015] In one possible implementation, before filtering out a target disassembly code from at least one second disassembly code based on the number of instructions contained in each code in at least one second disassembly code and / or the second target address of the jump instruction jump, the method further includes: filtering out a second disassembly code that does not contain garbled characters from at least one second disassembly code to obtain k second disassembly codes, where k is a natural number greater than or equal to 1; and filtering out the target disassembly code from the k second disassembly codes.

[0016] That is to say, the second disassembly code containing garbled characters is definitely not the target disassembly code. When selecting the target disassembly code from the at least one second disassembly code obtained, the second disassembly code containing garbled characters is excluded, so that when determining the target disassembly code, the scope of selecting the target disassembly code (correct disassembly code) can be narrowed, thereby improving efficiency.

[0017] In one possible implementation, filtering out a target disassembly code from at least one second disassembly code includes: assembling the at least one second disassembly code to obtain at least one third assembly code; selecting a target third assembly code from the at least one third assembly code based on the obtained machine code / binary file, and using the second disassembly code corresponding to the target third assembly code as the target disassembly code; or executing the at least one third assembly code, selecting a target third assembly code from the at least one third disassembly code based on the execution result of the at least one third assembly code, and using the second disassembly code corresponding to the target third assembly code as the target disassembly code, so that the target third assembly code can be correctly executed.

[0018] That is to say, when selecting the target disassembly code from at least one second disassembly code, the at least one second disassembly code obtained can be directly assembled, the assembly code obtained by assembling is compared with the machine code / binary file, and the second disassembly code corresponding to the assembly code that matches the machine code / binary file is selected as the target disassembly code. Wherein, the assembly code obtained by assembling matches the machine code / binary file, which means that the assembly code obtained by assembling has the same function as the machine code / binary file. When selecting the target disassembly code from at least one second disassembly code, the assembly code obtained by assembling can also be executed, and the second disassembly code corresponding to the assembly code that can be successfully executed is used as the target disassembly code. Directly assembling the at least one second disassembly code and determining the target disassembly code based on the assembly result reduces the difficulty of selecting the target disassembly code and also ensures the accuracy of the selected target disassembly code.

[0019] In one possible implementation, backing off one byte or multiple different bytes from the first target address in the machine code / binary file and re-disassembling one by one to obtain at least one second disassembly code includes: starting from the first target address and backing off different bytes respectively to generate multiple disassembly tasks; executing the multiple disassembly tasks in parallel to obtain at least one second disassembly code.

[0020] That is, after determining the first target address (the xth byte) in the target machine code / binary file, the process can be backed off by 1 byte, 2 bytes, ..., m bytes starting from the first target address. Then, the machine code / binary file is disassembled starting from the x-1th byte, the x-2th byte, ..., the xmth byte, respectively, to obtain m disassembly tasks. By executing the m disassembly tasks, m versions of the disassembly code are generated. When executing the m disassembly tasks, one or more processors can execute them in parallel to improve the efficiency of the processors in executing the disassembly tasks.

[0021] In a possible implementation, the method further includes: when no garbled characters exist in the first disassembly code, using the first disassembly code as the target disassembly code.

[0022] That is to say, in the process of disassembling the machine code / binary file, if there is no garbled code in the obtained disassembly code, the disassembly will be continued in the current order until the machine code / binary file is disassembled and a disassembly code is obtained, which is the target disassembly code.

[0023] In a second aspect, the present application provides a disassembly device, comprising: at least one processor and a memory,

[0024] Memory is used to store computer instructions;

[0025] At least one processor is configured to perform the following steps according to computer instructions:

[0026] Disassembling the obtained machine code / binary file line by line to obtain a first disassembly code;

[0027] If there is garbled code in the first disassembled code, rewind one byte or multiple different bytes from the first target address in the machine code / binary file and re-disassemble the code line by line to obtain at least one second disassembled code, where the first target address is the address corresponding to the last byte in the garbled code or the address corresponding to the current disassembly address in the machine code / binary file;

[0028] Filtering a target disassembled code from at least one second disassembled code;

[0029] Output target disassembly code.

[0030] In one possible implementation, at least one processor is configured to:

[0031] Target disassembly code is filtered out from the at least one second disassembly code according to the number of instructions contained in each code in the at least one second disassembly code and / or the second target address of the jump instruction jump.

[0032] In one possible implementation, at least one processor is configured to:

[0033] A second disassembly code with the largest number of instructions is selected from the at least one second disassembly code as a target disassembly code.

[0034] In one possible implementation, at least one processor is configured to:

[0035] A second disassembly code having a correct second target address of the jump instruction is selected from the at least one second disassembly code as the target disassembly code.

[0036] In one possible implementation, at least one processor is configured to:

[0037] Selecting n second disassembly codes with the largest number of instructions from at least one second disassembly code, where n is a natural number greater than or equal to 2;

[0038] A second disassembly code having a correct second target address of the jump instruction is selected from the n second disassembly codes as a target disassembly code.

[0039] In one possible implementation, at least one processor is configured to:

[0040] Filtering out second disassembly codes that do not contain garbled characters from the at least one second disassembly code to obtain k second disassembly codes, where k is a natural number greater than or equal to 1;

[0041] Filter out the target disassembly code from the k second disassembly codes.

[0042] In one possible implementation, at least one processor is configured to:

[0043] Assembling the at least one second disassembled code to obtain at least one third assembly code;

[0044] Selecting a target third assembly code from at least one third assembly code according to the obtained machine code / binary file, and using the second disassembly code corresponding to the target third assembly code as the target disassembly code; or

[0045] Execute at least one third assembly code, select a target third assembly code from at least one third disassembly code based on the execution result of the at least one third assembly code, and use the second disassembly code corresponding to the target third assembly code as the target disassembly code, so that the target third assembly code can be correctly executed.

[0046] In one possible implementation, at least one processor is configured to:

[0047] Starting from the first target address, different bytes are respectively retreated to generate multiple disassembly tasks;

[0048] The multiple disassembly tasks are executed in parallel to obtain at least one second disassembly code.

[0049] In one possible implementation, the at least one processor is further configured to:

[0050] When no garbled characters exist in the first disassembled code, the first disassembled code is used as the target disassembled code.

[0051] In a third aspect, an embodiment of the present application provides a computer storage medium, in which instructions are stored. When the instructions are executed on a computer, the computer executes the method described in the first aspect or any possible implementation of the first aspect.

[0052] In a fourth aspect, an embodiment of the present application provides a computer program product comprising instructions, which, when executed on a computer, enables the computer to execute the method described in the first aspect or any possible implementation of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0054] FIG1 is a schematic diagram of an assembly code provided in an embodiment of the present application;

[0055] FIG2 is a schematic diagram of an assembly code provided in an embodiment of the present application;

[0056] FIG3 is a schematic diagram of the structure of a disassembly tool provided in an embodiment of the present application;

[0057] FIG4 is a schematic diagram of a process flow of a disassembly method provided in an embodiment of the present application;

[0058] FIG5 is a schematic diagram of a disassembly process provided by an embodiment of the present application;

[0059] FIG6 is a schematic diagram of a disassembly process provided by an embodiment of the present application;

[0060] FIG7 is a schematic diagram of a disassembly result provided in an embodiment of the present application. DETAILED DESCRIPTION

[0061] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described below with reference to the accompanying drawings.

[0062] In the description of the embodiments of this application, any embodiment or design scheme using "exemplary," "for example," or "for example" should not be understood as being more preferred or advantageous than other embodiments or designs. Rather, the use of words such as "exemplary," "for example," or "for example" is intended to present the relevant concepts in a concrete manner.

[0063] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly identifying the technical features being referred to. Thus, features specified as "first" or "second" may explicitly or implicitly include one or more of such features. The terms "include," "comprising," "having," and their variations all mean "including but not limited to," unless otherwise specifically emphasized.

[0064] Disassembly algorithms often use two modes, namely linear disassembly and recursive disassembly. Linear disassembly starts from the first byte of a code segment, scans the entire code segment in a linear mode, and disassembles each instruction one by one until the entire code segment is completed. For example, Figure 1 shows a segment of assembly code. As shown in Figure 1, the segment of assembly code includes instructions BB0, BB1, BB2 and inline data. Taking the assembly code shown in Figure 1 as an example, linear disassembly can cover all code segments of the program as much as possible, but linear disassembly can easily identify inline data as instructions. As a result, two types of errors will be caused. One is that the numerical value is converted into an invalid machine instruction, and the other is that the numerical value just corresponds to a certain machine instruction, which will cause greater interference to subsequent analysis.

[0065] The basic idea behind recursive disassembly is to find the program's control flow, starting with program entry points like main, then performing linear disassembly. If a jump instruction, such as a jump, is encountered during linear disassembly, the code jumps to the corresponding address and continues disassembling. Because program control flow is difficult to track, many jumps are invisible. These invisible jumps do not have specific addresses in the binary file, requiring verification at runtime. This can prevent parts of the program from being disassembled.

[0066] In one possible example, referring to Figure 2, in BB0, BB1 in Jmp BB1 and BB2 in JI BB2 are explicit addresses, while call eax in BB1 and BB2 are implicit addresses. [func_address] in BB1 corresponds to the address of f1, and [func_address] in BB2 corresponds to the address of f2. Because call eax is an implicit address, the value of eax is only known when the code segments corresponding to BB1 and BB2 are executed. In a static state, disassembly tools cannot determine the specific value corresponding to eax and, therefore, cannot jump to the given address for disassembly, resulting in the inability to disassemble f1 and f2.

[0067] In view of this, an embodiment of the present application provides a disassembly method. In the process of linear disassembly using a disassembly tool, if an error is encountered, one or more bytes are rolled back at the location where the error occurred for reassembly to obtain multiple versions of disassembly code. After obtaining multiple versions of disassembly code, a correct version of disassembly code can be determined from multiple versions of disassembly code by comparing the number of instructions in the multiple versions of disassembly code or judging whether the target address of the unconditional jump instruction jump in the multiple versions of disassembly code is correct. This solves the problem that existing disassembly technology requires more computing power and memory and has low accuracy.

[0068] For example, Figure 3 shows a schematic diagram of the structure of a disassembly tool. As shown in Figure 3, the disassembly tool includes: a processor 310, a network interface 320, and a memory 330. The processor 310, the network interface 320, and the memory 330 can be connected via a bus or other means.

[0069] Memory 330 is used to store target files. A target file is a binary or machine code generated by a compiler from source code that is directly recognizable by a processor. A target file includes code and data used by the code at runtime, such as relocation information, program symbols used for linking or debugging, and other debugging information.

[0070] The processor 310 obtains the target file from the memory 330 and performs a disassembly operation on the obtained target file. If garbled characters appear in the disassembled code obtained during the disassembly of the target file, the processor 310 can use byte offset technology to re-disassemble the target file to obtain multiple versions of the disassembled code. Among them, garbled characters refer to incorrect instructions or unrecognizable codes in the disassembled code. Then, the processor 310 selects a correct version from the multiple versions of the disassembled code obtained. Among them, the byte offset fallback technology means that if garbled characters appear in the disassembled code obtained during the disassembly of the processor 310, the processor 310 can fall back one byte or multiple bytes from the place where the garbled characters occur and re-disassemble the target file. For example, the processor 310 can fall back 1 byte, 2 bytes, and 3 bytes from the place where the garbled characters occur to obtain three versions of the disassembly code.

[0071] In one possible example, the processor 310 may be a multi-core processor, and multiple threads in the processor may be used to execute n disassembly tasks in parallel, where n is a natural number greater than or equal to 1. For example, the processor 310 includes four physical cores, and hyperthreading technology is used to simulate two virtual cores with one physical core. That is, the processor 310 is a processor with four cores and eight threads. During disassembly, the processor 210 may execute eight disassembly tasks simultaneously.

[0072] In one possible example, the disassembly tool may include two processors (not shown in FIG3 ), namely a first processor and a second processor, and the first processor and the second processor together execute n disassembly tasks. Specifically, both the first processor and the second processor may be multi-core processors, and when executing n disassembly tasks, the n disassembly tasks may be executed in parallel by multiple threads in the first processor and the second processor.

[0073] After the processor 310 obtains multiple versions of disassembly codes, the processor 310 also needs to select a correct version of the disassembly code from the obtained multiple versions of the disassembly code. Specifically, the processor 310 can obtain the number of instructions in each version of the disassembly code and select the version with the largest number of instructions as the correct disassembly code. Alternatively, the processor 310 can also select the version with the correct address pointed to by the jump instruction from the multiple versions of the disassembly code as the correct disassembly code. Alternatively, when the processor 310 determines the correct disassembly code from the multiple versions of the disassembly code, it can simultaneously consider the number of instructions in the disassembly code and whether the address pointed to by the jump instruction is correct. For example, the processor 310 sorts the multiple versions of the disassembly code from most to least according to the number of instructions, and selects n disassembly code versions with the highest number of instructions, where n is a natural number greater than or equal to 2. Then, the processor 310 selects the version with the correct address pointed to by the jump instruction from the n versions of the disassembly code as the correct disassembly code version.

[0074] In one possible example, while the processor 310 is re-disassembling the target file, the processor 310 can detect in real time whether garbled characters appear in the generated disassembly code. If garbled characters appear in the currently generated disassembly code, the processor 310 does not save the disassembly code that contains the garbled characters. By retaining only the versions of the disassembly code that do not contain garbled characters among the multiple obtained versions, the range of selecting the correct disassembly code is narrowed, and the efficiency of the disassembly tool is improved.

[0075] The network interface 320 is used to receive the target file and store the target file in the memory 330. Alternatively, the network interface 320 can also be used to send the disassembled code obtained by the processor 310.

[0076] It is understandable that in the embodiments of the present application, the disassembly tool can be a separate disassembler or a tool that integrates a disassembler. Among them, the tool that integrates a disassembler can be a complete tool chain that includes a disassembler. For example, a program analyzer, a binary translator, and a code security & vulnerability detector. Taking the code security & vulnerability detector as an example, the code security & vulnerability detector first calls the disassembler to disassemble the target binary code into assembly code, and then processes the assembly code, for example, converting it into a certain intermediate language & representation. Then, security & vulnerability analysis and detection are performed based on the obtained intermediate language & representation.

[0077] The structure shown in FIG3 of the embodiment of the present application does not constitute a specific limitation on the disassembly tool. The disassembly tool may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.

[0078] For example, based on the disassembly tool in the above embodiment, the present application also provides a disassembly method, which can be executed by the processor in the disassembly tool shown in Figure 3. Figure 4 shows a flow chart of a disassembly method. As shown in Figure 4, the method includes: steps 401 to 407.

[0079] Step 401: Obtain the target file.

[0080] In this embodiment, the target file refers to the binary code or machine code generated by the source code through the compiler and can be directly recognized by the central processing unit (CPU). The target file includes the machine code that can be directly processed by the CPU and the data used by the code at runtime. For example, relocation information, program symbols used for linking or debugging, and other debugging information. The obtained target file can be the binary code or machine code input by the user, or it can be the binary code or machine code obtained from a given program.

[0081] It's understandable that when a computer runs a piece of code, it needs to convert it into an executable file. This process involves converting code written in a high-level language into assembly code, converting the assembly code into machine code (target file), and linking multiple target files containing machine code into a single executable file. Disassembly refers to the process of converting machine code into assembly code.

[0082] Step 402: disassemble the target file line by line to obtain a first disassembled code.

[0083] In this embodiment, after the target file is acquired, linear disassembly may be performed on the target file, that is, instructions in the target file may be disassembled line by line.

[0084] Step 403 , determining whether there is garbled code in the first disassembled code, if so, executing step 404 , otherwise executing step 406 .

[0085] In this embodiment, the assembly code can be composed of instructions and embedded data. For example, in the process of opening a PPT through program A, the data used to execute the instruction "open PPT" can be called instructions, and the content of the "PPT" can be called embedded data.

[0086] During the disassembly process, the processor cannot distinguish between instructions and embedded data (both instructions and embedded data are represented as binary "0" and "1"). Therefore, when disassembling embedded data, the processor may obtain a combination of instructions and garbled code. If the binary representation of some embedded data is identical to that of an instruction, this embedded data may be disassembled into instructions. Instructions include both correct and incorrect instructions. Incorrect instructions are instructions that conform to the correct assembly syntax but cannot be executed correctly or produce incorrect results. Incorrect execution can mean that the program address or data corresponding to the instruction is incorrect during execution. For embedded data that may be disassembled into data other than instructions, garbled code is generated during disassembly—code that is, chaotic and unreadable by the machine. Garbled code refers to the inability to generate the corresponding binary code or machine code of the instruction by searching the assembly manual during the disassembly process.

[0087] In one possible example, as shown in Figure 5, the processor disassembles the target file in Figure 4 to obtain the corresponding disassembled code. After obtaining the disassembled code contained in BBO in Figure 5, the processor continues to disassemble the target file. At this time, the disassembled data in the target file is inline data. Because instructions and data are displayed in binary form in the target file, inline data will also be disassembled into assembly instructions during the disassembly process, causing the BB1 ​​instruction to fail abnormally.

[0088] Step 404 : Go back one byte or multiple different bytes from the first target address in the target file and re-disassemble line by line to obtain at least one second disassembled code.

[0089] In this embodiment, the first target address is the address corresponding to the last byte in the garbled code or the address corresponding to the current address in the disassembled machine code / binary file.

[0090] In a possible example, during the disassembly of a target file, after each assembly instruction is generated, the processor determines whether the assembly instruction is a correct assembly instruction or is garbled. If the processor determines that the generated assembly instruction is garbled, the processor needs to obtain the address corresponding to the last byte in the garbled code or the address corresponding to the address of the current disassembly in the target file (the first target address), and record it.

[0091] It can be understood that during the disassembly of the target file, the processor can make a judgment every time an assembly instruction is obtained, or can make a judgment after obtaining multiple assembly instructions. The embodiments of the present application do not limit this.

[0092] In a possible example, as shown in FIG. 6, assume that "08AF 176B BBAC BBFF ACAA" in FIG. 6 is the target code to be disassembled. Among them, every four bytes identify an instruction or a character. For example, "08AF" identifies an instruction, and "BBFF ACAA" is the internal data. "BBFF" identifies the character "I", and "ACAA" identifies the character "love".

[0093] During the disassembly of the target file, the processor sequentially obtains 4-byte data and performs disassembly. As shown in ① in FIG. 6, when the processor obtains "BBFF", it disassembles "BBFF" and judges the disassembly result. In the case where the disassembly result corresponding to "BBFF" is garbled, it is necessary to determine the starting address (the first target address) for backward movement in the target file. As shown in ② in FIG. 6, the processor determines that the first target address is the 16th byte in the target file, that is, the last byte in "BBFF". At this time, the processor can backward move one or more different bytes from the first target address.

[0094] Exemplarily, ③ in FIG. 6 shows that after backward moving 3 bytes from the first target address, disassembly starts again. At this time, starting from the 14th byte, 4 bytes are sequentially obtained for disassembly.

[0095] ④ in FIG. 6 shows that after backward moving 2 bytes from the first target address, disassembly starts again. At this time, starting from the 15th byte, 4 bytes are sequentially obtained for disassembly.

[0096] ⑤ in FIG. 6 shows that after backward moving 1 byte from the first target address, disassembly starts again. At this time, starting from the 16th byte, 4 bytes are sequentially obtained for disassembly.

[0097] In one possible example, assume that the identification information of the binary code in the target file is: apple is abcdef lag. Among them, "apple", "is", and "lag" are all correct assembly words, and "abcdef" is a piece of data. However, during the disassembly process, the identification information of the obtained disassembly code may be: apple is abcde flag, among which "apple", "is", and "flag" are all correct assembly words, and "abcde" is a piece of data. Obviously, the assembly instruction corresponding to "flag" has been garbled, and the correct assembly instruction should be the assembly instruction corresponding to "lag". At this time, it can be determined that the address where the garbled code occurs in the target file (the first target address) is the address where the binary code corresponding to "flag" is located.

[0098] In one possible example, after determining the first target address (the xth byte) in the target file, the processor can backtrack 1 byte, 2 bytes, ..., m bytes from the first target address. Then, the target file is disassembled starting from the x-1th byte, the x-2th byte, ..., the xmth byte, respectively, to obtain m disassembly tasks, where m is a natural number greater than 1. By executing these m disassembly tasks, m versions of disassembly code are generated. For example, after obtaining the first target address in the target file, the processor backtracks 1 byte, 2 bytes, and 3 bytes from the first target address in the target file to perform re-disassembly. Therefore, when re-disassembling the same target file, the processor needs to execute three disassembly tasks. These three disassembly tasks can be executed in parallel or sequentially. For example, if the processor is a multi-core processor, these three disassembly tasks can be executed in parallel by multiple threads in the processor to obtain three versions of disassembly code. The processor then selects the version without garbled characters from the three obtained disassembly codes and saves it.

[0099] When the processor re-disassembles from the first target address of the target file, the processor may first roll back 1 byte from the first target address in the target file to re-disassemble, thereby obtaining a first version of the disassembled code. Then, the processor re-disassembles from the first target address in the target file by 2 bytes and 3 bytes, respectively, to obtain a second version of the disassembled code and a third version of the disassembled code. After obtaining each version of the disassembled code, the processor needs to determine whether there is garbled code in the disassembled code and save the disassembled code that does not contain garbled code.

[0100] In one possible example, when the processor rewinds 1 byte, 2 bytes, and 3 bytes in sequence from the first target address in the target file to re-disassemble, if garbled characters appear in the disassembled code obtained during each re-disassembly process, the processor can stop the disassembly and discard the disassembly result obtained by the disassembly. For example, when the processor rewinds 1 byte from the first target address in the target file to disassemble, if it is determined that garbled characters appear in the disassembled code obtained, the processor stops the disassembly. Then, the processor rewinds two bytes from the first target address in the target file to re-disassemble, and obtains the second version of the disassembly code corresponding to the rewind two bytes. If there is no garbled character in the second version of the disassembly code, the processor saves the second version of the disassembly code. Then, the processor continues to rewind 3 bytes from the first target address in the target file and re-disassembles to obtain the third version of the disassembly code corresponding to the rewind 3 bytes, and if there is no garbled character in the third version of the disassembly code, the processor saves the third version of the disassembly code.

[0101] In the embodiment of the application, during the re-disassembly process, the processor only retains the version without garbled characters from the multiple disassembly code versions obtained, thereby narrowing the range of selecting the correct disassembly code and improving the efficiency of the disassembly tool. It is understandable that in the embodiment of the present application, the disassembly code version without garbled characters is not equivalent to the correct disassembly code corresponding to the target file. The indication of the disassembly code without garbled characters indicates that there is no data or instructions in the disassembly code that affect the correctness of the disassembly.

[0102] In one possible example, the processor re-disassembles by rolling back 1 byte, 2 bytes, and 3 bytes starting from the first target address in the target file, and three versions of the disassembly results can be obtained. For example, Figure 7 shows the three versions of the disassembly results obtained after the processor rolls back 1 byte, 2 bytes, and 3 bytes from the first target address in the target file. Referring to Figure 7, version 1 corresponds to the disassembly result obtained by rolling back 1 byte, version 2 corresponds to the disassembly result obtained by rolling back 2 bytes, and version 3 corresponds to the disassembly result obtained by rolling back 3 bytes. Among them, the boxes labeled "1" in each version represent instructions generated by disassembly, the boxes labeled 2 represent data or instructions that do not affect the correctness of disassembly, and the boxes labeled "3" represent abnormal disassembly codes. Abnormal disassembly codes include garbled codes or erroneous instructions generated by disassembly. In the embodiment of the present application, "erroneous instructions" refer to instructions that conform to the correct assembly syntax, but the instructions happen to be composed of "data" plus "bytes of a portion of the real instruction".

[0103] In one possible example, inline data may include "dummy code." Generally, even if disassembled, this "dummy code" in the inline data does not affect the correctness of the generated disassembled code. The correctness of the generated disassembled code is affected only when the "dummy code" occurs at the boundary of the inline data.

[0104] Step 405 : Filter out target disassembly code from at least one second version of disassembly code.

[0105] In this embodiment, after the processor obtains at least one version of the disassembled code, the processor also needs to determine a correct disassembled code version from the obtained at least one version of the disassembled code.

[0106] In one possible example, the processor may determine the correct disassembly code version by comparing the number of instructions in each version of the disassembly code, for example, selecting the disassembly code version with the largest number of instructions from at least one version of the disassembly code as the correct disassembly code version.

[0107] In one possible example, when a jump instruction exists in all of the obtained multiple disassembly code versions, since the correct version of the disassembly code will certainly contain the second target address of the jump instruction, the disassembly code version in which the target address pointed to by the jump instruction is correct can be selected as the correct disassembly code version. Wherein, the second target address pointed to by the jump instruction is correct means that the second target address pointed to by the jump instruction exists, and by executing the disassembly code, the second target address can be correctly jumped to.

[0108] In one possible example, in order to improve the accuracy of selecting the correct disassembly code version from multiple disassembly code versions, the processor can simultaneously consider the number of instructions in the disassembly code of each version and whether the second target address of the jump instruction is correct when selecting the correct disassembly code version from multiple disassembly code versions. For example, the processor can select k disassembly code versions with the largest number of instructions from multiple disassembly code versions, where k is a natural number greater than or equal to 2. Then, the processor selects the disassembly version with the correct second target address pointed to by the jump instruction from the k disassembly code versions.

[0109] In one possible example, when selecting a target disassembly code from the at least one second disassembly code, the at least one second disassembly code may be directly assembled, the assembled code obtained may be compared with the machine code / binary file, and the second disassembly code corresponding to the assembly code that matches the machine code / binary file may be selected as the target disassembly code. The assembled assembly code matching the machine code / binary file means that the assembled assembly code and the machine code / binary file implement the same function.

[0110] When selecting a target disassembly code from the at least one second disassembly code, the assembled assembly code may be executed, and the second disassembly code corresponding to the successfully executed assembly code may be used as the target disassembly code. Directly assembling the at least one second disassembly code and determining the target disassembly code based on the assembly result reduces the difficulty of selecting the target disassembly code while also ensuring the accuracy of the selected target disassembly code.

[0111] Step 406: Use the first disassembled code as the target disassembled code.

[0112] In this embodiment, the processor needs to make a real-time judgment on the disassembly result during the process of disassembling the target code. If there is no garbled code in the disassembly result, the processor will disassemble the target file according to the current disassembly order and obtain the disassembly result (first disassembly code). It is understandable that if the disassembly code obtained during the disassembly process of the target file does not contain garbled code, then only one version of the disassembly code will be obtained in the end, which is the target disassembly code.

[0113] Step 407: Output the target disassembly code.

[0114] In the embodiment of the present application, when disassembling the target file, if garbled characters appear, the code is rewound one or more bytes and disassembled again to obtain one or more versions of the disassembled code. Then, the only correct version is selected from the multiple versions of the disassembled code. This achieves the goal of achieving higher code coverage with less computing power and memory.

[0115] It is understood that the order of execution of the steps in the above embodiments does not necessarily imply a specific order of execution. The order of execution of each process should be determined by its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. In addition, in some possible implementations, the steps in the above embodiments can be selectively executed according to actual circumstances, and can be executed partially or completely, which is not limited here.

[0116] Based on the method in the above embodiment, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a processor, the processor executes the method in the above embodiment.

[0117] Based on the method in the above embodiment, an embodiment of the present application provides a computer program product, characterized in that when the computer program product runs on a processor, the processor executes the method in the above embodiment.

[0118] The method steps in the embodiments of the present application can be implemented by hardware or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, mobile hard disks, CD-ROMs or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an ASIC.

[0119] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted via the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid state drive (SSD)).

[0120] It will be understood that the various numerical numbers involved in the embodiments of the present application are merely distinctions for the convenience of description and are not intended to limit the scope of the embodiments of the present application.

Claims

1. A disassembly method, characterized in that: include: Disassembling the obtained machine codes / binary files one by one to obtain a first disassembled code; In the case where there are garbled codes in the first disassembled code, back off one byte or multiple different bytes from the first target address in the machine code / binary file and re-disassemble one by one to obtain at least one second disassembled code, wherein the first target address is the address corresponding to the last byte in the garbled code or the address corresponding to the current address of the disassembly in the machine code / binary file; Filtering out a target disassembled code from the at least one second disassembled code; Output the target disassembled code.

2. The method according to claim 1, characterized in that The step of filtering out a target disassembly code from the at least one second disassembly code comprises: According to the number of instructions contained in each code in the at least one second disassembled code and / or the second target address of the jump instruction jump, a target disassembled code is screened out from the at least one second disassembled code.

3. The method according to claim 2, characterized in that The step of filtering out a target disassembly code from the at least one second disassembly code according to the number of instructions contained in each code in the at least one second disassembly code comprises: A second disassembly code with the largest number of instructions is selected from the at least one second disassembly code as the target disassembly code.

4. The method according to claim 2, characterized in that: The step of selecting a target disassembly code from the at least one second disassembly code according to a second target address of a jump instruction jump included in each code in the at least one second disassembly code comprises: A second disassembly code having a correct second target address of the jump instruction is selected from the at least one second disassembly code as the target disassembly code.

5. The method according to claim 2, characterized in that: The step of selecting a target disassembly code from the at least one second disassembly code according to the number of instructions contained in each code in the at least one second disassembly code and the second target address of the jump instruction jump comprises: Select n second disassembly codes with the largest number of instructions from the at least one second disassembly code, where n is a natural number greater than or equal to 2; A second disassembly code having a correct second target address of the jump instruction is selected from the n second disassembly codes as the target disassembly code.

6. The method according to any one of claims 1 to 5, characterized in that: Before filtering out a target disassembled code from the at least one second disassembled code according to the number of instructions contained in each code in the at least one second disassembled code and / or the second target address of the jump instruction jump, the method further includes: Filter out the second disassembly code that does not contain garbled code from the at least one second disassembly code to obtain k second disassembly codes, where k is a natural number greater than or equal to 1; A target disassembly code is screened out from the k second disassembly codes.

7. The method according to claim 1, characterized in that The step of filtering out a target disassembly code from the at least one second disassembly code comprises: Assembling the at least one second disassembled code to obtain the at least one third assembly code; According to the obtained machine code / binary file, a target third assembly code is selected from the at least one third assembly code, and the second disassembly code corresponding to the target third assembly code is used as the target disassembly code; or Execute the at least one third assembly code, select a target third assembly code from the at least one third disassembly code according to the execution result of the at least one third assembly code, and use the second disassembly code corresponding to the target third assembly code as the target disassembly code, so that the target third assembly code can be correctly executed.

8. The method according to any one of claims 1 to 7, characterized in that: The step of retracing a plurality of different bytes from the first target address in the machine code / binary file and re-disassembling the code line by line to obtain at least one second disassembly code includes: Starting from the first target address, different bytes are respectively rolled back to generate multiple disassembly tasks; The multiple disassembly tasks are executed in parallel to obtain at least one second disassembly code.

9. The method according to any one of claims 1 to 8, characterized in that: The method further comprises: When there is no garbled code in the first disassembled code, the first disassembled code is used as the target disassembled code.

10. A disassembly device, characterized in that: include: at least one processor and memory, The memory is used to store computer instructions; The at least one processor is configured to perform the following steps according to the computer instructions: Disassembling the obtained machine codes / binary files one by one to obtain a first disassembled code; In the case where there are garbled codes in the first disassembled code, back off one byte or multiple different bytes from the first target address in the machine code / binary file and re-disassemble one by one to obtain at least one second disassembled code, wherein the first target address is the address corresponding to the last byte in the garbled code or the address corresponding to the current address of the disassembly in the machine code / binary file; Filtering out a target disassembled code from the at least one second disassembled code; Output the target disassembled code.

11. The device according to claim 10, characterized in that The at least one processor is configured to: According to the number of instructions contained in each code in the at least one second disassembled code and / or the second target address of the jump instruction jump, a target disassembled code is screened out from the at least one second disassembled code.

12. The device according to claim 11, characterized in that The at least one processor is configured to: A second disassembly code with the largest number of instructions is selected from the at least one second disassembly code as the target disassembly code.

13. The device according to claim 11, characterized in that The at least one processor is configured to: A second disassembly code having a correct second target address of the jump instruction is selected from the at least one second disassembly code as the target disassembly code.

14. The device according to claim 11, characterized in that The at least one processor is configured to: Select n second disassembly codes with the largest number of instructions from the at least one second disassembly code, where n is a natural number greater than or equal to 2; A second disassembly code having a correct second target address of the jump instruction is selected from the n second disassembly codes as the target disassembly code.

15. The device according to any one of claims 10 to 14, characterized in that: The at least one processor is configured to: Filter out the second disassembly code that does not contain garbled code from the at least one second disassembly code to obtain k second disassembly codes, where k is a natural number greater than or equal to 1; A target disassembly code is screened out from the k second disassembly codes.

16. The device according to claim 10, characterized in that The at least one processor is configured to: Assembling the at least one second disassembled code to obtain the at least one third assembly code; According to the obtained machine code / binary file, selecting a target third assembly code from the at least one third assembly code, and using the second disassembly code corresponding to the target third assembly code as the target disassembly code; or, Execute the at least one third assembly code, select a target third assembly code from the at least one third disassembly code according to the execution result of the at least one third assembly code, and use the second disassembly code corresponding to the target third assembly code as the target disassembly code, so that the target third assembly code can be correctly executed.

17. The device according to any one of claims 10 to 16, characterized in that: The at least one processor is configured to: Starting from the first target address, different bytes are respectively rolled back to generate multiple disassembly tasks; The multiple disassembly tasks are executed in parallel to obtain at least one second disassembly code.

18. The device according to any one of claims 10 to 17, characterized in that: The at least one processor is further configured to: When there is no garbled code in the first disassembled code, the first disassembled code is used as the target disassembled code.

19. A computer-readable medium, wherein instructions are stored in the computer storage medium, and when the instructions are executed on a computer, the computer executes the method according to any one of claims 1 to 9.

20. A computer program product comprising instructions, which, when executed on a computer, causes the computer to perform the method according to any one of claims 1 to 9.