A co-processor firmware parsing method and apparatus

The Hoffman dictionary table is constructed through Gaussian elimination method and heuristic method, combined with the Hoffman decompression algorithm, and automatically analyzing the coprocessor firmware module, solving the problem of coprocessor firmware analysis and realizing the automated analysis and content recovery of the latest version of coprocessor firmware.

CN115756964BActive Publication Date: 2025-07-08Chinese People's Liberation Army Cyberspace Force Information Engineering University
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211469737.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-22
Publication Date
2025-07-08
Estimated Expiration
2042-11-22

AI Technical Summary

Technical Problem

It is difficult for the prior art to effectively parse coprocessor firmware in modern microprocessors, especially Intel ME firmware of V11 or above, which has problems such as large structural changes and difficulty in parsing information.

Method used

The Hoffman dictionary table is constructed using Gaussian elimination method and heuristic method, combined with the Hoffman decompression algorithm, and automatically analyzing the coprocessor firmware module. The Hoffman encoded value is solved through the Gaussian elimination method, and the unknown encoded value is restored with heuristic strategy to realize the information extraction and analysis of the coprocessor firmware.

Benefits of technology

实现了对最新版本协处理器固件的自动化解析,恢复压缩模块的文件内容,为进一步的逆向分析和模拟仿真提供支持。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115756964B_ABST
    Figure CN115756964B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of firmware reverse analysis and security, and particularly relates to a method and device for parsing coprocessor firmware. The method includes: first, using Gaussian elimination method and heuristic method to construct a Huffman dictionary table used for recovering the compression algorithm of coprocessor firmware; then extracting the BIOS firmware from the target motherboard, automatically decompressing the BIOS firmware and extracting the coprocessor firmware therefrom as the firmware to be parsed; secondly, scanning the coprocessor firmware to obtain the firmware partition table, traversing the firmware partition table, and respectively extracting modules from the code type partition and the data type partition; finally, traversing all the extracted modules, decompressing according to the compression algorithms used by the modules respectively and storing the decompression results. Based on the method for constructing the Huffman dictionary table by Gaussian elimination method and heuristic method, the present invention realizes the decompression of coprocessor firmware modules, and finally completes the information extraction and parsing of coprocessor firmware.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of firmware reverse analysis and security, and particularly relates to a method and device for parsing coprocessor firmware, which are mainly applied to firmware reverse analysis, firmware vulnerability mining, etc. Background Art

[0002] With the development of trusted computing technology, in the X86 platform, in addition to the x86 core, the CPU also has multiple embedded microcontrollers. Among them, the security coprocessor has been increasingly used. Currently, most modern microprocessors include such processors, such as Intel's ME (Management Engine), AMD's PSP (Platform Security Processor), and Apple's T2 processor. These processors are generally used for system initialization or assisting the main operating system to execute power management tasks during operation. In addition, they also act as a TPM to provide a trusted execution environment or act as a system trust anchor and other functions. However, the structures of these coprocessors and the firmware they run are not open to the public (a coprocessor is a chip used to relieve the specific processing tasks of the system microprocessor). According to the research results of relevant researchers, Intel ME can fully access and control the PC, start and shut down the computer, read open files, check all running programs, track key presses and mouse movements, and even capture screenshots. It also has a network interface that has been proven to be insecure, allowing attackers to implant rootkit programs to control and invade the computer.

[0003] Currently, most research on ME focuses on the lower versions of ME. With the continuous development of the processor architecture and manufacturing process, the supporting ME chips and firmware are also changing. The ME firmware version has evolved from V11 to V16. Moreover, through testing, it is found that the firmware structure of versions above V11 has changed greatly, and there are also significant differences between different versions, which poses great difficulties and higher requirements for conducting reverse analysis of ME firmware. In December 2020, Intel officially released the security white paper of Intel CSME for the first time, systematically introducing the security mechanisms and countermeasure technologies introduced by CSME, but not disclosing details such as ME firmware, which makes it a great challenge to process encrypted or compressed coprocessor firmware, especially to restore the dictionary or reference table used for compressing content. Summary of the Invention

[0004] Aiming at the problems existing in the prior art, the present invention proposes a method and device for parsing coprocessor firmware, which realizes decompression of coprocessor firmware modules based on a Huffman dictionary table construction method of Gaussian elimination method and heuristic method, and finally completes information extraction and parsing of coprocessor firmware.

[0005] To achieve the above object, the present invention adopts the following technical solutions:

[0006] The present invention provides a method for parsing coprocessor firmware, comprising the following steps:

[0007] Use Gaussian elimination method and heuristic method to construct the Huffman dictionary table used in the recovery coprocessor firmware compression algorithm;

[0008] Extract the BIOS firmware from the target motherboard, automatically decompress the BIOS firmware and extract the coprocessor firmware therefrom as the firmware to be parsed;

[0009] Scan the coprocessor firmware, obtain the firmware partition table, traverse the firmware partition table, and perform module extraction on the code type partition and data type partition respectively;

[0010] Traverse all the extracted modules, and decompress and store the decompression results respectively according to the compression algorithms used by the modules.

[0011] Further, constructing the Huffman dictionary table used in the recovery coprocessor firmware compression algorithm specifically includes:

[0012] Step 11, determine the range of the codeword lengths of the Huffman dictionary table used by the coprocessor firmware;

[0013] Step 12, extract the modules with the same name in the coprocessor firmware that adopt different compression algorithms, calculate the hash of the module after decompression using LZMA, and compare the hash values with the modules that adopt Huffman compression in the modules with the same name, and store the modules with the same hash values in the paired module table;

[0014] Step 13, construct the range of the codeword values under the same length respectively according to the length range in Step 11;

[0015] Step 14, according to the paired module table in Step 12, create a system of linear equations, each page corresponds to a system of equations, the coefficients of the system of equations are the number of occurrences of the code sequences in the compressed page, the length of the coding value is the unknown, and the free term is the page size;

[0016] Step 15, perform elementary transformation on the system of linear equations in Step 14, use Gaussian elimination method to solve the transformed matrix, and obtain the length values of each coding value;

[0017] Step 16, referring to the plaintext values in the paired module table, according to the length and sequence order of the coding values, obtain the corresponding relationship between the code sequences and the coding values;

[0018] Step 17, for the sequences that cannot be restored in Steps 11 - 16, use heuristic strategies to determine the coding values of the unknown code sequences;

[0019] Step 18, store the codeword value and the corresponding coding value into the Huffman dictionary table in sequence.

[0020] Furthermore, the following heuristic strategy is used in Step 17 to determine the coding value of the unknown code sequence:

[0021] Step 171, since there are two Huffman dictionary tables in the unpacker, use these two tables to compress the data, compare the sizes of the compressed data, and retain the Huffman dictionary table that occupies less space; and view different versions of the same module, find the same segment packed in the other table, and restore the unknown byte after comparison.

[0022] Step 172, find the code sequences that appear multiple times in the same or different code and data modules, and determine the constraints imposed on the unknown values according to the sequence changes.

[0023] Step 173, extract the text strings or function constants and offsets in all modules and store them. For modules of the same version, obtain the code or data segments to which they are applied according to the offset values, and restore the coding value corresponding to the code sequence after comparison.

[0024] Step 174, analyze the string constants of the open source library, and restore the coding value corresponding to the text string segment through context and source code information.

[0025] Step 175, analyze the source code of the open library, find the code text corresponding to the function, then compile the source code into a binary file, and restore the coding value corresponding to the function by comparing the function binary information.

[0026] Step 176, compare different versions of the same module, find the equivalent functions, and restore the coding value of the unknown module through the coding values that have been restored in the module.

[0027] Furthermore, scan the coprocessor firmware with the keyword string as the feature, obtain the firmware partition table, extract the partition number, offset, size, and type information of the firmware partition table, and store them in the partition table data structure.

[0028] Furthermore, traverse all the partitions in the firmware partition table, judge the partition type. If the partition type is code, execute Steps 21 to 26.

[0029] Step 21, parse the header information of the code partition directory and store it in the partition directory table data structure.

[0030] Step 22, parse the data information of the code partition directory, extract the structure information of each module under this partition, and stop extracting until the number of analyzed modules is the same as the number of modules in Step 21.

[0031] Step 23: Traverse all modules and determine the compression algorithm type in the module structure information according to the value of the compression algorithm field;

[0032] Step 24: If the value of the compression algorithm field is None, directly store the file content in binary format locally according to the module size;

[0033] Step 25: If the value of the compression algorithm field is LZMA, extract the file content in binary format according to the module size, and then call the LZMA decompression algorithm to parse the file content and store it locally;

[0034] Step 26: If the value of the compression algorithm field is Huffman, extract the file content in binary format according to the module size, and then call the Huffman decompression algorithm to parse the file content and store it locally.

[0035] Furthermore, the header information of the code partition directory includes the number of partition modules and the partition name; the data information of the code partition directory includes the compression algorithm, offset, compressed module size, and decompressed module size.

[0036] Furthermore, the Huffman decompression algorithm in Step 26 further includes:

[0037] Step 261: Calculate the number of pages occupied by the compressed module;

[0038] Step 262: Establish a module offset information table;

[0039] Step 263: Traverse the offset information table, determine whether it is the last item. If so, execute Step 267; otherwise, execute Step 264;

[0040] Step 264: Extract the information of the offset parts of the current item and the next item, locate the position of the compressed page, and use the information of the offset part of the next item as the page compression size;

[0041] Step 265: Extract the compressed module content bit by bit and store it in a string. As the size of the module to be decompressed decreases bit by bit, determine whether the string value is consistent with the codeword in the Huffman dictionary table. If it is consistent, extract the encoded value corresponding to the codeword from the Huffman dictionary table and append it to the decompressed file, and clear the string value; execute Step 266;

[0042] Step 266: Determine whether the size of the module to be decompressed is 0. If so, stop decompressing this page and execute Step 263; otherwise, execute Step 265;

[0043] Step 267: Extract the content of the compression module bit by bit and store it in a string. Then, determine whether the string value at this time is the same as the code word in the Huffman dictionary table. If they are the same, extract the coding value corresponding to the code word from the Huffman dictionary table and append it to the decompressed file, and clear the string value; then execute Step 268;

[0044] Step 268: Determine whether the number of bytes occupied by the decompressed code is the page size value. If it is, stop the decompression of this module; otherwise, execute Step 267.

[0045] Furthermore, traverse all partitions in the firmware partition table, determine the partition type. If the partition type is data, execute Steps 31 to 38;

[0046] Step 31: Calculate the number of pages in the data partition and locate the starting position of the first page of the data partition;

[0047] Step 32: Analyze the directory header information of the data partition. According to the flag byte, determine whether the page is the last page of the partition. If it is, execute Step 38; otherwise, extract the page type field value and execute Step 33;

[0048] Step 33: Determine whether the page type field value is 0. If it is, set the page type to the system page and execute Step 34; otherwise, set the page type to the data page and execute Step 35;

[0049] Step 34: Locate the position of the next page according to the partition offset and execute Step 32;

[0050] Step 35: Calculate the number of page blocks and the position of the first block;

[0051] Step 36: Analyze the file allocation table in sequence according to the block size, and determine whether there is a file currently. If there is, further recursively analyze the field information according to the file type and store it; the number of blocks decreases;

[0052] Step 37: Locate the position of the next page or block. Determine whether the number of blocks is 0. If it is, locate the position of the next page according to the offset and execute Step 32; otherwise, enter the starting position of the next block according to the index and execute Step 36;

[0053] Step 38: Extract the header information, set the page type, and record the parsing result.

[0054] The present invention also provides a co-processor firmware parsing device, including:

[0055] A Huffman dictionary table construction module, which is used to construct the Huffman dictionary table for restoring the co-processor firmware compression algorithm by using the Gaussian elimination method and the heuristic method;

[0056] A firmware decompression module, which is used to extract the BIOS firmware from the target mainboard, automatically decompress the BIOS firmware and extract the coprocessor firmware therefrom as the firmware to be parsed;

[0057] A partition table traversal module, which scans the coprocessor firmware to obtain the firmware partition table, traverses the firmware partition table, and separately extracts modules from the code type partition and the data type partition;

[0058] A decompression module, which is used to traverse all the extracted modules, decompress and store the decompression results respectively according to the compression algorithms used by the modules.

[0059] Furthermore, the Huffman dictionary table construction module includes:

[0060] A hash value comparison module, which is used to calculate the hash values of the same-name modules using different compression algorithms. Among them, the module using the LZMA compression algorithm is decompressed first and then the hash value is calculated, and then the calculated hash values are compared to obtain a paired module table;

[0061] A Gaussian elimination module, which is used to create a system of linear equations, perform elementary transformations on the system of equations, and then use the Gaussian elimination method to solve the matrix, construct the length of the encoded value, and restore the encoded value of the code sequence with reference to the paired module table;

[0062] A heuristic module, which is used to further restore the encoded value of the unknown code sequence.

[0063] Compared with the prior art, the present invention has the following advantages:

[0064] The Huffman dictionary table construction method combining the Gaussian elimination method and the heuristic method proposed by the present invention is applicable to the restoration of the Huffman compressed content in the latest version of the coprocessor firmware, can automatically analyze the module composition of the coprocessor firmware and restore the file content corresponding to each compressed module, and provides important support for further reverse analysis and simulation. Description of the Drawings

[0065] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required to be used in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0066] Figure 1 It is a schematic flowchart of the coprocessor firmware parsing method with the partition type of code in the embodiment of the present invention;

[0067] Figure 2It is a schematic flowchart of a coprocessor firmware parsing method with the partition type being data in an embodiment of the present invention;

[0068] Figure 3 It is a schematic flowchart of the Huffman decompression algorithm in an embodiment of the present invention;

[0069] Figure 4 It is a schematic flowchart of restoring the Huffman dictionary table in an embodiment of the present invention. Detailed implementation manners

[0070] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0071] As Figure 1 and Figure 2 shown, this embodiment provides a firmware parsing method for an Intel coprocessor, aiming to analyze the composition of the coprocessor firmware, including the following steps:

[0072] Step S1, construct a Huffman dictionary table used for restoring the coprocessor firmware compression algorithm.

[0073] Step S2, obtain the coprocessor firmware.

[0074] The coprocessor firmware is generally stored together with the BIOS firmware. The coprocessor firmware is extracted by extracting the target motherboard firmware and using a firmware parsing tool (such as UEFITOOL) to parse the firmware.

[0075] Step S3, obtain the firmware partition table.

[0076] Scan the firmware binary file with the keyword "$FPT" or the hexadecimal value 0x24465054 as the feature to obtain the firmware Flash partition table, extract information such as the number of partition table partitions, partition offsets, sizes, and types, and store them in the partition table data structure.

[0077] Step S4, traverse all the partitions in the firmware partition table, judge the partition type. If the partition type is code, execute steps S5 - S7; if the partition type is data, execute steps S8 - S15.

[0078] Step S5, parse the code partition directory header information, extract information such as "$CPD", the number of partition modules, and the partition name, and store them in the partition directory table data structure.

[0079] Step S6, analyze the code partition directory data information, and extract the module structure information such as the compression algorithm, offset, compressed module size, decompressed module size, etc. of each module in this partition in units of 24 bytes. For each module structure information extracted, the module count is incremented by 1. When the module count is equal to the number of partition modules extracted in Step S5, stop and execute Step S7.

[0080] Step S7, traverse all modules, and judge the compression algorithm type in the module structure information according to the value of the compression algorithm field.

[0081] If the value of the compression algorithm field is None, calculate the actual position of the module in the firmware according to the offset in the module structure information, and then extract the module content starting from the actual position in byte order according to the module size information to form a module file; if the value of the compression algorithm field is LZMA, directly call the open-source LZMA decompression algorithm. Similarly, extract the module content starting from the actual position in byte order according to the compressed module size information as the input of the LZMA decompression algorithm, and then store the content output by the LZMA algorithm as a module file; if the value of the compression algorithm field is Huffman, call the Huffman decompression algorithm to decompress this module; execute Step S4.

[0082] Step S8, calculate the number of partition pages, divide the partition size in the partition table data structure by 0x2000 to solve the number of partition pages, and locate the position of the first page according to the partition offset.

[0083] Step S9, analyze the data partition directory header information, and judge whether the flag byte is "AA557887". If so, extract the first chunk field information and execute Step S10; otherwise, execute Step S15.

[0084] Step S10, judge whether the value of the first chunk field is 0. If so, set the page type to "system" and execute Step S11; otherwise, set the page type to "data" and execute Step S12.

[0085] Step S11, verify the crc integrity of each chunk in this page, locate the position of the next page according to the partition offset, and execute Step S9.

[0086] Step S12, calculate the number of chunks in the page and the offset of the first chunk.

[0087] Step S13, judge whether there is a file in the current directory according to the file allocation table value. If so, further judge the file type, recursively analyze all field information according to the defined structure and store it; otherwise, skip it; decrement the number of chunks by 1 and execute Step S14.

[0088] In step S14, it is determined whether the number of chunks is equal to 0. If so, the next page position is located according to the partition offset, and step S9 is executed; otherwise, it jumps to the starting position of the next chunk according to the index, and step S13 is executed.

[0089] In step S15, the header information is extracted, the page type is set to "Scratch", the parsing result is recorded, and the parsing work is ended.

[0090] As Figure 4 shown, step S1 further includes:

[0091] In step S101, the composition of the Huffman dictionary table of the lower version is analyzed to determine the codeword length range in the Huffman dictionary table used by the higher version firmware.

[0092] In step S102, the same-name modules using different compression algorithms in the coprocessor firmware are extracted, the modules compressed using the LZMA algorithm are decompressed, then the SHA256 of the decompressed modules is calculated, and the SHA256 values are compared with the modules compressed using Huffman in the same-name modules. The modules with the same values are stored in the paired module table.

[0093] In step S103, the range (boundary) of codeword values with the same length is constructed respectively according to the length range determined in step S101.

[0094] In step S104, according to the paired module table in step S102, a system of linear equations is created. Each page corresponds to a system of equations. The coefficients of this system of equations are the number of occurrences of the code sequence in the compressed page, the length of the coding value is the unknown, and the free term takes the value of 4096.

[0095] In step S105, elementary transformations are performed on the system of linear equations in step S104, and the Gaussian elimination method is used to solve the transformed matrix to obtain the length values of each coding value.

[0096] In step S106, by referring to the plaintext values in the paired module table, according to the length and sequence order of the coding values, the corresponding relationship between the code sequence and the coding value is obtained.

[0097] In step S107, for the sequences that cannot be restored in steps S101 - S106, the following heuristic strategy is used to determine the coding values of the unknown code sequences.

[0098] In step S1071, since there are two Huffman dictionary tables in the unpacker, these two tables are used to compress the data, the sizes of the compressed data are compared, and the Huffman dictionary table that occupies less space is retained; and the different versions of the same module are viewed, the same fragments packed in the other table are searched, and the unknown bytes are restored after comparison.

[0099] Step S1072: Search for code sequences that appear multiple times in the same or different code and data modules, and determine the constraints imposed on the unknown values based on the sequence variations.

[0100] Step S1073: Extract constants such as text strings or functions and offsets in all modules and store them. For modules of the same version, obtain the code or data segments to which they apply based on the offset values, and restore the encoded values corresponding to the code sequences after comparison.

[0101] Step S1074: Analyze the string constants of open-source libraries such as WPA, and restore the encoded values corresponding to the text string segments through context and source code information.

[0102] Step S1075: Analyze the source code of the open library, search for the code text corresponding to the function, then compile the source code into a binary file, and restore the encoded value corresponding to the function by comparing the binary information of the function.

[0103] Step S1076: Compare different versions of the same module, find equivalent functions, and restore the encoded values of the unknown module through the encoded values already restored in the module.

[0104] Step S108: Store the encoded values restored in steps S101 - 107 into the Huffman dictionary table in sequence according to the length value, together with the codeword values and the corresponding encoded values.

[0105] As Figure 3 shown, the Huffman decompression algorithm in step S7 further includes:

[0106] Step S701: Calculate the number of pages occupied by the compressed module based on the decompressed size in the module structure information and the fixed size of each independent page.

[0107] Step S702: Establish a module offset information table, then calculate the actual position of the module in the firmware based on the offset in the module structure information, read data in 4-byte sizes from the actual position, store the first 2 bytes in the offset part of the offset information table, and judge the values of the last two bytes. If the value is 0x0040, it means this page is a code page, and store the code attribute in the attribute part of the offset information table. If the value is 0x00C0, it means this page is a data page, and store the data attribute in the attribute part of the offset information table.

[0108] Step S703: Traverse the offset information table, judge whether it is the last item. If so, execute step S707; otherwise, execute step S704.

[0109] Step S704: Extract the information of the offset parts of the current item and the next item, locate the position of the compressed page based on the information of the offset part of the current item, and use the information of the offset part of the next item as the page compression size.

[0110] Step S705, extract the content of the compression module one by one according to the bit size and store it in a string. Decrease the size of the compression module by 1, and determine whether the string value at this time is the same as the code word in the Huffman dictionary table. If they are the same, extract the encoding value corresponding to the code word from the Huffman dictionary table, and store it in the decompressed file in an appended form. Clear the string value; execute Step S706.

[0111] Step S706, determine whether the page compression size value at this time is 0. If it is, stop decompressing this page and execute Step S703; otherwise, execute Step S705.

[0112] Step S707, extract the content of the compression module one by one according to the bit size and store it in a string, and determine whether the string value at this time is the same as the code word in the Huffman dictionary table. If they are the same, extract the encoding value corresponding to the code word from the Huffman dictionary table, and store it in the decompressed file in an appended form. Clear the string value; execute Step S708.

[0113] Step S708, determine whether the number of bytes occupied by the decompressed encoding is 4096 bytes. If it is, stop decompressing this module; otherwise, execute Step S707.

[0114] Corresponding to the above method for parsing co-processor firmware, this embodiment also proposes a co-processor firmware parsing device, including a Huffman dictionary table construction module, a firmware decompression module, a partition table traversal module, and a decompression module.

[0115] The Huffman dictionary table construction module is used to construct the Huffman dictionary table used to restore the co-processor firmware compression algorithm by using the Gaussian elimination method and the heuristic method;

[0116] The firmware decompression module is used to extract the BIOS firmware from the target motherboard, automatically decompress the BIOS firmware, and extract the co-processor firmware from it as the firmware to be parsed;

[0117] The partition table traversal module scans the co-processor firmware, obtains the firmware partition table, traverses the firmware partition table, and extracts modules from the code type partition and the data type partition respectively;

[0118] The decompression module is used to traverse all the extracted modules, decompress and store the decompression results respectively according to the compression algorithms used by the modules.

[0119] Further, the Huffman dictionary table construction module includes a hash value comparison module, a Gaussian elimination module, and a heuristic module.

[0120] The hash value comparison module is used to calculate the hash values of modules with the same name using different compression algorithms. For the module using the LZMA compression algorithm, it is first decompressed and then the hash value is calculated. Then, the calculated hash values are compared to obtain the paired module table;

[0121] The Gaussian elimination module is used to create a system of linear equations, perform elementary transformations on this system of equations, then use the Gaussian elimination method to solve the matrix, construct the length of the coded value, and restore the coded value of the code sequence with reference to the paired module table;

[0122] The heuristic module is used to further restore the coded value of the unknown code sequence.

[0123] The present invention utilizes the Huffman dictionary table restoration and construction technology combining the Gaussian elimination method and the heuristic method to automatically analyze the coprocessor firmware, can decompress the compressed content, and is applicable to the new version of the coprocessor firmware.

[0124] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.

[0125] Finally, it should be noted that the above are only the preferred embodiments of the present invention, which are only used to illustrate the technical solutions of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are all included within the protection scope of the present invention.

Claims

1. A co-processor firmware parsing method, characterized in that, It includes the following steps: Step 1: Use Gaussian elimination and heuristic methods to construct the Huffman dictionary table used in the recovery co-processor firmware compression algorithm, specifically including: Step 11: Determine the range of Huffman dictionary table codeword lengths used by the co-processor firmware; Step 12: Extract the same-name modules in the co-processor firmware that use different compression algorithms, calculate the hash of the module after decompression using LZMA, and compare the hash values with the modules compressed using Huffman in the same-name modules. Store the modules with the same hash values in the paired module table; Step 13: Construct the range of codeword values for the same length according to the length range in Step 11; Step 14: Create a system of linear equations based on the paired module table in Step 12. Each page corresponds to a system of equations. The coefficients of this system of equations are the number of occurrences of the code sequence in the compressed page, the encoding value length is the unknown, and the free term is the page size; Step 15: Perform elementary transformations on the system of linear equations in Step 14, use Gaussian elimination to solve the transformed matrix, and obtain the length values of each encoding value; Step 16: According to the plaintext values in the paired module table, obtain the correspondence between the code sequence and the encoding value based on the encoding value length and sequence order; Step 17: For the sequences that cannot be restored in Steps 11-16, use heuristic methods to determine the encoding values of the unknown code sequences; Step 18: Store the codeword values and the corresponding encoding values into the Huffman dictionary table in sequence; Step 2: Extract the BIOS firmware from the target mainboard, automatically decompress the BIOS firmware and extract the co-processor firmware from it as the firmware to be parsed; Step 3: Scan the co-processor firmware to obtain the firmware partition table, traverse the firmware partition table, and extract modules from the code-type partition and data-type partition respectively; Step 4: Traverse all the extracted modules, decompress and store the decompression results according to the compression algorithms used by the modules; 2. The co-processor firmware parsing method according to claim 1, wherein The following heuristic method is used in Step 17 to determine the encoding values of the unknown code sequences: Step 171: Since there are two Huffman dictionary tables in the unpacker, use these two tables to compress data, compare the sizes of the compressed data, and retain the Huffman dictionary table that occupies less space; and view different versions of the same module, find the same segment packed in the other table, and restore the unknown bytes after comparison; Step 172: Find the code sequences that appear multiple times in the firmware code and data modules, and determine the constraints imposed on the unknown values according to the sequence changes; Step 173: Extract and store the constants and offsets in all modules. The constants are text strings or functions. For modules of the same version, obtain the code or data segments to which they apply according to the offset values, and restore the encoding values corresponding to the code sequences after comparison; Step 174: Analyze the string constants of the open source library, and restore the encoding values corresponding to the text string segments through context and source code information; Step 175: Analyze the source code of the open library, find the code text corresponding to the function, then compile the source code into a binary file, and restore the encoding value corresponding to the function by comparing the function binary information; Step 176: Compare different versions of the same module, find equivalent functions, and restore the encoding values of unknown modules through the encoding values restored by the module.

3. The co-processor firmware parsing method according to claim 1, wherein Scan the co-processor firmware with the keyword string as the feature, obtain the firmware partition table, extract the partition quantity, offset, size, and type information of the firmware partition table, and store them in the partition table data structure.

4. The co-processor firmware parsing method according to claim 3, wherein Traverse all partitions in the firmware partition table, determine the partition type. If the partition type is code, execute Steps 21 to 26. Step 21: Parse the header information of the code partition directory and store it in the partition directory table data structure. Step 22: Parse the data information of the code partition directory, extract the structure information of each module under this partition until the number of analyzed modules is the same as the number of modules in Step 21, then stop extraction. Step 23: Traverse all modules, and judge the compression algorithm type in the module structure information according to the value of the compression algorithm field. Step 24: If the value of the compression algorithm field is None, directly store the file content in binary form according to the module size to the local. Step 25: If the value of the compression algorithm field is LZMA, extract the file content in binary form according to the module size, and then call the LZMA decompression algorithm to parse the file content and store it to the local. Step 26: If the value of the compression algorithm field is Huffman, extract the file content in binary form according to the module size, and then call the Huffman decompression algorithm to parse the file content and store it to the local.

5. The co-processor firmware parsing method according to claim 4, characterized in that, The header information of the code partition directory includes the number of partition modules and the partition name; the data information of the code partition directory includes the compression algorithm, offset, compressed module size, and decompressed module size.

6. The co-processor firmware parsing method according to claim 4, wherein The Huffman decompression algorithm in Step 26 further includes: Step 261: Calculate the number of pages occupied by the compressed module. Step 262: Establish a module offset information table. Step 263: Traverse the offset information table, judge whether it is the last item. If so, execute Step 267; otherwise, execute Step 264. Step 264: Extract the information of the offset part of the current item and the next item, locate the position of the compressed page, and use the information of the next item's offset part as the page compression size. Step 265: Extract the compressed module content bit by bit and store it in a string. As the size of the module to be decompressed decreases bit by bit, judge whether the string value is the same as the code word in the Huffman dictionary table. If they are the same, extract the encoding value corresponding to the code word from the Huffman dictionary table and append it to the decompressed file, and clear the string value; execute Step 266. Step 266: Judge whether the size of the module to be decompressed is 0. If so, stop decompressing this page, execute Step 263; otherwise, execute Step 265. Step 267: Extract the compressed module content bit by bit and store it in a string, and judge whether the current string value is the same as the code word in the Huffman dictionary table. If they are the same, extract the encoding value corresponding to the code word from the Huffman dictionary table and append it to the decompressed file, and clear the string value; execute Step 268. Step 268, determine whether the number of bytes occupied by the decompressed code is the page size value. If so, stop decompressing this module; otherwise, execute Step 267.

7. The co-processor firmware parsing method according to claim 3, wherein Traverse all partitions in the firmware partition table, determine the partition type. If the partition type is data, execute Steps 31 to 38; Step 31, calculate the number of data partition pages and locate the starting position of the first page of the data partition; Step 32, parse the directory header information of the data partition, and determine whether the page is the last page of the partition according to the flag byte. If so, execute Step 38; otherwise, extract the page type field value and execute Step 33; Step 33, determine whether the page type field value is 0. If so, set the page type as a system page and execute Step 34; otherwise, set the page type as a data page and execute Step 35; Step 34, locate the position of the next page according to the partition offset and execute Step 32; Step 35, calculate the number of page blocks and the position of the first block; Step 36, parse the file allocation table in sequence according to the block size, determine whether there is a file currently. If so, further recursively analyze the field information according to the file type and store it; the number of blocks decreases; Step 37, locate the next page position or block position, and determine whether the number of blocks is 0. If so, locate the next page position according to the offset and execute Step 32; otherwise, enter the starting position of the next block according to the index and execute Step 36; Step 38, extract the header information, set the page type, and record the parsing result.

8. A co-processor firmware parsing device, characterized in that A device for implementing the co-processor firmware parsing method according to any one of claims 1-7, the device comprising: A Huffman dictionary table construction module for constructing a Huffman dictionary table used to restore the co-processor firmware compression algorithm using Gaussian elimination and heuristic methods; A firmware decompression module for extracting the BIOS firmware from the target motherboard, automatically decompressing the BIOS firmware and extracting the co-processor firmware therefrom as the firmware to be parsed; A partition table traversal module for scanning the co-processor firmware, obtaining the firmware partition table, traversing the firmware partition table, and separately extracting modules for code-type partitions and data-type partitions; A decompression module for traversing all the extracted modules, decompressing and storing the decompression results respectively according to the compression algorithms used by the modules.

9. The co-processor firmware parsing apparatus according to claim 8, wherein The Huffman dictionary table construction module includes: A hash value comparison module for calculating the hash values of modules with the same name using different compression algorithms, where the module using the LZMA compression algorithm is first decompressed and then the hash value is calculated, and then the calculated hash values are compared to obtain a paired module table; A Gaussian elimination module for creating a system of linear equations, performing elementary transformations on the system of equations, and then using Gaussian elimination to solve the matrix, construct the length of the encoded value, and restore the encoded value of the code sequence with reference to the paired module table; A heuristic module for further restoring the encoded value of the unknown code sequence.

Citation Information

Patent Citations

  • System and method for manipulating data with a plurality of processors

    CN1601511A

  • Dual mode data compression for operating code

    WO2002015408A2