An Open-Source Component Version Identification Method and Device for Binary Code Reuse

By selecting five version sensitive features and designing two-stage matching process, the problem of inaccurate identification of open source component versions in the existing technology is solved, which improves the recognition accuracy rate and enhances the reliability of security risk assessment.

CN114035794BActive Publication Date: 2025-06-10INSTITUTE OF INFORMATION ENGINEERING CHINESE ACADEMY OF SCIENCES
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111119984.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-07-30
Filing Date
2021-09-24
Publication Date
2025-06-10
Estimated Expiration
2041-09-24

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify open source component versions of closed source binary files reusable, especially when distinguishing different versions, resulting in inaccurate security risk assessment.

Method used

Five version sensitive features are adopted, including string features, export function features, constant features in assignment statements, function constant parameter features, and output and in-degree features in function call diagrams, and a two-stage matching process is designed to improve the accuracy of version recognition.

Benefits of technology

By selecting version-sensitive code characteristics, the recognition accuracy between binary code and different versions of source code is significantly improved, helping to evaluate the security risks of closed source binary software.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114035794B_ABST
    Figure CN114035794B_ABST
Patent Text Reader

Abstract

The present invention discloses an open-source component version identification method and device for binary code reuse, including: respectively extracting sensitive features of the target binary code and each version of the source code, where the sensitive features include: global constant features and function-level features; based on the sensitive features, calculating the similarity between the target binary code and each version of the source code, and obtaining the open-source component version identification result. Through the original setting rules, the present invention selects the code features sensitive to the version, thereby effectively improving the identification accuracy between the binary code and the source code of different versions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of code similarity detection, and in particular to a method and device for identifying the version of an open-source component for binary code reuse, specifically a method for identifying the version information of an open-source component for binary code reuse based on version-sensitive features for a given closed-source binary file. Based on the version information of the open-source component reused by the binary file, it is possible to know whether the open-source component is outdated and whether it contains public vulnerabilities, which helps to evaluate whether there are security risks in the closed-source binary software. Background Art

[0002] Given a certain closed-source binary file, identifying whether the binary file has code segments similar to a version with vulnerabilities in an open-source library is an important method for evaluating the security of closed-source software. Taking the open-source component Openssl as an example, 66% of the versions in Openssl have historical vulnerabilities. Identifying whether a certain closed-source software reuses the vulnerable version of Openssl can analyze whether there are security risks in the software. To identify the version of the open-source component reused by the binary file, code similarity analysis technology can be used, which can also be called code clone detection. Code clone refers to two or more identical or similar source code segments existing in a code library. Code similarity analysis technology generally consists of three parts, namely code feature extraction, code feature matching, and similarity calculation. Code similarity detection mainly focuses on the similarity detection between source code and source code or between binary and binary.

[0003] There are relatively few tools for detecting the similarity between binary code and source code, and the code features used by existing tools cannot effectively distinguish different versions of open-source components. For example, in "CN111045670A - A Method and Device for Identifying the Reuse Relationship between Binary Code and Source Code" and "CN111078227A - A Method and Device for Analyzing the Similarity between Binary Code and Source Code Based on Code Features", when comparing the similarity between binary code and source code, the features used are mainly program-global-level features, such as string features, function name features, global array features, constant features in switch cases, and constant features in if-else statements. These features help calculate the similarity degree between the target binary file and different open-source components, aiming to determine the set of open-source components reused by the target binary file. However, when further determining which specific version of an open-source component the target binary file has reused, some of these features have little discrimination between different versions, that is, they cannot be effectively used to distinguish different versions of open-source components. Taking the open-source component Openssl as an example, existing tools "CN111045670A" and "CN111078227A" can calculate whether a target binary file has reused Openssl, but they do not analyze which specific version of Openssl the target binary file has reused, and this is exactly the problem to be solved by the technical solution of the present invention.

[0004] Specifically, we extracted the above features from three open-source projects, freetype, sqlite, and libtiff, calculated the corresponding discrimination degrees, and presented them in Table 1. It can be seen from Table 1 that the constant features in switch cases and if-else statements have insufficient discrimination between different versions. The global constant array feature has relatively good discrimination, but since the global constant array exists in the form of a string of bitstreams in the binary file, the front and back boundaries of the bitstreams in the binary file cannot be effectively distinguished. Using the global constant array will cause multiple arrays of different versions to be found in the same binary file, resulting in the highest similarity between the binary code and the source code of the wrong version, leading to false alarms. For example, Figure 1 shows the array ft_extra_glyph_unicodes in version 2-6 of the open-source library freetype and its representation in the binary file freetype-VER-2-6.so. Figure 2 shows that a certain global array in two versions, 2-3-6 and 2-6, of the open-source library freetype can be found in the binary file freetype-VER-2-6.so. Finally, Table 1 shows that the string feature and the function name feature have relatively good discrimination between different versions, but relying solely on these two features is not sufficient to distinguish all versions.

[0005]

[0006] Table 1

[0007] To effectively identify the open-source component versions reused in closed-source binary software, aiming at the limitations of current tools in feature selection, the present invention proposes a method for selecting version-sensitive features, and based on the selected version-sensitive features, designs an effective similarity comparison scheme to calculate the similarity between binary code and source code of different versions. Furthermore, a two-stage matching process is designed to identify the open-source component versions reused in binary code. Summary of the Invention

[0008] To overcome the limitations of the current solution in the accuracy of version identification, the present invention discloses a method and device for identifying open-source component versions reused in binary code. Five version-sensitive features are selected, including string features, exported function features, constant features in assignment statements, constant parameters of function calls, and out-degree and in-degree features in function call graphs. They are further divided into two categories: global features and function-level features. Correspondingly, a two-stage matching process is designed to improve the accuracy of version identification.

[0009] The technical content of the present invention includes:

[0010] A method for identifying open-source component versions reused in binary code, the steps of which include:

[0011] 1) Extract the sensitive features of the target binary code and the source code of each version respectively, where the sensitive features include: global constant features and function-level features;

[0012] 2) Based on the sensitive features, calculate the similarity between the target binary code and the source code of each version, and obtain the open-source component version identification result.

[0013] Furthermore, the global constant features include: string features and exported function features.

[0014] Furthermore, the function-level features include: function-level constant features and function out-degree and in-degree features.

[0015] Furthermore, the function-level constant features include: constant features in assignment statements and constant parameters in function calls.

[0016] Furthermore, the string features of any version of the source code are extracted through the following steps:

[0017] 1) Use the compiler front-end clang to perform lexical analysis on the source code to generate tokens corresponding to each code fragment;

[0018] 2) Parse the tokens to generate string features.

[0019] Furthermore, the export function features and function-level features of any version of the source code are extracted through the following steps:

[0020] 1) Based on the grammar parser generator Antlr, generate a syntax tree corresponding to the code snippet;

[0021] 2) Parse the nodes of the syntax tree to obtain the export function features and function-level features.

[0022] Furthermore, the method for extracting sensitive features of the target binary code includes: parsing based on the interactive disassembler IDA.

[0023] Furthermore, the similarity includes: global constant feature similarity match score , function-level constant feature similarity constants similarity and function in-degree and out-degree feature similarity callgraph similarity .

[0024] Furthermore, the global constant feature similarity where BIN represents the set of global constant features extracted from the target binary code, OSS represents the set of global constant features extracted from a version of the source code, N src represents the total number of global constant features extracted from this version of the source code, N f represents the number of matching global constant features, and n(f) represents the number of times a global constant feature appears in all versions of the source code.

[0025] Furthermore, the function-level constant feature similarity where # represents the total number of items in the constant list, Binfunc_const represents the function-level constant features extracted from the target binary code, and Srcfunc_const represents the function-level constant features extracted from a version of the source code.

[0026] Furthermore, the function in-degree and out-degree feature similarity callgraphsimilarit y = distEclud(BinFunc in , SrcFunc in ) + distEclud(BinFunc_out, SrcFunc_out), where distEclud represents calculating the Euclidean distance between vectors, in represents the in-degree feature vector, and out represents the out-degree feature vector.

[0027] Furthermore, the open source component version recognition result is obtained through the following steps:

[0028] 1) Construct the version set GlobalMatch based on the highest numerical global constant feature similarity match score , and if there is only one version in this version set, take this version as the open-source component version recognition result; otherwise, proceed to step 2);

[0029] 2) Add the function-level constant feature similarity constants similarity and the function in-degree and out-degree feature similarity callgraph similarity to obtain the function-level feature similarity;

[0030] 3) In the version set GlobalMatch, construct the version set FunctionMatch based on the highest numerical function-level feature similarity, and take the version set FunctionMatch as the open-source component version recognition result.

[0031] A storage medium stores a computer program, wherein the computer program is configured to execute the above method when running.

[0032] An electronic device includes a memory and a processor, wherein the memory stores a program for executing the above method.

[0033] Compared with the prior art, the present invention selects version-sensitive code features through an original setting rule, thereby effectively improving the recognition accuracy between binary code and source code of different versions. Description of the Drawings

[0034] Figure 1 Diagram showing the representation forms of arrays in source files and binary files.

[0035] Figure 2 Arrays in multiple different versions of source files appear in the same binary file.

[0036] Figure 3 Flowchart for calculating the similarity between binary code and source code of different versions.

[0037] Figure 4 Example of constant features in assignment statements.

[0038] Figure 5 Example of function constant parameter features in source code.

[0039] Figure 6 Example of function constant parameter features in binary files.

[0040] Figure 7 Examples of global and function code features.

[0041] Figure 8 The data structure corresponding to the function code feature in ctree. Detailed implementation manners

[0042] The present invention will be described in detail below with reference to the accompanying drawings and embodiments. It should be noted that the described embodiments are only intended to facilitate the understanding of the present invention and do not impose any limitation on it.

[0043] The open-source component version identification method proposed by the present invention has a specific process as Figure 3 shown. First, version-sensitive features are extracted from source code and binary code. Then, different feature matching schemes are designed for different types of features. Finally, based on the matching scores, it is analyzed which version of the source code is used by the target binary file. The detailed implementation manners are described as follows:

[0044] I. Extraction of sensitive features

[0045] The present invention evaluates the discrimination ability of different code features when applied to open-source library version identification. That is, for a given open-source library, the present invention calculates the proportion of the number of versions that can be uniquely identified at the source code level using these features in the total number of versions, and this proportion is called the discrimination degree of the feature.

[0046] The version-sensitive features selected by the present invention are divided into global constant features and function-level features. The global constant features include string features and exported function features. The function-level features include function-level constant features and function in-degree and out-degree features. The function-level constant features include constant features in assignment statements and constant parameters in function calls.

[0047] The present invention calls the code features that meet the following two criteria version-sensitive features. First, the feature needs to exist in both binary code and source code, and the differences of these features before and after compilation are not significant. Second, in order to use these features for version-level similarity detection, the selected features need to have a high discrimination degree on open-source libraries of different versions, and at the same time, the selected features can be distinguished on binary code. We use the following formula to calculate the discrimination degree of the feature at the source code level and take the string feature as an example for introduction.

[0048]

[0049] Given an open-source library, first extract string features from the source code of each version of this open-source library, sort these features in descending order according to the alphabet, and store these features in a file. Next, calculate the hash value of the entire file, and represent this value as feature_hash. If the feature_hash of one version is different from that of the other versions, it means that this version has string features that are not present in the other versions. At the source code level, this version can be distinguished from the other versions by identifying this string. We use #(distinct(feature_hash)) to represent the total number of unique feature_hash in all versions, and the symbol n represents the total number of versions in this project. As can be seen from Table 1, the discrimination degrees corresponding to string features and exported function features are relatively high, and these two features can be used as version-sensitive features.

[0050] Considering that the version change of the open-source library will bring changes to functions, the present invention heuristically adopts function-level features, specifically including 3 types: a) The present invention takes the constants in the assignment statement arranged in ascending order as features. Figure 4 An example of this feature is given. Figure 4 The code snippet on the left is intercepted from the binary file freetype-VER-2-6.so. Check Figure 4 It can be clearly found from the code snippet on the right that the constants extracted from the source code of the same version (freetype-VER-2-6) are the same as those extracted from the binary file. The constants extracted by both parties are [0,0,0,1,2,6] after being arranged in ascending order. The constants extracted from the functions with the same name in the two versions VER-2-6-1 and VER-2-6-2 are [0,0,1,1,2,6] and [0,0,1,2,2,6] respectively after being arranged in ascending order, and both of these sequences are different from those in VER-2-6; b) The present invention takes the constant parameters in the function call as features. Figure 5 And Figure 6 An example of this feature is shown. Specifically, the openDatabase function calls the sqlite3MisuseError function, and the function parameter passed in is a constant "113824"; c) The last feature selected by the present invention is the function call graph. Specifically, the present invention statically analyzes the function call relationship, and takes the number of times a function calls other functions and the number of times this function is called by other functions, that is, the out-degree and in-degree values of this function on the function call graph, as a kind of code feature.

[0051] The following introduces the specific extraction implementation steps of source code features and binary features.

[0052] 1. Source code feature extraction scheme

[0053] The string features of the present invention are mainly extracted based on Clang. Clang is the front end of the LLVM compiler, and this tool uses the LLVM compiler infrastructure framework as the back end. Clang can perform lexical analysis on code snippets and generate tokens corresponding to the code snippets. The present invention extracts the string features in the source file by parsing these tokens.

[0054] The exported function features and function-level features of the present invention are extracted based on Antlr. Antlr is a parser generator implemented based on the LL algorithm and is written in the Java language. Compared with the tool based on Clang, Antlr can generate the syntax tree corresponding to the code snippet without adding compilation options, and the required features can be extracted more accurately by parsing the nodes of the syntax tree. For the exported function features, the set of function names in the source code without static or inline modifiers is extracted as the exported function list. Figure 7 Shows an example of the code features extracted by the present invention and the corresponding source code snippets.

[0055] It is easy to understand that the extraction tool used to extract the source code features can be replaced by other similar compiler front-end tools that support syntax parsing functions.

[0056] 2. Binary feature extraction scheme

[0057] Compared with the source code, information at the code semantic level such as data types and data structures is lost in the binary code. Binary files exist in the form of bytes before disassembly, and it is impossible to directly locate information such as functions from binary files. Therefore, this system uses a disassembly tool for analysis. The interactive disassembly software IDA (The Interactive Disassembler) is a mature commercial disassembly tool that can parse binary files and identify function fragments in binary files. This tool provides a Python SDK called IDAPython, which encapsulates the main functions of IDA and can extract the features contained in the binary files parsed by IDA by writing IDAPython scripts. IDA parses the list of strings used in the binary and the list of exported functions and displays them in the string windows and export windows of the IDA workspace sub-window respectively. These two types of features can be directly obtained through the corresponding API interfaces. IDA currently releases the intermediate language microcode used in decompilation and the ctree with an abstract syntax tree structure similar to the source code. The present invention obtains the required constant type features by parsing the ctree and identifying layer by layer downward. In addition, the function call relationship can be extracted by obtaining the operator (opcode) of the idaapi.cot_call type. Based on the call relationship between functions, the out-degree and in-degree values of each binary function can be extracted. The corresponding structures of the above code features in the ctree are shown in Figure 8 in.

[0058] It is easy to understand that the extraction tool used to extract binary code features can be replaced by other similar tools that support binary disassembly, post-disassembly syntax parsing, and provide interfaces convenient for feature extraction.

[0059] II. Matching of Sensitive Feature Instances

[0060] The matching of binary code feature instances and source code feature instances is the second key link of the present invention. When judging whether a binary code feature instance matches a source code feature instance, two schemes are adopted, namely, exact matching and matching based on semantic equivalence judgment. Among them, the global features include string features and exported function features, and these features adopt the exact matching scheme to check whether the code feature instances extracted from the source code and the binary are exactly the same. If they are exactly the same, it is considered that this pair of instances matches. Function matching adopts the method of judging based on semantic equivalence, extracts function-level features from the source code and the binary respectively, and calculates the similarity. Given a certain binary function, the source code function with the highest similarity value is considered to match.

[0061] The present invention designs a suitable feature matching scheme to calculate the similarity between binary codes and source codes of different versions. Table 2 summarizes the matching schemes for each type of feature.

[0062]

[0063] Table 2

[0064] The detailed calculation method of the similarity is as follows:

[0065] 1. Global features: The present invention uses two code features, namely strings and exported functions, as global features. Considering that global features remain consistent before and after compilation. Therefore, for global code feature instances extracted from source code and binary respectively, only when the values extracted from both sides are the same, this pair of code feature instances is considered to be matched. Further, according to the total number of string feature and exported function feature instances that are matched, the overall similarity between the source code and the binary code under the global features is calculated.

[0066] There are many duplicate global features between different versions. Intuitively, for version identification, the information carried by these duplicate global features is less than the information brought by strings that only exist in specific versions. Therefore, the present invention considers weighting the feature instances, and the feature instances with higher information content have the highest corresponding weights. Imitating the idea of the TF-IDF algorithm, the present invention weights the feature instances according to the frequency, and the obtained weight calculation formula is represented by weighted_matched_features. This method uses the following formula to calculate the similarity score match between the binary code and the source code under the global feature matching score , when match score is greater than the threshold set according to experience, it can be considered that both sides are matched. The present invention uses BIN and OSS respectively to represent the global code feature instances extracted from the binary and the source code, that is, string feature and exported function feature instances. Where n(f) represents the number of times a global code feature instance appears in all versions of the open-source library source code. In addition to the weighted value weighted_matched_features, this method also considers the similarity of the global code features extracted from both sides. The loss term in the equation is used to calculate the proportion of the number of matched feature instances to the total number of source code feature instances in the current version. For each version, N src represents the total number of global constant code feature instances extracted from the source code of the current version, and N f represents the number of global code feature instances matched with the target binary file.

[0067]

[0068]

[0069] 2. Function - level constant features: Given a binary function Binfunc, this method uses the following formula to calculate its similarity constants with a source - code function Srcfunc similarity . Specifically, Binfunc_const represents two features: function - assigned constants and constants in function parameters extracted from the function in the target binary file, Srcfunc_const represents the above - mentioned two features extracted from the source - code function, and the marker # represents the total number of items in the obtained constant list.

[0070]

[0071] 3. Function out - degree and in - degree features: Based on this feature, the present invention uses the following formula to calculate the similarity of a pair of functions. Specifically, the present invention defines the out - degree vector and in - degree vector of a function as the number of times this function is called by different functions and the number of times this function calls different functions, respectively. For example, if on the function - call graph, function A is called once by function B and twice by function C. At the same time, this function calls function D three times and function E once. Then the in - degree vector of this function is (1, 2) and the out - degree vector of this function is (1, 3). If the dimensions of the two vectors are inconsistent, we pad 0s in front to ensure that the vector dimensions of the two functions are the same. Then use the following formula to calculate the similarity of the two function - call graphs. In the following formula, in / out represent the out - degree vector and in - degree vector of the function respectively, and distEclud represents the Euclidean distance between vectors.

[0072] callgraph similarity (Binfunc, Srcfunc)=distEclud(BinFunc_in, SrcFunc_in)+distEclud(BinFunc_out, SrcFunc_out)

[0073] III. Version identification

[0074] Based on the foregoing feature matching, three kinds of similarities can be obtained. One is the global - level similarity match score , and the other two are function - level similarities constants similarity and callgraph similarity . Based on the foregoing similarity calculation method, the present invention designs a two - stage identification method to obtain the source - code version of the specific reuse of the binary code. The two - stage identification process is as follows:

[0075] 1. Global matching stage: Calculate the match between the target binary file and each version of the open - source component it reuses scoreThe value, record the set GlobalMatch corresponding to the version with the highest value = {w 1 , w 2 , …, w n}. If this set contains only one version, i.e., n = 1, then this version is the open-source component version reused by this target binary file. Otherwise, if n > 1, proceed to the next step 2 to further screen GlobalMatch using function-level features.

[0076] 2. Function-level matching phase: First, calculate the set of unique functions in the source code of each version in GlobalMatch. A unique function is a function that exists only in the source code of this version. It can be achieved by normalizing each function in GlobalMatch and calculating the hash value of the normalized function text. If the function name + function hash value is unique, then this function appears only in the source code of one version and is a unique function. The detailed normalization method can be found in Section 3.2.2 of the article "MVP: Detecting Vulnerabilities using Patch-Enhanced Vulnerability Signatures". Secondly, compare the target binary file with the unique functions in the source code of all versions in GlobalMatch, and calculate the matching situation between all functions of the target binary file and the unique functions of each version in GlobalMatch, that is, the function pair with the highest constants similarity +callgraph similarity value is considered a match. Record the set of open-source component versions FunctionMatch = {v1, v2, …, vm} in GlobalMatch that match the unique functions of the target binary file. Obviously, m ≤ n. The set FunctionMatch is the list of reused versions finally identified by the present invention.

[0077] If the set FunctionMatch contains only one version, i.e., m = 1, it is considered that the target binary file only reuses one version of this open-source component. If m > 1, it means that the binary file matches the unique functions of multiple versions, and it is considered that the target binary file reuses multiple versions of this open-source component. Through subsequent experimental analysis, the appearance of this multi-version reuse situation is because the target binary file may update the version of the originally reused open-source component by patching to ensure software compatibility.

[0078] Experimental data

[0079] The present invention manually marks the versions of 585 binary files and uses these binary files as a test set. This test set covers multiple versions of 10 open-source libraries. Using the solution of the present invention, the open-source component versions reused by 89% (522) of the binary files can be correctly identified, proving the effectiveness of the present invention.

[0080] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Those skilled in the art should understand that any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention, and the protection scope shall be defined by the claims.

Claims

1. An open-source component version identification method using binary code reuse, the steps of which include: 1) Extract the sensitive features of the target binary code and each version of the source code respectively, where the sensitive features include: global constant features and function-level features; among them, the global constant features include: string features and exported function features; the function-level features include: function-level constant features and function in-degree and out-degree features; the function-level constant features include: constant features in assignment statements and constant parameters in function calls. 2) Calculate the similarity between the target binary code and the source code of each version based on the sensitive features to obtain the open-source component version recognition result; wherein, the similarity includes: the global constant feature similarity match score , the function-level constant feature similarity constants similarity and the function in-degree and out-degree feature similarity callgraph similarity , and the obtaining of the open-source component version recognition result includes: Based on the global constant feature similarity match with the highest value score , construct the version set GlobalMatch: If there is only one version in this version set, then use this version as the open-source component version recognition result; otherwise, proceed to add the function-level constant feature similarity constants similarity and the function in-degree and out-degree feature similarity callgraph similarity to obtain the function-level feature similarity; In the version set GlobalMatch, construct a version set FunctionMatch based on the highest function-level feature similarity, and use the version set FunctionMatch as the open-source component version identification result.

2. The method according to claim 1, characterized in that the string features of any version of the source code are extracted through the following steps: 1) Use the compiler front-end clang to perform lexical analysis on the source code to generate tokens corresponding to each code snippet; 2) Parse the tokens to generate string features.

3. The method according to claim 1, characterized in that the exported function features and function-level features of any version of the source code are extracted through the following steps: 1) Generate a syntax tree corresponding to the code snippet based on the syntax parser generator Antlr; 2) Parse the nodes of the syntax tree to obtain the exported function features and function-level features.

4. The method according to claim 1, characterized in that the method for extracting the sensitive features of the target binary code includes: parsing based on the interactive disassembler IDA.

5. The method according to claim 1, characterized in that Global constant feature similarity where BIN represents the set of global constant features extracted from the target binary code, OSS represents the set of global constant features extracted from a version of the source code, N src represents the total number of global constant features extracted from this version of the source code, N f represents the number of matching global constant features, and n(f) represents the number of times a global constant feature appears in all versions of the source code; Function-level constant feature similarity where # represents the total number of items in the constant list, Binfunc_const represents the function-level constant features extracted from the target binary code, and Srcfunc_const represents the function-level constant features extracted from a version of the source code; Function out-degree and in-degree feature similarity callgraph similarity = distEclud(BinFunc in , SrcFunc in ) + distEclud(BinFunc_out, SrcFunc_out), where distEclud represents calculating the Euclidean distance between vectors, in represents the in-degree feature vector, and out represents the out-degree feature vector.

6. A storage medium, in which a computer program is stored, wherein the computer program is set to execute any one of the methods described in claims 1-5 when running.

7. An electronic device, including a memory and a processor, a computer program is stored in the memory, and the processor is set to run the computer program to execute any one of the methods described in claims 1-5.

Citation Information

Patent Citations

  • Method and device for identifying multiplexing relationship between binary code and source code

    CN111045670A

  • Binary code and source code similarity analysis method and device based on code features

    CN111078227A

  • Firmware homology detection method based on multi-dimensional features

    CN112084146A