C / C + + software dependency analysis method in cross-platform transplantation scene

By establishing offline dependency subtree knowledge base and abstract syntax tree technology, the accuracy and performance problems in cross-platform C/C++ software dependency analysis are solved, and efficient and accurate analysis of binary and source code packages is achieved, especially the processing of variables being replaced with real values.

CN120276738APending Publication Date: 2025-07-08NORTHEASTERN UNIV CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510452620.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The prior art has insufficient accuracy and performance in C/C++ software dependency analysis in cross-platform transplant scenarios, especially when processing variables as parameters, and does not support binary form dependency analysis.

Method used

Using methods based on offline dependency subtree knowledge base and abstract syntax tree, we perform dependency analysis on C/C++ software in binary and source code forms respectively. For binary forms, by establishing an offline dependency subtree knowledge base, parsing the dynamic segment information of the shared library file and matching it with the knowledge base to build a dependency subtree; for source code forms, converting the software bill of materials file through an abstract syntax tree, extracting and replacing the variables with the real value.

Benefits of technology

It improves the accuracy and efficiency of C/C++ software dependency analysis in cross-platform transplant scenarios, supports the analysis of binary and source code packages, solves the problem of variable replacement, and improves the accuracy and performance of analysis results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120276738A_ABST
    Figure CN120276738A_ABST
Patent Text Reader

Abstract

The invention provides a C / C + + software dependency analysis method in a cross-platform transplantation scene, relates to the technical field of computer software, and respectively adopts different methods to carry out dependency analysis on C / C + + software in a source code form and a binary form. When the to-be-detected software is in a binary form, analyzing dynamic segment information of a shared library in the to-be-detected software in real time by pre-establishing a dependency sub-tree knowledge base, matching with an offline dependency sub-tree knowledge base, gradually constructing a dependency sub-tree, and completing the construction of a complete dependency tree; when the to-be-detected software is in a source code form, text content of a software bill of material file is converted through the abstract syntax tree, mapping of a temporary variable and a real value and obtaining of parameters in a dependent introduction statement are completed by traversing nodes of the abstract syntax tree, and the parameter value existing in a variable form is replaced with the real value; the problem that a regular expression can only complete extraction of parameter content and cannot obtain a true value of a variable is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer software, and in particular, to a method for analyzing C / C++ software dependencies in a cross-platform porting scenario. Background Art

[0002] In recent years, the instruction set architecture and operating system have shown a trend of diversified development. The mainstream software originally oriented to private instruction set architectures and operating systems needs to be ported to domestic instruction set architectures and operating systems. Since C / C++ is closely related to the instruction set architecture and operating system, when porting the software itself, its dependencies also need to be ported. Therefore, it is important to analyze the dependencies of C / C++ software in a cross-platform porting scenario.

[0003] The dependency analysis of C / C++ software is mainly carried out from two perspectives: binary packages and source code packages. Among them, the source code packages of C / C++ software introduce dependencies through the software bill of materials files of build tools and package management tools. There are various build tools and package management tools used by C / C++ software. For example, the software bill of materials file of the build tool CMake is CMakeLists.txt, and the software bill of materials file of the package management tool Conan is conanfile.py. Since C / C++ does not have a unified build tool or package management tool similar to maven or pypi in Java or Python, researchers have investigated and summarized the common build tools and package management tools for C / C++ software, and then summarized the syntax for dependency introduction in the software bill of materials files corresponding to these tools, and designed corresponding regular expressions based on these syntaxes.

[0004] In Proceedings of the 37th IEEE / ACM International Conference on Automated Software Engineering (ASE'22). Association for Computing Machinery, New York, NY, USA, Article 106, 1–12. The paper "Towards Understanding Third-party Library Dependency in C / C++ Ecosystem" presents a method for parsing third-party library dependencies in the C language ecosystem based on software bill of materials. The paper proposes a tool called CCScanner to parse the software bill of materials files of build tools or package management tools used by C / C++ software. CCScanner defines parsers for software bill of materials files of various build tools and package management tools through regular expressions. When performing dependency analysis, CCScanner traverses all source files of the software to be detected, calls the corresponding parser for the software bill of materials file according to the file format of the source file for parsing. The parser scans the text content of the software bill of materials file, completes the detection of dependency introduction statements through regular expressions, obtains the parameters in the dependency introduction statements, and outputs the parameter content as the result of dependency analysis. However, this way of dependency analysis has certain deficiencies in terms of accuracy and performance.

[0005] In terms of accuracy, regular expressions are difficult to handle the situation where variables are passed as parameters into dependency analysis statements and cannot replace these variables with the actual values of the dependencies. This is because developers will define variables through its syntax in the software bill of materials file, and these defined variables will also be passed as parameters into the dependency introduction statements. When performing dependency analysis, regular expressions will match the dependency introduction statements and extract the parameters in these statements as the results. Therefore, the dependency results output by existing tools contain the temporary variables defined in the software bill of materials file, rather than the actual values of the dependencies corresponding to these variables. This results in a reduction in the accuracy of the dependency analysis of existing tools.

[0006] In terms of performance, the matching process of regular expressions usually relies on the backtracking mechanism, especially when using greedy quantifiers (such as * and +). If the regular expression is designed improperly, it may lead to a large number of backtracking operations, resulting in performance problems and causing the program to wait for a long time or crash. In addition, the technical solution of this paper does not support the dependency analysis of binary C / C++ software. Summary of the Invention

[0007] The technical problem to be solved by the present invention is to provide a C / C++ software dependency analysis method in a cross-platform porting scenario, which effectively improves the accuracy and efficiency of dependency analysis through technical means such as an offline dependency subtree knowledge base and an abstract syntax tree.

[0008] To solve the above technical problems, the technical solution adopted by the present invention is as follows:

[0009] The present invention provides a C / C++ software dependency analysis method in a cross-platform porting scenario, including the following steps:

[0010] Obtain the software package to be detected, determine whether the software to be detected is a C / C++ software in source code form or binary package form, and adopt different methods for dependency analysis for C / C++ software in source code form and binary form respectively; if the software to be detected is in binary form, use the C / C++ software dependency analysis method for binary form to perform dependency analysis on the software to be detected; if the software to be detected is in source code form, use the C / C++ software dependency analysis method for source code form to perform dependency analysis on the software to be detected.

[0011] The C / C++ software dependency analysis method for binary form is a C / C++ software dependency analysis method based on feature extraction.

[0012] The C / C++ software dependency analysis method for source code form is a C / C++ software dependency analysis method based on abstract syntax tree and pattern matching.

[0013] The C / C++ software dependency analysis method for binary form includes the following steps:

[0014] Step 1: Establish an offline dependency subtree knowledge base.

[0015] Obtain knowledge base data, including software packages and software package metadata in the software package repository of Linux distributions.

[0016] Obtain the knowledge parameters included in the knowledge base data, including project name, version, supported instruction set architecture, file list, source code package link, direct dependencies, version constraint range of direct dependencies.

[0017] Based on the knowledge base data and the knowledge parameters included in the knowledge base data, establish an offline dependency subtree knowledge base.

[0018] Step 2: Obtain the software package to be detected in binary form, decompress the software package to be detected in binary form, traverse the files in the software package to be detected in binary form, and obtain all shared library files included in the software.

[0019] The specific method for obtaining all shared library files included in the software is:

[0020] Traverse the files in the software package to be detected in binary form. For the currently accessed file, judge the file type of the file according to the file name suffix of the file; if the file name of the file has a suffix and the suffix is not.so, skip the file; if the file name suffix of the file is.so or there is no suffix, further judge whether the file is a shared library file. If the file is a shared library file, obtain the file. If the file is not a shared library file, skip the file;

[0021] The specific method for judging whether a file is a shared library file is as follows:

[0022] Pass the file path of the file into the os.path.islink() function in python, and use the readelf function and the islink() function to judge whether the file is a shared library file. If the output result of the readelf function is "Type: DYN (Shared object file)" and the output of the islink() function is False, then the file is a shared library file; otherwise, the file is not a shared library file;

[0023] Step 3: Traverse all the shared library files in the software package to be detected in binary form, obtain the direct dependencies included in each shared library file, and generate a direct dependency list for each shared library file;

[0024] Traverse all the shared library files in the software package to be detected in binary form. For the currently accessed shared library file, obtain the header file information of the shared library file, and judge whether the shared library file is a soft link file. If the shared library file is a soft link file, skip the shared library file; if the shared library file is not a soft link file, parse the dynamic segment of the shared library file, obtain the direct dependencies of the shared library file, and get the direct dependency list of the shared library file;

[0025] Step 4: Generate a direct dependency subtree for each shared library file;

[0026] Step 4.1: Based on the direct dependency list of each shared library file, extract the characteristics of the direct dependencies included in each list, including string literals, export symbol tables, and function call graphs;

[0027] For the direct dependency list of any shared library file, the specific method for extracting the characteristics of the i-th direct dependency in it is as follows:

[0028] Extract the string literal of the i-th direct dependency from the.rodata segment of the shared library file through the LIEF library;

[0029] Extract the export symbol table of the i-th direct dependency using the readelf tool; the export symbol table includes the symbol information publicly available by the shared library.

[0030] Use the disassembly framework Ghidra to disassemble the i-th direct dependency and extract the function call graph of the i-th direct dependency.

[0031] Step 4.2: Based on the name of the shared library file, perform a preliminary filter on the offline dependency subtree knowledge base to obtain the shared library file with the same name as the i-th direct dependency in the offline dependency subtree knowledge base and its features.

[0032] Step 4.3: Use the relevant features of the i-th direct dependency to match with the shared library file with the same name as the i-th direct dependency in the offline dependency subtree knowledge base to generate the dependency subtree of the i-th direct dependency. The specific method is as follows:

[0033] Successively determine whether the features of the i-th direct dependency are consistent with the features of the shared library file with the same name as the i-th direct dependency in the offline dependency subtree knowledge base, including:

[0034] 1) Compare whether the hash values of the string literals of the i-th direct dependency and the string literals of the shared library file with the same name as the i-th direct dependency in the offline dependency subtree knowledge base are the same;

[0035] 2) Compare whether the function names and quantities included in the export symbol table of the i-th direct dependency and the export symbol table of the shared library file with the same name as the i-th direct dependency in the offline dependency subtree knowledge base are consistent;

[0036] 3) Compare whether the difference in the number of nodes and edges between the function call graph of the i-th direct dependency and the function call graph of the shared library file with the same name as the i-th direct dependency in the offline dependency subtree knowledge base is less than the set threshold;

[0037] If the features of the i-th direct dependency are consistent with the features of the shared library file with the same name as the i-th direct dependency in the offline dependency subtree knowledge base, then the i-th direct dependency can be matched with the shared library with the same name as the i-th direct dependency in the offline dependency subtree knowledge base from the feature perspective, and then generate the dependency subtree of the i-th direct dependency.

[0038] Step 5: Concatenate the direct dependency subtrees of each shared library file in the binary form of the software package to be detected to generate the complete dependency tree of the binary form of the software to be detected.

[0039] Step 6: Based on the complete dependency tree of the binary form of the software to be detected, perform dependency analysis on the binary form of the C / C++ software to obtain the set of shared libraries on which the binary form of the software to be detected depends.

[0040] The C / C++ software dependency analysis method in source code form includes the following steps:

[0041] S1: Collect build tools and package management tools, and generate software bill of materials files for the build tools and package management tools respectively;

[0042] S2: Obtain the software package to be detected in source code form, initialize the dependency relationship linear list, decompress the software package to be detected in source code form, traverse the files in the software package to be detected in source code form, and judge the file type of the j-th file currently accessed. If the suffix name of the j-th file is within the range of the software bill of materials files of the build tools and package management tools generated in S1, obtain the software bill of materials file of the j-th file; otherwise, skip this file;

[0043] S3: Based on the software bill of materials file of the j-th file, call the corresponding software bill of materials file parser, extract the dependencies introduced in the software bill of materials file of the j-th file, and store them in the linear list for saving dependency relationships;

[0044] S3.1: Perform abstract syntax tree transformation on the software bill of materials file of the j-th file through an abstract syntax tree transformation tool to generate the abstract syntax tree of the j-th file;

[0045] S3.2: Perform a depth-first traversal on the abstract syntax tree of the j-th file, traverse the nodes of the abstract syntax tree of the j-th file. For the k-th node currently traversed, if the k-th node has no child nodes, end the recursion and return to the upper-level node; if the k-th node has child nodes, perform variable extraction and dependency analysis on the k-th node;

[0046] S3.3: Judge whether the k-th node is a statement for variable definition. If the k-th node is not a statement for variable definition, skip the k-th node; if the k-th node is a statement for variable definition, perform variable extraction on the k-th node, write the variable name and the corresponding real value into a hash table, and obtain the hash table for saving the mapping between variable names and variable real values;

[0047] S3.4: Judge whether the k-th node is a statement for dependency introduction. If the k-th node is not a statement for dependency introduction, skip the k-th node; if the k-th node is a statement for dependency introduction, parse and judge the parameter value in the k-th node. If the type of the parameter value of the k-th node is in variable form, obtain the real value corresponding to the current parameter value according to the key in the hash table, and write the real value of the introduced dependency into the dependency relationship linear list; if the type of the parameter value of the k-th node is not in variable form, write the real value of the introduced dependency into the dependency relationship linear list;

[0048] S4: After completing the traversal of all files in the software package to be detected in source code form, a dependency relationship linear list including the dependency results of all files in the software package to be detected in source code form is obtained, and all the results in the dependency relationship linear list are output to obtain all dependencies of the source code package.

[0049] The beneficial effects of adopting the above technical solution are as follows: A C / C++ software dependency analysis method in a cross-platform porting scenario provided by the present invention not only supports the analysis of both binary packages and source code packages of software, but also effectively improves the accuracy and efficiency of dependency analysis. When the software to be detected is in binary form, by pre-establishing a dependency subtree knowledge base, dynamically parsing the dynamic segment information of shared libraries in the software to be detected in real time, and matching it with the offline dependency subtree knowledge base, a dependency subtree is gradually constructed, and finally the construction of a complete dependency tree is completed to realize the dependency analysis of binary-form C / C++ software; when the software to be detected is in source code form, the text content of the software bill of materials file is transformed through an abstract syntax tree, and by traversing the nodes of the abstract syntax tree, the mapping between temporary variables and real values and the work of obtaining parameters in dependency introduction statements are completed, and the parameter values existing in the form of variables are replaced with real values; due to the variable substitution in the software bill of materials file, when the parameters in the dependency introduction statement exist in the form of variables, the abstract syntax tree is used for dependency analysis to solve the problem that regular expressions can only complete the extraction of parameter content and cannot obtain the real values of variables. Description of the Drawings

[0050] Figure 1 It is a flowchart of the C / C++ software dependency analysis method in a cross-platform porting scenario provided by an embodiment of the present invention;

[0051] Figure 2 It is a schematic diagram of the construction of a dependency subtree knowledge in the C / C++ software dependency analysis method in binary form provided by an embodiment of the present invention. Detailed Embodiments

[0052] The following combines the drawings and embodiments to further describe in detail the specific embodiments of the present invention. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.

[0053] A C / C++ software dependency analysis method in a cross-platform porting scenario of this embodiment can not only complete the dependency analysis of binary-form C / C++ software, but also complete the dependency analysis of source code-form C / C++ software, and at the same time improve the accuracy and efficiency of dependency analysis. A C / C++ software dependency analysis method, as Figure 1 shown, includes:

[0054] Obtain the software package to be detected, and determine whether the software to be detected is a C / C++ software in source code form or binary package form. Different methods are used for dependency analysis of C / C++ software in source code form and binary form respectively. If the software to be detected is in binary form, use the dependency analysis method for binary-form C / C++ software to perform dependency analysis on the software to be detected. If the software to be detected is in source code form, use the dependency analysis method for source-code-form C / C++ software to perform dependency analysis on the software to be detected.

[0055] The dependency analysis method for binary-form C / C++ software is the C / C++ software dependency analysis method based on feature extraction.

[0056] The dependency analysis method for source-code-form C / C++ software is the C / C++ software dependency analysis method based on abstract syntax tree and pattern matching.

[0057] The dependency analysis method for binary-form C / C++ software includes the following steps:

[0058] Step 1: Establish an offline dependency subtree knowledge base.

[0059] Obtain knowledge base data, including software packages and software package metadata in the software package repository distributed by Linux.

[0060] Obtain the knowledge parameters included in the knowledge base data, including project name, version, supported instruction set architecture, file list, source code package link, direct dependencies, and version constraint ranges of direct dependencies.

[0061] Based on the knowledge base data and the knowledge parameters included in the knowledge base data, establish an offline dependency subtree knowledge base.

[0062] Step 2: Obtain the software package to be detected in binary form, decompress the software package to be detected in binary form, traverse the files in the software package to be detected in binary form, and obtain all shared library files included in the software.

[0063] The specific method for obtaining all shared library files included in the software is as follows:

[0064] Traverse the files in the software package to be detected in binary form. For the currently accessed file, judge the file type of the file according to the file name suffix of the file. If the file has a suffix and the suffix is not.so, skip the file. If the file name suffix of the file is.so or there is no suffix, further judge whether the file is a shared library file. If the file is a shared library file, obtain the file. If the file is not a shared library file, skip the file.

[0065] The specific method for judging whether a file is a shared library file is as follows:

[0066] Pass the file path of the file into the os.path.islink() function in Python. Use the readelf function and the islink() function to determine whether the file is a shared library file. If the output result of the readelf function is "Type: DYN (Shared object file)" and the output of the islink() function is False, then the file is a shared library file; otherwise, the file is not a shared library file;

[0067] Step 3: Traverse all the shared library files in the binary form of the software package to be detected, obtain the direct dependencies included in each shared library file, and generate a direct dependency list for each shared library file;

[0068] Traverse all the shared library files in the binary form of the software package to be detected. For the currently accessed shared library file, obtain the header file information of the shared library file, and determine whether the shared library file is a soft link file. If the shared library file is a soft link file, then skip the shared library file; if the shared library file is not a soft link file, then parse the dynamic segment of the shared library file, obtain the direct dependencies of the shared library file, and get the direct dependency list of the shared library file;

[0069] Step 4: Generate a direct dependency subtree for each shared library file;

[0070] Step 4.1: Based on the direct dependency list of each shared library file, extract the features of the direct dependencies included in each list, including string literals, export symbol tables, and function call graphs;

[0071] For the direct dependency list of any shared library file, the specific method for extracting the features of the i-th direct dependency is as follows:

[0072] Extract the string literal of the i-th direct dependency from the.rodata segment of the shared library file through the LIEF library;

[0073] Extract the export symbol table of the i-th direct dependency through the readelf tool; the export symbol table includes the symbol information publicly exposed by the shared library;

[0074] Use the disassembling framework Ghidra to disassemble the i-th direct dependency and extract the function call graph of the i-th direct dependency;

[0075] Step 4.2: Based on the name of the shared library file, perform a preliminary filtering on the offline dependency subtree knowledge base, and obtain the shared library file with the same name as the i-th direct dependency in the offline dependency subtree knowledge base and its features;

[0076] Step 4.3: Use the relevant features of the i-th direct dependency and the shared library file with the same name as the i-th direct dependency in the offline dependency subtree knowledge base for matching to generate the dependency subtree of the i-th direct dependency. The specific method is as follows:

[0077] Successively determine whether the features of the i-th direct dependency are consistent with the features of the shared library file with the same name as the i-th direct dependency in the offline dependency subtree knowledge base, including:

[0078] 1) Compare whether the hash values of the string literals of the i-th direct dependency and the shared library file with the same name as the i-th direct dependency in the offline dependency subtree knowledge base are the same;

[0079] 2) Compare whether the function names and quantities included in the export symbol table of the i-th direct dependency and the export symbol table of the shared library file with the same name as the i-th direct dependency in the offline dependency subtree knowledge base are consistent;

[0080] 3) Compare whether the difference in the number of nodes and edges in the function call graph of the i-th direct dependency and the function call graph of the shared library file with the same name as the i-th direct dependency in the offline dependency subtree knowledge base is less than the set threshold;

[0081] If the features of the i-th direct dependency are consistent with the features of the shared library file with the same name as the i-th direct dependency in the offline dependency subtree knowledge base, then the i-th direct dependency can be matched with the shared library with the same name as the i-th direct dependency in the offline dependency subtree knowledge base from the feature perspective, and then the dependency subtree of the i-th direct dependency is generated;

[0082] Step 5: Concatenate the direct dependency subtrees of each shared library file in the binary-form software package to be detected to generate the complete dependency tree of the binary-form software to be detected;

[0083] Step 6: Perform dependency analysis on the binary-form C / C++ software based on the complete dependency tree of the binary-form software to be detected to obtain the set of shared libraries on which the binary-form software to be detected depends;

[0084] The method for dependency analysis of source-code-form C / C++ software includes the following steps:

[0085] S1: Collect build tools and package management tools, and generate software bill of materials files for the build tools and package management tools respectively;

[0086] In this embodiment, before performing C / C++ software dependency analysis in source code form, common build tools and package management tools in the C / C++ ecosystem are collected in advance. Among them, the build tools include CMake, and the package management tools include Conan. A software bill of materials file for the build tool CMake is generated as CMakeLists.txt, and a software bill of materials file for the package management tool Conan is conanfile.py.

[0087] S2: Obtain the software package to be detected in source code form, initialize the dependency relationship linear list, decompress the software package to be detected in source code form, traverse the files in the software package to be detected in source code form, and determine the file type of the j-th file currently accessed. If the suffix name of the j-th file is within the range of the software bill of materials files of the build tools and package management tools generated in S1, obtain the software bill of materials file of the j-th file; otherwise, skip the file.

[0088] S3: Based on the software bill of materials file of the j-th file, call the corresponding software bill of materials file parser to extract the dependencies introduced in the software bill of materials file of the j-th file and store them in the linear list for saving dependency relationships.

[0089] After completing the type judgment of the software bill of materials files in the source code package, determine the package management tool or build tool used by the software package to be detected according to the name of the software bill of materials file, and then call the corresponding software bill of materials file parser to extract the dependencies introduced in the software bill of materials file, including the following steps:

[0090] S3.1: Perform abstract syntax tree transformation on the software bill of materials file of the j-th file through an abstract syntax tree transformation tool to generate the abstract syntax tree of the j-th file.

[0091] S3.2: Perform a depth-first traversal of the abstract syntax tree of the j-th file, traverse the nodes of the abstract syntax tree of the j-th file. For the k-th node currently traversed, if the k-th node has no child nodes, end the recursion and return to the upper-level node; if the k-th node has child nodes, perform variable extraction and dependency analysis on the k-th node.

[0092] S3.3: Determine whether the k-th node is a statement for variable definition. If the k-th node is not a statement for variable definition, skip the k-th node; if the k-th node is a statement for variable definition, perform variable extraction on the k-th node, write the variable name and the corresponding real value into the hash table, and obtain the hash table that saves the mapping between the variable name and the variable real value.

[0093] S3.4: Determine whether the k-th node is a statement introducing a dependency. If the k-th node is not a statement introducing a dependency, skip the k-th node; if the k-th node is a statement introducing a dependency, parse and determine the parameter value in the k-th node. If the type of the parameter value of the k-th node is in variable form, obtain the true value corresponding to the current parameter value according to the key in the hash table, and write the true value of the introduced dependency into the dependency relationship linear list; if the type of the parameter value of the k-th node is not in variable form, write the true value of the introduced dependency into the dependency relationship linear list.

[0094] S4: After completing the traversal of all files in the software package to be detected in source code form, obtain a dependency relationship linear list including the dependency results of all files in the software package to be detected in source code form, output all the results in the dependency relationship linear list, and obtain all the dependencies of the source code package.

[0095] Since the efficiency of real-time binary package dependency analysis is relatively low, in this embodiment, an offline method of constructing a dependency subtree knowledge base is adopted. A large number of shared libraries and their dependency subtrees are stored in the knowledge base. A schematic diagram of the construction of the offline dependency subtree knowledge base in binary package dependency analysis is as Figure 2 shown. First, crawl the software packages in the software package repositories of mainstream Linux distributions, such as Debian, CentOS, and Loongnix. Extract the shared library files from the crawled software packages, recursively analyze the dynamic segments of each shared library file, obtain the next-level direct dependencies, and record the analyzed shared library files in a hash table. During the recursive analysis, synchronously determine whether the currently analyzed shared library exists in the hash table. If it exists, it means that the current file has been completed, perform a pruning operation, end the recursion, and return to the upper layer. In this way, the construction of the knowledge base is completed and used in the direct dependency matching process.

[0096] Compared with the prior art, the technical solution proposed in this embodiment not only supports the analysis of software in both binary package and source code package forms, but also effectively improves the accuracy and efficiency of dependency analysis. In terms of the dependency analysis of binary packages, under the experimental conditions of a 12-core processor and 16GB of memory, this tool performed dependency analysis on 15,339 C / C++ software in binary form, and the average construction time was only 0.44s, showing good performance in the analysis of binary packages. In terms of the dependency analysis of source code packages, by using the abstract syntax tree for dependency analysis, the problem of variable substitution in the software bill of materials file is solved, and at the same time, the time-consuming of dependency analysis can be reduced. After performing dependency analysis on 20 open-source projects, dependencies were manually extracted from the software bill of materials files of the open-source projects and compared with the output results of the present invention. Dependency analysis was successfully completed on 19 open-source projects, a total of 372 dependencies were detected, and the correctness of the dependency analysis reached 98%.

[0097] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope defined by the claims of the present invention.

Claims

1. A method for C / C++ software dependency analysis in a cross-platform porting scenario, characterized in that: It includes the following steps: Obtain the software package to be detected. Determine whether the software to be detected is a C / C++ software in source code form or binary package form, and adopt different methods for dependency analysis on C / C++ software in source code form and binary form respectively. If the software to be detected is in binary form, use the dependency analysis method for binary C / C++ software to perform dependency analysis on the software to be detected. If the software to be detected is in source code form, use the dependency analysis method for source code C / C++ software to perform dependency analysis on the software to be detected. The dependency analysis method for binary C / C++ software is a C / C++ software dependency analysis method based on feature extraction. The dependency analysis method for source code C / C++ software is a C / C++ software dependency analysis method based on abstract syntax tree and pattern matching.

2. The method for analyzing C / C++ software dependencies in a cross-platform porting scenario according to claim 1, wherein: The dependency analysis method for binary C / C++ software includes the following steps: Step 1: Establish an offline dependency subtree knowledge base. Step 2: Obtain the software package to be detected in binary form, decompress the software package to be detected in binary form, traverse the files in the software package to be detected in binary form, and obtain all shared library files included in the software. Step 3: Traverse all shared library files in the software package to be detected in binary form, obtain the direct dependencies included in each shared library file, and generate a direct dependency list for each shared library file. Step 4: Generate the direct dependency subtree for each shared library file. Step 5: Concatenate the direct dependency subtrees of each shared library file in the software package to be detected in binary form to generate a complete dependency tree for the software package to be detected in binary form. Step 6: Based on the complete dependency tree of the software package to be detected in binary form, perform dependency analysis on the binary C / C++ software to obtain the set of shared libraries on which the software package to be detected in binary form depends.

3. A method for analyzing C / C++ software dependencies in a cross-platform porting scenario according to claim 2, characterized in that: The specific method of the above Step 1 is as follows: Obtain knowledge base data, including software packages and software package metadata in the software package repository of Linux distribution. Obtain the knowledge parameters included in the knowledge base data, including project name, version, supported instruction set architecture, file list, source code package link, direct dependency, and version constraint range of direct dependency. Based on the knowledge base data and the knowledge parameters included in the knowledge base data, establish an offline dependency subtree knowledge base.

4. The C / C++ software dependency analysis method in a cross-platform porting scenario according to claim 3, wherein: The above Step 2 includes: The specific method for obtaining all shared library files included in the software is as follows: Traverse the files in the software package to be detected in binary form. For the currently accessed file, judge the file type according to the file name suffix of the file. If the file has a file name suffix and the suffix is not.so, skip the file. If the file name suffix of the file is.so or there is no suffix, further judge whether the file is a shared library file. If the file is a shared library file, obtain the file. If the file is not a shared library file, skip the file. The specific method for judging whether a file is a shared library file is as follows: Pass the file path of the file into the os.path.islink() function in Python. Use the readelf function and the islink() function to determine whether the file is a shared library file. If the output of the readelf function is "Type: DYN (Shared object file)" and the output of the islink() function is False, then the file is a shared library file; otherwise, the file is not a shared library file.

5. A method for analyzing C / C++ software dependencies in a cross-platform porting scenario according to claim 4, characterized in that: The specific method of step 3 is as follows: Traverse all the shared library files in the software package to be detected in binary form. For the currently accessed shared library file, obtain the header file information of the shared library file, and determine whether the shared library file is a soft link file. If the shared library file is a soft link file, skip the shared library file; if the shared library file is not a soft link file, parse the dynamic segment of the shared library file to obtain the direct dependencies of the shared library file, and obtain the direct dependency list of the shared library file.

6. The C / C++ software dependency analysis method in a cross-platform porting scenario according to claim 5, wherein: Step 4 includes: Step 4.1: Based on the direct dependency lists of each shared library file, extract the features of the direct dependencies included in each list, including string literals, export symbol tables, and function call graphs; For the direct dependency list of any shared library file, the specific method of extracting the features of the i-th direct dependency therein is: Extract the string literal of the i-th direct dependency from the.rodata segment of the shared library file through the LIEF library; Extract the export symbol table of the i-th direct dependency through the readelf tool; the export symbol table includes the symbol information publicly exposed by the shared library; Use the disassembly framework Ghidra to disassemble the i-th direct dependency and extract the function call graph of the i-th direct dependency; Step 4.2: Based on the name of the shared library file, perform a preliminary filter on the offline dependency subtree knowledge base to obtain the shared library file and its features in the offline dependency subtree knowledge base that have the same name as the i-th direct dependency; Step 4.3: Use the relevant features of the i-th direct dependency to match with the shared library file in the offline dependency subtree knowledge base that has the same name as the i-th direct dependency to generate the dependency subtree of the i-th direct dependency.

7. A method for analyzing C / C++ software dependencies in a cross-platform porting scenario according to claim 6, characterized in that: The specific method of step 4.3 is as follows: Successively determine whether the features of the i-th direct dependency are consistent with the features of the shared library file in the offline dependency subtree knowledge base that has the same name as the i-th direct dependency, including: 1) Compare whether the hash values of the string literal of the i-th direct dependency and the string literal of the shared library file in the offline dependency subtree knowledge base that has the same name as the i-th direct dependency are the same; 2) Compare whether the function names and quantities included in the export symbol table of the i-th direct dependency and the export symbol table of the shared library file in the offline dependency subtree knowledge base that has the same name as the i-th direct dependency are consistent; 3) Compare whether the difference in the number of nodes and edges between the function call graph of the i-th direct dependency and the function call graph of the shared library file in the offline dependency subtree knowledge base that has the same name as the i-th direct dependency is less than the set threshold; If the features of the i-th directly dependent feature are consistent with the features of the shared library file in the offline dependent subtree knowledge base that has the same name as the i-th direct dependency, then the i-th direct dependency can be matched with the shared library in the offline dependent subtree knowledge base that has the same name as the i-th direct dependency from the perspective of features, and then the dependent subtree of the i-th direct dependency is generated.

8. A method for analyzing C / C++ software dependencies in a cross-platform porting scenario according to claim 1, characterized in that: The C / C++ software dependency analysis method in the form of source code includes the following steps: S1: Collect build tools and package management tools, and generate software bill of materials files for the build tools and package management tools respectively; S2: Obtain the software package to be detected in the form of source code, initialize the linear list of dependency relationships, decompress the software package to be detected in the form of source code, traverse the files in the software package to be detected in the form of source code, and judge the file type of the j-th file currently accessed. If the suffix name of the j-th file is within the range of the software bill of materials files of the build tools and package management tools generated in S1, obtain the software bill of materials file of the j-th file; otherwise, skip this file. S3: Based on the software bill of materials file of the j-th file, call the corresponding software bill of materials file parser to extract the dependencies introduced in the software bill of materials file of the j-th file, and store them in the linear list for saving dependency relationships. S4: After completing the traversal of all files in the software package to be detected in the form of source code, obtain the linear list of dependency relationships including the dependency results of all files in the software package to be detected in the form of source code, and output all the results in the linear list of dependency relationships to obtain all the dependencies of the source code package.

9. A method for C / C++ software dependency analysis in a cross-platform porting scenario according to claim 8, characterized in that: The said S3 includes the following steps: S3.1: Perform abstract syntax tree transformation on the software bill of materials file of the j-th file through an abstract syntax tree transformation tool to generate the abstract syntax tree of the j-th file; S3.2: Perform a depth-first traversal on the abstract syntax tree of the j-th file, traverse the nodes of the abstract syntax tree of the j-th file. For the k-th node currently traversed, if the k-th node has no child nodes, end the recursion and return to the upper-level node; if the k-th node has child nodes, perform variable extraction and dependency analysis on the k-th node. S3.3: Judge whether the k-th node is a statement for variable definition. If the k-th node is not a statement for variable definition, skip the k-th node; if the k-th node is a statement for variable definition, perform variable extraction on the k-th node, and write the variable name and the corresponding real value into the hash table to obtain the hash table that saves the mapping between the variable name and the variable real value. S3.4: Judge whether the k-th node is a statement for dependency introduction. If the k-th node is not a statement for dependency introduction, skip the k-th node; if the k-th node is a statement for dependency introduction, parse and judge the parameter value in the k-th node. If the type of the parameter value of the k-th node is in the form of a variable, obtain the real value corresponding to the current parameter value according to the key in the hash table, and write the real value of the introduced dependency into the linear list of dependency relationships; if the type of the parameter value of the k-th node is not in the form of a variable, write the real value of the introduced dependency into the linear list of dependency relationships.