Software component analysis method oriented to C / C + + source code

By building a cross-platform TPL feature library and multiple threshold judgment, the problems of incomplete features and insufficient dependencies in existing tools are solved, efficient and accurate software component analysis of C/C++ source code is realized, the risk of misjudgment is reduced, and library granularity multiplexing detection and dependency analysis are supported.

CN120295890APending Publication Date: 2025-07-11NORTHEASTERN UNIV CHINA
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510356202.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing software component analysis tools for C/C++ source code have problems such as incomplete features, inaccurate identification of library granularity reuse, and insufficient dependency analysis, resulting in increased misjudgment and security risks.

Method used

Build a cross-platform TPL feature library, remove noise functions through TLSH hash value and clustering algorithm, combine multiple thresholds to judge the existence of TPL, analyze the multiplexing ratio at the file and function level, identify the TPL module and parse the dependencies.

Benefits of technology

Improve the accuracy and efficiency of software component analysis, reduce false positives, ensure the accuracy of library granularity reuse detection, and support security vulnerabilities and license compliance analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120295890A_ABST
    Figure CN120295890A_ABST
Patent Text Reader

Abstract

The invention provides a software component analysis method oriented to C / C + + source codes, relates to the technical field of software component analysis, and provides a flexible and efficient data collection mechanism, comprehensively constructs a TPL feature library to improve the quality of the feature library, preliminarily filters non-specific TPL functions in combination with a directory structure, reduces noise interference and improves the software component analysis efficiency. The method comprises the following steps: firstly, classifying similar TPLs into the same family by using a clustering algorithm, sharing a certain number of common functions in each family, and meanwhile, including a unique marker function of each TPL, further analyzing a multiplexing relationship between the TPLs sharing the common functions, and removing the common functions which do not belong to a specific TPL range, thereby overcoming the defects related to the birth time of the functions; according to the method, the multiplexing proportion of the target code at the function level and the file level is further comprehensively analyzed, the TPL multiplexing situation of the code is accurately judged, the target code is divided into TPL modules, the dependency relationship between the modules is analyzed through include and extran statements, and therefore more accurate component recognition and dependency analysis are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of software composition analysis, and particularly to a software composition analysis method for C / C++ source code. Background Art

[0002] In the prior art, many tools and methods have been proposed to achieve similar goals. For example, CENTRIS infers the dependencies of reused third-party libraries (TPLs) by analyzing the birth time of functions, so as to distinguish the external code introduced by TPLs from its own code and eliminate redundant code. However, this tool relies on the feature of the function birth time, which is vulnerable to version control tools and management mechanisms, and may be difficult to accurately obtain or even unable to obtain, thus limiting the accuracy of the analysis. In addition, CENTRIS determines the matching relationship between the target code and TPL by setting a fixed function number threshold (such as when the number of the same functions reaches 10%). However, in the reuse scenario at the software library granularity, this method is prone to misjudgment, inferring some identical functions as reusing third-party libraries, thus increasing the false positives of the analysis.

[0003] Another tool, OSSFP, attempts to divide functions into cloned functions, supporting functions, general functions, and core functions, and removes the first three types of functions by setting thresholds and function birth time, and finally generates TPL features only based on core functions. However, the proportion of common functions and supporting functions in different TPLs varies greatly, and using static thresholds may cause valuable feature functions to be misdeleted, further reducing the reliability of the analysis results. In addition, OSSFP also relies on the feature of function birth time and is limited by the same defect.

[0004] In terms of TPL dependency detection, CNEPS divides the target code into modules, and each module consists of a file set composed of a header file and several source files. Then, component analysis is carried out for each module. From the analysis results, a dominant component is selected. However, there are certain deviations in the results of component analysis. On this basis, further determining the dominant component of the module will amplify the uncertainty of component analysis. Even if there is only one function in the module that is the same as a certain component, CNEPS will determine this component as the main component of the module. However, in this case, the same function is likely to be caused by chance, which will result in node redundancy and subsequent incorrect dependency edges. In addition, the component analysis operation is time-consuming. If component analysis is to be carried out on multiple modules, the efficiency of the tool will be reduced. CNEPS uses the C / C++ syntax as the criterion for judging dependencies and establishes dependencies by scanning include instructions and the corresponding.h files. However, in case of.h file name conflicts, the accuracy of the tool will be reduced.

[0005] In summary, the existing SCA tools for C / C++ source code have the following defects:

[0006] (1) The existing SCA tools lack comprehensive TPL features. C / C++ lacks a unified package management tool, and its TPL resources are scattered across multiple platforms, such as package managers like GitHub, GitLab, Debian repositories, and Conan. However, most SCA tools only rely on collecting popular TPLs from Github, ignoring the actual situation that developers may introduce TPLs from multiple channels, which directly limits the comprehensiveness of the SCA results.

[0007] (2) The existing SCA tools cannot be directly applied to the scenario of software library - level reuse detection. Software library - level reuse involves introducing a complete library or module as a whole to provide comprehensive functional support for a project, rather than just using individual functions or fragments. In contrast, function - level reuse focuses more on the precise invocation of a single function, while library - level reuse emphasizes integrity and systematicness. However, the existing SCA tools pay more attention to the reuse detection of functions or code fragments and are difficult to accurately identify library - level reuse. This limitation may lead to misjudging the use of some functions as the reuse of the entire library, thus misleading developers' understanding of the library usage in the project.

[0008] (3) The existing SCA tools have deficiencies in the analysis of TPL dependency relationships. In a complex software development environment, dependency tracking is becoming increasingly important. Inaccurate dependency analysis may trigger serious security problems. Especially in a multi - license environment, this deficiency will hinder license compatibility checks and increase legal risks. However, mainstream SCA tools (such as CENTRIS) lack the ability to analyze the dependency relationships between TPLs, which limits the comprehensiveness of vulnerability impact scope assessment and license compliance checks. Summary of the Invention

[0009] Aiming at the deficiencies of the prior art, the object of the present invention is to propose a software component analysis method for C / C++ source code, including:

[0010] Step 1: Collect TPL source code and build a final TPL feature library;

[0011] Step 2: According to the final TPL feature library, determine the TPL libraries reused by the target code, as well as the reused functions and reuse paths corresponding to each TPL library. The reused functions are the functions in the TPL library that are the same as those in the target code. The reused functions are written in the reused files, and the reuse paths are the file paths in the target code corresponding to the reused files;

[0012] Step 3: According to the reuse function and reuse path determined in Step 2, identify the files corresponding to the TPL library in the target code, and divide the files corresponding to the TPL library into the TPL modules corresponding to the TPL library to obtain multiple divided TPL modules;

[0013] Step 4: Analyze the dependency relationships between the reused files in the TPL module and analyze the dependency relationships between the TPL modules.

[0014] Optionally, Step 1 specifically includes:

[0015] Step 1.1: Collect the TPL source code, and all the TPL source code forms a data set;

[0016] Step 1.2: For each TPL source code in the data set, use the Ctags tool to parse the TPL source code to obtain function information, where the function information includes the function name, the file location where the function is located, and the line numbers at the start and end of the function; according to the incremental storage strategy, store the information of each TPL source code to obtain a TPL library, where the TPL library includes a source code summary, version information, a repository address, and a repository directory. The source code summary includes the TLSH hash value of each function in the TPL source code except for the test functions, the version of the TPL source code to which the function belongs, and the storage path of the TPL source code to which the function belongs; the version information includes the time and name of the TPL source code; the repository address includes the download address of the TPL source code; the repository directory includes a string of the directory structure of the TPL source code. The TPL libraries corresponding to all the TPL source codes in the data set form a TPL feature library;

[0017] Step 1.3: Perform screening and removal processing on the functions in the TPL feature library to obtain the final TPL feature library.

[0018] Optionally, Step 1.1 specifically includes:

[0019] Obtain TPL metadata in the application-level package manager, the operating system central repository, and the code repository. The TPL metadata includes the source code name, source, and download address. Remove duplicate data in the TPL metadata to obtain the deduplicated TPL metadata. The deduplicated TPL metadata constitutes an open-source software index library. Determine multiple open-source software according to actual needs. For each open-source software, use the CCScanner tool to parse it to obtain the TPL names referenced by the open-source software. Search in the open-source software index library according to the TPL names. If the source and download address can be found in the open-source software index library, download the TPL source code according to the source and download address. If the source and download address cannot be found in the open-source software index library, search and download the TPL source code from GitHub, GitLab, and Gitee according to the TPL names. The downloaded TPL source code is similarly parsed using the CCScanner tool to obtain the newly downloaded TPL source code. Repeat the operation of using the CCScanner tool to parse and download the source code until the number of downloaded TPL source codes reaches the preset threshold. Then all the TPL source codes form a dataset.

[0020] Optionally, step 1.3 specifically includes:

[0021] Step 1.3.1: Remove the external library functions and duplicate functions in the TPL feature library to obtain a target feature library;

[0022] Step 1.3.1.1: For the storage path of the TPL source code to which the functions in the source code summary of the TPL feature library belong, identify the storage path containing the path identifier. The path identifier includes 3rdParty, external, and deps. Delete the storage path containing the path identifier in the source code summary, as well as the TLSH hash value of its corresponding function and the version of the TPL source code to which the function belongs, to obtain a first feature library;

[0023] Step 1.3.1.2: Perform clustering analysis on the repository directory, screen out the TPL library with the highest update frequency, and use it as the representative TPL library. For the other TPL libraries in the first feature library except the representative TPL library, remove the functions in the other TPL libraries that are the same as those in the representative TPL library to obtain a target feature library;

[0024] Step 1.3.2: Remove the common functions in the target feature library to obtain the final TPL feature library;

[0025] Step 1.3.2.1: For two TPL libraries in the target feature library, calculate the similarity between the two TPL libraries, which is specifically implemented through the following formula:

[0026]

[0027] Among them, A is the set of TLSH hash values of all functions in a TPL library, B is the set of TLSH hash values of all functions in another TPL library, and J(A, B) is the similarity between the two TPL libraries;

[0028] Based on the similarity of TPL libraries, all TPL libraries are clustered through a hierarchical clustering algorithm to obtain multiple families, and each family contains multiple TPL libraries;

[0029] Step 1.3.2.2: Calculate the inverse library frequency of each function in the family, which is specifically implemented through the following formula:

[0030]

[0031] Among them, FILF(f) is the inverse library frequency of function f, N is the number of TPL libraries in the family, and d(f) is the number of TPL libraries in the family that contain function f;

[0032] Step 1.3.2.3: For each TPL library in the family, calculate the uniqueness index of the TPL library according to the inverse library frequency of each function in the TPL library, which is specifically implemented through the following formula:

[0033]

[0034] Among them, DI(TPL) is the uniqueness index of the TPL library, Totals is the total number of functions in the TPL library, and FILF(f i ) is the inverse library frequency of the i-th function f in the TPL library i ;

[0035] Step 1.3.2.4: Sort the uniqueness indices of all TPL libraries in the family in ascending order to obtain the first list. Starting from the first TPL library in the first list, mark all functions in the first TPL library as flag functions, and then obtain the next TPL library. Remove the marked functions contained in this TPL library, and mark the functions in this TPL library except for the marked functions, and then obtain the next TPL library. Repeat the above steps until the removal of the marked functions in the last TPL library in the first list is completed, so as to obtain the TPL library with duplicate functions removed, and then obtain the final TPL feature library.

[0036] Optionally, the TLSH hash value of each function in the TPL source code described in step 1.2 is obtained through the following method:

[0037] Extract the TPL source code to obtain the code of each function. For the code of each function, remove spaces, line breaks, and comments to obtain the processed function code. Then calculate the TLSH hash value of the processed function code and use it as the TLSH hash value of the function. Thus, obtain the TLSH hash value of each function in the TPL source code.

[0038] Optionally, step 2 specifically includes:

[0039] Step 2.1: Match the functions in the target code to be detected with the functions in each TPL library in the final TPL feature library to obtain the number Mfuncs of functions in each TPL library that match the target code. Then, for each TPL library, calculate the average number Afuncs of functions included in all versions of the TPL library. Further, calculate the ratio of Mfuncs to Afuncs. In the final TPL feature library, obtain the TPL libraries for which the ratio of Mfuncs to Afuncs is greater than or equal to the threshold α, and form the first candidate database.

[0040] Step 2.2: For each file in each TPL library in the first candidate database, match the functions included in the target code with the functions included in the file to obtain the number MFuncsFile of functions in the target code that match the file. Calculate the ratio of MFuncsFile to the total number TFuncsFile of functions in the file. In each TPL library in the first candidate database, determine the number Rfiles of files for which the ratio of MFuncsFile to TFuncsFile is greater than or equal to the threshold β.

[0041] Step 2.3: Calculate the ratio of Rfiles to the total number Tfiles of files in the TPL library. In the first candidate database, obtain the TPL libraries for which the ratio of Rfiles to Tfiles is greater than or equal to the threshold θ to obtain the TPL libraries reused by the target code. Then, determine the reused functions and reuse paths corresponding to each TPL library.

[0042] Optionally, step 3 specifically includes:

[0043] Step 3.1: For each reuse path, determine whether the reuse path contains a TPL name. If the reuse path contains a TPL name, divide the reuse files under the reuse path into the TPL module corresponding to the TPL name; if the reuse path does not contain a TPL name, perform step 3.2.

[0044] Step 3.2: Determine whether the directory where the reusable files are located contains the header files of TPL. If the directory contains the header files of TPL, divide all the reusable files in this directory into the TPL module corresponding to the header files of TPL. If the directory does not contain the header files of TPL, execute Step 3.3;

[0045] Step 3.3: Determine whether the directory where the reusable files are located contains TPL metadata files. The TPL metadata files include README, LICENSE, and COPYING. If the directory does not contain TPL metadata files, do nothing. If the directory contains TPL metadata files, determine whether the TPL metadata files contain the TPL name. If the TPL metadata files contain the TPL name, divide all the reusable files in this directory into the TPL module corresponding to the TPL name; if the TPL metadata files do not contain the TPL name, do nothing.

[0046] Optionally, after Step 3.3, it further includes:

[0047] Step 3.4: For all the reusable files corresponding to each TPL library, if there are reusable files among all the reusable files that can be divided through Steps 3.1 to 3.3, the division ends. If all the reusable files cannot be divided through Steps 3.1 to 3.3, obtain the total number of reusable functions in the directory where the reusable files are located, and then calculate the ratio of it to the functions in the directory where the reusable files are located to obtain the reuse frequency of the directory where the reusable files are located, and then obtain the reuse frequency of each directory where the reusable files are located. Select the directory corresponding to the maximum value among all the reuse frequencies, and divide all the reusable files in this directory into the module corresponding to this TPL library.

[0048] Optionally, Step 4 specifically includes:

[0049] Step 4.1: Analyze two reusable files for each TPL module. If one reusable file directly depends or indirectly depends on another reusable file, at this time, ignore the dependency relationship between the two reusable files;

[0050] Among them, that one reusable file directly depends on another reusable file means that one reusable file directly references another reusable file through the #include instruction or the extern keyword declaration. That one reusable file indirectly depends on another reusable file means that one reusable file references another reusable file through other files;

[0051] Step 4.2: For the dependency relationships between TPL modules, determine whether there are direct dependencies or indirect dependencies between the reused files in two TPL modules. If there are direct dependencies or indirect dependencies between the reused files in two TPL modules, it indicates that there is a dependency relationship between the two TPL modules. If there are no direct dependencies or indirect dependencies between the reused files in two TPL modules, it indicates that there is no dependency relationship between the two TPL modules.

[0052] The beneficial effects of adopting the above technical solutions are as follows:

[0053] The present invention proposes a flexible and efficient data collection mechanism to comprehensively construct a TPL feature library to improve the quality of the feature library. At the same time, it initially filters functions that are not specific TPLs in combination with the directory structure to reduce noise interference. Then, a clustering algorithm is used to classify similar TPLs into the same family. A certain number of common functions are shared within each family, and at the same time, it includes the signature functions unique to each TPL. Based on the assumption that the proportion of signature functions in the reused library is relatively low, the reuse relationships of the shared common functions among various TPLs are analyzed, and the common functions that do not belong to the scope of a specific TPL are removed, thereby overcoming the defects related to the function birth time. The present invention also comprehensively analyzes the reuse ratios of the target code at the function and file levels, accurately determines the TPL reuse situation of the code, divides the target code into TPL modules, and uses include and extern statements to analyze the dependency relationships between modules, thereby achieving more accurate component identification and dependency analysis. Brief Description of the Drawings

[0054] Figure 1 It is a schematic flowchart of a software component analysis method for C / C++ source code in an embodiment of the present invention;

[0055] Figure 2 It is a schematic structural diagram of the storage method in an embodiment of the present invention, where Figure (a) is a schematic structural diagram of the storage method of a traditional SCA tool, and Figure (b) is a schematic structural diagram of an incremental storage strategy;

[0056] Figure 3 It is a dependency detection framework diagram in an embodiment of the present invention. Detailed Embodiments

[0057] The following combines the drawings and embodiments to further describe in detail the specific embodiments of the present invention. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.

[0058] In view of the problems existing in the prior art, the present invention provides a software component analysis method for C / C++ source code, which is used to analyze the source code, modules, frameworks, libraries, etc. of open-source software and third-party commercial software. The main purpose of SCA technology is to identify the external components used in a software project and their dependencies, detect known security vulnerabilities, outdated patches, and license compliance risks, so as to ensure the security and manageability of the software supply chain.

[0059] According to the different inputs and targets of software component analysis, SCA can be divided into Source-to-Source (S2S) analysis and Binary-to-Source (B2S) analysis. The present invention focuses on the S2S analysis technology in the C / C++ field. S2S analysis directly parses the source code to analyze the dependencies, imported library files, header files, etc. in the project, and identifies the referenced external components and libraries. This method is commonly used in scenarios such as open-source software management and license compliance analysis.

[0060] Combined with Figure 1 , the present invention specifically may include the following steps:

[0061] Step 1: Collect TPL source code and build a final TPL feature library;

[0062] Step 1.1: Collect TPL source code, and all the TPL source code forms a data set;

[0063] In the application-level package manager, the operating system central repository, and the code repository, obtain TPL metadata, where the TPL metadata includes the source code name, source, and download address. In a specific implementation, the number of TPL metadata collected in the present invention is shown in Table 1.

[0064] Table 1 TPL metadata collected

[0065]

[0066]

[0067] Remove duplicate data in the TPL metadata to obtain the deduplicated TPL metadata. The deduplicated TPL metadata forms an open-source software index library. Determine multiple open-source software according to actual needs. For each open-source software, use the CCScanner tool to parse and obtain the TPL names cited by the open-source software. Search in the open-source software index library according to the TPL names. If the source and download address can be found in the open-source software index library, then download the TPL source code according to the source and download address. If the source and download address cannot be found in the open-source software index library, search and download the TPL source code from GitHub, GitLab, and Gitee according to the TPL names. Similarly, use the CCScanner tool to parse the downloaded TPL source code, and then obtain the newly downloaded TPL source code, that is, use the CCScanner tool to parse and obtain the TPL names cited by the downloaded TPL source code. Search in the open-source software index library according to the TPL names. If the source and download address can be found in the open-source software index library, then download the TPL source code, that is, the newly downloaded TPL source code. Repeat the operation of using the CCScanner tool to parse and download the source code until the number of downloaded TPL source codes reaches the preset threshold. Then all the TPL source codes form a dataset.

[0068] This method efficiently constructs a widely covered C / C++ TPL dataset, which not only covers popular TPLs but also introduces more relevant resources through dependency mining.

[0069] Step 1.2: For each TPL source code in the dataset, use the Ctags tool to parse the TPL source code to obtain function information. The function information includes the function name, the file location where the function is located, and the line numbers at the start and end of the function. According to the incremental storage strategy, store the information of each TPL source code and the function information to obtain a TPL library.

[0070] Since there are usually a large number of duplicate codes between different versions in the version evolution of existing open-source software. Traditional SCA tools store function records separately for each version. As shown in Figure 2 (a), Version 1, Version k, and Version n are different versions, and the corresponding function a, function b, and function c behind are the functions included in the versions. Such a storage method leads to waste of storage space and the problem of repeated comparison of the common parts of multiple versions and the target code. Therefore, the present invention adopts an incremental storage strategy. For functions that are repeated across versions, only store a single record and append version and path information; reduce storage redundancy and at the same time reduce the computational complexity. The storage structure refers to Figure 2 Figure (b) in.

[0071] Combined with Figure 2 in Figure (b), the TPL library includes source code summaries, version information, repository addresses, and repository directories. The source code summaries include the TLSH hash values of each function in the TPL source code except for the test functions, that is Figure 2 Hash(function a), Hash(function b), and Hash(function c) in Figure (b) in , where the TLSH hash values of each function in the TPL source code are obtained by the following method:

[0072] Extract the TPL source code to obtain the code of each function. For the code of each function, remove spaces, line breaks, and comments to obtain the processed function code, and then calculate the TLSH hash value of the processed function code and use it as the TLSH hash value of the function. In this way, the TLSH hash values of each function in the TPL source code are obtained.

[0073] The source code summary also includes the version of the TPL source code to which the function belongs, that is Figure 2 [version1,...,version k], [version 1,...,version k], and [version 1,...,version n] in Figure (b) in , and the source code summary also includes the storage path of the TPL source code to which the function belongs, that is Figure 2 [path1,...,path k], [path1,...,path k], and [path1,...,path n] in Figure (b) in .

[0074] The version information includes the time and name of the TPL source code, that is Figure 2 "Version 1: version1_time" and "Version n: version n_time" in Figure (b) in . The version information is used to evaluate the activity level of the repository;

[0075] The repository address includes the download address of the TPL source code. Specifically, it may include the URL of the TPL source code. The repository address supports cloning and updating via Git commands.

[0076] The repository directory includes a string of the directory structure of the TPL source code, which is convenient for identifying whether repositories with similar directories belong to the same TPL.

[0077] The TPL libraries corresponding to all TPL source codes in the dataset constitute the TPL feature library.

[0078] It should be noted that in order to improve the representativeness and accuracy of TPL features, test functions will be excluded during the construction of the source code summary. These functions are usually used to verify functions, performance, and stability, and are located in specific test directories (such as ". / test") and do not directly participate in the implementation of core functions. Removing test functions can effectively avoid interference and improve the accuracy of feature matching. Therefore, in the present invention, when constructing the TPL library, test functions can be removed first in the source code summary part, and then the source code summary part can be constructed.

[0079] Since there are multiple TPL libraries in the TPL feature library that contain the same functions, the widespread distribution of such functions blurs the TPL boundary and affects the accuracy of SCA. For this reason, the following multi-stage processing is required. For details, see Step 1.3.

[0080] Step 1.3: Screen and remove the functions in the TPL feature library to obtain the final TPL feature library.

[0081] Step 1.3.1: Remove the external library functions and duplicate functions in the TPL feature library to obtain the target feature library;

[0082] Step 1.3.1.1: For the storage path of the TPL source code to which the functions in the source code summary of the TPL feature library belong, identify the storage paths containing path identifiers. The path identifiers include 3rdParty, external, and deps. Delete the storage paths containing path identifiers in the source code summary, as well as the TLSH hash values of their corresponding functions and the versions of the TPL source code to which the functions belong, to obtain the first feature library;

[0083] Multiple repositories of the same TPL may be maintained due to cross-platform or specific requirements, but the directory structures and function implementations are highly similar. Therefore, further analysis is performed. For details, refer to Step 1.3.1.2.

[0084] Step 1.3.1.2: Perform clustering analysis on the repository directories. Specifically, clustering analysis can be performed through the K-means algorithm to screen out the TPL library with the highest update frequency and use it as the representative TPL library. For other TPL libraries in the first feature library except the representative TPL library, remove the functions in other TPL libraries that are the same as those in the representative TPL library to obtain the target feature library;

[0085] Developers copy external library code without indicating the source or independently implement basic algorithms (such as sorting and searching), resulting in the widespread distribution of common functions and increasing the difficulty of differentiation. For this reason, Step 1.3.2 is performed for further analysis and removal.

[0086] Step 1.3.2: Remove the common functions in the target feature library to obtain the final TPL feature library;

[0087] Step 1.3.2.1: Calculate the similarity between two TPL libraries in the target feature library, which is specifically implemented through the following formula:

[0088]

[0089] Among them, A is the set of TLSH hash values of all functions in one TPL library, B is the set of TLSH hash values of all functions in another TPL library, and J(A, B) is the similarity between the two TPL libraries;

[0090] Based on the similarity of TPL libraries, use the Hierarchical Clustering algorithm to cluster all TPL libraries to obtain multiple families. Each family contains multiple TPL libraries. Among them, a certain number of common functions are shared within each family, and each TPL has its own unique signature function. By comparing the proportion of signature functions, the common functions are preferentially assigned to the TPL with a lower proportion of signature functions.

[0091] Step 1.3.2.2: Calculate the inverse library frequency of each function in the family, which is specifically implemented through the following formula:

[0092]

[0093] Among them, FILF(f) is the inverse library frequency of function f, N is the number of TPL libraries in the family, and d(f) is the number of TPL libraries in the family that contain function f;

[0094] Among them, FILF is a statistic that measures the universality of a function in multiple TPLs. It is based on the concept of inverse document frequency (IDF). The higher the FILF value, the fewer TPLs the function appears in, and thus it has higher uniqueness.

[0095] Step 1.3.2.3: For each TPL library in the family, calculate the uniqueness index of the TPL library according to the inverse library frequency of each function in the TPL library, which is specifically implemented through the following formula:

[0096]

[0097] Among them, DI(TPL) is the uniqueness index of the TPL library, Totals is the total number of functions in the TPL library, and FILF(f i ) is the inverse library frequency of the i-th function f i in the TPL library;

[0098] Among them, DI is a quantitative indicator that measures the proportion of signature functions in a specific family for a TPL. The higher the DI value, the more unique the functions in the TPL are and the less dependent on common functions.

[0099] Step 1.3.2.4: Sort the uniqueness indices of all TPL libraries in the family in ascending order to obtain a first list. Starting from the first TPL library in the first list, mark all functions in the first TPL library as flagged functions, then obtain the next TPL library, remove the flagged functions included in this TPL library, and mark the functions other than the flagged functions in this TPL library. Then obtain the next TPL library and repeat the above steps until the removal of the flagged functions in the last TPL library in the first list is completed, thus obtaining the TPL libraries with duplicate functions removed, and further obtaining the final TPL feature library.

[0100] To better adapt to the software library granularity reuse detection scenario, the Ctags and TLSH algorithms are used to represent the target code as a feature set, and a multi-level and multi-threshold evaluation system is designed to perform similarity matching between the target code features and the TPLs in the feature library. Only when a specific TPL meets the triple thresholds simultaneously can it be confirmed as an effective component of the target code. For the specific execution steps, refer to Step 2.

[0101] Step 2: According to the final TPL feature library, determine the TPL libraries reused by the target code, as well as the reused functions and reuse paths corresponding to each TPL library. The reused functions are the functions in the TPL library that are the same as those in the target code. The reused functions are written in the reused file, and the reuse path is the path of the reused file in the file corresponding to the target code;

[0102] Step 2.1: Match the functions in the target code to be detected with the functions in each TPL library in the final TPL feature library to obtain the number Mfuncs of functions in each TPL library that match the target code. Then, for each TPL library, calculate the average number Afuncs of functions included in all versions of this TPL library, and then calculate the ratio of Mfuncs to Afuncs. In the final TPL feature library, obtain the TPL libraries whose ratio of Mfuncs to Afuncs is greater than or equal to the threshold α (Function-Level Reuse Metric) to form a first candidate database;

[0103] Step 2.2: For each file in each TPL library in the first candidate database, match the functions included in the target code with the functions included in this file to obtain the number MFuncsFile of functions in the target code that match this file, calculate the ratio of MFuncsFile to the total number TFuncsFile of functions in this file, and in each TPL library in the first candidate database, determine the number Rfiles of files for which the ratio of MFuncsFile to TFuncsFile is greater than or equal to the threshold β (File-Level Metric);

[0104] Step 2.3: Calculate the ratio of Rfiles to the total number Tfiles of files in the TPL library. In the first candidate database, obtain the TPL libraries for which the ratio of Rfiles to Tfiles is greater than or equal to the threshold θ (Library-Level Metric) to get the TPL libraries for target code reuse, and further determine the reusable functions and reuse paths corresponding to each TPL library.

[0105] Step 3: According to the reusable functions and reuse paths determined in Step 2, identify the files in the target code corresponding to the TPL library, and divide the files corresponding to this TPL library into the TPL modules corresponding to this TPL library to obtain multiple divided TPL modules;

[0106] Combined with Figure 3 , the TPL libraries for target code reuse obtained in Step 2 of the present invention and the reusable functions corresponding to each TPL library correspond to the target code TPL list. Further, it is necessary to perform TPL component module division through the reuse path and reuse files, that is Figure 3 the module division performed according to the directory structure and metadata in

[0107] The obtained TPLA Module and TPLB Module are the divided TPL modules. The specific division includes: For each reuse path, determine whether the reuse path contains the TPL name. If the reuse path contains the TPL name, divide the reuse files under this reuse path into the TPL module corresponding to this TPL name. For example, the reuse path of spirv-tools is "skia / third_party / externals / spirv-tools / source / opt / optimizer.cpp", and the reuse path contains "spirv-tools", so optimizer.cpp is divided into the spirv-tools module; if the reuse path does not contain the TPL name, perform Step 3.2;

[0108] Step 3.2: Determine whether the directory where the reused files are located contains the header files of TPL, i.e., .h files. If the directory contains the header files of TPL, all the reused files in this directory are classified into the TPL module corresponding to the TPL header file. For example, if the directory contains the libfoo.h file and there is code reuse between the reused files in this directory and the libfoo library, the reused files in this directory should belong to the libfoo module; if the directory does not contain the header files of TPL, execute Step 3.3;

[0109] Taking the reuse path A / B / C / D.cpp as an example, where D.cpp is the reused file and the C directory is the directory where the reused file is located. It should be noted that in the present invention, the reuse directory may contain multiple reused files. If the directory where a certain reused file is located contains a TPL header file, all the reused files in this directory are directly classified into the corresponding TPL module, and there is no need to repeatedly judge other reused files in this directory.

[0110] Step 3.3: Determine whether the directory where the reused file is located contains TPL metadata files. TPL metadata files include README, LICENSE, and COPYING. If the directory does not contain TPL metadata files, no processing is performed. If the directory contains TPL metadata files, determine whether the TPL metadata files contain the TPL name. If the TPL metadata files contain the TPL name, all the reused files in this directory are classified into the TPL module corresponding to the TPL name; if the TPL metadata files do not contain the TPL name, no processing is performed.

[0111] After the above judgments, there may be a situation where all the reused files corresponding to the TPL library cannot be further judged through Steps 3.1 to 3.3. In this case, the reuse frequency of functions can be relied on to determine whether the reused file is an actual component of TPL. The specific execution steps refer to Step 3.4.

[0112] Step 3.4: For all the reused files corresponding to each TPL library, if there are reused files that can be classified through Steps 3.1 to 3.3 among all the reused files, the classification ends. If all the reused files cannot be classified through Steps 3.1 to 3.3, obtain the total number of reused functions in the directory where the reused file is located, and then calculate the ratio of it to the functions in the directory where the reused file is located to obtain the reuse frequency of the directory where the reused file is located, and then obtain the reuse frequency of each directory where the reused file is located. Select the directory corresponding to the maximum value among all the reuse frequencies, and classify all the reused files in the directory into the module corresponding to this TPL library.

[0113] For example, in the target code, TPL_A is reused, and the corresponding reused file paths are: A / B / C / d.cpp, A / B / E / f.cpp, and A / B / G / h.cpp. Taking the C directory as an example, there are 5 functions in d.cpp that are the same as TPL_A, and there are a total of 10 functions in all source code files in the C directory. Therefore, the reuse frequency of the C directory is 5 / 10 = 0.5. Similarly, the reuse frequency of the E directory is 0.4, and the reuse frequency of the G directory is 0.3. According to the principle of the highest reuse frequency, only the reused files related to TPL_A in the C directory (such as d.cpp) are divided into the TPL_A module, while the reused files related to TPL_A in the E and G directories (such as f.cpp and h.cpp) are not divided into the TPL_A module.

[0114] Step 4: Analyze the dependency relationships between the reused files in the TPL module and the dependency relationships between the TPL modules.

[0115] Step 4.1: Analyze two reused files for each TPL module. If one reused file directly depends or indirectly depends on another reused file, at this time, ignore the dependency relationship between the two reused files;

[0116] Combined with Figure 3 , Step 4.1 corresponds to the part of eliminating internal dependencies in the figure, detecting the internal file reference relationships in the module, hiding the reference relationships that only exist internally, and avoiding internal dependencies from interfering with the analysis results.

[0117] Among them, the situation where one reused file directly depends on another reused file means that one reused file directly references another reused file through an #include instruction declaration or an extern keyword declaration.

[0118] Next, a detailed description will be given for the direct reference through an #include instruction declaration:

[0119] If file A includes file B (.h or.hpp) through an #include instruction, so that it can access the variables, functions, or type definitions declared in file B, at this time, it means that A directly depends on file B.

[0120] For the case where there are multiple header files with the same name, follow the path resolution rules of GNU GCC, select the file with the closest path, that is, the file closest to the path of file A, or select the header file with a clear specification. If the path conflict cannot be resolved, leave such conflicts for subsequent indirect dependency analysis.

[0121] Next, a detailed description will be given for the direct reference through an extern keyword declaration:

[0122] The present invention generates a code index through the Ctags tool, extracts all statements related to extern, records the names of variables and functions and their definition and declaration locations. By parsing this information, it can be confirmed whether File A references external resources in File B. If File A declares and references an external variable or function in File B, thereby allowing File A to use the global variables or functions in File B, it indicates that A directly depends on File B.

[0123] Saying that one reusable file indirectly depends on another reusable file means that one reusable file references another reusable file through other files. That is to say, if File A references File C and File C references File B, then File A indirectly references File B.

[0124] For indirect dependencies, the present invention can analyze the class, structure, and enumeration names defined and used in File A and File B. If the same identifiers are found and their namespaces and scopes are consistent, it can be determined that File A indirectly depends on File B.

[0125] Step A2: For the dependency relationships between TPL modules, determine whether there are direct dependencies or indirect dependencies between the reusable files in two TPL modules. If there are direct dependencies or indirect dependencies between the reusable files in two TPL modules, it indicates that there is a dependency relationship between the two TPL modules. If there are no direct dependencies or indirect dependencies between the reusable files in two TPL modules, it indicates that there is no dependency relationship between the two TPL modules.

[0126] That is to say, the present invention obtains a reusable file in one TPL module and a reusable file in another TPL module, and determines the dependency relationship between the two modules by analyzing the dependency relationship between these two reusable files. Among them, to improve efficiency, once the dependency relationship between modules is established, there is no need to repeatedly analyze the dependencies of other files in the same module pair. That is to say, if there is a direct dependency or an indirect dependency between two modules, there is no need to analyze the dependency relationship between other reusable files in the two modules.

[0127] Combined with Figure 3 , Step A2 corresponds to the part of establishing external dependencies in the figure, checking the dependency relationships between different modules. After performing the dependency analysis between modules, combined with Figure 3 the dependency analysis results can be obtained. The TPLA module directly depends on TPLB and TPLC, the TPLB module directly depends on TPLD, and the TPLA module indirectly depends on TPLD.

[0128] The present invention provides more precise TPL location and dependency analysis capabilities for the scenario of library-level reuse. To verify its effectiveness, the present invention conducts experiments in 100 actual projects, compares with mainstream tools CENTRIS, TPLite, and OSSFP, and simultaneously conducts actual development and application verification.

[0129] Among them, the reference for CENTRIS is: "Seunghoon Woo, Sunghan Park, Seulbae Kim, HeejoLee, and Hakjoo Oh. 2021. CENTRIS: A Precise and Scalable Approach for Identifying Modified Open-Source Software Reuse. In 2021 IEEE / ACM 43rd International Conference on Software Engineering (ICSE). IEEE, 860–872"; the reference for TPLite is: "Ling Hengchen Yuan, Qiyi Tang, Sen Nie, Shi Wu, and Yuqun Zhang*. 2023. Third-Party Library Dependency for Large-Scale SCA in the C / C++ Ecosystem: How Far Are We?. In Proceedings of the 32nd ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA’23), July 17–21, 2023, Seattle, WA, United States. ACM, New York, NY, USA, 13 pages. https: / / doi.org / 10.1145 / 3597926.3598143”; OSSFP reference: “J. Wu et al, "OSSFP: Precise and Scalable C / C++ Third-Party Library Detection using Fingerprinting Functions," 2023 IEEE / ACM 45th International Conference on Software Engineering (ICSE), Melbourne, Australia, 2023, pp. 270-282, doi: 10.1109 / ICSE48619.2023.00034.”.

[0130] Based on this, the experimental results of the present invention are as follows:

[0131] The present invention conducts detections on 201 TPL components involved in 100 standard set projects with CENTRIS, TPLite, and OSSFP, analyzes the experimental results, and explores the reasons for false positives and false negatives.

[0132] Comparison with CENTRIS. Given that CENTRIS has released a feature library and source code, the present invention directly conducts detections on 100 projects based on its feature library. CENTRIS determines the attribution of common functions by the birth time of functions, believing that the library with an earlier birth time is more likely to be the “origin” of common functions. Therefore, when constructing the feature library, the common functions in the later-born libraries are regarded as non-critical features and excluded. When the proportion of functions in the target code that are the same as those in the TPL reaches 10%, it is considered that the TPL is an effective component of the target code.

[0133] As shown in Table 2, CENTRIS identified 184 components from 100 projects, with a precision of 51.63% and a recall rate of 47.26%. CENTRIS still relies on the function birth time to handle the attribution problem of common functions. The limitation of the function birth time makes the attribution judgment of common functions inaccurate, thus causing false positives.

[0134] The present invention introduces file - granularity analysis and a more stringent threshold criterion to judge the existence of TPL, thus performing excellently in reducing false positives, with a precision reaching 90.63%. In addition, since the feature library of CENTRIS only contains 10,294 TPLs, some real components are omitted, resulting in a large number of false negatives (FN) and reducing the recall rate. While the present invention uses a more comprehensive feature library, with a recall rate reaching 86.57%, higher than that of CENTRIS.

[0135] Table 2 Comparison experimental results of the present invention and related tools in the ability to identify TPL components

[0136]

[0137] Comparison with TPLite. TPLite has published its source code but not its feature library. Therefore, the present invention re - executed the process of constructing the feature library of TPLite and used the same TPL data set as the present invention, and finally constructed a feature library containing 32,706 TPLs. Based on CENTRIS's derivation of TPL dependencies based on function birth time, TPLite constructed the feature library into a directed acyclic graph and combined the PageRank algorithm and the centrality filter of node in - degree to correct the inaccurate dependencies derived from function birth time. Similar to CENTRIS, TPLite also uses a 10% threshold to judge whether a TPL is a valid component of the target code.

[0138] As shown in Table 2, the present invention is significantly superior to TPLite in terms of precision. Most of its false - positive (FP) cases stem from the fact that the reported components reuse part of the code of real components. This indicates that although TPLite uses a centrality filter to make up for the limitations of function birth time, it still cannot accurately attribute common functions, thus leading to the problem of false positives.

[0139] TPLite focuses on the centrality of TPLs in the global graph and determines which dependencies to retain by evaluating the importance of TPLs in the graph. If a certain TPL has insufficient connectivity in the global graph, related dependencies may be erroneously deleted, resulting in functions being misclassified as other TPLs. For example, a core TPL that is only used in a specific domain and is reused by only a few other TPLs may have a low importance in the global graph, and the dependencies pointing to this TPL may be deleted, leading to these functions being misclassified as other TPLs.

[0140] In contrast, the present invention adopts a local analysis method to avoid global noise interference. It divides the feature library into multiple "families" with high code similarity and analyzes the dependencies between TPLs within each family to ensure that the functions of each TPL itself are retained and non-TPL functions are removed. Therefore, when calculating the overlap ratio between the TPL and the target code, the denominator can accurately reflect the actual number of functions of each TPL, avoiding the false negative (FN) problem caused by an overly high preset threshold or an overly small matching ratio. The experimental results further verify that although the present invention uses strict multiple thresholds, the recall rate can still be retained.

[0141] Comparison with OSSFP. Since the commercial tool has not released its feature library and source code, the present invention uploads 100 projects to the OSSFP website for online detection and uses the standard set and the results generated by OSSFP to manually check its accuracy. OSSFP divides the functions of TPLs into cloned functions, supporting functions, general functions, and core functions, and removes the first three types of functions by setting thresholds and the function birth time, and finally generates TPL features only based on the core functions to solve the internal cloning problem between TPLs. OSSFP believes that as long as the target code is the same as any function in the TPL, then this TPL is a valid component of the target code.

[0142] 126 correct cases were confirmed from the results generated by OSSFP, with the precision and recall rates being 61.46% and 62.69% respectively, as shown in Table 2. There are a large number of components in the form of special variables such as git, lib, and similar ${library}, and single letters (such as m) in the OSSFP results. These components may be the results of analyzing the directory structure of the target code, so the false positive (FP) cases increase significantly. In addition, since OSSFP considers that the target code reuses the TPL as long as a single function is the same, this does not meet the requirements of library-level reuse detection, which further leads to an increase in FPs.

[0143] Compared with the present invention, the recall rate of OSSFP is lower, probably because the scope of the feature library is incomplete. OSSFP constructs a TPL feature library based on 23,427 C / C++ repositories on GitHub, but its scale is still smaller than the 33,100 TPL feature libraries collected from 15 platforms in the present invention. In addition, since OSSFP does not disclose more details about TPL reuse, the accuracy of its results can only be compared by comparing the TPL names with the component names in the standard set, and errors may be introduced in the checking process, thus affecting the judgment of precision and recall rate.

[0144] The method proposed in the present invention is superior to current mainstream tools in both accuracy and adaptability. By preprocessing the feature library and verifying the existence of TPL through a triple-threshold strategy, the present invention effectively reduces the false positive problem and ensures the precision of component analysis. Although the recall rate is slightly lower than that of some tools due to the influence of the threshold, the threshold can be dynamically adjusted according to requirements in actual applications to achieve a balance between accuracy and recall rate and meet the component analysis requirements of different scenarios. When evaluating the TPL dependency detection ability, we used 1,784 pairs of file dependencies in the OpenHarmony project as the standard set and achieved a recall rate of 99% and a precision of 95%, as shown in Table 3. It can be seen from this that the present invention can accurately detect the dependencies between file sets. On the premise that the TPL module division is accurate, each TPL module can be regarded as an independent file set. Therefore, it can be inferred that the present invention can also efficiently identify and extract the dependencies between TPLs.

[0145] Table 3 Experimental results of the present invention in detecting dependencies between file sets

[0146]

[0147] To verify the application value of the present invention in actual development, 689 code repositories in the OpenHarmony community were analyzed, and 166 external components were successfully identified, among which 61 external components were labeled for the first time and were recognized by community developers. The author team was awarded the "Outstanding Contribution Team Award for Security Governance" by the OpenHarmony community, further demonstrating the important value of the present invention in open source community management.

[0148] Based on this, in response to the detection requirements of C / C++ software library granularity reuse, the present invention proposes a component analysis technology, and the main innovation points and technical advantages include:

[0149] (1) Comprehensive construction of the TPL feature library: A cross-platform data collection mechanism is designed, covering 15 representative platforms (such as Debian, GitHub, Conan, etc.). The function information of 33,100 software is extracted to construct a TPL feature library containing 30,047,290 functions, providing comprehensive data support for component analysis in the C / C++ ecosystem.

[0150] (2) Preprocessing of the feature library: Noise functions are identified and filtered through algorithms, shared common functions are cleaned, and external dependencies and TPL's own features are accurately distinguished, thus effectively solving the problem of fuzzy TPL boundaries and improving the accuracy of component analysis.

[0151] (3) Precise TPL component localization: Through the preprocessing of the feature library and the TPL existence judgment criteria with multiple thresholds, in the scenario of reuse detection at the software library granularity, the precise localization of software components is achieved, providing reliable support for developers to understand and analyze the software composition structure.

[0152] (4) Automatically constructing TPL dependency relationships: Based on the results of component analysis, the target code is divided into different TPL component modules. The direct dependency relationships are parsed based on the include and extern statements in the code, and the indirect dependency relationships are inferred by combining identifiers such as classes and structures, automatically generating a component dependency view to support the analysis of the propagation path of security vulnerabilities and license compatibility analysis.

[0153] (5) Experimental verification and community recognition: In 100 actual projects, the accuracy of component analysis reaches 90.63%, which is significantly better than existing tools (such as CENTRIS, TPLite, OSSFP). The accuracy of dependency relationship detection reaches 95%, fully meeting the requirements of automatically constructing TPL dependency relationships. This method has been adopted by the OpenHarmony community (https: / / gitee.com / openharmony-sig / compliance_composition_analysis), helping its 689 code repositories accurately identify 166 external components, 61 of which are first marked by the tool and recognized by developers. The team thus won the "Outstanding Contribution Team Award for Security Governance" of the OpenHarmony community.

[0154] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) disclosed in the embodiments of the present disclosure that have similar functions.

Claims

1. A software component analysis method for C / C++ source code, characterized in that Including: Step 1: Collect TPL source code and construct the final TPL feature library; Step 2: According to the final TPL feature library, determine the TPL libraries for target code reuse, as well as the reuse functions and reuse paths corresponding to each TPL library. The reuse functions are the functions in the TPL library that are the same as those in the target code. The reuse functions are written in the reuse file, and the reuse path is the path of the reuse file in the file corresponding to the target code; Step 3: According to the reuse functions and reuse paths determined in Step 2, identify the files in the target code corresponding to the TPL library, and divide the files corresponding to the TPL library into the TPL modules corresponding to the TPL library to obtain multiple divided TPL modules; Step 4: Analyze the dependency relationships between the TPL modules by analyzing the dependency relationships between the reuse files in the TPL modules.

2. The software component analysis method for C / C++ source code according to claim 1, characterized in that Step 1 specifically includes: Step 1.1: Collect TPL source code, and all TPL source code forms a dataset; Step 1.2: For each TPL source code in the dataset, use the Ctags tool to parse the TPL source code to obtain function information, where the function information includes the function name, the file location where the function is located, and the line numbers at the start and end of the function; according to the incremental storage strategy, store the information of each TPL source code to obtain a TPL library. The TPL library includes a source code summary, version information, repository address, and repository directory. The source code summary includes the TLSH hash value of each function in the TPL source code except for the test functions, the version of the TPL source code to which the function belongs, and the storage path of the TPL source code to which the function belongs; the version information includes the time and name of the TPL source code; the repository address includes the download address of the TPL source code; the repository directory includes a string of the directory structure of the TPL source code. The TPL libraries corresponding to all TPL source codes in the dataset form the TPL feature library; Step 1.3: Perform screening and removal processing on the functions in the TPL feature library to obtain the final TPL feature library.

3. The software component analysis method for C / C++ source code according to claim 2, characterized in that, Step 1.1 specifically includes: Obtain TPL metadata in the application-level package manager, the operating system central repository, and the code repository. The TPL metadata includes source code names, sources, and download addresses. Remove duplicate data in the TPL metadata to obtain the deduplicated TPL metadata. The deduplicated TPL metadata constitutes an open-source software index library. Determine multiple open-source software according to actual needs. For each open-source software, use the CCScanner tool to parse it to obtain the TPL names referenced by the open-source software. Search in the open-source software index library according to the TPL names. If the source and download address can be found in the open-source software index library, download the TPL source code according to the source and download address. If the source and download address cannot be found in the open-source software index library, search and download the TPL source code from GitHub, GitLab, or Gitee according to the TPL name. The downloaded TPL source code is similarly parsed using the CCScanner tool to obtain the newly downloaded TPL source code. Repeat the operation of using the CCScanner tool to parse and download the source code until the number of downloaded TPL source codes reaches the preset threshold. Then all the TPL source codes form a dataset.

4. The software component analysis method for C / C++ source code according to claim 2, characterized in that, Step 1.3 specifically includes: Step 1.3.1: Remove external library functions and duplicate functions in the TPL feature library to obtain the target feature library; Step 1.3.1.1: For the storage path of the TPL source code to which the function in the source code abstract of the TPL feature library belongs, identify the storage path containing the path identifier. The path identifier includes 3rdParty, external, and deps. Delete the storage path containing the path identifier in the source code abstract, as well as the TLSH hash value of its corresponding function and the version of the TPL source code to which the function belongs, to obtain the first feature library; Step 1.3.1.2: Perform clustering analysis on the repository directory, screen out the TPL library with the highest update frequency as the representative TPL library. For other TPL libraries in the first feature library except the representative TPL library, remove the functions in other TPL libraries that are the same as those in the representative TPL library to obtain the target feature library; Step 1.3.2: Remove the common functions in the target feature library to obtain the final TPL feature library; Step 1.3.2.1: Calculate the similarity between two TPL libraries in the target feature library, which is specifically implemented through the following formula: where A is the set of all function TLSH hash values in one TPL library, B is the set of all function TLSH hash values in another TPL library, and J(A,B) is the similarity between the two TPL libraries; Based on the similarity of the TPL libraries, perform clustering on all TPL libraries through the hierarchical clustering algorithm to obtain multiple families, and each family contains multiple TPL libraries; Step 1.3.2.2: Calculate the inverse library frequency of each function in the family, which is specifically implemented through the following formula: where FILF(f) is the inverse library frequency of function f, N is the number of TPL libraries in the family, and d(f) is the number of TPL libraries in the family that contain function f; Step 1.3.2.3: For each TPL library in the family, calculate the uniqueness index of the TPL library according to the inverse library frequency of each function in the TPL library. The specific implementation is through the following formula: Among them, DI(TPL) is the distinctiveness index of the TPL library, Totals is the total number of functions in the TPL library, and FILF(f i ) is the inverse library frequency of the i-th function f i in the TPL library; Step 1.3.2.4: Sort the uniqueness indices of all TPL libraries in the family in ascending order to obtain the first list. Starting from the first TPL library in the first list, mark all functions in the first TPL library as flagged functions, then obtain the next TPL library, remove the flagged functions included in this TPL library, and mark the functions other than the flagged functions in this TPL library. Then obtain the next TPL library, and repeat the above steps until the removal of the flagged functions in the last TPL library in the first list is completed, thereby obtaining the TPL library with duplicate functions removed, and thus obtaining the final TPL feature library.

5. The software component analysis method for C / C++ source code according to claim 2, wherein The TLSH hash value of each function in the TPL source code described in Step 1.2 is obtained through the following method: Extract the TPL source code to obtain the code of each function. For the code of each function, remove spaces, line breaks, and comments to obtain the processed code of the function, and then calculate the TLSH hash value of the processed code of the function, and use it as the TLSH hash value of the function. Thus, the TLSH hash value of each function in the TPL source code is obtained.

6. A software component analysis method for C / C++ source code according to claim 1, characterized in that Step 2 specifically includes: Step 2.1: Match the functions in the target code to be detected with the functions in each TPL library in the final TPL feature library to obtain the number Mfuncs of functions in each TPL library that match the target code. Then, for each TPL library, calculate the average number Afuncs of functions included in all versions of this TPL library, and then calculate the ratio of Mfuncs to Afuncs. In the final TPL feature library, obtain the TPL libraries whose ratio of Mfuncs to Afuncs is greater than or equal to the threshold α, and form the first candidate database; Step 2.2: For each file in each TPL library in the first candidate database, match the functions included in the target code with the functions included in this file to obtain the number MFuncsFile of functions in the target code that match this file, and calculate the ratio of MFuncsFile to the total number TFuncsFile of functions in this file. In each TPL library in the first candidate database, determine the number Rfiles of files whose ratio of MFuncsFile to TFuncsFile is greater than or equal to the threshold β; Step 2.3: Calculate the ratio of Rfiles to the total number Tfiles of files in the TPL library. In the first candidate database, obtain the TPL libraries whose ratio of Rfiles to Tfiles is greater than or equal to the threshold θ, to obtain the TPL libraries reused by the target code, and then determine the reused functions and reuse paths corresponding to each TPL library.

7. A software component analysis method for C / C++ source code according to claim 1, characterized in that, Step 3 specifically includes: Step 3.1: For each reuse path, determine whether the reuse path contains a TPL name. If the reuse path contains a TPL name, divide the reuse files under the reuse path into the TPL module corresponding to the TPL name; if the reuse path does not contain a TPL name, execute Step 3.2; Step 3.2: Determine whether the directory where the reuse file is located contains a TPL header file. If the directory contains a TPL header file, divide all the reuse files under the directory into the TPL module corresponding to the TPL header file; if the directory does not contain a TPL header file, execute Step 3.3; Step 3.3: Determine whether the directory where the reuse file is located contains a TPL metadata file. The TPL metadata file includes README, LICENSE, and COPYING. If the directory does not contain a TPL metadata file, do nothing. If the directory contains a TPL metadata file, determine whether the TPL metadata file contains a TPL name. If the TPL metadata file contains a TPL name, divide all the reuse files under the directory into the TPL module corresponding to the TPL name; if the TPL metadata file does not contain a TPL name, do nothing.

8. A software component analysis method for C / C++ source code according to claim 7, characterized in that After Step 3.3, it further includes: Step 3.4: For all the reuse files corresponding to each TPL library, if there are reuse files among all the reuse files that can be divided through Steps 3.1 to 3.3, the division ends. If all the reuse files cannot be divided through Steps 3.1 to 3.3, obtain the total number of reuse functions in the directory where the reuse file is located, and then calculate its ratio to the functions in the directory where the reuse file is located to obtain the reuse frequency of the directory where the reuse file is located, and then obtain the reuse frequency of each directory where the reuse file is located. Select the directory corresponding to the maximum value among all the reuse frequencies, and divide all the reuse files under the directory into the module corresponding to the TPL library.

9. A software component analysis method for C / C++ source code according to claim 1, characterized in that Step 4 specifically includes: Step 4.1: Analyze two reuse files for each TPL module. If one reuse file directly depends or indirectly depends on the other reuse file, at this time, ignore the dependency relationship between the two reuse files; Among them, one reuse file directly depending on another reuse file means that one reuse file directly references another reuse file through the #include directive or the extern keyword declaration, and one reuse file indirectly depending on another reuse file means that one reuse file references another reuse file through other files; Step 4.2: For the dependency relationship between TPL modules, determine whether there is a direct dependency or an indirect dependency between the reuse files in two TPL modules. If there is a direct dependency or an indirect dependency between the reuse files in two TPL modules, it indicates that there is a dependency relationship between the two TPL modules. If there is no direct dependency or an indirect dependency between the reuse files in two TPL modules, it indicates that there is no dependency relationship between the two TPL modules.

Citation Information

Cited By

  • Open source component dependency identification method

    CN121166187A

  • An open source component dependency identification method

    CN121166187B

  • Open source component tracing method based on composite fingerprint and large model variant judgment

    CN122451857A