Method and device for processing byte code file and storage medium

By obtaining multiple function call diagrams and splitting the large bytecode file into multiple small files, each small file containing the function content of the same function call diagram, the problem of time-consuming and low precision of static analysis of large bytecode files is solved, and more efficient and accurate static analysis is achieved.

CN120029873APending Publication Date: 2025-05-23HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311514749.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-14
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

Large bytecode files take a long time to analyze statically, which may lead to memory exhaustion, and splitting files in a physically partitioned manner will reduce the accuracy of static analysis.

Method used

By obtaining multiple function call graphs, the large bytecode file is split into multiple small files, each small file contains the function content belonging to the same function call graph, thereby avoiding interruption of data flow between functions.

Benefits of technology

Improves the accuracy and efficiency of static analysis, avoids the problem of memory exhaustion, and reduces the number of split files.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029873A_ABST
    Figure CN120029873A_ABST
Patent Text Reader

Abstract

The invention discloses a method and device for processing a byte code file and a storage medium, and belongs to the field of computers. The method comprises the steps that a first BC file is obtained, wherein the first BC file is in an intermediate representation form in the source program compiling process; based on the first BC file, multiple first function call graphs are obtained, the first function call graphs comprise multiple functions in the source program, and any two functions in the multiple functions have a direct call relation or an indirect call relation; based on the plurality of first function call graphs, splitting the first BC file into a plurality of second BC files, the second BC files corresponding to at least one first function call graph in the plurality of first function call graphs, the second BC file comprises the content, belonging to each function in the at least one first function call graph, in the first BC file. The precision of static analysis can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computers, and in particular to a method, device and storage medium for processing bytecode files. Background Art

[0002] Bytecode (BC) files are used to save the intermediate form of the C / C++ language source program during compilation. They are another form of the source program. Program analysis tools can obtain the BC file of the source program and perform precise data flow static analysis on the source program based on the BC file.

[0003] Usually BC files contain a lot of information and a large amount of data. Static analysis based on large BC files takes a long time. Sometimes, the large amount of data in the large BC files causes memory exhaustion and static analysis fails. For this reason, the large BC file is physically separated into multiple small BC files, and then static analysis is performed based on each small BC file.

[0004] Splitting a large BC file into multiple small BC files by physically separating them will cause data flow interruption, which will reduce the accuracy of static analysis when static analysis is performed based on each small BC file. Summary of the invention

[0005] The present application provides a method, device and storage medium for processing bytecode files to improve the accuracy of static analysis. The technical solution is as follows:

[0006] In a first aspect, the present application provides a method for processing a bytecode file, in which a first bytecode BC file is obtained, and the first BC file is an intermediate representation form in the source program compilation process. Based on the first BC file, multiple first function call graphs are obtained, and the first function call graph includes multiple functions in the source program, and there is a direct call relationship or an indirect call relationship between any two functions in the multiple functions. Based on the multiple first function call graphs, the first BC file is split into multiple second BC files, and the second BC file corresponds to at least one first function call graph in the multiple first function call graphs, and the second BC file includes the content of each function in the first BC file belonging to the at least one first function call graph.

[0007] Since multiple first function call graphs are obtained, there are call relationships between multiple functions included in the first function call graphs. Based on the multiple first function call graphs, the first BC file is split into multiple second BC files. For at least one first function call graph corresponding to the second BC file, the second BC file includes the content of each function in the first BC file belonging to at least one first function call graph. In this way, when splitting the first BC file into multiple second BC files, the contents of functions with call relationships in the same first function call graph are split into the same second BC file, avoiding data flow interruption between functions. Static analysis is performed based on each second BC file, which can improve the accuracy of static analysis.

[0008] In a possible implementation, multiple function pairs are obtained based on the first BC file, each function pair includes two functions with a direct calling relationship or an indirect calling relationship, and the functions included in each function pair are functions in the source program. Based on the multiple function pairs, multiple first function call graphs are obtained, and the two functions connected by a calling edge included in the first function call graph are functions in the same function pair. In this way, the first function call graph is composed of multiple functions with a calling relationship, and when splitting the first BC file, the contents of the functions with a calling relationship can be split into the same second BC file.

[0009] In another possible implementation, the plurality of function pairs include one or more of the following function pairs: a first function pair, a second function pair, and a third function pair. The first function pair includes a callback function and a first calling function, the first calling function calls a global variable pointer of the callback function, and the global variable pointer is used to indicate the callback function. The second function pair includes a first function and a second calling function, the first function is a template function or a template class, and the second calling function is a function that calls the first function. The third function pair includes a second function and a third function, and the second function calls the third function through a local function pointer variable.

[0010] In another possible implementation, multiple second function call graphs are obtained based on the multiple function pairs, the root node of the second function call graph is a function in the source program that is not called by other functions, and the nodes other than the root node in the second function call graph are functions that have a direct call relationship or an indirect call relationship with the function. Multiple first function call graphs are obtained by splitting at the first call edge included in the second function call graph, there is no data flow transmission between the two functions connected by the first call edge, or the first call edge includes at least one call edge that has the least impact on the data flow integrity of the second function call graph. In this way, the number of first function call graphs can be increased, and more first function call graphs can be obtained.

[0011] In another possible implementation, at least one first function call graph is a function call graph whose function overlap satisfies the first condition. In this way, the content of each function included in at least one first function call graph with high function overlap is split into the same second BC file, so as to avoid splitting the first BC file too small, thereby reducing the number of split second BC files.

[0012] In another possible implementation, the first condition includes:

[0013] The function overlap between any two first function call graphs in the at least one first function call graph exceeds a first overlap threshold; or,

[0014] The function overlap between a first function call graph included in the at least one first function call graph and each other first function call graph exceeds a second overlap threshold.

[0015] In another possible implementation, the second BC file further includes a correspondence between functions and global variable pointers, and each record in the correspondence includes a function and a global variable pointer called by the function, and the function is a function in the at least one first function call graph.

[0016] In a second aspect, the present application provides a device for processing a bytecode file, for executing the method in the first aspect or any possible implementation of the first aspect. Specifically, the device includes a unit for executing the method in the first aspect or any possible implementation of the first aspect.

[0017] In a third aspect, the present application provides a computing device cluster, the cluster comprising at least one computing device, each device in the at least one computing device comprising at least one processor and at least one memory, the at least one memory storing computer-readable instructions; the at least one processor executes the computer-readable instructions so that the cluster implements the method in the first aspect or any possible implementation manner of the first aspect.

[0018] In a fourth aspect, the present application provides a computer program product, which includes a computer program stored in a computer-readable storage medium, and the computer program is loaded by a processor to implement the method in the first aspect or any possible implementation manner of the first aspect.

[0019] In a fifth aspect, the present application provides a computer-readable storage medium for storing a computer program, wherein the computer program is loaded by a processor to execute the method in the first aspect or any possible implementation manner of the first aspect.

[0020] In a sixth aspect, the present application provides a chip comprising a memory and a processor, wherein the memory is used to store computer instructions, and the processor is used to call and run the computer instructions from the memory to execute the method in the first aspect or any possible implementation of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 This is a schematic diagram of an application scenario provided by an embodiment of the present application;

[0022] Figure 2 This is another application scenario schematic diagram provided by an embodiment of the present application;

[0023] Figure 3 It is a flow chart of a method for processing bytecode files provided in an embodiment of the present application;

[0024] Figure 4 is a second function call graph provided by an embodiment of the present application;

[0025] Figure 5 is a first function call graph provided by an embodiment of the present application;

[0026] Figure 6 It is a schematic diagram of the structure of a device for processing bytecode files provided in an embodiment of the present application;

[0027] Figure 7 is a schematic diagram of the structure of a computing device provided in an embodiment of the present application;

[0028] Figure 8 It is a schematic diagram of the structure of a cluster provided in an embodiment of the present application;

[0029] Fig. 9 It is a schematic diagram of the structure of another cluster provided in an embodiment of the present application. DETAILED DESCRIPTION

[0030] The following is an introduction to the relevant concepts that appear in this application.

[0031] A function call graph, also known as a call graph in English, is a control flow graph that represents the direct call relationship of a computer program. The function call graph includes a first function call graph and a second function call graph.

[0032] The first function call graph is obtained by splitting the second function call graph.

[0033] The root node of the second function call graph is a function in the source program that is not called by other functions, and other nodes of the second function call graph are functions that have a direct calling relationship or an indirect calling relationship with the root node.

[0034] BC file is an intermediate representation during the compilation process of the source program.

[0035] See also Figure 1 In the application scenario shown, a source program is written in C / C++ programming language, the source program is then compiled, and the compiled source program is linked to obtain a first BC file.

[0036] The source program may be a program with a large amount of code, many functions, and complex code logic, resulting in a large amount of data in the first BC file. Static analysis of the first BC file takes a long time, reducing the efficiency of static analysis, and / or there is not enough memory to save the first BC file, resulting in the inability to perform static analysis on the source program.

[0037] Therefore, it is necessary to split the first BC file into multiple second BC files, the data volume of each second BC file is smaller than the data volume of the first BC file, and then perform static analysis on the source program based on each second BC file.

[0038] Linking means compiling each source code unit in the source program independently, and then "assembling" each compiled source code unit as needed. This assembly process is linking. The main content of linking is to handle the places where each source code module directly references each other (including functions and variables), so that each source code module can be directly and correctly connected.

[0039] The so-called static analysis refers to a code analysis technology that scans the program code of the source program through lexical analysis, syntax analysis, control flow, data flow analysis and other technologies without running the source program to verify whether the code meets indicators such as standardization, security, reliability, and maintainability.

[0040] The first BC file is an intermediate code, that is, it is neither the source code included in the source program nor the machine code. From the code organization structure, the first BC file is closer to the machine code, but it uses many high-level programming language features at the function and instruction levels.

[0041] See also Figure 2 , a source program static analysis tool can be used to perform static analysis on the source program. A function of splitting BC files is added to the source program static analysis tool, so that after receiving the first BC file generated by linking, the source program static analysis tool splits the first BC file to obtain multiple second BC files. Then, static analysis of the source program is performed based on each second BC file.

[0042] Optionally, any of the following embodiments may be used to split a first BC file with a larger data volume into multiple second BC files with smaller data volumes. For detailed implementation, see the content of any of the following embodiments.

[0043] See also Figure 3 The present application embodiment provides a method 300 for processing a bytecode file, which can be applied to Figure 2 In the source program static analysis tool in the application scenario shown, the method 300 includes the following process.

[0044] Step 301: Obtain a first BC file, where the first BC file is an intermediate representation of a source program.

[0045] The source program is compiled, and then the compiled source program is linked to obtain a first BC file. The first BC file is an intermediate representation form in the process of the source program being compiled into binary machine code.

[0046] Step 302: Acquire multiple function pairs based on the first BC file, each function pair includes two functions that have a direct calling relationship or an indirect calling relationship, and the functions included in each function pair are functions in the source program.

[0047] In step 302, the function pair may be obtained in the following ways.

[0048] Method 1: Identify a global variable pointer and a first calling function that calls the global variable pointer in a first BC file, obtain a callback function indicated by the global variable pointer, and obtain a first function pair, wherein the first function pair includes the callback function and the first calling function.

[0049] Optionally, the first function pair may further include the global variable pointer.

[0050] For example, consider the following program:

[0051]

[0052]

[0053] int(*fun_ptr)(int,int)=add;

[0054] Among them, ptr is a global variable pointer, which is used to indicate the callback function add.

[0055] Assume that there is a function calculate that calls the global variable pointer fun_ptr, and the function calculate is the first calling function.

[0056]

[0057] The global variable pointer fun_ptr and the first call function calculate that calls the global variable pointer fun_ptr can be identified in the first BC file, and the callback function add indicated by the global variable pointer fun_ptr is obtained from the first BC file. In this way, the first function pair is obtained, and the first function pair includes the callback function add and the first call function calculate. Optionally, the first function pair may also include the global variable pointer fun_ptr.

[0058] The source program often includes multiple global variable pointers, so multiple first function pairs can be obtained.

[0059] Method 2: In the first BC file, a first function and a second calling function that calls the first function are identified to obtain a second function pair, wherein the first function is a template function or a template class, and the second function pair includes the first function and the second calling function.

[0060] For example, there is the following template function, the format of declaring the template function is:

[0061]

[0062] The instance of the template function tfunc is identified in the first BC file, and further analysis shows that the function pointer in its third parameter points to the function add by default. Thus, the second function pair is obtained, which includes the instance of the template function tfunc and the second calling function add.

[0063] A C++ source program often includes multiple template functions and template classes, so multiple second function pairs can be obtained in this case.

[0064] Method three, obtain the second function and the third function in the first BC file, the second function calls the third function through a local function pointer variable, the third function pair includes the second function and the third function, and the second function and the third function are two functions with an indirect calling relationship.

[0065] For example, see the following function:

[0066]

[0067]

[0068] Among them, fun_ptr is a local function pointer variable, calculate() is the second function, add() is the third function, and the second function calculate() calls the third function add() through the local function pointer variable fun_ptr.

[0069] In the third approach, the second function and the third function may be obtained in the first BC file based on a pointer analysis algorithm. Optionally, the pointer analysis algorithm may be an Anderson algorithm.

[0070] Step 303: Based on the multiple function pairs, multiple second function call graphs are obtained, where the root node of the second function call graph is a function in the source program that is not called by other functions, and the nodes other than the root node in the second function call graph are functions that have a direct calling relationship or an indirect calling relationship with the function.

[0071] In step 303, a function in the source program that is not called by other functions can be obtained based on the first BC file, and the function is used as the root node of the second function call graph, and at least one function pair including the function is searched from the multiple function pairs. At least one first node is added to the second function call graph, and a call edge is connected between the root node and each first node, and the at least one first node includes functions other than the function in the at least one function pair.

[0072] For each first node, at least one function pair including the function in the first node is searched from the plurality of function pairs. At least one second node is added to the second function call graph, the first node is connected to each second node by a call edge, and the at least one second node includes the function in the at least one function pair except the function in the first node.

[0073] For each second node, at least one function pair including the function in the second node is searched from the plurality of function pairs. At least one third node is added to the second function call graph, the second node is connected to each third node by a call edge, and the at least one third node includes the function in the at least one function pair except the function in the second node.

[0074] Repeat the above process until no function pair can be found, and obtain a second function call graph.

[0075] Based on the first BC file, multiple functions in the source program that are not called by other functions can be obtained, and the above process is performed on each function to obtain multiple second function call graphs.

[0076] For example, assume that in step 302, function pair 1 is obtained, which includes function 1 and function 2, function pair 2 includes function 1 and function 3, function pair 3 includes function 2 and function 3, function pair 4 includes function 3 and function 4, and function pair 5 includes function 4 and function 5.

[0077] Function 1 is a function in the source program that is not called by other functions. With function 1 as the root node, function pair 1, function pair 2, function pair 3, function pair 4 and function pair 5 are searched for function pair 1 and function pair 2 that include function 1.

[0078] See also Figure 4 , function pair 1 includes function 1 and function 2, and function pair 2 includes function 1 and function 3. In the second function call graph, add root node 1 including function 1, node 2 including function 2, and node 3 including function 3. There is a call edge connecting root node 1 and node 2, and there is a call edge connecting root node 1 and node 3.

[0079] For function 2 included in node 2, function pair 3 includes function 2 and function 3, and node 2 including function 2 and node 3 including function 3 are connected by a call edge.

[0080] For function 3 included in node 3, function pair 4 includes function 3 and function 4, node 4 including function 4 is added to the second function call graph, and node 3 and node 4 are connected by a call edge.

[0081] For function 4 included in node 4, function pair 5 includes function 4 and function 5. Node 5 including function 5 is added to the second function call graph, and a call edge is connected between node 4 and node 5. Since it is no longer possible to find a function pair including a function in the second function call graph, the following is obtained: Figure 4 The second function call graph is shown.

[0082] Step 304: Split the second function call graph at the first call edge included in the second function call graph to obtain multiple first function call graphs, where there is no data flow transmission between two functions connected by the first call edge, or the first call edge includes at least one call edge that has the least impact on the data flow integrity of the second function call graph.

[0083] In step 303, multiple second function call graphs are obtained, and each second function call graph can be split, or a larger second function call graph can be split. The so-called larger second function call graph refers to a second function call graph including a number of nodes exceeding a threshold.

[0084] For each second function call graph in the plurality of second function call graphs, or for each second function call graph of a larger scale selected from the plurality of second function call graphs, a first call edge is selected from the second function call graph, and the second function call graph is split into a plurality of first function call graphs at the first call edge. There is no data flow transfer between two functions connected by the first call edge, or the first call edge includes at least one call edge that has the least impact on the data flow integrity of the second function call graph.

[0085] In some embodiments, it is checked whether there is data flow transfer between any two functions with a calling relationship included in the second function call graph. If it is found that there is no data flow transfer between the two functions with a calling relationship, the call edge connecting the two functions is used as the first call edge, and the second function call graph is split into multiple first function call graphs at the first call edge.

[0086] If it is checked that there is data flow transmission between all two functions with a calling relationship, a directed graph segmentation algorithm is used to find at least one calling edge that has the least impact on the data flow integrity of the second function call graph in the second function call graph. The second function call graph is split into multiple first function call graphs at the at least one calling edge.

[0087] Optionally, a safe sub-callgraph (Sub-CG) edge trimming algorithm may be used to check whether there is data flow transmission between any two functions having a calling relationship included in the second function call graph.

[0088] Optionally, the directed graph segmentation algorithm may be a lossy Sub-CG edge trimming algorithm. That is, the lossy Sub-CG edge trimming algorithm may be used to find at least one call edge in the second function call graph that has the least impact on the data flow integrity of the second function call graph.

[0089] For example, for Figure 4 The second function call graph shown in the figure uses the secure Sub-CG edge trimming algorithm to check the second function call graph, and checks that there is no data flow between function 3 and function 4, which have a call relationship. Therefore, the first call edge is the call edge between function 3 and function 4, see Figure 5 , split the second function call graph into the first function call graph at this call edge Figure 1 and the first function call Figure 2 .

[0090] For the unsplit second function call graph, the unsplit second function call graph may be used as the first function call graph.

[0091] Step 305: split the first BC file into multiple second BC files based on the multiple first function call graphs.

[0092] Among them, each second BC file corresponds to at least one first function call graph among the multiple first function call graphs, and for each second BC file and at least one first function call graph corresponding to the second BC file, the second BC file includes the content of each function in the first BC file belonging to the at least one first function call graph.

[0093] For each second BC file, at least one first function call graph corresponding to the second BC file is a function call graph whose function overlap degree satisfies the first condition.

[0094] In some embodiments, the first condition includes: the function overlap between any two first function call graphs in at least one first function call graph exceeds a first overlap threshold; or, the function overlap between a first function call graph included in at least one first function call graph and each other first function call graph exceeds a second overlap threshold.

[0095] In step 305 , the first BC file may be split into multiple second BC files through the following operations 3051 - 3053 .

[0096] 3051: Obtain the function overlap degree between any two first function call graphs in the multiple first function call graphs.

[0097] For any two first function call graphs among the multiple first function call graphs, the function overlap between the two first function call graphs is obtained based on the function names of the functions included in the two first function call graphs and the identification information of the call edges included in the two first function call graphs.

[0098] Optionally, the function overlap between the two first function call graphs is equal to cosθ as shown in the following first formula.

[0099] First formula:

[0100] In the first formula, since the similarity of the first function call graphs is calculated, n=2, A i is the set of functions in the i-th first function call graph, B i is the set of call edges of the i-th first function call graph.

[0101] Among them, A i Including A 1 and A 2 , A 1 is a function set of the first first function call graph, the function set including the function name of each function in the first first function call graph, A 2 is a function set of the second first function call graph, and the function set includes the function name of each function in the second first function call graph.

[0102] B i Including B 1 and B 2 , B 1is a call edge set of the first first function call graph, the call edge set including identification information of each call edge in the first first function call graph, B 2 is a call edge set of the second first function call graph, and the call edge set includes identification information of each call edge in the second first function call graph.

[0103] In some embodiments, the function overlap between the two first function call graphs may include the similarity between the two first function call graphs, etc. For example, the function overlap between the two first function call graphs may include the cosine similarity between the two first function call graphs.

[0104] 3052: Based on the function overlap between any two first function call graphs in the multiple first function call graphs, cluster the multiple first function call graphs to obtain multiple clusters, each cluster including at least one first function call graph whose function overlap satisfies the first condition.

[0105] That is, for each cluster, the function overlap between any two first function call graphs in the cluster exceeds the first overlap threshold. Alternatively, the function overlap between a first function call graph included in the cluster and each other first function call graph exceeds the second overlap threshold.

[0106] For example, suppose the cluster includes the first function call Figure 1 , the first function call Figure 2 , the first function call Figure 3 and the first function call Figure 4 The first function call Figure 1 and the first function call Figure 2 The function overlap between them exceeds the first overlap threshold, and the first function call Figure 1 and the first function call Figure 3 The function overlap between them exceeds the first overlap threshold, and the first function call Figure 1 and the first function call Figure 4 The function overlap between them exceeds the first overlap threshold, and the first function call Figure 2 and the first function call Figure 3 The function overlap between them exceeds the first overlap threshold, and the first function call Figure 2 and the first function call Figure 4 The function overlap between them exceeds the first overlap threshold, and the first function call Figure 3 and the first function call Figure 4 The function overlap between exceeds the first overlap threshold. Or,

[0107] For the first function call Figure 1 , the first function call Figure 1and the first function call Figure 2 The function overlap between them exceeds the second overlap threshold, and the first function calls Figure 1 and the first function call Figure 3 The function overlap between them exceeds the second overlap threshold, and the first function calls Figure 1 and the first function call Figure 4 The function overlap between them exceeds the second overlap threshold.

[0108] 3053: Based on the multiple clusters, split the first BC file into multiple second BC files.

[0109] The multiple second BC files correspond one-to-one to the multiple clusters. For each second BC file and for the cluster corresponding to the second BC file, the second BC file includes the content of each function in the first BC file belonging to the cluster.

[0110] In 3053, for each cluster, the content of each function included in at least one first function call graph belonging to the cluster is obtained from the first BC file to obtain a second BC file, which includes the content of each function included in at least one first function call graph belonging to the cluster.

[0111] In some embodiments, at least one first function call graph in the cluster includes multiple functions that may call global variable pointers for some functions, and obtains the global variable pointers for the partial functions and the partial functions called by the partial functions. The second BC file also includes a correspondence between the function and the global variable pointer, and the correspondence includes the partial functions and the global variable pointers called by the partial functions.

[0112] The first function call graphs in the same cluster include a large number of identical functions, so clustering the multiple first function call graphs and splitting the first BC file into multiple second BC files based on the multiple clusters can reduce the number of split second BC files.

[0113] In an embodiment of the present application, multiple second function call graphs are obtained based on the first BC file, and the second function call graph with a larger scale is split to obtain multiple first function call graphs. Wherein, when splitting the second function call graph, the splitting is performed from the first call edge of the second function call graph, because there is no data flow transmission between the two functions connected by the first call edge, or the first call edge includes at least one call edge that has the least impact on the data flow integrity of the second function call graph, so that each split first function call graph is as complete and accurate as possible, and the first BC file can be split into a second BC file with a smaller data volume. When splitting the first BC file, at least one first function call graph with a higher function overlap can be clustered into a cluster, and the contents of each function in each first function call graph belonging to the cluster in the first BC file can be split into a second BC file, so that not only the number of split second BC files can be minimized, but also the data flow integrity of each function in the second BC file is retained, and static analysis is performed based on each split second BC file, which can improve the accuracy of static analysis.

[0114] See also Figure 6 The present application embodiment provides a device 600 for processing a bytecode file, and the device 600 can be deployed in the above Figure 2 The source program static analysis tool in the application scenario shown in the figure can be deployed in the above Figure 3 In the source program static analysis tool of the method 300 shown. The device 600 includes:

[0115] The acquisition unit 601 is used to acquire a first bytecode BC file, where the first BC file is an intermediate representation form in the source program compilation process;

[0116] The acquisition unit 601 is further used to acquire multiple first function call graphs based on the first BC file, where the first function call graphs include multiple functions in the source program, and there is a direct call relationship or an indirect call relationship between any two functions in the multiple functions;

[0117] The splitting unit 602 is used to split the first BC file into multiple second BC files based on the multiple first function call graphs, where the second BC file corresponds to at least one first function call graph among the multiple first function call graphs, and the second BC file includes the content of each function in the first BC file belonging to the at least one first function call graph.

[0118] Optionally, the detailed implementation process of obtaining the first BC file by the obtaining unit 601 is shown in Figure 3 Step 301 of the method 300 is not described in detail here.

[0119] Optionally, the detailed implementation process of the acquisition unit 601 acquiring multiple first function call graphs based on the first BC file can be found in Figure 3 Steps 302 - 304 of the method 300 are not described in detail herein.

[0120] Optionally, the splitting unit 602 splits the first BC file into multiple second BC files based on the multiple first function call graphs. Figure 3 Step 305 of the method 300 is not described in detail here.

[0121] Optionally, the acquiring unit 601 is configured to:

[0122] Acquire multiple function pairs based on the first BC file, each function pair includes two functions that have a direct calling relationship or an indirect calling relationship, and the functions included in each function pair are functions in the source program;

[0123] Based on the multiple function pairs, the multiple first function call graphs are obtained, wherein two functions connected by a call edge included in the first function call graph are functions in the same function pair.

[0124] Optionally, the detailed implementation process of the acquisition unit 601 acquiring multiple function pairs based on the first BC file can be found in Figure 3 Step 302 of the method 300 is not described in detail here.

[0125] Optionally, the detailed implementation process of obtaining the plurality of first function call graphs based on the plurality of function pairs by the obtaining unit 601 is shown in Figure 3 Steps 303 - 304 of the method 300 are not described in detail here.

[0126] Optionally, the plurality of function pairs include one or more of the following function pairs: a first function pair, a second function pair and a third function pair;

[0127] The first function pair includes a callback function and a first calling function, the first calling function calls a global variable pointer of the callback function, and the global variable pointer is used to indicate the callback function;

[0128] The second function pair includes a first function and a second calling function, the first function is a template function or a template class, and the second calling function is a function that calls the first function;

[0129] The third function pair includes a second function and a third function, and the second function calls the third function through a local function pointer variable.

[0130] Optionally, the acquiring unit 601 is configured to:

[0131] Acquire multiple second function call graphs based on multiple function pairs, wherein the root node of the second function call graph is a function in the source program that is not called by other functions, and the nodes other than the root node in the second function call graph are functions that have a direct calling relationship or an indirect calling relationship with the function;

[0132] The second function call graph is split at the first call edge included in the second function call graph to obtain multiple first function call graphs, there is no data flow transmission between the two functions connected by the first call edge, or the first call edge includes at least one call edge that has the least impact on the data flow integrity of the second function call graph.

[0133] Optionally, the detailed implementation process of the acquisition unit 601 acquiring multiple second function call graphs based on multiple function pairs can be found in Figure 3 Step 303 of the method 300 is not described in detail here.

[0134] Optionally, the acquisition unit 601 splits the first call edge included in the second function call graph to obtain multiple first function call graphs. Figure 3 Step 304 of the method 300 is not described in detail here.

[0135] Optionally, the at least one first function call graph is a function call graph whose function overlap degree satisfies a first condition.

[0136] Optionally, the first condition includes:

[0137] The function overlap between any two first function call graphs in the at least one first function call graph exceeds a first overlap threshold; or,

[0138] The function overlap between a first function call graph included in the at least one first function call graph and each other first function call graph exceeds a second overlap threshold.

[0139] Optionally, the second BC file further includes a correspondence between functions and global variable pointers, and each record in the correspondence includes a function and a global variable pointer called by the function, and the function is a function in the at least one first function call graph.

[0140] In the embodiment of the present application, since the acquisition unit acquires multiple first function call graphs, and there are call relationships between multiple functions included in the first function call graph, the splitting unit splits the first BC file into multiple second BC files based on the multiple first function call graphs, and for at least one first function call graph corresponding to the second BC file, the second BC file includes the content of each function in the first BC file belonging to at least one first function call graph. In this way, when the splitting unit splits the first BC file into multiple second BC files, the content of the functions with call relationships in the same first function call graph is split into the same second BC file, thereby avoiding data flow interruption between functions, and performing static analysis based on each second BC file can improve the accuracy of static analysis.

[0141] See also Figure 7 , the present application embodiment provides a computing device 700. For example, the computing device 700 may be Figure 1 The cloud computing platform in the network architecture shown includes devices, or Figure 2 A device in a cloud computing platform of the method 200 is shown.

[0142] like Figure 7 As shown, computing device 700 includes: bus 702, processor 704, memory 706 and communication interface 708. Processor 704, memory 706 and communication interface 708 communicate through bus 702. Computing device 700 can be a server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in computing device 700.

[0143] The bus 702 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 7 The bus 702 may include a path for transmitting information between various components of the computing device 700 (eg, the memory 706, the processor 704, and the communication interface 708).

[0144] The processor 704 may include any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0145] The memory 706 may include a volatile memory, such as a random access memory (RAM). The processor 704 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD).

[0146] See also Figure 7 , the memory 706 stores executable program codes, and the processor 704 executes the executable program codes to respectively implement Figure 6 The functions of the acquisition unit 601 and the splitting unit 602 in the device 600 shown in the figure are implemented to realize the method provided by any of the above embodiments. That is, the memory 706 stores instructions for executing the method provided by any of the above embodiments. Or,

[0147] The communication interface 708 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 700 and other devices or communication networks.

[0148] The embodiment of the present application also provides a cluster. The cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.

[0149] like Figure 8 As shown, the cluster includes at least one computing device 700. The memory 706 in one or more computing devices 700 in the cluster may store the same instructions for executing the method provided by any of the above embodiments.

[0150] In some possible implementations, the memory 706 of one or more computing devices 700 in the cluster may also store partial instructions for executing the above-mentioned method for processing bytecode files. In other words, the combination of one or more computing devices 700 may jointly execute instructions for executing the method for processing bytecode files provided in any of the above-mentioned embodiments.

[0151] In some possible implementations, one or more computing devices in the cluster may be connected via a network, which may be a wide area network or a local area network. Fig. 9 A possible implementation is shown. Fig. 9 As shown, two computing devices 700A and 700B are connected via a network. Specifically, they are connected to the network via a communication interface in each computing device.

[0152] In this type of possible implementation, the memory 706 in the computing device 700A stores the execution Figure 6 Instructions for obtaining the functions of the unit 601 in the embodiment shown. Meanwhile, the memory 706 in the computing device 700B stores instructions for executing the following Figure 6 Instructions for the functionality of the split unit 602 in the illustrated embodiment.

[0153] It should be understood that Fig. 9 The functions of the computing device 700A shown in FIG. 7 may also be completed by multiple computing devices 700. Similarly, the functions of the computing device 700B may also be completed by multiple computing devices 700.

[0154] The present application embodiment also provides another cluster. The connection relationship between the computing devices in the cluster can be similar to that of Fig. 9 The connection mode of the cluster for processing bytecode files is different in that the memory 706 in one or more computing devices 700 in the cluster may store the same instructions for executing the method provided in any of the above embodiments.

[0155] In some possible implementations, the memory 706 of one or more computing devices 700 in the cluster may also store partial instructions for executing the method provided in any of the above embodiments. In other words, the combination of one or more computing devices 700 may jointly execute instructions for executing the method provided in any of the above embodiments.

[0156] The present application also provides a computer program product including instructions. The computer program product may be software or a program product including instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device executes the method provided in any of the above embodiments.

[0157] The embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium may be any available medium that can be stored by a computing device or a data storage device such as a data center including one or more available media. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive), etc. The computer-readable storage medium includes instructions that direct the computing device to execute the method provided in any of the above embodiments.

[0158] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium. The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disk, etc.

[0159] The above are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the principles of the present application shall be included within the protection scope of the present application.

Claims

1. A method for processing bytecode files, It is characterized in that The method comprises: Obtaining a first bytecode BC file, where the first BC file is an intermediate representation form in a source program compilation process; Acquire a plurality of first function call graphs based on the first BC file, wherein the first function call graphs include a plurality of functions in the source program, and a direct call relationship or an indirect call relationship exists between any two functions in the plurality of functions; Based on the multiple first function call graphs, the first BC file is split into multiple second BC files, the second BC file corresponds to at least one first function call graph among the multiple first function call graphs, and the second BC file includes the content of each function in the first BC file belonging to the at least one first function call graph.

2. The method according to claim 1, It is characterized in that The acquiring a plurality of first function call graphs based on the first BC file includes: Acquire multiple function pairs based on the first BC file, each function pair includes two functions that have a direct calling relationship or an indirect calling relationship, and the functions included in each function pair are functions in the source program; Based on the multiple function pairs, the multiple first function call graphs are obtained, wherein two functions connected by a call edge included in the first function call graph are functions in the same function pair.

3. The method according to claim 2, It is characterized in that The plurality of function pairs include one or more of the following function pairs: a first function pair, a second function pair, and a third function pair; The first function pair includes a callback function and a first calling function, the first calling function calls a global variable pointer of the callback function, and the global variable pointer is used to indicate the callback function; The second function pair includes a first function and a second calling function, the first function is a template function or a template class, and the second calling function is a function that calls the first function; The third function pair includes a second function and a third function, and the second function calls the third function through a local function pointer variable.

4. The method according to claim 2 or 3, It is characterized in that The obtaining the plurality of first function call graphs based on the plurality of function pairs comprises: Acquire multiple second function call graphs based on the multiple function pairs, wherein a root node of the second function call graph is a function in the source program that is not called by other functions, and nodes other than the root node in the second function call graph are functions that have a direct calling relationship or an indirect calling relationship with the function; The second function call graph is split at the first call edge included in the second function call graph to obtain multiple first function call graphs, where there is no data flow transmission between the two functions connected by the first call edge, or the first call edge includes at least one call edge that has the least impact on the data flow integrity of the second function call graph.

5. The method according to any one of claims 1 to 4, It is characterized in that The at least one first function call graph is a function call graph whose function overlap degree satisfies a first condition.

6. The method according to claim 5, It is characterized in that The first condition includes: The function overlap between any two first function call graphs in the at least one first function call graph exceeds a first overlap threshold; or, The function overlap between a first function call graph included in the at least one first function call graph and each other first function call graph exceeds a second overlap threshold.

7. The method according to any one of claims 1 to 6, It is characterized in that The second BC file also includes a correspondence between functions and global variable pointers, each record in the correspondence includes a function and a global variable pointer called by the function, and the function is a function in the at least one first function call graph.

8. A device for processing bytecode files, It is characterized in that The device comprises: An acquisition unit, used to acquire a first bytecode BC file, where the first BC file is an intermediate representation form in a source program compilation process; The acquisition unit is further used to acquire a plurality of first function call graphs based on the first BC file, wherein the first function call graphs include a plurality of functions in the source program, and there is a direct call relationship or an indirect call relationship between any two functions in the plurality of functions; A splitting unit is used to split the first BC file into multiple second BC files based on the multiple first function call graphs, wherein the second BC file corresponds to at least one first function call graph among the multiple first function call graphs, and the second BC file includes the content of each function in the at least one first function call graph in the first BC file.

9. The device as claimed in claim 8, It is characterized in that The acquisition unit is used to: Acquire multiple function pairs based on the first BC file, each function pair includes two functions that have a direct calling relationship or an indirect calling relationship, and the functions included in each function pair are functions in the source program; Based on the multiple function pairs, the multiple first function call graphs are obtained, wherein two functions connected by a call edge included in the first function call graph are functions in the same function pair.

10. The device according to claim 9, It is characterized in that The plurality of function pairs include one or more of the following function pairs: a first function pair, a second function pair, and a third function pair; The first function pair includes a callback function and a first calling function, the first calling function calls a global variable pointer of the callback function, and the global variable pointer is used to indicate the callback function; The second function pair includes a first function and a second calling function, the first function is a template function or a template class, and the second calling function is a function that calls the first function; The third function pair includes a second function and a third function, and the second function calls the third function through a local function pointer variable.

11. The device according to claim 9 or 10, It is characterized in that The acquisition unit is used to: Acquire multiple second function call graphs based on the multiple function pairs, wherein a root node of the second function call graph is a function in the source program that is not called by other functions, and nodes other than the root node in the second function call graph are functions that have a direct calling relationship or an indirect calling relationship with the function; The second function call graph is split at the first call edge included in the second function call graph to obtain multiple first function call graphs, where there is no data flow transmission between the two functions connected by the first call edge, or the first call edge includes at least one call edge that has the least impact on the data flow integrity of the second function call graph.

12. The device according to any one of claims 8 to 11, It is characterized in that The at least one first function call graph is a function call graph whose function overlap degree satisfies a first condition.

13. The device according to claim 12, It is characterized in that The first condition includes: The function overlap between any two first function call graphs in the at least one first function call graph exceeds a first overlap threshold; or, The function overlap between a first function call graph included in the at least one first function call graph and each other first function call graph exceeds a second overlap threshold.

14. The device according to any one of claims 8 to 13, It is characterized in that The second BC file also includes a correspondence between functions and global variable pointers, each record in the correspondence includes a function and a global variable pointer called by the function, and the function is a function in the at least one first function call graph.

15. A computing device cluster, It is characterized in that The cluster includes at least one computing device, each of the at least one computing device includes at least one processor and at least one memory, and the at least one memory stores computer-readable instructions; the at least one processor executes the computer-readable instructions so that the cluster executes the method according to any one of claims 1 to 7.

16. A computer storage medium having a computer program stored thereon, It is characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

17. A computer program product comprising a computer program, It is characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.