A method and device for process difference location
By executing instrumentation operations in open source code and recording operation data, the process differences are automatically positioned, and the functional process activation problem caused by incomplete initialization of open source code fragments is solved, which improves the code debugging efficiency.
Patent Information
- Application Number
- CN202111340478.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-12
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-11-12
AI Technical Summary
During the process of open source code vulnerability mining, the extracted code snippets cannot activate the functional flow in the code or the static variable is not initialized incompletely, resulting in the inability to activate the functional flow in the code or the activate the error flow, and thus the expected results cannot be obtained. In the face of a huge number of open source code snippets, it takes a lot of time to read the code repeatedly, which is inefficient.
Provide a process differential positioning method. By obtaining source code, extracting source code snippets, initializing and compiling the run, obtaining the run results, performing instrumentation operations at different results, finding process control key points, and recording the operation data after insertion to assist in determining the process differential position.
Through the automated process difference positioning method, the process difference content and location of source code fragments relative to source code is quickly determined, the code debugging efficiency is improved, the time for manual reading and analysis is reduced, and it is suitable for use in a large number of code debugging processes.
Smart Images

Figure CN114003507B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of computer technology, and particularly to a method and device for locating process differences. Background Art
[0002] In the process of open-source code vulnerability mining, it is often necessary to extract key code from open-source code for testing. However, the extracted code often fails to activate the functional process in the code or activates the wrong process due to incomplete parameter settings or incomplete initialization of static variables, resulting in the inability to obtain the expected result for the running code.
[0003] Traditionally, to solve such problems, it is necessary to manually and repeatedly read the open-source source code in detail to find out the places where the initialization is imperfect and complete them. For example, when the extracted open-source code snippet is run alone, the desired result is often not obtained. Programmers need to find out the underlying causal relationship for the failure to activate a specific process, so they need to read a large amount of source code to find the key points for modification. The specific process is as follows Figure 1 as Figure 1 shown. Programmers first read the open-source code to understand its meaning, and then extract specific functions from the open-source code. After that, programmers will complete the initialization part of the extracted code to make it run properly, and then compare whether the running result meets the expectation. When the running result does not meet the expectation, they will re-understand the meaning of the open-source code and complete the imperfect parts in the initialization stage.
[0004] Using the above method to repair open-source code, when faced with a huge number of open-source code snippets, it often takes a lot of time to repeatedly read the code, which is not only inefficient, but also difficult to determine the key process difference points. Summary of the Invention
[0005] The present invention provides a method and device for locating process differences, which are used to assist staff in quickly determining the process difference content and location between the extracted source code snippet and the source code.
[0006] To solve the above technical problems, an embodiment of the present invention provides a method for locating process differences, including:
[0007] Obtain the source code;
[0008] Extract a source code snippet from the source code;
[0009] Initialize the source code snippet and compile and run the initialized source code snippet;
[0010] Obtain the running result;
[0011] In the case where the running result is different from the result when the source code snippet runs in the source code, perform instrumentation operations at the corresponding positions of the source code snippet and the source code;
[0012] Run the instrumented source code snippet and the source code to at least respectively find the flow control key points in the source code snippet and the source code, and obtain a log file recording the running data of the instrumented source code snippet and the source code;
[0013] At least determine the flow difference position of the source code snippet relative to the source code based on the log file.
[0014] As an optional embodiment, running the instrumented source code snippet and the source code to at least respectively find the flow control key points in the source code snippet and the source code includes:
[0015] Perform block processing on the source code snippet and the source code respectively to obtain a plurality of code blocks;
[0016] At least determine, based on the context of each code block in the corresponding source code snippet or source code, a first code block as a candidate flow control key point from the plurality of code blocks;
[0017] Perform key point filtering processing on the first code block to obtain a second code block corresponding to the flow control key points in the source code snippet and the source code respectively.
[0018] As an optional embodiment, at least determining, based on the context of each code block in the corresponding source code snippet or source code, a first code block as a candidate flow control key point from the plurality of code blocks includes:
[0019] Determine the context of each code block in the corresponding source code snippet or source code;
[0020] Determine whether the context has a target keyword or a function header. If so, determine the corresponding code block as the first code block.
[0021] As an optional embodiment, running the instrumented source code snippet and the source code to obtain a log file recording the running data of the instrumented source code snippet and the source code includes:
[0022] Insert log printing functions at each of the flow key points in the source code snippet and the source code, so as to respectively generate and output the log file when the instrumented source code snippet and the source code run.
[0023] As an optional embodiment, the log files each include the running data of multiple threads, wherein the log file corresponding to the source code is the standard log file, and the log file corresponding to the source code segment is the log file to be tested;
[0024] The determining of the process difference position of the source code segment relative to the source code based at least on the log files includes:
[0025] Respectively block the standard log file and the log file to be tested based on the thread ID to obtain multiple log blocks corresponding to each thread respectively;
[0026] Compare the log block corresponding to the standard log file with the log block corresponding to the log file to be tested to determine the process difference position of the source code segment.
[0027] As an optional embodiment, the comparing of the log block corresponding to the standard log file with the log block corresponding to the log file to be tested includes:
[0028] Compare each log block in the log file to be tested with all the log blocks in the corresponding standard log file;
[0029] During the comparison process, determine the length of each log block being compared. If the length difference between the two compared log blocks meets the threshold, then extract the valid data from the log block with the longer length to form a third log block, and the length of the third log block and the length of the shorter log block satisfy a matching relationship;
[0030] Compare the third log block with the shorter log block.
[0031] As an optional embodiment, the comparing of the log block corresponding to the standard log file with the log block corresponding to the log file to be tested includes:
[0032] Process each log block based on a feature generation algorithm to determine the features of each log block;
[0033] Based on the features of each log block, compare the frequency similarity of the log block corresponding to the standard log file with the log block corresponding to the log file to be tested;
[0034] If the determined frequency similarity meets the first matching threshold, then compare the edit distance similarity of the log block corresponding to the standard log file with the log block corresponding to the log file to be tested.
[0035] As an optional embodiment, comparing the edit distance similarity of the log blocks corresponding to the standard log file and the log blocks corresponding to the log file to be tested includes:
[0036] Calculating the shortest edit distance of the two log blocks to be compared respectively;
[0037] Determining the differential edit distance based on the shortest edit distance;
[0038] Determining the edit distance similarity based on the differential edit distance.
[0039] As an optional embodiment, it further includes:
[0040] Counting the number of repetitions of each log block of the log file to be tested whose edit distance similarity does not meet the second matching threshold in the source code segment respectively;
[0041] Calculating the standard deviation of all statistical values;
[0042] Calculating the Euclidean distance value between each statistical value and the standard deviation;
[0043] Sorting in descending order based on multiple Euclidean distance values and outputting the sorting result, where the Euclidean distance value is proportional to the difference points of the log block occurrence process of the corresponding log file to be tested and the degree of difference of the process difference points.
[0044] Another embodiment of the present invention provides a process difference positioning device, including:
[0045] A first acquisition module, configured to acquire source code;
[0046] An extraction module, configured to extract a source code segment from the source code;
[0047] An initialization module, configured to initialize the source code segment and compile and run the initialized source code segment;
[0048] A second acquisition module, configured to acquire a running result;
[0049] An instrumentation module, configured to perform an instrumentation operation at corresponding positions in the source code segment and the source code when the running result is different from the result of running the source code segment in the source code;
[0050] A processing module, configured to run the instrumented source code segment and the source code to at least respectively find the process control key points in the source code segment and the source code, and obtain a log file recording the running data of the instrumented source code segment and the source code;
[0051] A determination module, configured to assist in determining a process difference position of the source code segment relative to the source code at least according to the log file.
[0052] Based on the disclosure of the above embodiments, the beneficial effects of the embodiments of the present invention include that when it is known that the running result after the initialization of the source code segment is inconsistent with the running result in the source code, instrumentation operations are performed in both the source code and the source code segment. Then, by running the instrumented source code segment and the source code, process control key points in the source code segment and the source code are found at least respectively, and a log file recording the running data of the instrumented source code segment and the source code is obtained. The system can at least assist the user in determining the process difference position of the source code segment relative to the source code based on the log file, that is, determining the process difference point of the source code segment and its position in the source code segment. The above methods are all executed by the device without manual operation, and both the repeatability and the execution efficiency are relatively high, and they can be applied to the debugging process of a large number of codes to assist in accelerating the user's code debugging. Description of the Drawings
[0053] Figure 1 It is a flowchart for determining the process difference point of the source code segment relative to the source code in the prior art.
[0054] Figure 2 It is a flowchart of the process difference location method in the embodiments of the present invention.
[0055] Figure 3 It is an application flowchart of the process difference location method in the embodiments of the present invention.
[0056] Figure 4 It is an application flowchart of the process difference location method in another embodiment of the present invention.
[0057] Figure 5 It is an application flowchart of the process difference location method in another embodiment of the present invention.
[0058] Figure 6 It is an application flowchart of the process difference location method in another embodiment of the present invention.
[0059] Figure 7 It is an application flowchart of the process difference location method in another embodiment of the present invention.
[0060] Figure 8 It is an application flowchart of the process difference location method in another embodiment of the present invention.
[0061] Figure 9 It is an actual application process diagram of the process difference location method in another embodiment of the present invention.
[0062] Figure 10It is a diagram of the actual application process of the process difference localization method in another embodiment of the present invention.
[0063] Figure 11 It is a diagram of the actual application process of the process difference localization method in another embodiment of the present invention.
[0064] Figure 12 It is a diagram of the actual application process of the process difference localization method in another embodiment of the present invention.
[0065] Figure 13 It is a block diagram of the structure of the process difference localization device in an embodiment of the present invention. Detailed implementation manners
[0066] Next, specific embodiments of the present invention will be described in detail with reference to the accompanying drawings, but it is not a limitation of the present invention.
[0067] It should be understood that various modifications can be made to the embodiments disclosed herein. Therefore, the following description should not be regarded as a limitation, but only as an example of the embodiments. Those skilled in the art will think of other modifications within the scope and spirit of the present disclosure.
[0068] The accompanying drawings included in the specification and forming a part of the specification illustrate the embodiments of the present disclosure, and together with the general description of the present disclosure given above and the detailed description of the embodiments given below are used to explain the principles of the present disclosure.
[0069] These and other features of the present invention will become apparent from the following description of the preferred forms of the embodiments given by way of non-limiting example with reference to the accompanying drawings.
[0070] It should also be understood that although the present invention has been described with reference to some specific examples, those skilled in the art can surely implement many other equivalent forms of the present invention, which have the features as described in the claims and thus are all within the protection scope defined thereby.
[0071] When combined with the accompanying drawings, the above and other aspects, features, and advantages of the present disclosure will become more apparent in view of the following detailed description.
[0072] Hereinafter, specific embodiments of the present disclosure will be described with reference to the accompanying drawings; however, it should be understood that the disclosed embodiments are merely examples of the present disclosure and can be implemented in various ways. Well-known and / or repetitive functions and structures are not described in detail to avoid unnecessary or redundant details from obscuring the present disclosure. Therefore, the specific structural and functional details disclosed herein are not intended to be limiting, but merely as a basis for the claims and a representative basis for teaching those skilled in the art to use the present disclosure in substantially any suitable detailed structure in a variety of ways.
[0073] This specification may use the phrases "in one embodiment", "in another embodiment", "in yet another embodiment", or "in other embodiments", all of which may refer to one or more of the same or different embodiments according to the present disclosure.
[0074] Next, embodiments of the present invention will be described in detail with reference to the accompanying drawings.
[0075] As Figure 2 and Figure 3 shown, an embodiment of the present invention provides a method for locating process differences, including:
[0076] Obtain the source code;
[0077] Extract source code fragments from the source code;
[0078] Initialize the source code fragments and compile and run the initialized source code fragments;
[0079] Obtain the running result;
[0080] In the case where the running result is different from the result when the source code fragment runs in the source code, perform instrumentation operations at the corresponding positions of the source code fragment and the source code;
[0081] Run the instrumented source code fragment and the source code to at least respectively find the process control key points in the source code fragment and the source code, and obtain a log file recording the running data of the instrumented source code fragment and the source code;
[0082] At least based on the log file, assist in determining the process difference position of the source code fragment relative to the source code.
[0083] The method in this embodiment can be applied to users for open-source code detection and debugging processes, especially suitable for processing open-source code with a huge volume.
[0084] For example, after the first comparison of the running results, if the initialization running result of the source code fragment does not meet the expectation, that is, the independent running result of the source code fragment is inconsistent with its running result in the source code, then enter the instrumentation process of the control instruction, that is, find the process controllable instructions in the code, which is equivalent to finding the process key points described below, and at the same time, log output code can be inserted into the above code to obtain the subsequent log file. Since the log file records the running parameters, that is, the running data, generated when the source code fragment and the source code are instrumented and run, the log file can effectively assist the system in quickly locking the possible positions of the process difference points, enabling the user to quickly determine the actual process difference points based on the results given by the system.
[0085] Based on the disclosure of the above embodiments, the beneficial effects of this embodiment include that when it is determined that the running result after the initialization of the source code fragment is inconsistent with the running result in the source code, instrumentation operations are performed in both the source code and the source code fragment. Then, by running the instrumented source code fragment and the source code, at least the flow control key points in the source code fragment and the source code are found respectively, and a log file recording the running data of the instrumented source code fragment and the source code is obtained. The system can at least assist the user in determining the flow difference position of the source code fragment relative to the source code based on the log file, that is, determine the flow difference points of the source code fragment and their positions in the source code fragment. The above methods are all executed by the device without manual operation. That is, there is no need for manual reading and understanding of all code fragments to analyze and determine the flow difference points by oneself. The method of this embodiment replaces manual review of a large amount of code and locks out the areas and positions where flow difference points are most likely to occur, quickly assisting the user in streamlining the code analysis scope and avoiding the user from repeatedly reading invalid code. Moreover, the repeatability and execution efficiency of this method are high, especially suitable for use in the debugging process of a large amount of code, assisting to speed up the user's code debugging and reducing the user's workload.
[0086] Continuing with Figure 2 As shown, the instrumentation process in this embodiment is different from the currently common instrumentation process. The common instrumentation process is to perform instrumentation based on the semantic analysis of the C language or pre-compile the source code and perform instrumentation on the basis of the pre-compiled block. However, the implementation of these two instrumentation processes is too difficult, the parsing algorithm is complex, the instrumentation code is huge, the efficiency is relatively low, and there is a certain usage threshold. Therefore, this embodiment provides a process that can achieve lightweight instrumentation and positioning.
[0087] Specifically, running the instrumented source code fragment and the source code to at least respectively find the flow control key points in the source code fragment and the source code includes:
[0088] Perform block processing on the source code fragment and the source code respectively to obtain multiple code blocks;
[0089] Based at least on the context of each code block in the corresponding source code fragment or source code, determine the first code block as a candidate flow control key point from the multiple code blocks;
[0090] Perform key point filtering processing on the first code block to obtain the second code block corresponding to the flow control key points in the source code fragment and the source code respectively.
[0091] Among them, based at least on the context of each code block in the corresponding source code fragment or source code, determining the first code block as a candidate flow control key point from the multiple code blocks includes:
[0092] Determine the context of each code block in the corresponding source code snippet or source code;
[0093] Determine whether the target keyword exists in the context. If so, determine the corresponding code block as the first code block;
[0094] The method further includes:
[0095] Determine the function headers in the source code snippet and the source code as candidate process control key points.
[0096] For example, the instrumentation code in this embodiment is implemented by a self-written python (computer programming language) script. As Figure 4 shown, during the execution process, first, the source code snippet and the source code are segmented to form multiple code blocks. Since in C language code, code blocks are segmented in the form of curly brace pairs, the python script in this embodiment, that is, the above code, is block-segmented based on curly braces. Then, at least based on the context of each code block in the corresponding source code or source code snippet, the first code block as a candidate process control key point is determined from multiple code blocks. For example, in C language, process control statements include if, else, else if, while, for, switch, etc. (equivalent to the keywords in this embodiment). Therefore, in the context of each code block, the system can determine whether the current code block belongs to the above process control statements by judging whether the above keywords appear in the context of the code block. If so, it proves that the code block belongs to the above process control statements, and thus the code block is recorded as a candidate process control key point. In addition, for some functions, they will be started in some special ways, such as message handling callback functions, error handling functions, etc. These functions are very important for process control, but they usually do not appear in the above branch statements. Therefore, in this embodiment, the beginning part of each function, that is, the function header, also needs to be used as a candidate process control key point.
[0097] Furthermore, when recording process control key points, due to unexpected situations such as non-standard code writing, curly braces in strings, macro definitions, line breaks, etc., the positioning of process control key points may be inaccurate. Therefore, various situations need to be considered to filter each candidate process control key point. In this embodiment, when performing key point filtering processing on the first code block, that is, the candidate process key point, to obtain the second code block corresponding to the process control key points in the source code snippet and the source code respectively, it includes:
[0098] (1) Keyword conflict points. For example, for a function named func_while(), during statistics, the function name conflicts with the keyword while. Therefore, before instrumentation, it is necessary to strictly judge the context relationship to exclude such errors. That is, the above function should be regarded as a separate process control key point.
[0099] (2) Some special symbols may appear inside double quotes or single quotes (such as curly braces in string constants). When counting curly brace pairs, it may cause statistical errors. Therefore, each time statistics are entered, this situation needs to be analyzed and excluded. That is, remove the code block determined based on the curly brace pairs included in the string constant, and remove the candidate process control points determined based on this code block.
[0100] (3) In the source code, it is necessary to consider the process issues controlled by macro definitions (such as #if code blocks). Therefore, special processing needs to be performed on the control flow statements in macro definitions.
[0101] (4) When defining functions, variables, and macro definitions, it is allowed to use the \' character for line breaks to make the code beautiful and concise. When calculating code blocks, it is necessary to fully consider the impact of line break characters on the code content, and make supplements or phased processing in combination with the context. For example, based on the above characters, when dividing code blocks, some statement contents are not split into the same code block. At this time, the missing content needs to be re-added to the code block.
[0102] Furthermore, after running the instrumented source code fragment and the source code, a log file recording the running data of the instrumented source code fragment and the source code is obtained, including:
[0103] Insert log printing functions at each process key point in the source code fragment and the source code, so that when the instrumented source code fragment and the source code are running, log files are generated and output respectively.
[0104] For example, different log files can be printed according to different requirements. In the C language, the __FILE__, __func__, and __LINE__ macros can be used to obtain which file, which function, and which line the current code block comes from. After instrumentation, each piece of code will generate a log file during compilation and running. This log file will record the data that appears during the running process, its source, location information, and what operations are involved, etc. At the same time, in a multi-threaded environment, the log printing function also needs to print the current thread ID for subsequent filtering. That is, the log file will contain the thread IDs of each thread.
[0105] Specifically, such as Figure 5As shown, the log files in this embodiment all include the running data of multiple threads. Among them, the log file corresponding to the source code is the standard log file, that is, the normal process log in the figure, and the log file corresponding to the source code segment is the log file to be tested, which corresponds to the problem process log in the figure;
[0106] At least assist in determining the process difference position of the source code segment relative to the source code based on the log file, including:
[0107] Chunk the standard log file and the log file to be tested respectively based on the thread ID to obtain multiple log chunks corresponding to each thread respectively;
[0108] Compare the log chunks corresponding to the standard log file with the log chunks corresponding to the log file to be tested to determine the process difference position of the source code segment.
[0109] For example, chunk the standard log file and the log file to be tested respectively based on the thread ID, and classify the log information using the thread ID to obtain multiple log chunks corresponding to each thread respectively. That is, each log chunk contains the log information of one thread. Then, as Figure 6 shown, compare the log chunks corresponding to the standard log file with each log chunk corresponding to the log file to be tested to determine the process difference position of the source code segment.
[0110] Specifically, when performing the comparison of the log chunks corresponding to the standard log file with the log chunks corresponding to the log file to be tested, it includes:
[0111] Compare each log chunk in the log file to be tested with all the log chunks in the standard log file;
[0112] During the comparison process, determine the length of each log chunk being compared. If the length difference between the two log chunks being compared meets the threshold, then extract the valid data from the log chunk with the longer length to form a third log chunk, and the length of the third log chunk satisfies a matching relationship with the length of the shorter log chunk;
[0113] Compare the third log chunk with the shorter log chunk.
[0114] Such as Figure 7As shown in the figure, if one log block is too long and the other is too short during the comparison process, that is, when comparing long and short log blocks, the long log block can be traversed cyclically to extract the valid data therein, so as to form a third log block with a shorter length. For example, the length of the third log block is about 1.3 times, or about 1.2 times, etc. of the length of the shorter log block to be compared. As long as the length difference between the two log blocks to be compared is not large and they match each other. Because only when the lengths match can more effective comparison be carried out, especially for the subsequent comparison of the compilation distance, the error is smaller.
[0115] Further, comparing the log blocks of the corresponding standard log file with the log blocks of the corresponding log file to be tested includes:
[0116] Processing each log block based on the feature generation algorithm to determine the features of each log block;
[0117] Comparing the frequency similarity of the log blocks of the corresponding standard log file with the log blocks of the corresponding log file to be tested based on the features of each log block;
[0118] If the determined frequency similarity meets the first matching threshold, then compare the edit distance similarity of the log blocks of the corresponding standard log file with the log blocks of the corresponding log file to be tested.
[0119] Specifically, before performing the similarity matching between log blocks, feature generation needs to be performed on each log block. The features can be extracted by using the feature generation algorithm. For example, the valid information of the log block is calculated by hash, and the smallest byte in the hash result is obtained as the feature record, etc. As Figure 9 shown, in this embodiment, when performing the similarity matching of log blocks, the comparison of the frequency similarity is relatively fast. Because the features for comparison are of fixed length, fast similarity matching can be supported for fast filtering to obtain log blocks with similarity meeting the requirements. Then, the edit distance similarity matching is performed on this log block. This matching is mainly to compare the similarity of the feature structures of the log blocks. The execution code of the above comparison process can be referred to Figure 10 shown. When performing the frequency similarity matching, in this embodiment, the first matching threshold is set to 0.68. If the frequency similarity is less than 0.68, it is considered that there is no similarity between the two feature data. And when performing the edit distance similarity matching, if the similarity is less than 0.8, it is considered that there is no similarity between the two log blocks.
[0120] In practical applications, when calculating the frequency similarity, the frequencies of the feature data of the two log blocks to be compared in the corresponding code can be counted, the frequencies of the two data sets are compared, and their Euclidean distance is calculated. The greater the Euclidean distance, the lower the similarity. Since the element frequency is a value between 0 and 1 (inclusive) and the sum of the frequencies of all elements is 1, the calculated Euclidean distance ranges from 0 to 1. Subtracting the Euclidean distance from 1 gives the frequency similarity of the current two data sets.
[0121] When comparing the edit distance similarity between the log blocks of the corresponding standard log file and the log blocks of the corresponding log file to be tested, it includes:
[0122] Calculate the shortest edit distance of the two log blocks to be compared respectively;
[0123] Determine the differential edit distance based on the shortest edit distance;
[0124] Determine the edit distance similarity based on the differential edit distance.
[0125] For example, calculate the shortest edit distance of the data of the two log blocks to be compared, calculate the difference between the edit distances of the two data sets (i.e., the differential edit distance), and count the similarity based on the edit distance difference. The relevant implementation can call the Levenshtein.ratio(str1, str2) function in Python. Through the shortest edit distance algorithm, the shortest edit operations for changing from sample A to sample B can be obtained. This means that two similar samples can be transformed into each other through the shortest edit operations (such as insertion, replacement, deletion) given by the algorithm. Therefore, these shortest compilation operations can be recorded as the difference points between sample A and sample B for subsequent process operations. If the above-mentioned edit distance similarity meets the preset range, it can be determined that the data in the current log block probably has process difference points.
[0126] As an optional embodiment, the method in this embodiment further includes:
[0127] Count the number of times each log block of the log file to be tested with an edit distance similarity not meeting the second matching threshold appears repeatedly in the source code segment;
[0128] Calculate the standard deviation of all statistical values;
[0129] Calculate the Euclidean distance value between each statistical value and the standard deviation;
[0130] Sort the multiple Euclidean distance values from largest to smallest and output the sorting result. Among them, the Euclidean distance value is proportional to the occurrence of process difference points in the log block of the corresponding log file to be tested and the degree of difference of the process difference points.
[0131] For example, the number of repetitions of each log block in the log file to be tested whose edit distance similarity does not meet the second matching threshold in the source code snippet is counted respectively. This is equivalent to counting the number of repetitions of all log blocks where process difference points may occur in the corresponding source code snippet or source code, and can be denoted as a = {...}. Calculate the standard deviation of a, and this standard deviation is denoted as std. Then calculate the Euclidean distance between each value in a and std, and denote this distance as b = {...}. Sort the obtained b from largest to smallest, and it is considered that the greater the deviation from the standard deviation of the difference point, the greater the possibility of being a key difference point. That is, the log block corresponding to the feature data with a greater deviation from the standard deviation, the degree of the process difference point and the difference degree of this process difference point, or in other words, the greater the role it plays in whether the code snippet runs correctly. The system can output the above sorting result and description to the user accordingly, so that the user can analyze and determine the location of the actual process difference point based on this sorting result.
[0132] Specifically, when the method in this embodiment is actually applied, for example, when applied to quickly locate the process difference point in the extracted code, it can specifically use the code in FreeRTOS on windows as an example. FreeRTOS is an open-source operating system based on specific IOT hardware. Now the user wants to extract the TCP / IP part and use it in their own product. In actual work, the extracted code fails to be initialized successfully, that is, the running result after the initialization of this section of code is inconsistent with the running result in the actual source code. Therefore, it is necessary to find the reason for the difference. Refer to Figure 10 、 11 、Figure 12, first modify the parameters of the automated script for instrumentation, use the script to parse the key points in the C language source code, and mark the process control statements, that is, mark the key points of process control. Then automatically add log printing code at the process control statements in the script. The log printing code can refer to the following code content: printf("%s;%s;%d;%ul;\n",__FILE__,__func__,__LINE__,GetCurrentThreadId()). Then use the obtained script to insert log printing code into the FreeRTOS source code and the extracted code respectively. The operation effect is as Figure 11 shown. Then, as Figure 12As shown, compile the FreeRTOS source code and the product project source code containing the above script, run and compare the generated Log files. By comparing the script with the log file and checking the results, the key points of process differences can be easily found. In this example, through comparison, it is found that the problem point that may cause abnormal process differences is on line 1563 of the FreeRTOS_IP.c file. Based on this, the user can view the context of the FreeRTOS code and the product project source code, and find the actual problem in the product project source code, that is, find the actual process difference point.
[0133] As Figure 13 shown, another embodiment of the present application also provides a device for locating process differences, including:
[0134] A first acquisition module, configured to acquire source code;
[0135] An extraction module, configured to extract source code fragments from the source code;
[0136] An initialization module, configured to initialize the source code fragments and compile and run the initialized source code fragments;
[0137] A second acquisition module, configured to obtain the running result;
[0138] An instrumentation module, configured to perform an instrumentation operation at corresponding positions in the source code fragment and the source code when the running result is different from the result when the source code fragment runs in the source code;
[0139] A processing module, configured to run the instrumented source code fragment and the source code to at least respectively find the process control key points in the source code fragment and the source code, and obtain a log file recording the running data of the instrumented source code fragment and the source code;
[0140] A determination module, configured to at least assist in determining the process difference position of the source code fragment relative to the source code according to the log file.
[0141] As an optional embodiment, running the instrumented source code fragment and the source code to at least respectively find the process control key points in the source code fragment and the source code includes:
[0142] Performing block processing on the source code fragment and the source code respectively to obtain a plurality of code blocks;
[0143] Determining, at least based on the context of each code block in the corresponding source code fragment or source code, a first code block as a candidate process control key point from the plurality of code blocks;
[0144] Perform key point filtering on the first code block to obtain a second code block corresponding to the source code segment and the process control key points in the source code respectively.
[0145] As an optional embodiment, determining the first code block as a candidate process control key point from multiple code blocks based at least on the context of each code block in the corresponding source code segment or source code includes:
[0146] Determine the context of each code block in the corresponding source code segment or source code;
[0147] Determine whether there is a target keyword in the context. If so, determine the corresponding code block as the first code block;
[0148] The method further includes:
[0149] Determine the function headers in the source code segment and the source code as the candidate process control key points.
[0150] As an optional embodiment, running the instrumented source code segment and source code to obtain a log file recording the running data of the instrumented source code segment and source code includes:
[0151] Insert log printing functions at each process key point in the source code segment and source code, so as to generate and output the log file respectively when the instrumented source code segment and source code are running.
[0152] As an optional embodiment, the log file includes the running data of multiple threads. Among them, the log file corresponding to the source code is the standard log file, and the log file corresponding to the source code segment is the log file to be tested;
[0153] The at least based on the log file to assist in determining the process difference position of the source code segment relative to the source code includes:
[0154] Block the standard log file and the log file to be tested based on the thread ID respectively to obtain multiple log blocks corresponding to each thread;
[0155] Compare the log block corresponding to the standard log file with the log block corresponding to the log file to be tested to assist in determining the process difference position of the source code segment.
[0156] As an optional embodiment, the comparing the log block corresponding to the standard log file with the log block corresponding to the log file to be tested includes:
[0157] Compare each log block in the log file to be tested with all log blocks in the corresponding standard log file;
[0158] During the comparison process, determine the length of each log block being compared. If the length difference between the two log blocks being compared meets the threshold, extract the valid data from the log block with the longer length to form a third log block, and the length of the third log block and the length of the shorter log block meet the matching relationship;
[0159] Compare the third log block with the shorter log block.
[0160] As an optional embodiment, the comparison of the log blocks in the corresponding standard log file with the log blocks in the corresponding log file to be tested includes:
[0161] Process each log block based on a feature generation algorithm to determine the features of each log block;
[0162] Compare the frequency similarity of the log blocks in the corresponding standard log file with the log blocks in the corresponding log file to be tested based on the features of each log block;
[0163] If the determined frequency similarity meets the first matching threshold, compare the edit distance similarity of the log blocks in the corresponding standard log file with the log blocks in the corresponding log file to be tested.
[0164] As an optional embodiment, the comparison of the edit distance similarity of the log blocks in the corresponding standard log file with the log blocks in the corresponding log file to be tested includes:
[0165] Calculate the shortest edit distance between the two log blocks being compared respectively;
[0166] Determine the differential edit distance based on the shortest edit distance;
[0167] Determine the edit distance similarity based on the differential edit distance.
[0168] As an optional embodiment, the device in this embodiment further includes:
[0169] A statistics module for respectively counting the number of times each log block in the log file to be tested with an edit distance similarity not meeting the second matching threshold appears repeatedly in the source code segment;
[0170] A first calculation module for calculating the standard deviation of all statistical values;
[0171] A second calculation module for calculating the Euclidean distance value between each statistical value and the standard deviation;
[0172] A sorting module, configured to perform sorting from large to small according to the plurality of Euclidean distance values and output a sorting result, wherein the Euclidean distance value is proportional to the difference point of the log block occurrence process of the corresponding log file to be tested and the degree of difference of the difference point.
[0173] Another embodiment of the present invention further provides an electronic device, including:
[0174] One or more processors;
[0175] A memory configured to store one or more programs;
[0176] When the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the above method.
[0177] One embodiment of the present invention further provides a storage medium, on which a computer program is stored, and when the program is executed by a processor, the above-mentioned method is implemented. It should be understood that each of the solutions in this embodiment has the corresponding technical effects in the above method embodiment, and will not be described in detail here.
[0178] The embodiment of the present invention further provides a computer program product, the computer program product is tangibly stored on a computer-readable medium and includes computer-readable instructions, and the computer-executable instructions, when executed, cause at least one processor to execute the method in the above-mentioned embodiment. It should be understood that each of the solutions in this embodiment has the corresponding technical effects in the above method embodiment, and will not be described in detail here.
[0179] It should be noted that the computer storage medium of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable medium can, for example but is not limited to, be an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access storage medium (RAM), a read-only storage medium (ROM), an erasable programmable read-only storage medium (EPROM or flash memory), an optical fiber, a portable compact disk read-only storage medium (CD-ROM), an optical storage medium, a magnetic storage medium, or any suitable combination of the above. In the present application, a computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device. And in the present application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program configured to be used by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, antenna, optical cable, RF, etc., or any suitable combination of the above.
[0180] It should be understood that although the present application is described according to various embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
[0181] The above embodiments are only exemplary embodiments of the present invention and are not used to limit the present invention. The protection scope of the present invention is defined by the claims. Those skilled in the art can make various modifications or equivalent replacements within the essence and protection scope of the present invention, and such modifications or equivalent replacements should also be regarded as falling within the protection scope of the present invention.
Claims
1. A method for locating process differences, characterized in that, Including: Obtain the source code; Extract source code fragments from the source code; Initialize the source code fragments and compile and run the initialized source code fragments; Obtain the running result; In the case where the running result is different from the result when the source code fragment runs in the source code, perform a stamping operation at the corresponding positions of the source code fragment and the source code; Run the instrumented source code fragment and source code to at least respectively find the process control key points in the source code fragment and source code, and obtain a log file recording the running data of the instrumented source code fragment and source code; Among them, running the instrumented source code fragment and source code to at least respectively find the process control key points in the source code fragment and source code includes: performing block processing on the source code fragment and source code respectively to obtain a plurality of code blocks; at least based on the context of each code block in the corresponding source code fragment or source code, determining a first code block as a candidate process control key point from the plurality of code blocks; performing key point filtering processing on the first code block to obtain a second code block corresponding to the process control key points in the source code fragment and the source code respectively; Among them, the at least based on the context of each code block in the corresponding source code fragment or source code, determining a first code block as a candidate process control key point from the plurality of code blocks includes: determining the context of each code block in the corresponding source code fragment or source code; determining whether there is a target keyword in the context, if so, determining the corresponding code block as the first code block; the method further includes: determining the function headers in the source code fragment and source code as the candidate process control key points; At least based on the log file, assist in determining the process difference position of the source code fragment relative to the source code.
2. The method according to claim 1, wherein Running the instrumented source code fragment and source code to obtain a log file recording the running data of the instrumented source code fragment and source code includes: Insert log printing functions at the corresponding positions of each process control key point in the source code fragment and source code, so as to respectively generate and output the log file when the instrumented source code fragment and source code run.
3. The method according to claim 1, wherein The log files all include the running data of multiple threads, where the log file corresponding to the source code is the standard log file, and the log file corresponding to the source code fragment is the log file to be tested; The at least based on the log file, assist in determining the process difference position of the source code fragment relative to the source code includes: Respectively perform block processing on the standard log file and the log file to be tested based on the thread ID to obtain a plurality of log blocks corresponding to each thread respectively; Compare the log blocks corresponding to the standard log file with the log blocks corresponding to the log file to be tested to assist in determining the process difference position of the source code fragment.
4. The method according to claim 3, wherein The comparison of the log blocks corresponding to the standard log file with the log blocks corresponding to the log file to be tested includes: Comparing each log block in the log file to be tested with all log blocks in the corresponding standard log file; During the comparison process, determine the length of each log block being compared. If the length difference between the two log blocks being compared meets the threshold, extract the valid data from the log block with the longer length to form a third log block, and the length of the third log block satisfies a matching relationship with the length of the shorter log block; Compare the third log block with the shorter log block.
5. The method according to claim 3, characterized in that, The comparison of the log blocks corresponding to the standard log file with the log blocks corresponding to the log file to be tested includes: Processing each log block based on a feature generation algorithm to determine the features of each log block; Comparing the frequency similarity of the log blocks corresponding to the standard log file with the log blocks corresponding to the log file to be tested based on the features of each log block; If the determined frequency similarity meets the first matching threshold, compare the edit distance similarity of the log blocks corresponding to the standard log file with the log blocks corresponding to the log file to be tested.
6. The method according to claim 5, wherein The comparison of the edit distance similarity of the log blocks corresponding to the standard log file with the log blocks corresponding to the log file to be tested includes: Calculating the shortest edit distance between the two log blocks being compared respectively; Determining the differential edit distance based on the shortest edit distance; Determining the edit distance similarity based on the differential edit distance.
7. The method according to claim 6, wherein It further includes: Counting the number of repetitions of each log block of the log file to be tested whose edit distance similarity does not meet the second matching threshold in the source code segment respectively; Calculating the standard deviation of all statistical values; Calculating the Euclidean distance value between each statistical value and the standard deviation; Sorting the multiple Euclidean distance values from largest to smallest and outputting the sorting result, where the Euclidean distance value is proportional to the difference points in the occurrence process of the log blocks of the corresponding log file to be tested and the degree of difference in the difference points of the process.
8. A process difference positioning device, characterized in that It includes: A first acquisition module for acquiring the source code; An extraction module for extracting the source code segment from the source code; An initialization module for initializing the source code segment and compiling and running the initialized source code segment; A second acquisition module for obtaining the running result; An instrumentation module for performing an instrumentation operation at the corresponding positions in the source code segment and the source code when the running result is different from the result of running the source code segment in the source code; A processing module for running the instrumented source code segment and the source code to at least respectively find the process control key points in the source code segment and the source code, and obtaining a log file recording the running data of the instrumented source code segment and the source code. Among them, running the source code segment and the source code after instrumentation to respectively find at least the flow control key points in the source code segment and the source code, including: respectively performing chunking processing on the source code segment and the source code to obtain a plurality of code chunks; determining, at least based on the context of each code chunk in the corresponding source code segment or source code, a first code chunk as a candidate flow control key point from the plurality of code chunks; performing key point filtering processing on the first code chunk to obtain a second code chunk corresponding to the flow control key point in the source code segment and the source code respectively; Among them, the determining, at least based on the context of each code chunk in the corresponding source code segment or source code, a first code chunk as a candidate flow control key point from the plurality of code chunks includes: determining the context of each code chunk in the corresponding source code segment or source code; determining whether there is a target keyword in the context, and if so, determining the corresponding code chunk as the first code chunk; further including: determining the function headers in the source code segment and the source code as the candidate flow control key points; A determining module is configured to at least assist in determining the flow difference position of the source code segment relative to the source code according to the log file.
Citation Information
Patent Citations
A test device and method for performing white-box testing on coverage calculation visualization
CN104331361A
Monitoring method and device for JAVA application, server and storage medium
CN109542444A