Software code defect detection method and system based on program code feature fusion
By combining static and dynamic analysis tools to extract defect features, perform preprocessing and weighted fusion, and using decision trees and density clustering algorithms to identify and locate stealth defects, the problem of difficulty in detecting software code stealth defects in the existing technology is solved, and efficient defect repair and quality improvement is achieved.
Patent Information
- Application Number
- CN202510421827.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
It is difficult for the prior art to effectively detect and locate invisible defects in software code, especially problems such as confusion in control flow, improper exception handling and resource leakage. Traditional tools have not had the best detection effect.
The combination of static analysis and dynamic analysis is used to extract defect features and preprocess and weighted fusion, and the decision tree model is used for classification, and the defect type and distribution are determined by combining density clustering algorithms, and repair suggestions are generated, and the test case library is optimized for detection through periodic updates.
Effectively identify and locate invisible defects in code, improve software quality and reliability, provide developers with timely and accurate defect repair guidance, and improve detection coverage and accuracy.
Smart Images

Figure CN120256273A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information technology, and in particular, to a software code defect detection method and system based on program code feature fusion. Background Art
[0002] In the process of software development, stealth defects are a type of defects that are difficult to discover and locate. They often hide in complex code logics and execution paths. These stealth defects may stem from problems such as chaotic control flows, improper exception handling, and resource leaks in the code. Although traditional static analysis tools and dynamic analysis tools can detect some common defects, their detection effects for stealth defects are not ideal. Static analysis tools mainly focus on the syntax and structure of the code and are difficult to deeply analyze complex execution paths and runtime states; although dynamic analysis tools can capture runtime exceptions and errors, they are limited by the coverage rate of test cases and cannot comprehensively reveal hidden defects.
[0003] Therefore, how to effectively converge and aggregate the dispersed stealth defect features in the code has become an urgent technical problem to be solved. On the one hand, we need to comprehensively utilize static analysis and dynamic analysis technologies to mine the stealth defect features in the code from multiple dimensions; on the other hand, we need to design a unified detection framework to fuse and integrate these dispersed features to form a complete stealth defect detection solution. During the feature fusion process, we need to consider the correlation and complementarity between different features to avoid feature redundancy and conflicts; when designing the detection framework, we need to take into account the scalability and adaptability of the framework to cope with the constantly changing software development scenarios.
[0004] In short, the convergence and aggregation of stealth defects is a complex technical problem that requires in-depth exploration and innovation at the theoretical and practical levels. Only by continuously mining new defect features, optimizing the feature fusion algorithm, and improving the detection framework design can we truly achieve the accurate detection and efficient repair of stealth defects, and improve software quality and reliability. Summary of the Invention
[0005] On the one hand, the present invention provides a software code defect detection method based on program code feature fusion, mainly including: Using a static analysis tool to check the syntax and structure of the code, and extracting defect features related to chaotic control flow, exception handling, and resource leakage; Running preset test cases through a dynamic analysis tool to capture abnormal behaviors in the execution path and runtime state, and obtaining defect features not covered by static analysis; Preprocess the features extracted by static analysis and dynamic analysis, remove redundant features, retain the key features related to code logic and execution paths, and generate a preprocessed defect feature set. The preprocessed defect feature set is represented in vector form, with each element corresponding to a specific defect feature; Design a feature fusion algorithm. Based on a pre-established weight model, perform weighted fusion on the control flow chaos, exception handling, and resource leakage features in the preprocessed defect feature set to generate a comprehensive defect feature vector. The weight model is constructed based on historical data and expert experience, reflecting the importance of different features for defect detection; Construct a unified detection framework. Input the comprehensive defect feature vector into a pre-established decision tree model, classify the features, and determine whether there are stealth defects in the code. The decision tree model is trained with labeled historical defect data for identifying new defect patterns; If the output result of the decision tree model indicates the existence of a defect, use a density-based clustering algorithm to group the comprehensive defect feature vector, determine the type and distribution range of the defect. The density clustering algorithm is selected to handle defect features with non-spherical distributions; Generate defect repair suggestions according to the clustering results, and combine code logic and execution path information to determine the specific location and impact range of the defect; Periodically feedback the repair suggestions and location results to the code repository, update the test case libraries of the static analysis tool and the dynamic analysis tool, and optimize the coverage and accuracy of subsequent defect detection for newly added defect types and features. This step is executed after multiple detection iterations as part of continuous improvement.
[0006] On the other hand, the present invention provides a software code defect detection system based on program code feature fusion, mainly including: A static analysis module for performing syntax and structure checks on the code using a static analysis tool, and extracting defect features related to control flow chaos, exception handling, and resource leakage; A dynamic analysis module for running preset test cases through a dynamic analysis tool, capturing abnormal behaviors in the execution path and runtime state, and obtaining defect features not covered by static analysis; A feature preprocessing module for preprocessing the features extracted by static analysis and dynamic analysis, removing redundant features, retaining the key features related to code logic and execution paths, and generating a preprocessed defect feature set. The preprocessed defect feature set is represented in vector form, with each element corresponding to a specific defect feature; A feature fusion module for designing a feature fusion algorithm to perform weighted fusion on the control flow chaos, exception handling, and resource leakage features in the preprocessed defect feature set based on a pre-established weight model, generating a comprehensive defect feature vector. The weight model is constructed according to historical data and expert experience, reflecting the importance of different features for defect detection; A defect detection module for constructing a unified detection framework, inputting the comprehensive defect feature vector into a pre-established decision tree model to classify the features and determine whether there are stealth defects in the code. The decision tree model is trained with labeled historical defect data for identifying new defect patterns; A defect clustering module for, if the output result of the decision tree model is that there are defects, using a density-based clustering algorithm to group the comprehensive defect feature vector, determining the type and distribution range of the defects. The density clustering algorithm is selected to handle defect features with non-spherical distributions; A repair suggestion generation module for generating defect repair suggestions according to the clustering results, and determining the specific location and influence range of the defects in combination with code logic and execution path information; A continuous optimization module for periodically feeding back the repair suggestions and location results to the code library, updating the test case libraries of static analysis tools and dynamic analysis tools, and optimizing the coverage rate and accuracy of subsequent defect detection for newly added defect types and features. This step is executed after multiple detection iterations as part of continuous improvement.
[0007] The technical solution provided by the embodiments of the present invention may include the following beneficial effects: The present invention discloses a method and system for detecting stealth defects in code. Defect features in the code are extracted through a combination of static analysis and dynamic analysis. After preprocessing and weighted fusion of the features, a comprehensive defect feature vector is generated. A decision tree model is used to classify the features to determine whether there are stealth defects. If there are defects, a density clustering algorithm is used to determine the defect type and distribution range, and repair suggestions are generated. The present invention also continuously optimizes the detection coverage rate and accuracy by periodically feeding back and updating the test case library. This method can effectively identify and locate stealth defects in the code, improve software quality and reliability, and provide timely and accurate defect repair guidance for developers. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Figure 1 It is a flowchart of a method for detecting software code defects based on program code feature fusion according to the present invention.
[0009] Figure 2 It is a schematic diagram of a method and system for detecting software code defects based on program code feature fusion according to the present invention.
[0010] Figure 3Another schematic diagram of a software code defect detection method and system based on program code feature fusion according to the present invention.
[0011] Figure 4 Structural schematic diagram of a software code defect detection method and system based on program code feature fusion according to the present invention. Specific implementation manner
[0012] Next, the technical solutions in the embodiments of the present invention will be clearly and detailedly described in conjunction with the accompanying drawings in the embodiments of the present invention. The described embodiments are only a part of the embodiments of the present invention.
[0013] Such as Figures 1-4 , a software code defect detection method and system based on program code feature fusion in this embodiment may specifically include: S101. Use a static analysis tool to perform syntax and structure checks on the code, and extract defect features related to control flow chaos, exception handling, and resource leakage.
[0014] Use a static analysis tool to perform a syntax check on the code to obtain the syntax error information of the code. Use a static analysis tool to perform a structure check on the code to obtain the structure error information of the code. Extract the control flow chaos feature from the syntax error information and the structure error information to determine the code segment with control flow chaos. Extract the exception handling feature from the syntax error information and the structure error information to determine the code segment with improper exception handling. Extract the resource leakage feature from the syntax error information and the structure error information to determine the code segment with resource leakage. If the control flow chaos feature exists, use a control flow optimization algorithm to reconstruct the code segment. If the exception handling feature exists, use an exception handling optimization algorithm to reconstruct the code segment.
[0015] Exemplarily, a static analysis tool such as FindBugs is used to scan Java code. A rule set is set up to cover common syntax errors and bad programming habits, such as unused variables, null pointer references, etc., and the warning threshold is set to medium level and above, that is, defects with a severity level of medium and above in the scan results need to be concerned about and fixed. After the scan is completed, the tool will generate a defect report, which contains detailed defect location and type information. Taking a specific code snippet as an example, assume that FindBugs detects an unused private field `private int tempValue = 100;` in a certain class. The report will point out that this field has never been read in the class and suggest deleting it to optimize the code structure. Next, a control flow analysis tool such as JDeodorant is used to construct a control flow graph of the code to detect situations of chaotic control flow. For example, JDeodorant will analyze the control flow inside each method and calculate the cyclomatic complexity. Assume that the cyclomatic complexity of a certain method is as high as 25, far exceeding the recommended value of 10. The tool will mark this method as a high-risk area. By visualizing the control flow graph, it can be found that there are multiple nested `if-else` branches and loop structures inside this method, resulting in a complex logical path, which is difficult to understand and maintain and needs to be refactored. In terms of exception handling, a static analysis tool such as SonarQube can be used to check for defects related to exception handling. For example, SonarQube can be configured with rules to detect issues such as empty `catch` blocks and over-catching exceptions. Assume that there is an empty exception handling block `catch (Exception e) {}` in a certain piece of code. SonarQube will mark it as a defect and give the specific code location and recommended repair measures, such as adding logging or throwing a more specific exception. In terms of resource leak detection, tools such as LeakCanary can be used to perform static and dynamic analysis on Android applications. For example, LeakCanary can monitor the object reference situation in the application and issue a warning when a potential memory leak is detected. Assume that in a certain Activity, a static `List <view>If an object holds a reference to a View in a destroyed Activity, LeakCanary will detect this situation and generate a detailed reference chain report to indicate the source of the leak, helping developers locate and fix the problem.
[0016] S102: Run the preset test cases through a dynamic analysis tool, capture the abnormal behaviors in the execution path and runtime state, and obtain the defect features not covered by static analysis.
[0017] According to the execution path captured by the dynamic analysis tool, extract the control flow chaos features and determine the code segments with control flow chaos. Extract the exception handling features from the runtime state and determine the code segments with improper exception handling. Through the analysis of abnormal behaviors, extract the resource leakage features and determine the code segments with resource leakage. If the control flow chaos features exist, use the control flow optimization algorithm to reconstruct the code segments. If the exception handling features exist, use the exception handling optimization algorithm to reconstruct the code segments. If the resource leakage features exist, use the resource management optimization algorithm to reconstruct the code segments. Generate new test cases based on the reconstructed code segments, run the dynamic analysis tool again, and obtain the optimized defect features.
[0018] Exemplarily, combining the results of static analysis, we can further utilize dynamic analysis tools to verify and discover deeper defects. For example, use Memcheck in the Valgrind toolchain to detect memory errors in C / C++ programs, such as illegal memory access and memory leaks. Suppose there is a C++ program where static analysis reveals a potential risk of memory out-of-bounds, but the specific scenario cannot be determined. By writing test cases with various boundary conditions, such as an array index of 0, one less than the array length, and the array length, and running these test cases using Memcheck, when a function in the program processes an array index equal to the array length and wrongly writes to memory outside the array boundary, Memcheck will capture this illegal write operation and report detailed error information, including the memory address where the error occurred, the call stack, and the specific code line that caused the error. For example, `Invalid write of size 4 at 0x4C329EB:testFunction(test.cpp:55)` indicates that a 4-byte illegal write operation occurred in the `testFunction` function on line 55 of the test.cpp file, accessing the memory address 0x4C329EB. For Java programs, performance analysis tools such as JProfiler can be used, combined with stress testing, to monitor the performance and resource usage of the program during long-term operation or under high load. For example, conduct a stress test on a web server, simulating 1000 concurrent users continuously making requests for 1 hour, and JProfiler monitors the usage of resources such as the memory, CPU, and threads of the JVM in real time. After the test, a performance report is generated, and it is found that the CPU occupancy rate of a certain business processing thread continuously reaches as high as 95%, and there is a loop in the method executed by this thread. Temporary objects are frequently created inside the loop, causing the garbage collector to work frequently. Further analysis reveals that the object creation inside this loop can be optimized using object pool technology to reduce the overhead of object creation and destruction. In Android application development, the built-in Profiler tool in Android Studio can be utilized, combined with Monkey testing, to comprehensively analyze the performance of the application. For example, use the Monkey tool to perform random event testing on the application, simulating various irregular operations of users, and at the same time enable Profiler to monitor the usage of CPU, memory, network, and battery power.After a 2-hour Monkey test, the Profiler report shows that the memory usage of the application continues to rise with multiple peaks. By analyzing the memory snapshots, it is found that a certain background service fails to release resources in a timely manner after completing tasks, resulting in memory leakage. Further examination of the code reveals that the service holds a static collection that continuously adds data without clearing, leading to continuous memory growth.
[0019] S103. Preprocess the features extracted by static analysis and dynamic analysis, remove redundant features, retain the key features related to code logic and execution paths, and generate a preprocessed defect feature set. The preprocessed defect feature set is represented in vector form, with each element corresponding to a specific defect feature.
[0020] Use a feature selection algorithm to perform dimensionality reduction on the preprocessed defect feature set, remove redundant features, and retain key features. Based on the dimensionality-reduced feature set, construct a defect feature vector space model and map each defect feature to a specific dimension in the vector space. Classify the defect feature vectors through a clustering algorithm to determine the defect categories and their distribution patterns. If the defect category belongs to chaotic control flow, use a control flow optimization algorithm to reconstruct the code segment. If the defect category belongs to improper exception handling, use an exception handling optimization algorithm to reconstruct the code segment. If the defect category belongs to resource leakage, use a resource management optimization algorithm to reconstruct the code segment. Generate new test cases based on the reconstructed code segment, run the dynamic analysis tool again, and obtain the optimized defect features.
[0021] Exemplarily, through static analysis and dynamic analysis, we have collected a series of defect features. For example, uninitialized variables, null pointer dereferences, out-of-bounds array accesses identified by static analysis, and memory leaks, concurrent access conflicts, performance bottlenecks, etc. captured by dynamic analysis. Suppose static analysis outputs 100 features, including variable types, variable scopes, function call relationships, etc. For example, the variable `int a` is uninitialized, the function `foo()` has a risk of null pointer dereference, and the array `arr
[10] ` may have out-of-bounds access when accessing index `i`. Dynamic analysis outputs 50 features, including memory allocation / deallocation addresses, thread IDs, function execution times, etc. For example, 128 bytes of memory at address `0x7ffeea7b8` are not released after being allocated in the function `bar()`, and threads `T1` and `T2` simultaneously access the shared variable `count`. We integrate these features and use feature selection algorithms such as information gain to evaluate the importance of each feature. For example, by calculating the information gain of each feature, it is found that features such as "uninitialized variables", "null pointer dereferences", and "memory leaks" have higher information gains, while features such as "variable types" and "thread IDs" have lower information gains. According to the information gain ranking, we eliminate 20 redundant features with information gains lower than the threshold of 1 and retain 130 key features. Then, these features are transformed into vector form. For example, "uninitialized variable" is represented as `[1, 0, 0,...]`, "null pointer dereference" is represented as `[0, 1, 0,...]`, "memory leak" is represented as `[0, 0, 1,...]`, where each dimension of the vector corresponds to a specific feature, 1 indicates the presence of the feature, and 0 indicates the absence. Finally, we obtain a set of 130-dimensional defect feature vectors for subsequent defect prediction and analysis. This process fuses the features at the code logic level obtained from static analysis with the features at the runtime state level obtained from dynamic analysis to generate feature vectors, providing high-quality input data for subsequent defect prediction and analysis models.
[0022] S104. Design a feature fusion algorithm. Based on a pre-established weight model, perform weighted fusion on the control flow chaos, exception handling, and resource leak features in the preprocessed defect feature set to generate a comprehensive defect feature vector. The weight model is constructed based on historical data and expert experience, reflecting the importance of different features for defect detection.
[0023] The feature fusion algorithm is used to perform weighted fusion on the control flow chaos, exception handling, and resource leakage features to generate a comprehensive defect feature vector. According to the pre-established weight model, the weight value of each feature is determined to reflect its importance for defect detection. The weight model is constructed through historical data and expert experience to ensure the accuracy and reliability of the weight values. The weighted fusion method is adopted to perform weighted summation on the control flow chaos, exception handling, and resource leakage features according to the weight values to generate a comprehensive defect feature vector. Based on the comprehensive defect feature vector, the defect type and its severity are judged. If the defect type is control flow chaos, the control flow optimization algorithm is used to reconstruct the code segment. If the defect type is improper exception handling, the exception handling optimization algorithm is used to reconstruct the code segment. If the defect type is resource leakage, the resource management optimization algorithm is used to reconstruct the code segment.
[0024] Exemplarily, assume that we have obtained a set of 130-dimensional defect feature vectors from the previous step, which contain features related to control flow chaos, exception handling, and resource leakage. We now need to perform weighted fusion on these features. First, we consult historical defect data and expert experience to construct a weight model. For example, we have counted the occurrence frequencies and severities of three types of defects, namely control flow chaos, improper exception handling, and resource leakage, in the past 1000 software projects. The results show that defects caused by control flow chaos account for 30% with an average repair cost of 10 person-hours; defects caused by improper exception handling account for 25% with an average repair cost of 8 person-hours; and defects caused by resource leakage account for 20% with an average repair cost of 12 person-hours. Based on these data and combined with expert experience, we assign weights 35, 3, and 35 to these three types of features respectively. These weights reflect the importance of different features for defect detection, and the larger the value, the more important it is. Next, we perform weighted fusion on each defect feature vector. For example, for a specific code snippet, among the preprocessed feature vectors, there are 5 features related to control flow chaos with a value of 1 (indicating existence), 3 features related to exception handling with a value of 1, and 2 features related to resource leakage with a value of 1. We multiply these feature values by their corresponding weights and then sum them up. The weighted sum of the control flow chaos features is 5 * 35 = 75, the weighted sum of the exception handling features is 3 * 3 = 9, and the weighted sum of the resource leakage features is 2 * 35 = 7. Finally, we perform normalization on these three weighted sums to obtain the comprehensive defect scores of the code snippet in the three dimensions of control flow chaos, exception handling, and resource leakage. For example, by dividing 75, 9, and 7 by their sum (75 + 9 + 7 = 35) respectively, we get the normalized scores approximately as 52, 27, and 21. These scores can be used as the defect tendency indicators of the code snippet in the three dimensions. To further integrate this information, we can combine the scores of these three dimensions into a new 3D vector [52, 27, 21] as the comprehensive defect feature vector of the code snippet. This vector comprehensively considers the information of control flow chaos, exception handling, and resource leakage and is weighted according to their importance. In this way, we perform the above-mentioned weighted fusion operation on the defect feature vectors of each code snippet, and finally obtain a brand-new set of comprehensive defect feature vectors, providing more representative inputs for the subsequent defect prediction model.
[0025] S105. Construct a unified detection framework, input the comprehensive defect feature vector into a pre-established decision tree model, classify the features, and determine whether there are stealth defects in the code. The decision tree model is trained with labeled historical defect data and is used to identify new defect patterns.
[0026] The feature fusion algorithm is used to perform weighted fusion on the control flow chaos, exception handling, and resource leakage features to generate a comprehensive defect feature vector. According to the pre-established weight model, the weight value of each feature is determined to reflect its importance for defect detection. The weight model is constructed through historical data and expert experience to ensure the accuracy and reliability of the weight values. The weighted fusion method is adopted to perform weighted summation on the control flow chaos, exception handling, and resource leakage features according to the weight values to generate a comprehensive defect feature vector. Based on the comprehensive defect feature vector, the defect type and its severity are judged. If the defect type is control flow chaos, the control flow optimization algorithm is used to reconstruct the code segment. If the defect type is improper exception handling, the exception handling optimization algorithm is used to reconstruct the code segment. If the defect type is resource leakage, the resource management optimization algorithm is used to reconstruct the code segment.
[0027]
[0028] Er represents the resource management optimization effect, R_0 represents the degree of resource leakage before optimization, and R_1 represents the degree of resource leakage after optimization. This formula is used to evaluate the effect of the resource management optimization algorithm.
[0029]
[0030] Ee represents the exception handling optimization effect, E0 represents the degree of improper exception handling before optimization, and E1 represents the degree of improper exception handling after optimization. This formula is used to evaluate the effect of the exception handling optimization algorithm.
[0031]
[0032] Ec represents the control flow optimization effect, C0 represents the degree of control flow chaos before optimization, and C1 represents the degree of control flow chaos after optimization. This formula is used to evaluate the effect of the control flow optimization algorithm.
[0033]
[0034] S represents the severity of the defect, F represents the size of the comprehensive defect feature vector, T represents the severity coefficient of the defect type, and α and β are adjustment parameters. This formula is used to evaluate the overall severity of the defect.
[0035]
[0036] wi represents the weight value of the i-th feature, hj represents the importance of the j-th sample in the historical data, ej represents the expert's score for the j-th sample, and n represents the total number of samples. This formula describes how to construct a weight model based on historical data and expert experience.
[0037]
[0038] Let \(F\) represent the comprehensive defect feature vector, \(C\), \(E\), and \(R\) represent the control flow chaos, exception handling, and resource leakage features respectively, and \(w_1\), \(w_2\), \(w_3\) be the corresponding weight values. This formula describes the process of feature fusion.
[0039] Exemplarily, assume that we have obtained the set of comprehensive defect feature vectors generated in the previous step. For example, the comprehensive defect feature vector of a code snippet is \[52, 27, 21], representing the weighted scores of three dimensions: control flow chaos, exception handling, and resource leakage. Now, we input these vectors into a pre-trained decision tree model for defect prediction. The decision tree model is constructed based on historical annotation data. For example, we used 5000 annotated software projects, among which 3000 projects contain defects and 2000 projects have no defects, and extracted the comprehensive defect feature vectors of each project as training data. There are various algorithms for building decision trees, such as the ID3 algorithm, C5 algorithm, etc. In this embodiment, the C5 algorithm is used for training, and this algorithm selects the best splitting attribute by calculating the information gain ratio of each feature. For example, when constructing a certain node of the decision tree, a feature needs to be selected for splitting. Assume that the information gain ratio of the control flow chaos feature is 25, the information gain ratio of the exception handling feature is 18, and the information gain ratio of the resource leakage feature is 31. Then, the resource leakage feature is selected as the splitting attribute for this node. During the splitting process, the data is divided into different subsets according to the feature values. For example, the samples with a resource leakage score greater than 20 are divided into one subset, and those less than or equal to 20 are divided into another subset. By recursively constructing the decision tree until the stopping condition is met, such as all samples in the node belong to the same category or reach the preset maximum depth of 5. To avoid overfitting, pre-pruning and post-pruning techniques are used. For example, the minimum number of samples for each leaf node is set to 10, or the constructed decision tree is pruned and optimized using a validation set, and the pruning is based on whether the error rate on the validation set decreases after pruning. The finally trained decision tree model can classify new comprehensive defect feature vectors. For example, for the input vector \[52, 27, 21], starting from the root node, according to the resource leakage score of 21 being greater than 20, it enters the right subtree; then according to the control flow chaos score of 52 being greater than 45, it enters the right subtree; and finally reaches a leaf node, and the class label of this leaf node is "defective". Therefore, this code snippet is predicted to have a stealth defect. The model output result is a binary label, "defective" or "non-defective". To evaluate the performance of the model, we can use a test set for testing, such as another 1000 annotated software projects, and calculate indicators such as the accuracy, recall rate, and F1 value of the model. For example, the test results show that the accuracy of the model is 85%, the recall rate is 80%, and the F1 value is 82%, indicating that this decision tree model can effectively identify stealth defects in the code.
[0040] S106. If the output result of the decision tree model indicates the existence of defects, the density-based clustering algorithm is used to group the comprehensive defect feature vectors to determine the types and distribution ranges of the defects. The density clustering algorithm is selected to handle defect features with non-spherical distributions.
[0041] The density-based clustering algorithm is used to group the comprehensive defect feature vectors to obtain the clustering results of the defects. Based on the clustering results, the types and distribution ranges of the defects are determined. If the defect type is chaotic control flow, the control flow optimization algorithm is used to reconstruct the code segment. If the defect type is improper exception handling, the exception handling optimization algorithm is used to reconstruct the code segment. If the defect type is resource leakage, the resource management optimization algorithm is used to reconstruct the code segment. According to the pre-established weight model, the weight value of each feature is determined to reflect its importance for defect detection. The weight model is constructed through historical data and expert experience to ensure the accuracy and reliability of the weight values. The weighted fusion method is used to perform weighted summation of the chaotic control flow, exception handling, and resource leakage features according to the weight values to generate the comprehensive defect feature vectors.
[0042] Exemplarily, assume that the decision tree model determines that a batch of code snippets have defects. We use the comprehensive defect feature vector set of this batch of code snippets as the input. For example, it contains the feature vectors of 100 code snippets determined to be "defective", and each vector contains scores in three dimensions: control flow chaos, exception handling, and resource leakage. We choose the DBSCAN algorithm for clustering analysis. There are two core parameters of this algorithm: the neighborhood radius Eps and the minimum number of samples MinPts. First, a distance matrix is constructed by calculating the Euclidean distance between each vector. For example, the Euclidean distance between the vector \[52, 27, 21] and the vector \[48, 25, 23] is approximately 39. Then, we need to determine the appropriate Eps and MinPts. We assist in selecting Eps by drawing a k-distance graph, where k takes the value of the feature dimension plus 1, that is, 4. Observe the inflection point of the k-distance graph. For example, if the inflection point appears near a distance of 5, then set Eps to 5. The selection of MinPts is usually related to the dataset size and domain knowledge. For example, we set MinPts to 5 according to experience. Next, the DBSCAN algorithm starts to iterate. Starting from any unvisited vector, for example, the vector \[52, 27, 21], check the number of vectors in its Eps neighborhood. It is found that there are 6 vectors in its neighborhood, which is greater than MinPts. Therefore, mark this vector as a core object and create a new cluster. Then, the algorithm recursively checks all unvisited vectors in its neighborhood, for example, the vector \[48, 25, 23]. If this vector is also a core object, then add the vectors in its neighborhood to this cluster. If this vector is a border point (the number of vectors in the neighborhood is less than MinPts but in the neighborhood of a core object), for example, the vector \[45, 20, 18], then temporarily mark it as a member of this cluster but do not expand it further. If this vector is a noise point (neither a core object nor a border point), for example, the vector \[5, 2, 3], then temporarily mark it as noise. When all vectors have been visited, the clustering process ends. Assume that finally 3 clusters are formed, containing 40, 35, and 20 vectors respectively, and another 5 vectors are marked as noise points. We can further analyze the characteristics of each cluster. For example, the vectors in the first cluster generally have a higher score in the control flow chaos dimension, with an average value of 55, while the scores in the exception handling and resource leakage dimensions are lower, with average values of 20 and 15 respectively, indicating that the defect type represented by this cluster is mainly control flow chaos problems. The vectors in the second cluster have relatively balanced scores in the three dimensions, with average values of 35, 30, and 32 respectively, which may represent a mixed type of defect pattern.The vector of the third cluster scores significantly higher on the resource leakage dimension than the other two dimensions, with an average value of 45, while the scores on the control flow chaos and exception handling dimensions are lower, with average values of 25 and 20 respectively, indicating that the defect type represented by this cluster is mainly resource leakage problems. In this way, we can classify the defective code segments and identify the main defect types and distribution ranges, providing more refined guidance for subsequent defect repair.
[0043] S107. Generate defect repair suggestions based on the clustering results, and determine the specific location and scope of influence of the defect by combining code logic and execution path information.
[0044] Use the density clustering algorithm to group the comprehensive defect feature vectors to obtain the clustering results of the defects. Combine code logic and execution path information to determine the specific location and scope of influence of the defects. If the defect type is control flow chaos, use the control flow optimization algorithm to reconstruct the code segment. If the defect type is improper exception handling, use the exception handling optimization algorithm to reconstruct the code segment. If the defect type is resource leakage, use the resource management optimization algorithm to reconstruct the code segment. According to the pre-established weight model, determine the weight value of each feature, reflecting its importance for defect detection. Use the weighted fusion method to perform weighted summation of the control flow chaos, exception handling, and resource leakage features according to the weight values to generate the comprehensive defect feature vector.
[0045] Exemplarily, based on the aforementioned clustering analysis, for the three identified main defect types, targeted repair suggestions can be generated by combining static analysis of the code and dynamic execution information. For clusters with chaotic control flow, a control flow graph analysis tool is used, such as generating an abstract syntax tree and constructing a control flow graph, and calculating the cyclomatic complexity of each code block. Cyclomatic complexity is a metric for measuring code complexity. For example, code blocks with a cyclomatic complexity greater than 10 are usually difficult to understand and maintain. Suppose the cyclomatic complexity of the code snippet "A" in the cluster is 15, which is significantly higher than the average. Further analyzing its control flow graph, it is found that there are multiple nested if-else statements and loop statements in this code block, and there are multiple exits. Combining the dynamic execution path information, the instrumentation technique is used, such as inserting logging code before and after each branch statement and loop statement to record the actual path of program execution. It is observed that this code block will execute a path with 7 branch jumps under specific input conditions, which is extremely error-prone. Therefore, for the code snippet "A", it is recommended to refactor the control logic, such as rewriting the multiple nested if-else statements as switch-case statements or using the strategy pattern, and reducing the number of exits to reduce the cyclomatic complexity to below 10. For clusters with mixed defects, since the defect characteristics in its three dimensions are relatively balanced, specific code needs to be analyzed in combination. For example, the control flow chaos, exception handling, and resource leakage scores of the code snippet "B" are 38, 33, and 35 respectively. Using a static code analysis tool, such as FindBugs or PMD, to scan the code and combining manual review, it is found that there is a try-catch-finally block in this code snippet, where the catch block catches the exception but does not perform any processing, and there is a situation where file resources are not closed in the finally block. Combining the dynamic execution information, it is found that when specific data is input, the program will throw an exception, resulting in resource leakage in the finally block. Therefore, for the code snippet "B", it is recommended to add exception handling logic in the catch block, such as logging or throwing a custom exception, and add code to close the file resources in the finally block. For clusters with resource leakage, a heap memory analysis tool, such as Valgrind or Eclipse Memory Analyzer, is used to detect memory leakage during program operation. Suppose the code snippet "C" in the cluster has a resource leakage dimension score as high as 48. Using the memory analysis tool, it is found that there is a database connection object in this code snippet that is not closed after each use. Combining the static code analysis, it is found that this connection object is created in a loop and not closed after the loop ends, resulting in a new connection object being created each time the loop runs, while the old connection object is not released. Therefore, for the code snippet "C", it is recommended to move the creation of the database connection object outside the loop and add code to close the connection after the loop ends, ensure that the same connection object is used each time the loop runs, and release the object before the program ends.
[0046] S108. Periodically feedback the repair suggestions and location results to the code library, update the test case libraries of the static analysis tool and the dynamic analysis tool, and optimize the coverage rate and accuracy of subsequent defect detection for the newly added defect types and characteristics. This step is executed after multiple detection iterations and is part of continuous improvement.
[0047] Use a static analysis tool to scan the code library to obtain potential defect information in the code. Combine the execution path information of the dynamic analysis tool to determine the specific location and scope of influence of the defect. Generate repair suggestions and location results according to the defect types and characteristics. Feedback the repair suggestions and location results to the code library, and update the test case libraries of the static analysis tool and the dynamic analysis tool. Optimize the coverage rate and accuracy of subsequent defect detection for the newly added defect types and characteristics. Execute this step after multiple detection iterations and is part of continuous improvement. Determine the weight value of each defect feature according to the pre-established weight model, reflecting its importance to defect detection. Use a weighted fusion method to perform weighted summation of the control flow chaos, exception handling, and resource leakage features according to the weight values to generate a comprehensive defect feature vector.
[0048] Exemplarily, after multiple iterations of detection, for example, 5 rounds of defect detection and repair have been carried out, accumulating a large amount of defect data and repair experience. At this time, this step can be executed for continuous improvement. Suppose in the most recent iteration, for the code snippet "D", the static analysis tool reports a potential null pointer dereference risk. The defect localization system locates it to line 35 and gives a repair suggestion: "Add null value check". The repair suggestion and the localization result, that is, the source code of the code snippet "D", the information that there is a null pointer dereference risk at line 35, and the suggestion of "Add null value check", are stored together in the version control system of the code library, such as Git, and note in the commit message "Defect repair: Add null value check for null pointer dereference risk". At the same time, update the rule library of the static analysis tool. For example, in the rule configuration file of FindBugs, add a new rule for null pointer dereference. This rule is based on the pattern of the code snippet "D" and set its priority to "high". For the dynamic analysis tool, based on the characteristics of the code snippet "D", generate a new test case. For example, construct an input data so that when the program executes to line 35, the value of the relevant variable is null, and observe whether the program will throw a null pointer exception. Add this test case to the test case library of the dynamic analysis tool, such as the test case set of Junit. As the number of iterations increases, suppose after the 10th iteration, a new type of defect is discovered: "Concurrent access conflict". This type was not recognized in previous iterations. Its characteristic is that in a multi-threaded environment, multiple threads access and modify the same shared resource simultaneously without proper synchronization control, resulting in data inconsistency. For this newly added type of defect, it is necessary to optimize the coverage and accuracy of subsequent defect detection. First, update the rule library of the static analysis tool. For example, add a detection rule for concurrent access conflict in PMD. This rule can be judged based on whether there are access and modification operations on shared resources in the code and whether there are synchronization control mechanisms such as synchronized and Lock. Second, update the test case library of the dynamic analysis tool. For example, use JMH (Java Microbenchmark Harness) to construct concurrent test cases, simulate the scenario where multiple threads access and modify the same shared resource simultaneously, and detect whether the program will have data inconsistency. In this way, the coverage and accuracy of the defect detection system can be continuously improved to achieve continuous improvement.
[0049] The present invention provides a software code defect detection system based on program code feature fusion, mainly including: A static analysis module for performing syntax and structure checks on the code using a static analysis tool and extracting defect features related to control flow chaos, exception handling, and resource leakage; A dynamic analysis module, which is used to run preset test cases through a dynamic analysis tool, capture abnormal behaviors in the execution path and runtime state, and obtain defect features not covered by static analysis; A feature preprocessing module, which is used to preprocess the features extracted by static analysis and dynamic analysis, remove redundant features, retain key features related to code logic and execution path, generate a preprocessed defect feature set, and the preprocessed defect feature set is represented in vector form, with each element corresponding to a specific defect feature; A feature fusion module, which is used to design a feature fusion algorithm, and based on a pre-established weight model, perform weighted fusion on the control flow chaos, exception handling, and resource leakage features in the preprocessed defect feature set to generate a comprehensive defect feature vector. The weight model is constructed based on historical data and expert experience and reflects the importance of different features for defect detection; A defect detection module, which is used to build a unified detection framework, input the comprehensive defect feature vector into a pre-established decision tree model, classify the features, and determine whether there are stealth defects in the code. The decision tree model is trained with labeled historical defect data and is used to identify new defect patterns; A defect clustering module, which is used to, if the output result of the decision tree model is that there are defects, use a density-based clustering algorithm to group the comprehensive defect feature vectors, determine the types and distribution ranges of the defects, and the density clustering algorithm is selected to handle defect features with non-spherical distributions; A defect repair suggestion generation module, which is used to generate defect repair suggestions according to the clustering results, and combine code logic and execution path information to determine the specific location and influence range of the defects; A continuous optimization module, which is used to periodically feedback the repair suggestions and positioning results to the code library, update the test case libraries of the static analysis tool and the dynamic analysis tool, and optimize the coverage rate and accuracy of subsequent defect detection for newly added defect types and features. This step is executed after multiple detection iterations and is part of continuous improvement.
[0050] The present invention discloses a method and system for detecting stealth defects in code. Defect features in the code are extracted by combining static analysis and dynamic analysis. After preprocessing and weighted fusion of the features, a comprehensive defect feature vector is generated. A decision tree model is used to classify the features to determine whether there are stealth defects. If there are defects, a density clustering algorithm is used to determine the defect types and distribution ranges, and repair suggestions are generated. The present invention also continuously optimizes the detection coverage rate and accuracy by periodically feedbacking and updating the test case library. This method can effectively identify and locate stealth defects in the code, improve software quality and reliability, and provide timely and accurate defect repair guidance for developers.
[0051] It should also be understood that, in the embodiments herein, the term "and / or" is merely a description of the relationship between associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. Additionally, in this text, the character " / " generally indicates that the associated objects before and after are in an "or" relationship.
[0052] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this text.
[0053] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0054] In the several embodiments provided herein, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Additionally, the displayed or discussed couplings or direct couplings or communication connections to each other can be indirect couplings or communication connections through some interfaces, devices, or units, and can also be in electrical, mechanical, or other forms of connection.
[0055] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiments herein.
[0056] In addition, the functional units in each embodiment herein can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0057] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution herein, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments herein. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0058] The above content is only an example and explanation of the structure of the present invention. Those skilled in the art of this technology can make various modifications or supplements to the specific embodiments described, or use similar methods for substitution, as long as they do not deviate from the structure of the invention or exceed the scope defined by this claim book, they should all fall within the protection scope of the present invention.< / view>
Claims
1. A software code defect detection method based on program code feature fusion, characterized in that The method includes: Using a static analysis tool to perform syntax and structure checks on the code, and extracting defect features related to control flow chaos, exception handling, and resource leakage; Running preset test cases through a dynamic analysis tool to capture abnormal behaviors in the execution path and runtime state, and obtaining defect features not covered by static analysis; Preprocessing the features extracted by static analysis and dynamic analysis, removing redundant features, retaining key features related to code logic and execution path, generating a preprocessed defect feature set, and the preprocessed defect feature set is represented in vector form, with each element corresponding to a specific defect feature; Designing a feature fusion algorithm, based on a pre-established weight model, to perform weighted fusion on the control flow chaos, exception handling, and resource leakage features in the preprocessed defect feature set, generating a comprehensive defect feature vector, and the weight model is constructed based on historical data and expert experience, reflecting the importance of different features for defect detection; Constructing a unified detection framework, inputting the comprehensive defect feature vector into a pre-established decision tree model, classifying the features, and determining whether there are stealth defects in the code. The decision tree model is trained with labeled historical defect data for identifying new defect patterns; If the output result of the decision tree model is that there are defects, then using a density-based clustering algorithm to group the comprehensive defect feature vector, determining the type and distribution range of the defects, and the density clustering algorithm is selected to handle defect features with non-spherical distributions; Generating defect repair suggestions according to the clustering results, and combining code logic and execution path information to determine the specific location and influence range of the defects; Periodically feedbacking the repair suggestions and location results to the code library, updating the test case libraries of the static analysis tool and the dynamic analysis tool, and optimizing the coverage and accuracy of subsequent defect detection for newly added defect types and features. This step is executed after multiple detection iterations as part of continuous improvement.
2. The method according to claim 1, wherein The using a static analysis tool to perform syntax and structure checks on the code, and extracting defect features related to control flow chaos, exception handling, and resource leakage includes: Using a static analysis tool to perform syntax checks on the code to obtain syntax error information of the code; Using a static analysis tool to perform structure checks on the code to obtain structure error information of the code; Extracting control flow chaos features from the syntax error information and structure error information to determine the code segments with control flow chaos; Extracting exception handling features from the syntax error information and structure error information to determine the code segments with improper exception handling; Extracting resource leakage features from the syntax error information and structure error information to determine the code segments with resource leakage.
3. The method according to claim 2, wherein The running preset test cases through a dynamic analysis tool to capture abnormal behaviors in the execution path and runtime state, and obtaining defect features not covered by static analysis includes: Extracting control flow chaos features according to the execution path captured by the dynamic analysis tool to determine the code segments with control flow chaos; Extracting exception handling features from the runtime state to determine the code segments with improper exception handling; Extracting resource leakage features through abnormal behavior analysis to determine the code segments with resource leakage; If the control flow chaos feature exists, a control flow optimization algorithm is used to reconstruct the code segment; If the exception handling feature exists, an exception handling optimization algorithm is used to reconstruct the code segment; If the resource leakage feature exists, a resource management optimization algorithm is used to reconstruct the code segment; Based on the reconstructed code segment, new test cases are generated, and the dynamic analysis tool is run again to obtain the optimized defect features.
4. The method according to claim 3, characterized in that Preprocess the features extracted by static analysis and dynamic analysis, remove redundant features, retain the key features related to code logic and execution paths, and generate a preprocessed defect feature set; The preprocessed defect feature set is represented in vector form, and each element corresponds to a specific defect feature, including: Use a feature selection algorithm to perform dimensionality reduction on the preprocessed defect feature set, remove redundant features, and retain key features; Based on the dimensionality-reduced feature set, construct a defect feature vector space model and map each defect feature to a specific dimension in the vector space; Classify the defect feature vectors through a clustering algorithm to determine the defect categories and their distribution patterns.
5. The method according to claim 1, wherein Design a feature fusion algorithm. Based on a pre-established weight model, perform weighted fusion on the control flow chaos, exception handling, and resource leakage features in the preprocessed defect feature set to generate a comprehensive defect feature vector; The weight model is constructed based on historical data and expert experience and reflects the importance of different features for defect detection, including: Use a feature fusion algorithm to perform weighted fusion on the control flow chaos, exception handling, and resource leakage features to generate a comprehensive defect feature vector; Based on the pre-established weight model, determine the weight value of each feature, reflecting its importance for defect detection; Construct a weight model through historical data and expert experience to ensure the accuracy and reliability of the weight values; Use a weighted fusion method to perform weighted summation on the control flow chaos, exception handling, and resource leakage features according to the weight values to generate a comprehensive defect feature vector; Based on the comprehensive defect feature vector, judge the defect type and its severity.
6. The method according to claim 1, characterized in that, If the output result of the decision tree model is that there are defects, use a density-based clustering algorithm to group the comprehensive defect feature vectors to determine the type and distribution range of the defects; The density clustering algorithm is selected to handle defect features with non-spherical distributions, including: Use a density clustering algorithm to group the comprehensive defect feature vectors to obtain the clustering results of the defects; Based on the clustering results, determine the type and distribution range of the defects.
7. The method according to claim 6, wherein Generate defect repair suggestions based on the clustering results, and combine code logic and execution path information to determine the specific location and impact range of the defects, including: Use a density clustering algorithm to group the comprehensive defect feature vectors to obtain the clustering results of the defects; Combine code logic and execution path information to determine the specific location and impact range of the defects.
8. The method according to claim 1, characterized in that, Periodically feedback the repair suggestions and location results to the code library, update the test case libraries of the static analysis tool and the dynamic analysis tool, and optimize the coverage and accuracy of subsequent defect detection for newly added defect types and features; This step is executed after multiple detection iterations and is part of continuous improvement, including: Use a static analysis tool to scan the code library and obtain potential defect information in the code; Combine the execution path information of the dynamic analysis tool to determine the specific location and impact scope of the defect; Generate repair suggestions and localization results according to the defect type and characteristics; Feed back the repair suggestions and localization results to the code library, and update the test case libraries of the static analysis tool and the dynamic analysis tool; Optimize the coverage rate and accuracy of subsequent defect detection for newly added defect types and characteristics; Execute this step after multiple detection iterations as part of continuous improvement.
9. A software code defect detection system based on program code feature fusion, characterized in that The system includes: A static analysis module for using a static analysis tool to perform syntax and structure checks on the code, and extracting defect characteristics related to control flow chaos, exception handling, and resource leakage; A dynamic analysis module for running preset test cases through a dynamic analysis tool, capturing abnormal behaviors in the execution path and runtime state, and obtaining defect characteristics not covered by static analysis; A feature preprocessing module for preprocessing the characteristics extracted by static analysis and dynamic analysis, removing redundant characteristics, retaining key characteristics related to code logic and execution path, generating a preprocessed defect feature set, and representing the preprocessed defect feature set in vector form, with each element corresponding to a specific defect characteristic; A feature fusion module for designing a feature fusion algorithm, and based on a pre-established weight model, performing weighted fusion on the control flow chaos, exception handling, and resource leakage characteristics in the preprocessed defect feature set to generate a comprehensive defect feature vector. The weight model is constructed based on historical data and expert experience and reflects the importance of different characteristics for defect detection; A defect detection module for constructing a unified detection framework, inputting the comprehensive defect feature vector into a pre-established decision tree model, classifying the features, and determining whether there are hidden defects in the code. The decision tree model is trained with labeled historical defect data for identifying new defect patterns; A defect clustering module for, if the output result of the decision tree model is that there are defects, using a density-based clustering algorithm to group the comprehensive defect feature vectors, determining the type and distribution scope of the defects. The density clustering algorithm is selected to handle defect characteristics with non-spherical distributions; A repair suggestion generation module for generating defect repair suggestions according to the clustering results, and combining code logic and execution path information to determine the specific location and impact scope of the defects; A continuous optimization module for periodically feeding back the repair suggestions and localization results to the code library, updating the test case libraries of the static analysis tool and the dynamic analysis tool, and optimizing the coverage rate and accuracy of subsequent defect detection for newly added defect types and characteristics. This step is executed after multiple detection iterations as part of continuous improvement.
Citation Information
Cited By
Software defect collaborative detection method and system based on multi-agent dynamic adaptation
CN120560989A
Program performance analysis system, program performance analysis method, and computer program product
CN120821458A
Program performance analysis system, program performance analysis method, and computer program product
CN120821458B
Abnormity detection and repair method and system for big data task and medium
CN121743130A