Software defect positioning method based on context program reduction technology
Through stack trace and dependency analysis, combined with regular analysis and defect probability ranking, the defect positioning problem caused by noise data interference in the existing technology is solved, efficient defect positioning at the statement level is achieved, and software debugging efficiency is improved.
Patent Information
- Application Number
- CN202510774210.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-11
AI Technical Summary
Existing IR-based defect positioning technology is susceptible to noise data interference when processing defect reports, resulting in an expanded retrieval range and the inability to achieve statement-level defect positioning, especially inefficient in multi-threaded program activities.
The software defect positioning method based on context program reduction technology is adopted to generate stack traces through stack frames, extract suspicious program entities, use regular analysis and dependency analysis to calculate defect probability ranking, and split the code entities by line number, adjust detection priority and coverage to achieve accurate statement-level defect positioning.
有效减少了检索空间,提高了缺陷定位的精度和效率,能够在复杂软件系统中快速识别并定位缺陷代码,降低了开发成本。
Smart Images

Figure CN120295896A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of defect detection, and particularly to a software defect localization method based on context program reduction technology. Background Technique
[0002] Software defect localization is one of the most time-consuming and expensive activities in software debugging. As software systems become increasingly complex, these defects can significantly affect software performance and development, causing huge economic losses. Existing IR (information retrieval)-based defect localization techniques prioritize suspicious defect files based on the correlation between source code and defect reports, and provide defect repair guidelines for developers.
[0003] This technique generates defect reports based on source files and uses the structured information in the defect reports for information retrieval, such as matching with source files having or containing the same name, thereby retrieving suspicious program entities. However, because natural language processing technology is used to process defect reports, it will be interfered by noise data unrelated to defects in the defect reports. Especially when multiple program activities from different threads are provided in the defect report, a large amount of suspicious defect information expands the retrieval scope, resulting in an explosion of the retrieval space and making it impossible to achieve statement-level defect localization. Therefore, it is necessary to design a software defect localization method based on context program reduction technology for statement-level defect localization. Summary of the Invention
[0004] The purpose of the present invention is to provide a software defect localization method based on context program reduction technology to solve the problems raised in the above background technique.
[0005] To solve the above technical problems, the present invention provides the following technical solution: A software defect localization method based on context program reduction technology, which uses a defect localization system to work. The system includes a suspicious program extraction module, a code entity processing module, and a defect localization module. The suspicious program extraction module is used to generate a stack trace from the stack frames of program activities and extract suspicious program entities therefrom. The code entity processing module is used to parse the suspicious program entities and rank the defect probabilities of potentially defective code entities. The defect localization module is used to locate the defective code entities.
[0006] According to the above technical solution, the suspicious program extraction module includes a stack frame generation module and a stack trace analysis module, and the stack frame generation module is electrically connected to the stack trace analysis module; the stack frame generation module is used to record the stack frames of the program when a system failure occurs, and the stack trace analysis module is used to extract suspicious program entities from the stack frames; The code entity processing module includes a regular expression parsing module, a dependency analysis module, and a defect probability ranking module. The regular expression parsing module is electrically connected to the stack trace analysis module and the dependency analysis module, and the dependency analysis module is electrically connected to the defect probability ranking module. The regular expression parsing module is used to parse suspicious program entities. The dependency analysis module is used to analyze the dependencies of code entities in suspicious files. The defect probability ranking module is used to rank the defect probabilities of potentially defective code entities. The defect location module includes a defect detection module, a result output module, a priority adjustment module, and an entity splitting module. The defect detection module is electrically connected to the defect probability ranking module. The priority adjustment module is electrically connected to the result output module and the defect probability ranking module. The result output module is electrically connected to the defect detection module. The defect detection module is used to detect defective code. The result output module is used to output the detected defective code and the corresponding location information. The priority adjustment module is used to adjust the defect probability ranking during multiple detections. The entity splitting module is used to split code entities into multiple paragraphs.
[0007] According to the above technical solution, the following steps are included: S1. Record the program activities at the time of system failure, use the API provided by the system to obtain the stack frame information of the current execution. The stack frame records the order and location of function calls. Create a stack trace using the stack frame and extract suspicious program entities from it. S2. Use regular expressions to parse the extracted suspicious program entities, divide the code entities in the suspicious program entities into function names, file names, and line numbers, and analyze the dependencies of the code entities in the suspicious files to infer potentially defective code entities. S3. Rank the defect probabilities of relevant code entities, adjust the coverage rate of identifying defective statements to cover code entities with high defect probabilities, and minimize the retrieval space. S4. Detect defective code entities, output the detected defective code entities and their corresponding location information, count the number of defect detections for each detected defective code entity, and adjust the defect probabilities of each code entity. S5. Split the code entities according to line numbers, name the line number ranges where defects often occur as new code entities, and make the new code entities participate in the defect probability ranking.
[0008] According to the above technical solution, in S1, the specific method for creating a stack trace is: S1-1. Parse and format the stack frame data. According to the information in the stack frame, extract the program entity name and layer number of each layer, and connect these information into a string in the call order. S1-2. When defining a suspicious program entity, the definition conditions of the suspicious program entity include the program where the function with an abnormal call count in the code is located, the program where the called function has an error prompt in the error log, the program with high-complexity functions having loop nesting and multi-condition branches, the program with the recently modified code lines, and the program depending on a known defective module. The suspicious degree value of each program entity is obtained according to the number of definition conditions met by the current program entity. , where is the number of program entities in the stack trace, and the program entity with the number of definition conditions met greater than 0 is defined as a suspicious program entity.
[0009] S2-1. By the extracted file names, find the corresponding code entities. The dependency relationships include function mutual calls, functions using external function variables, data structure sharing between functions, and inheritance of function attributes from the parent class. Use the AST parser to perform static code analysis on the code entities in the non-running state, and use the debugger to perform dynamic trace analysis when the code entities are running. S2-2. Obtain the number of dependency relationships of each code entity in the program entity. , where is the number of code entities in the current suspicious program entity. Select the code entities with the number of dependency relationships greater than 0 and define them as potentially defective code entities.
[0010] According to the above technical solution, in S3, the specific method is as follows: S3-1. Perform a ranking of defect probabilities: Calculate the defect probabilities of all potentially defective code entities in each suspicious program entity. The higher the suspicious degree value of the program entity where the code entity is located and the higher the number of dependency relationships of the code entity, the greater the defect probability and the higher the ranking of the defect probability, and the higher the detection priority of this code entity. The specific calculation formula is: Defect probability , where is the weight coefficient of the suspicious degree value of the program entity, is the weight coefficient of the number of dependency relationships of the code entity, is the serial number of the program entity where the current code entity is located, is the serial number of the current code entity in the program entity; S3-2. Adjust the coverage rate of identifying defective statements: Before retrieval, adjust the coverage rate according to the average defect probability of all code entities in the program entity. The higher the average defect probability, the higher the coverage rate. The specific formula is: Coverage rate , where is the initial coverage rate of identifying defective statements, is the average defect probability influence coefficient, and .
[0011] According to the above technical solution, in S4, the specific method for adjusting the defect probability and detection priority is as follows: Each time the system fails and defect location is performed, the defect probability of each code entity is adjusted. Let the number of detections that have been performed be times, and the defect probability of the current code entity at the th detection is , the number of defect detections of the current code entity at the th detection is , the number of defect detections of the current code entity at the th detection is . If the number of detections of a code entity with a low defect probability is large, the defect probability at the next detection will be increased. If the number of detections of the code entity shows an upward trend, the defect probability at the next detection will also be increased. Then the defect probability of the current code entity at the next detection, where is the proportionality coefficient of the number of detections to the original defect probability, and is the weight coefficient of the change trend of the number of detections. The higher the defect probability, the higher the detection priority. It can be detected earlier in the next detection, and it is more likely to quickly detect defects. Also, code entities whose detection results do not match the theoretical defect probability calculation results can be adjusted to the appropriate detection priority faster, and code entities that gradually deteriorate the situation of increasing bugs due to data accumulation during system operation can be adjusted to the priority detection position faster.
[0012] According to the above technical solution, in S5, the specific method for splitting by line number is as follows: The code entity is split into several paragraphs, denoted as , where is the paragraph number, and the line number where the paragraph is located is denoted as . Record the line numbers where defects occur in the code entity, determine which line number interval they fall into, and count the number of defects accumulated in each line number interval . When the number of defects accumulated in a certain paragraph exceeds the cumulative threshold , the code at the line number where this paragraph is located is redefined as a new code entity. If multiple adjacent paragraphs all exceed the cumulative threshold , they are merged into one code entity. The new code entity is calculated for the defect probability separately, and when calculating the defect probability, the number of dependency relationships is converted according to the ratio of the line number interval of this paragraph to the line number interval of the entire code entity.
[0013] Compared with the prior art, the beneficial effects achieved by the present invention are as follows: In the present invention, first, the stack frames of program activities when a system failure occurs are recorded to generate a stack trace, suspicious program entities are extracted from the stack trace, the suspicious program entities are parsed using regular expressions, and the dependency relationships of code entities in the suspicious files are analyzed to trace potential defective code entities. The highly structured information in the stack trace can help identify code entities directly and indirectly related to the defect; Secondly, the defect probabilities of relevant code entities are ranked. By adjusting the coverage rate of identifying defect statements, it is made to cover as many code entities with high defect probabilities as possible, and as many defective codes as possible are detected on the premise of minimizing the retrieval space, realizing statement-level defect localization. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, and do not constitute a limitation to the present invention. In the drawings: Figure 1 is a schematic diagram of the overall module structure of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0015] Hereinafter, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative work shall fall within the protection scope of the present invention.
[0016] Please refer to Figure 1 , the present invention provides a technical solution: A software defect localization method based on context program reduction technology, which uses a defect localization system to work. The system includes a suspicious program extraction module, a code entity processing module, and a defect localization module. The suspicious program extraction module is used to generate a stack trace from the stack frames of program activities and extract suspicious program entities therefrom. The code entity processing module is used to parse the suspicious program entities and rank the defect probabilities of potential defective code entities. The defect localization module is used to locate the defective code entities; The suspicious program extraction module includes a stack frame generation module and a stack trace analysis module, and the stack frame generation module is electrically connected to the stack trace analysis module; the stack frame generation module is used to record the stack frames of the program when a system failure occurs, and the stack trace analysis module is used to extract suspicious program entities from the stack frames; The code entity processing module includes a regular parsing module, a dependency analysis module, and a defect probability ranking module. The regular parsing module is electrically connected to the stack trace analysis module and the dependency analysis module, and the dependency analysis module is electrically connected to the defect probability ranking module. The regular parsing module is used to parse suspicious program entities. The dependency analysis module is used to analyze the dependencies of code entities in suspicious files. The defect probability ranking module is used to rank the defect probabilities of potentially defective code entities. The defect location module includes a defect detection module, a result output module, a priority adjustment module, and an entity splitting module. The defect detection module is electrically connected to the defect probability ranking module. The priority adjustment module is electrically connected to the result output module and the defect probability ranking module. The result output module is electrically connected to the defect detection module. The defect detection module is used to detect defective code. The result output module is used to output the detected defective code and the corresponding location information. The priority adjustment module is used to adjust the defect probability ranking during multiple detections. The entity splitting module is used to split code entities into multiple paragraphs. It includes the following steps: S1. Record the program activities at the time of system failure, use the API provided by the system to obtain the stack frame information of the current execution. The stack frame records the order and location of function calls. Create a stack trace using the stack frame and extract suspicious program entities from it. S2. Use regular expressions to parse the extracted suspicious program entities, divide the code entities in the suspicious program entities into function names, file names, and line numbers, and analyze the dependencies of code entities in the suspicious files to infer potentially defective code entities. S3. Rank the defect probabilities of relevant code entities, adjust the coverage rate of identifying defective statements to cover code entities with high defect probabilities, and minimize the retrieval space. S4. Detect defective code entities, output the detected defective code entities and their corresponding location information, count the number of defect detections for each detected defective code entity, and adjust the defect probabilities of each code entity. S5. Split the code entities by line number, name the line number intervals where defects often occur as new code entities, and make the new code entities participate in the defect probability ranking. In S1, the specific method for creating a stack trace is: S1-1. Parse and format the stack frame data. According to the information in the stack frame, extract the program entity name and layer number of each layer, and connect these information into a string in the call order. S1-2. When defining a suspicious program entity, the definition conditions of the suspicious program entity include the program where the function with an abnormal number of calls in the code is located, the program where the called function has an error prompt in the error log, the program with high-complexity functions having loop nesting and multiple conditional branches, the program with the most recently modified code lines, and the program that depends on a known defective module. The suspicious degree value of each program entity is obtained according to the number of definition conditions met by the current program entity. , where is the number of program entities in the stack trace. The program entities with the number of definition conditions met greater than 0 are defined as suspicious program entities. In S2, the specific method for analyzing the dependency relationship of code entities in the suspicious file is as follows: S2-1. By the extracted file name, find the corresponding code entity. The dependency relationships include function mutual calls, functions using external function variables, data structure sharing between functions, and inheritance of function attributes from the parent class. Use the AST parser to perform static code analysis when the code entity is not in the running state, and use the debugger to perform dynamic tracking analysis when the code entity is running. S2-2. Obtain the number of dependency relationships of each code entity in the program entity, where is the number of code entities in the current suspicious program entity. Select the code entities with the number of dependency relationships greater than 0 and define them as potentially defective code entities. In S3, the specific method is as follows: S3-1. Rank the defect probabilities: Calculate the defect probabilities of all potentially defective code entities in each suspicious program entity. The higher the suspicious degree value of the program entity where the code entity is located and the higher the number of dependency relationships of the code entity, the greater the defect probability and the higher the ranking of the defect probability, and the higher the detection priority of this code entity. The specific calculation formula is: Defect probability , where is the weight coefficient of the suspicious degree value of the program entity, is the weight coefficient of the number of dependency relationships of the code entity, is the serial number of the program entity where the current code entity is located, is the serial number of the current code entity in the program entity; S3-2. Adjust the coverage rate of the identified defective statements: Before retrieval, adjust the coverage rate according to the average defect probability of all code entities in the program entity. The higher the average defect probability, the higher the coverage rate. The specific formula is: Coverage rate , where is the initial coverage rate of the identified defective statements, is the average defect probability influence coefficient, and ; In S4, the specific method for adjusting the defect probability and detection priority is as follows: Each time the system fails and defect location is performed, the defect probability of each code entity is adjusted. Suppose that tests have been performed, and the defect probability of the current code entity in the th test is . The number of detected defects of the current code entity in the th test is , and the number of detected defects of the current code entity in the th test is . If a code entity with a low defect probability has a large number of detected defects, the defect probability in the next test will be increased. If the number of detected defects of a code entity shows an upward trend, the defect probability in the next test will also be increased. Then the defect probability of the current code entity in the next test, where is the proportionality coefficient of the number of detected defects to the original defect probability, and is the weight coefficient of the change trend of the number of detected defects. The higher the defect probability, the higher the detection priority, and it can be ranked in the front for detection in the next test, which can more likely quickly detect defects, and enable code entities whose detection results do not conform to the theoretical defect probability calculation results to be adjusted to an appropriate detection priority faster, and can also adjust code entities that gradually deteriorate the situation of increasing bugs due to data accumulation during system operation to the priority detection position faster; In S5, the specific method for splitting by line number is as follows: The code entity is split into several paragraphs, denoted as , where is the paragraph number, and the line number where the paragraph is located is denoted as . Record the line numbers where defects occur in the code entity, determine which line number interval they fall into, and count the number of defects accumulated in each line number interval . When the number of defects accumulated in a certain paragraph exceeds the cumulative threshold , the code with the line number of this paragraph is redefined as a new code entity. If multiple adjacent paragraphs with line numbers all exceed the cumulative threshold , they are merged into one code entity. The new code entity is calculated for the defect probability separately, and when calculating the defect probability, the number of dependency relationships is converted according to the ratio of the line number interval of this paragraph to the line number interval of the entire code entity.
[0017] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device.
[0018] Finally, it should be noted that the above are only preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A software defect localization method based on context program reduction technology, characterized in that: This method works by using a defect localization system, which includes a suspicious program extraction module, a code entity processing module, and a defect localization module. The suspicious program extraction module is used to generate a stack trace from the stack frames of program activities and extract suspicious program entities therefrom. The code entity processing module is used to parse the suspicious program entities and rank the defect probabilities of potentially defective code entities. The defect localization module is used to locate the defective code entities.
2. The software defect localization method based on context program reduction technology according to claim 1, wherein: The suspicious program extraction module includes a stack frame generation module and a stack trace analysis module, which are electrically connected. The stack frame generation module is used to record the stack frames of the program during system failures, and the stack trace analysis module is used to extract suspicious program entities from the stack frames. The code entity processing module includes a regular expression parsing module, a dependency analysis module, and a defect probability ranking module. The regular expression parsing module is electrically connected to the stack trace analysis module and the dependency analysis module, and the dependency analysis module is electrically connected to the defect probability ranking module. The regular expression parsing module is used to parse the suspicious program entities, the dependency analysis module is used to analyze the dependencies of the code entities in the suspicious files, and the defect probability ranking module is used to rank the defect probabilities of potentially defective code entities. The defect localization module includes a defect detection module, a result output module, a priority adjustment module, and an entity splitting module. The defect detection module is electrically connected to the defect probability ranking module, the priority adjustment module is electrically connected to the result output module and the defect probability ranking module, and the result output module is electrically connected to the defect detection module. The defect detection module is used to detect defective code, the result output module is used to output the detected defective code and the corresponding location information, the priority adjustment module is used to adjust the defect probability ranking during multiple detections, and the entity splitting module is used to split the code entities into multiple paragraphs.
3. A software defect localization method based on context program reduction technology according to claim 2, characterized in that: It includes the following steps: S1. Record the program activities during system failures, use the API provided by the system to obtain the stack frame information of the current execution. The stack frames record the order and location of function calls. Create a stack trace using the stack frames and extract suspicious program entities therefrom. S2. Use regular expressions to parse the extracted suspicious program entities, divide the code entities in the suspicious program entities into function names, file names, and line numbers, and analyze the dependencies of the code entities in the suspicious files to infer potentially defective code entities. S3. Rank the defect probabilities of the relevant code entities, adjust the coverage rate of the identified defect statements to cover the code entities with high defect probabilities, and minimize the search space. S4. Detect the defective code entities, output the detected defective code entities and their corresponding location information, count the number of defect detections for each detected defective code entity, and adjust the defect probabilities of each code entity. S5. Split the code entities according to the line numbers, name the line number ranges where defects often occur as new code entities, and make the new code entities participate in the defect probability ranking.
4. A software defect localization method based on context program reduction technology according to claim 3, characterized in that: In S1, the specific method for creating a stack trace is as follows: S1-1. Parse and format the stack frame data. According to the information of the stack frame, extract the program entity name and layer number of each layer, and concatenate these information into a string in the call order. S1-2. When defining a suspicious program entity, the definition conditions of the suspicious program entity include the program where the function with an abnormal call count in the code is located, the program where the called function has an error prompt in the error log, the program with a highly complex function having loop nesting and multiple conditional branches, the program with the most recently modified code line, and the program depending on a known defective module. The suspicious degree value of each program entity is obtained according to the number of definition conditions met by the current program entity. , where is the number of program entities in the stack trace, and the program entity with the number of definition conditions met greater than 0 is defined as a suspicious program entity.
5. The software defect localization method based on context program reduction technology according to claim 4, characterized in that: In S2, the specific method for analyzing the dependency relationship of code entities in the suspicious file is as follows: S2-1. Through the extracted file name, find the corresponding code entity. The dependency relationships include function mutual calls, functions using external function variables, data structure sharing between functions, and inheritance of function attribute parent classes. Use the AST parser to perform static code analysis when the code entity is not running, and use the debugger to perform dynamic tracking analysis when the code entity is running. S2-2. Obtain the number of dependency relationships of each code entity in the program entity , where is the number of code entities in the current suspicious program entity. Select the code entities with the number of dependency relationships greater than 0 and define them as potentially defective code entities.
6. The software defect localization method based on context program reduction technology according to claim 5, characterized in that: In S3, the specific method is as follows: S3-1. Calculate the defect probability for all potentially defective code entities in each suspicious program entity. The higher the suspiciousness value of the program entity where the code entity is located, the greater the defect probability. The specific calculation formula is: Defect probability , where is the weight coefficient of the suspiciousness value of the program entity, is the weight coefficient of the number of dependency relationships of the code entity, is the serial number of the program entity where the current code entity is located, is the serial number of the current code entity in the program entity; S3-2. Adjust the coverage rate according to the average defect probability of all code entities in the program entity before retrieval. The higher the average defect probability, the higher the coverage rate. The specific formula is: Coverage rate , where is the coverage rate of the initially identified defective statements, is the average defect probability influence coefficient, and 7. A software defect localization method based on context program reduction technology according to claim 6, characterized in that: In S4, the specific method for adjusting the defect probability and detection priority is as follows: Each time the system fails and defect localization is performed, the defect probability of each code entity is adjusted. Let the current number of detections have been carried out, and the defect probability of the current code entity in the th detection is . The number of detected defects of the current code entity at the th detection is . The number of detected defects of the current code entity at the th detection is . Then the defect probability of the current code entity at the next detection is , where is the proportionality coefficient of the detected quantity to the original defect probability, and is the weight coefficient of the change trend of the detected quantity.
8. A software defect localization method based on context program reduction technology according to claim 7, characterized in that: In S5, the specific method of splitting according to line numbers is as follows: The code entity is split into several paragraphs, denoted as , where is the paragraph number, and the line number where the paragraph is located is denoted as . Record the line numbers with defects in the code entity, determine which line number interval they fall into, and count the number of defects accumulated in each line number interval . When the number of defects accumulated in a certain paragraph exceeds the cumulative threshold , redefine the code of the line number where this paragraph is located as a new code entity, and calculate the defect probability of the new code entity separately.
Citation Information
Patent Citations
Software defect positioning method based on relative redundant test set reduction
CN101866316A
Defect detection method for code covering device based on debugging information support
CN115422082A
Software defect positioning method based on feature crossing and structural semantic information matching
CN117851216A
Performance bug detection and code recommendation
US20250021460A1