An Automatic Classification Method for Interpreter Defects in CPython and PyPy

By using TF-IDF and cosine similarity calculation method in Python interpreter, keywords are extracted from issue reports and generated tags are solved, the problem of automatic classification of interpreter defects is improved, defect detection and repair efficiency is improved, and the healthy development of Python language is promoted.

CN112799960BActive Publication Date: 2025-07-08NANJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110213194.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-02-25
Publication Date
2025-07-08
Estimated Expiration
2041-02-25

AI Technical Summary

Technical Problem

The existing technology has failed to effectively automatically classify defects in Python interpreters, resulting in inefficient defect detection and repair, affecting the healthy development of the Python language ecosystem.

Method used

Using a method based on TF-IDF algorithm and cosine similarity calculation, keywords are extracted from the issue report, combined with quantitative analysis and manual classification results, defect record labels are generated, and defect categories are recommended through code and text similarity.

Benefits of technology

It realizes efficient automatic classification of Python interpreter defects, improves defect detection and repair efficiency, reduces the harm of difficult-to-identify defects to the interpreter, and promotes the healthy development of the Python language ecosystem.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112799960B_ABST
    Figure CN112799960B_ABST
Patent Text Reader

Abstract

The present invention discloses an automatic classification method for defects in two interpreters, CPython and PyPy, comprising the following steps: 1) extracting keywords from the title of the issue report based on syntactic components; 2) supplementing tags by combining quantitative analysis and the experimental results of manual classification to make up for the deficiencies in the extraction of syntactic components and improve the retrieval efficiency of defect records; 3) recommending the issue to the corresponding category based on code and text similarity and in combination with the vector space model VSM. The difficulty of the method of the present invention lies in automatically generating defect record tags by extracting keywords from the issue report. By using hybrid technology, the present invention fills the gap in the direction of automatic classification of interpreter defects and provides certain assistance to developers and maintainers of Python interpreters, developers of Python applications, and researchers in related fields.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention discloses an automatic classification method for defects in two interpreters, CPython and PyPy, belonging to the field of data analysis. Using a hybrid technique, the present invention fills the gap in the direction of automatically classifying interpreter defects, and provides certain help for developers and maintainers of Python interpreters, developers of Python applications, and researchers in related fields. Background Art

[0002] Python is an interpreted programming language. Applications written in the Python language do not need to be compiled into executable programs, but are directly interpreted and executed by the Python interpreter. As a low-level basic software, there are inevitably many software defects in the Python interpreter. Different from defects in ordinary applications, such defects existing in the Python interpreter have a wider scope of influence and a higher degree of harm.

[0003] Thus, by quantitatively analyzing 23,823 defect reports, 17,016 revisions, and 2,015 test cases in two mainstream Python interpreters, CPython and PyPy, we found that the distribution of defects in the Python interpreter among components and source files is extremely uneven, and the vast majority of defects are distributed in a few components and source files. At the same time, by manually analyzing the root causes of 400 defects in CPython and 110 defects in PyPy, a defect category distribution table and a label set for label classification are obtained.

[0004] In addition, we learned about the traditional keyword extraction algorithm - TF-IDF and the cosine similarity calculation formula - Sim(D). The TF-IDF algorithm is often used to evaluate the importance of a word to a certain document in a document set. The more important a word is to a document, the more likely it is to be a keyword of the document. TF counts the frequency of a word appearing in a document. IDF counts the number of documents in the document set in which a word appears. After representing a document as a word vector, the Sim(D) formula calculates the similarity Sim(D) between two documents using the cosine similarity between the two vectors.

[0005] Since there has been no previous research work on automatically classifying interpreter defects in China, we will, on the basis of the TF-IDF algorithm and the Sim(D) formula, combine the quantitative analysis results of a large number of defect reports in the two interpreters to give an automatic classification method for defects in two interpreters, CPython and PyPy. Summary of the Invention

[0006] Objective of the invention: To help interpreter developers detect and fix Python interpreter defects more efficiently, this paper proposes an automatic classification method for defects in two interpreters, CPython and PyPy. The difficulty of this method lies in automatically generating defect record tags by extracting keywords from issue reports.

[0007] Technical solution: To achieve the above objective, the technical solution adopted by the present invention includes the following steps:

[0008] Step 1): Extract keywords from the title of the issue report based on syntactic components;

[0009] Step 2): Supplement tags by combining quantitative analysis and the results of manual classification experiments;

[0010] Step 3): Recommend issues to corresponding categories based on code and text similarity;

[0011] Further, the specific steps of the above Step 1) are as follows:

[0012] Step 1.1): First, obtain the issue report links of CPython and PyPy, and extract the key information in the reports, including the number "ID", title "title", component "comp", content "msg", pull request "PR", and source file path "path";

[0013] Step 1.2): Propose a tag data meta-model for recording defects. This model stipulates the set of attributes that the tags for recording defects have. The set includes the following attributes:

[0014] Con (Condition): This attribute describes the situation under which the interpreter fails;

[0015] Sub (Subject): This attribute describes where the failure in the interpreter occurs;

[0016] Phe (Phenomena): This attribute describes what specific failure occurs in the interpreter;

[0017] Fix: This attribute describes how the defect in the interpreter is fixed;

[0018] Step 1.3): According to specific syntactic rules, divide the issue report title into two different sentence patterns, and extract keywords and their dimensions from them. The rules are as follows:

[0019] Text extraction rules for declarative sentences: For the issue titles in declarative sentence patterns, their subjects, objects, and conditional clauses have fixed formats, corresponding to the entity Sub where the fault occurs, the specific fault Phe that occurs, and the condition Con under which the fault occurs; according to this writing format, extract "Sub + Phe + Con";

[0020] Text extraction rules for imperative sentences: For the issue titles in imperative sentence patterns, there is no subject, only the predicate verb and the direct object, both of which are methods Fix used to express the repair of a certain fault. According to this writing format, extract "Fix + Fix";

[0021] Step 1.4), based on these rules, keywords can be extracted from the issue title as tags and filled into the corresponding tag attributes. The keyword extraction algorithm uses the TF-IDF metric model. Denote the i-th issue report as D i , and use a set of weighted words to represent the document D i , denoted as SE i = <W i,1 , W i,2 , W i,3 …W i,v >, where the weight W i,j represents the TF-IDF score of the j-th word in the document D i , and the specific calculation is as follows:

[0022]

[0023] Furthermore, the specific steps of the said Step 2) are as follows:

[0024] Step 2.1), by quantitatively analyzing the distribution of defects in components, obtain the top 15 components with higher defect occurrence frequencies in CPython and PyPy respectively. In CPython, there are Library (Lib), Documentation, InterpreterCore, and in PyPy, there are PyPy2 (running Python3.x), RPython. If these components appear in the report, fill them into the tag attribute table;

[0025] Step 2.2), By quantitatively analyzing the distribution of defects in the source files, 15 source files with higher defect occurrence frequencies are obtained in CPython and PyPy respectively. In CPython, they include Modules / posixmodule.c, Objects / unicodeobject.c, Python / ceval.c, and in PyPy, they include PyPy / module / cpyext / typeobject.py, PyPy / module / cpyext / api.py. If these files appear in the defect modification source file path path, they are filled into the label attribute table;

[0026] Step 2.3), Manually analyze the root cause of the defect to obtain the classification criteria. Then, match the keywords of the classification criteria in the msg of the issue, extract the code lines changed during defect repair in the PR field as the Fix attribute, and fill the extracted keywords and attributes into its label attribute table. The implementation of Step 2 can make up for the deficiencies in the extraction of syntactic components and improve the retrieval efficiency of defect records;

[0027] Further, the specific steps of the said Step 3) are as follows:

[0028] Step 3.1), Measure the text similarity between the label attribute table of an issue and the manually classified dataset. Based on the keyword vectors extracted in Steps 1 and 2, the user combines the vector space model VSM to classify the attributes in the two datasets into code attributes DC and non-code attributes DD;

[0029] In Step 3.2), after representing the non-code attribute DC as a word vector through the above steps, first let Then use the cosine similarity between two vectors SE m and SE n to calculate the similarity Sim(DD) of the two documents:

[0030]

[0031] Step 3.3), Measure the similarity between code snippets DC1 and DC1. First, extract the variable type names in the code snippets respectively to form string sets T1 and T2. Secondly, extract the called method names to form sets M1 and M2. Finally, the code similarity measurement work is transformed into measuring the similarity between sets T1 and T2 and M1 and M2;

[0032] The Jaccard coefficient is used to measure the set similarity, that is, the variable type similarity coefficient is:

[0033]

[0034] The method call similarity coefficient is:

[0035]

[0036] The similarity calculation formula for the complete code snippet is as follows:

[0037] sim(DC) = β·sim(T)+(1 - β)·sim(M)

[0038] Step 3.4), after obtaining the text similarity and code similarity, calculate the overall similarity. The similarity between each issue report tag vector and the set of manually classified sample vectors is denoted as score ID = <sim1, sim2, sim3…sim n >, select the maximum value (sim ID ) max from this set, and find the manually classified label tag of its corresponding sample and recommend it to this issue report. The calculation formula for the overall similarity is as follows:

[0039] sim = λ·sim(DD)+(1 - λ)·sim(DC). Description of the Drawings

[0040] Figure 1 is a flowchart for extracting keywords from the title of the issue report based on syntactic components

[0041] Figure 2 is a flowchart for supplementing labels by combining quantitative analysis and the results of manual classification experiments

[0042] Figure 3 is a flowchart for recommending categories based on code and text similarity Specific Implementation Modes

[0043] The present invention will be further clarified below in conjunction with the drawings and specific embodiments, so that those skilled in the art can make modifications to various equivalent forms of the present invention all fall within the scope defined by the appended claims of this application:

[0044] According to Figure 1 the process of extracting keywords from the title based on syntactic components, the specific steps are as follows:

[0045] Step 1.1), first obtain the issue report links of CPython and PyPy, and extract the keyword field information in the reports, including the number "ID", title "title", component "comp", and content "msg";

[0046] Step 1.2), according to the label data element model for recording defects proposed by this method and the specific syntactic rules of the issue report title, extract the keywords and their dimensions from the report and fill them into the label data element model;

[0047] Step 1.3), adopt the keyword extraction algorithm TF-IDF, and denote the i-th issue report as D i , and use a set of weighted words to represent the document D i , denoted as SE i = <W i,1 , W i,2 , W i,3 …W i,v >, where the weight W i,j represents the TF-IDF score of the j-th word in the document D i , and the calculation is as follows:

[0048]

[0049] According to Figure 2 the process of supplementing labels by combining quantitative analysis and manual classification results, the specific steps are as follows:

[0050] Step 2.1), by quantitatively analyzing the component fields of all fixed CPython and PyPy defect reports, obtain the distribution of defects in components, and screen out the 15 components with higher defect frequencies. If these components appear in the report, fill them into the label attribute table;

[0051] Step 2.2), by quantitatively analyzing the source files of all fixed CPython and PyPy defect reports, obtain the distribution of defects in source files, and screen out the 15 source files with higher defect frequencies. If these files appear in the defect modification source file path path, fill them into the label attribute table;

[0052] Step 2.3), manually analyze the root cause of the defect to obtain the classification criteria; the user matches the classification criteria keywords in the content field of the issue, extracts the code lines changed during defect repair from the PR field as the Fix attribute, and fills the extracted keywords and attributes into its label attribute table;

[0053] According to Figure 3 the process of recommending categories based on code and text similarity, the specific steps are as follows:

[0054] Step 3.1), based on the keyword vectors extracted in Step 1 and Step 2, the user combines the vector space model VSM to divide the attributes in the two datasets into code attributes DC and non-code attributes DD;

[0055] Step 3.2), after representing the non-code attribute DC as a word vector through the above steps, first let Then use the two vectors SE m and SE nCalculate the similarity Sim(DD) of two documents using the cosine similarity between them:

[0056]

[0057] Step 3.3), to measure the similarity between code snippets DC1 and DC1, first extract the variable type names in the code snippets respectively to form string sets T1 and T2, and then extract the called method names to form sets M1 and M2. Finally, the code similarity measurement work is transformed into measuring the similarity between sets T1 and T2 and between M1 and M2;

[0058] The Jaccard coefficient is used to measure the similarity of sets, that is, the variable type similarity coefficient is:

[0059]

[0060] The method call similarity coefficient is:

[0061]

[0062] The similarity calculation formula for the complete code snippet is as follows:

[0063] sim(DC) = β·sim(T) + (1 - β)·sim(M)

[0064] Step 3.4), calculate the overall similarity. The similarity between each issue report tag vector and the manual classification sample vector set is denoted as score ID = <sim1, sim2, sim3…sim n ( ID ) max , select the maximum value (sim

[0065] sim = λ·sim(DD) + (1 - λ)·sim(DC)

[0066] This paper proposes a method for the automatic classification of defects in two interpreters, CPython and PyPy, to help interpreter developers detect and repair Python interpreter defects more efficiently, and at the same time avoid the huge harm caused by some difficult-to-identify defects to the Python interpreter. Therefore, this method is of great significance for promoting the healthy development of the Python language ecosystem.

[0067] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. An automatic classification method for defects in two interpreters, CPython and PyPy, the main steps of which are as follows: Step 1): Extract keywords from the title of the issue report based on syntactic components; Step 2): Supplement tags by combining quantitative analysis and the results of manual classification experiments; Step 3): Recommend issues to corresponding categories based on code and text similarity; Furthermore, the specific steps of Step 1) are as follows: Step 1.1): First, obtain the issue report links of CPython and PyPy, and extract the key information in the reports, including the number "ID", title "title", component "comp", content "msg", pull request "PR", and source file path "path"; Step 1.2): Propose a tag data meta-model for recording defects, which stipulates the set of attributes of the tags for recording defects. The set includes the following attributes: Con (Condition): This attribute describes the situation under which the interpreter fails; Sub (Subject): This attribute describes where the failure in the interpreter occurs; Phe (Phenomena): This attribute describes what specific failure occurs in the interpreter; Fix: This attribute describes how the defect in the interpreter is fixed; Step 1.3): According to specific syntactic rules, divide the issue report title into two different sentence patterns, and extract keywords and their dimensions from them. The rules are as follows: Text extraction rule for declarative sentences: For the issue title in declarative sentence pattern, its subject, object, and conditional clause all have fixed formats, which respectively correspond to the subject Sub of the failure occurrence, the specific failure Phe that occurs, and the condition Con of the failure occurrence; extract "Sub+Phe+Con" according to this writing format; Text extraction rule for imperative sentences: For the issue title in imperative sentence pattern, there is no subject, only the predicate verb and the direct object, both of which are for expressing the method Fix for fixing a certain failure. Extract "Fix+Fix" according to this writing format; Step 1.4), according to these rules, keywords can be extracted from the title of the issue as tags and filled into the corresponding tag attributes. The keyword extraction algorithm uses the TF-IDF metric model. Denote the i-th issue report as D i , and represent the document D using a set of weighted words i , denoted as SE i = <W i,1 , W i,2 , W i,3 …W i,v >, where the weight W i,j represents the TF-IDF score of the j-th word in the document D i , and the calculation is as follows: Furthermore, the specific steps of Step 2) are as follows: Step 2.1): By quantitatively analyzing the distribution of defects in components, obtain the 15 components with higher defect occurrence frequencies in CPython and PyPy respectively. In CPython, there are Library (Lib), Documentation, InterpreterCore, and in PyPy, there are PyPy2 (running Python3.x), RPython. If these components appear in the report, fill them into the tag attribute table; Step 2.2): By quantitatively analyzing the distribution of defects in the source files, the top 15 source files with higher defect occurrence frequencies are obtained in CPython and PyPy respectively. In CPython, they include Modules / posixmodule.c, Objects / unicodeobject.c, Python / ceval.c, and in PyPy, they include PyPy / module / cpyext / typeobject.py, PyPy / module / cpyext / api.py. If these files appear in the defect modification source file path path, they are filled into the label attribute table; Step 2.3): Manually analyze the root causes of the defects to obtain the classification criteria. Then, match the keywords of the classification criteria in the msg of the issue, extract the code lines changed during defect repair from the PR field as the Fix attribute, and fill the extracted keywords and attributes into its label attribute table. The implementation of Step 2 can make up for the deficiencies in grammar component analysis and improve the retrieval efficiency of defect records; Further, the specific steps of step 3) are as follows: Step 3.1): Measure the text similarity between a certain issue label attribute table and the manually classified dataset. Based on the keyword vectors extracted in steps 1 and 2, and combined with the vector space model VSM, the attributes in the two datasets are divided into code attributes DC and non-code attributes DD; Step 3.2), after representing the non-code attribute DC as a word vector through the above steps, first let Then use two vectors SE m and SE n The cosine similarity between them is used to calculate the similarity Sim(DD) of two documents: Step 3.3): Measure the similarity between code snippets DC1 and DC1. First, extract the variable type names in the code snippets respectively to form string sets T1 and T2. Second, extract the called method names to form sets M1 and M2. Finally, the code similarity measurement work is transformed into measuring the similarity between sets T1 and T2 and between M1 and M2; The Jaccard coefficient is used to measure the set similarity, that is, the variable type similarity coefficient is: The method call similarity coefficient is: The similarity calculation formula for the complete code snippet is as follows: sim(DC) = β·sim(T) + (1 - β)·sim(M) Step 3.4), after obtaining the text similarity and code similarity, calculate the overall similarity. The similarity between each issue report tag vector and the set of manually classified sample vectors is denoted as score ID =<sim1, sim2, sim3…sim n (>, select the maximum value (sim ID ) max from this set, and find the manually classified label tag of its corresponding sample and recommend it to this issue report. The calculation formula for the overall similarity is as follows: sim = λ·sim(DD) + (1 - λ)·sim(DC).

Citation Information

Patent Citations

  • A similarity defect report recommendation method by combining weighted word vectors and latent semantic analysis

    CN109165382A

  • Software defect positioning method based on similarity integration

    CN112000802A