An App source code linking method based on user comments and developer wisdom
By adopting the intention classification and value comment extraction method of the BERT model in App development, combining GitHub's Issue and Commit information, an effective link between user comments and source code is established, and the problem of low efficiency and accuracy of user feedback information mining in the existing technology is solved, and valuable guidance for the continuous and rapid update and iteration of the App is achieved.
Patent Information
- Application Number
- CN202210393040.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-15
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-04-15
AI Technical Summary
The existing technology has shortcomings in identifying important user comments, solving the problem of weak natural language information in source code, and handling user comment version division, resulting in low efficiency and accuracy of user feedback information.
Using the intent classification and value comment extraction method based on the BERT model, combined with Issue and Commit related information in GitHub, we establish links from user comments to potentially change source codes and provide code modification suggestions through natural language processing and data mining technology.
It realizes rapid and accurate identification of important user feedback, narrows the semantic gap between user comments and source code, improves the efficiency and accuracy of user feedback information, and provides valuable guidance for the continuous and rapid updates and iterations of the App.
Smart Images

Figure CN114741088B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of software engineering, and particularly relates to an App source code linking method based on user comments and developer wisdom. Background Art
[0002] The continuous boom of the mobile application market has brought a large number of user comments to the app store. These comments can describe the problems encountered by users when using the App. By analyzing the internal information of these user comments, developers are driven to improve the applications they develop, mainly manifested as fixing problem loopholes or adding new functions. Useful user comments reflect the needs of the user group and contain valuable information for the subsequent update and maintenance of the App. However, due to the randomness and lack of threshold of user comments, a large number of useless comments have also been brought to the app store, making it difficult to monitor the behavior of user feedback in the app store. Directly obtaining user comments by developers not only cannot provide practical benefits to developers, but also brings the problem of wasting time for developers.
[0003] In addition to obtaining information from user comments in the app store, developers can also obtain useful information from open-source software hosting platforms. In the current popular background of social programming, there are numerous third-party open-source source code hosting platforms on the Internet, such as GitHub. Developers use GitHub to host application projects, open the source code, accept code contributions from other developers, use Issues to discuss problems in the development process, and use Commits to submit the modified code. These Issues and Commits gather the collective wisdom of developers, including semantic-level information and source code-level information, which have certain reference value for the future update and maintenance of the App.
[0004] In fact, the existing source code linking methods have at least the following deficiencies:
[0005] 1. The existing methods perform mediocrely in identifying important user comments, and both the production volume and accuracy rate of important user comments are relatively low.
[0006] 2. The existing methods do not solve the problem of weak natural language information of the source code. There is still a certain semantic gap between user feedback and the source code, which affects the accuracy of the link.
[0007] 3. The existing methods do not divide the versions of user comments, but directly put a large number of collected user comments into the experiment. Considering that a certain problem raised by user comments may have been solved by developers at a certain time point, it is meaningless to mine this part of user comments and feedback them to developers. Summary of the Invention
[0008] The object of the present invention is to overcome the deficiencies of the prior art and propose an App source code linking method based on user comments and developer wisdom, to mine valuable user feedback information from user comments, and at the same time consider the collective wisdom of the developer community, establish a link from user comments to the source code that potentially needs to be changed, and efficiently provide code modification suggestions for developers, so as to achieve the continuous and rapid update and iteration of the App.
[0009] The present invention solves its technical problems by adopting the following technical solutions:
[0010] An App source code linking method based on user comments and developer wisdom, comprising the following steps:
[0011] Step 1, crawl data information and preprocess the data information;
[0012] Step 2, use the BERT model to classify the intention of the data preprocessed in Step 1;
[0013] Step 3, extract valuable comments from the data with intention classification in the previous version in Step 2;
[0014] Step 4, use LDA to cluster the themes of the valuable comments in Step 3 and the data with intention classification in the current version in Step 2;
[0015] Step 5, use the text in the Issue data to enrich the semantics of the clustering themes in Step 4, and use the text in the Commit data to enrich the semantics of the source code components;
[0016] Step 6, through similarity calculation, calculate the similarity between the clustering themes with enriched semantics in Step 5 and the source code components, and perform source code linking through a potential source code recommendation algorithm.
[0017] Moreover, the crawling of data information in Step 1 includes: crawling the data information of the App from the Fdroid open source platform; crawling user comment information from Google Play; crawling Issue data and Commit data from GitHub, and storing the crawled data information in a database.
[0018] Moreover, the crawling of the data information of the App from the Fdroid open source platform includes the general description of the APP, the detailed description of the App, the GitHub address of the App, and the category to which the App belongs;
[0019] The crawling of user comment information from Google Play includes the APP to which the comment belongs, the number of likes of the comment, the score of the comment on the APP, the time of the comment, the user who made the comment, and the content of the comment;
[0020] The Issue data crawled from GitHub includes the GitHub address of the Issue, the title of the Issue, the status of the Issue, the comment content of the Issue, and the recording time of the Issue;
[0021] The Commit data crawled from GitHub includes the GitHub address of the Commit, the commit description of the Commit, the description of the Commit, the committer of the Commit, and the recording time of the Commit.
[0022] Moreover, the specific implementation method of the preprocessing in step 1 is: preprocess the crawled data information through non-English filtering, stop word removal, part-of-speech tagging, word correction, lemmatization, and short text removal of NLTK technology.
[0023] Moreover, the specific implementation method of step 2 is: add a CLS token and a SEP token to the user comments in the data preprocessed in step 1, obtain the input of the pre-trained language model BERT after the Embedding process, call the pre-trained language model BERT, select the feature vector at the CLS in its output, classify in the classification layer composed of the feed-forward neural network and the softmax function of the pre-trained language model BERT, and return the probability of each category to which the user comment belongs. Select the option with the largest probability value as the result of its intent classification, and the intent classification includes new feature requests, problem discovery, information prompt, information assistance, and others.
[0024] Moreover, the specific implementation method of step 3 is: use the Sentence-BERT model to vectorize the user comments classified in the previous version in the data preprocessed in step 2 and the text in the Issue data, and then calculate the similarity between the two through cosine similarity to extract valuable comments; as the similarity threshold between the user comments and the text in the Issue data gradually increases, the number of user comments that generate link pairs with the text in the Issue data gradually decreases, and at the same time, the number of valuable words gradually decreases; during the calculation process, compare the number of user comments that change with the reduction of unit valuable words under different similarity thresholds, and dynamically select the optimal similarity threshold.
[0025] Moreover, the specific implementation method of step 4 is: use LDA to cluster the topics of the valuable comments in step 3 and the data classified by intent in the current version in the data preprocessed in step 2, and select the optimal number of clustering topics of LDA in combination with PyLDAvis visualization and topic relevance metrics.
[0026] Moreover, the specific implementation method of step 5 is:
[0027] Semantically enrich the clustering themes of user comments using the text in the Issue data: Calculate the similarity between the clustering themes of user comments and the text in the Issue data using the asymmetric dice coefficient:
[0028]
[0029] Among them, is the set of words included in the clustering theme a of user comments, is the set of words included in the Issue text b. The role of the min function is to compare the number of words in the two sets of words, set the similarity threshold, regard the links with similarity results greater than the similarity threshold as valid links, and combine the two parts of data in each valid link to form the enriched user feedback information;
[0030] Semantically enrich the source code components using the text in the Commit data: Extract the Commit information related to the source code, and perform abstract syntax tree parsing on the source code. Combining the two together is the enriched source code component, where the abstract syntax tree parsing extracts the package name, class name, method name, variable name, and comments.
[0031] Moreover, the specific implementation method of step 6 is: Use the weighted asymmetric dice coefficient to calculate the similarity between the semantically enriched clustering theme and the source code component:
[0032]
[0033] Among them, RF i is the semantically enriched clustering theme of user comments, RC j is the semantically enriched source code component, k is the word that appears in both RF i and RC j . df represents the word frequency, is the set of words included in RF i , is the set of words included in RC j . The role of the min function is to compare the number of words in the two sets of words. After completing the similarity calculation, implement the link of the source code through the potential source code recommendation algorithm.
[0034] The advantages and positive effects of the present invention are:
[0035] The present invention uses related technologies such as natural language processing to connect the user comments after version division with the wisdom of the developer community. A method combining intent classification based on the BERT model and value comment extraction is adopted to identify important user feedback. By introducing information related to Issues and Commits in GitHub, the semantic gap between user comments and source code is narrowed, thereby achieving a better linking effect than existing methods. The present invention can quickly and accurately provide valuable user feedback to developers, and give a set of source code that may potentially be modified, providing guiding opinions for the continuous release of the App and improving the efficiency of developers in maintaining the App. In order to understand the degree of help of this tool to developers during actual use, this tool also has a function of collecting developer feedback information, so as to better optimize the method of the present invention and continuously improve the effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 is a flowchart of the present invention;
[0037] Figure 2 is a flowchart of intent classification based on the BERT model of the present invention;
[0038] Figure 3 is a flowchart of semantic enrichment of value comments for theme clustering of the present invention;
[0039] Figure 4 is an algorithm logic diagram of source code linking of the present invention;
[0040] Figure 5 is a diagram showing the effect of the application tool of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] The present invention will be further described in detail below with reference to the accompanying drawings.
[0042] An App source code linking method based on user comments and developer wisdom, in order to efficiently mine valuable user comments from user comments and further locate the App source code that potentially needs to be modified, so as to help developers better update and maintain the App. After considering the App version division, the present invention combines the content of two dimensions of user comments and developer wisdom to achieve accurate user feedback mining, so as to provide modification suggestions for developers. The whole process mainly involves technologies such as natural language processing and data mining, such as Figure 1 shown, including the following steps:
[0043] Step 1, crawl data information and preprocess the data information.
[0044] As shown in Table 1, the crawled data information includes: crawling the data information of apps from the Fdroid open-source platform; crawling user review information from Google Play; crawling Issue data and Commit data from GitHub, and storing the crawled data information into a database.
[0045] Among them, crawling the data information of apps from the Fdroid open-source platform includes the general description of the app, the detailed description of the app, the GitHub address of the app, and the category to which the app belongs.
[0046] Crawling user review information from Google Play includes the app to which the review belongs, the number of likes of the review, the score of the review for the app, the time of the review, the user who made the review, and the content of the review.
[0047] The Issue data crawled from GitHub includes the GitHub address of the Issue, the title of the Issue, the status of the Issue, the comment content of the Issue, and the recording time of the Issue.
[0048] The Commit data crawled from GitHub includes the GitHub address of the Commit, the commit description of the Commit, the description of the Commit, the committer of the Commit, and the recording time of the Commit.
[0049] Important information fields for each dimension in Table 1
[0050]
[0051]
[0052] Due to the informality of user evaluation behavior and the relatively casual spelling of words, the reviews submitted by users contain a large amount of noisy data, such as misspelled words, abbreviations, repeated-letter words, etc. These noisy data have a serious impact on extracting preference features from user reviews. Therefore, it is necessary to perform necessary data preprocessing on the crawled data. The present invention uses non-English filtering, stop-word removal, part-of-speech tagging, word correction, lemmatization, and short-text removal through NLTK technology to preprocess the crawled data information.
[0053] Step 2: Use the BERT model to perform intent classification on the preprocessed data in Step 1. To better help developers understand users' feedback on the App, classify users' comment intents into new feature requests, problem discovery, information prompts, information assistance, and others. Among them, new feature requests and problem discovery are most relevant to the update and maintenance of the App. The recognition accuracy of existing methods for these two types of user comments is between 60% and 70%. Therefore, this invention uses the pre-trained language model BERT with stronger feature extraction ability to perform intent classification on user comments:
[0054] As Figure 2 shown, add a CLS token and a SEP token to the user comments in the data preprocessed in Step 1. After the Embedding process (word embedding, segment embedding, position embedding, and summation operation), obtain the input of the pre-trained language model BERT. Call the pre-trained language model BERT, select the feature vector at the CLS position in its output, perform classification in the classification layer composed of the feed-forward neural network and the softmax function of the pre-trained language model BERT, and return the probability of each category to which the user comment belongs. Select the option with the largest probability value as the result of its intent classification.
[0055] Step 3: Extract valuable comments from the data classified by intent in the previous version in the data preprocessed in Step 2. The user comments of the current version can provide important information for the subsequent update and maintenance of the App. However, in actual development, there is a problem similar to "cold start" in obtaining user feedback. That is, after the release of the current version, most Apps cannot obtain enough user comments in the short term. To solve this problem, based on the division of user comment versions, select relevant comments from the user comments of the previous version to supplement the user comment information of the current version.
[0056] Use the Sentence-BERT model to vectorize the user comments classified in the previous version in the data preprocessed in Step 2 and the text in the Issue data, and then calculate the similarity between the two through cosine similarity to extract valuable comments; as the similarity threshold between the user comments and the text in the Issue data gradually increases, the number of user comments that generate link pairs with the text in the Issue data gradually decreases, and at the same time the number of valuable words gradually decreases; during the calculation process, compare the number of user comments that change with the decrease of the unit valuable word under different similarity thresholds, and dynamically select the optimal similarity threshold.
[0057] Step 4: Use LDA to cluster the topics of the valuable comments in Step 3 and the data classified by intent in the current version in the data preprocessed in Step 2. Achieve more fine-grained feature extraction, and combine PyLDAvis visualization and topic relevance metrics to select the optimal number of clustering topics for LDA.
[0058] Step 5: Semantically enrich the clustering topics in Step 4 using the text in the Issue data, and semantically enrich the source code components using the text in the Commit data.
[0059] Semantically enrich the user comment clustering topics using the text in the Issue data: Calculate the similarity between the user comment clustering topics and the text in the Issue data using the asymmetric dice coefficient:
[0060]
[0061] where is the set of words contained in the user comment clustering topic a, is the set of words contained in the Issue text b. The function of the min function is to compare the number of words in the two sets of words and return the smaller value. Set the similarity threshold to. Consider the links with similarity results greater than the 0.3 similarity threshold as valid links, and combine the two parts of data in each valid link to form the enriched user feedback information (abbreviation: RF).
[0062] Semantically enrich the source code components using the text in the Commit data: Extract the Commit information related to the source code, perform abstract syntax tree parsing on the source code, and combine the two to obtain the enriched source code components (abbreviation: RC). Among them, the abstract syntax tree parsing extracts package names, class names, method names, variable names, and comments:
[0063] As Figure 3 shown, first perform abstract syntax tree parsing and data preprocessing on the source code to obtain the bag-of-words model of the source code components. Then, use the file name of each java source code file as a parameter to splice the SQL statement. Next, query the Commit text with the same java file name in the database, extract and preprocess the semantic information (the content of the title and detailed description) of all the Commit text containing the java file name that is queried to obtain the bag-of-words model of the Commit information. Finally, combine the two.
[0064] Step 6: Calculate the similarity between the semantically enriched clustering topics and the source code components in Step 5 through similarity calculation, and link the source code through the potential source code recommendation algorithm.
[0065] To better reflect the importance of high-frequency words, use the weighted asymmetric dice coefficient to calculate the similarity between RF and RC:
[0066]
[0067] where RFi is the clustering theme of user comments after semantic enrichment, RC j is the source code component after semantic enrichment, and k is the word that appears in both RF i and RC j The word in, df represents the word frequency, is RF i The set of words contained in, is RC j The set of words contained in. The function of the min function is to compare the number of words in two sets of words and return the smaller value.
[0068] Such as Figure 4 As shown, after the similarity calculation is completed, this algorithm is used to implement the top-k link recommendation for locating user comments to the source code.
[0069] Such as Figure 5 As shown, according to the above method, an application tool for displaying the source code of user comment links is designed to provide valuable user feedback and source code modification suggestions for developers in an intuitive way. There is a developer feedback icon above each comment, and developers can judge the value of the comment, that is, whether the appearance of the comment is helpful to the developers. After the abbreviated comment is expanded, the complete user comment and the user evaluation time are displayed. The potential source code that may be linked is expanded to display the set of potential source codes of the user comment link, and the feedback of the developers on the recommended source code is collected to judge whether it is useful. The developer of this App is used as the audience to assist the developer in a more intuitive way and improve the efficiency of maintaining the App.
[0070] It should be emphasized that the embodiments described in the present invention are illustrative rather than restrictive. Therefore, the present invention includes but is not limited to the embodiments described in the specific embodiments. Any other embodiments obtained by those skilled in the art according to the technical solutions of the present invention also belong to the scope of protection of the present invention.
Claims
1. An App source code linking method based on user comments and developer wisdom, characterized in that: It includes the following steps: Step 1, crawl data information and preprocess the data information; Crawling data information includes: crawling the data information of the App from the Fdroid open source platform; crawling user comment information from Google Play; crawling Issue data and Commit data from GitHub, and storing the crawled data information into the database; Step 2, use the BERT model to perform intent classification on the preprocessed user comment data in Step 1; Intent classification includes new feature requests, problem discovery, information prompts, information assistance, and others; Step 3, extract valuable comments from the data after intent classification of the previous version of the APP in Step 2; Step 4, use LDA to cluster the themes of the valuable comments in Step 3 and the data after intent classification of the current version of the APP in Step 2; Step 5, use the text in the Issue data to enrich the semantics of the clustering themes in Step 4, and use the text in the Commit data to enrich the semantics of the source code components; Step 6, through similarity calculation, calculate the similarity between the clustering themes with enriched semantics in Step 5 and the source code components, and perform source code linking through a potential source code recommendation algorithm.
2. The App source code linking method based on user comments and developer wisdom according to claim 1, characterized in that: The crawling of the data information of the App from the Fdroid open source platform includes the general description of the App, the detailed description of the App, the GitHub address of the App, and the category to which the App belongs; Crawling user comment information from Google Play includes the APP to which the comment belongs, the number of likes of the comment, the score of the comment on the APP, the time of the comment, the user of the comment, and the content of the comment; The Issue data crawled from GitHub includes the GitHub address of the Issue, the title of the Issue, the status of the Issue, the comment content of the Issue, and the recording time of the Issue; The Commit data crawled from GitHub includes the GitHub address of the Commit, the commit description of the Commit, the description of the Commit, the committer of the Commit, and the recording time of the Commit.
3. The App source code linking method based on user comments and developer wisdom according to claim 1, characterized in that: The specific implementation method of the preprocessing in Step 1 is: preprocess the crawled data information through non-English filtering, stop word removal, part-of-speech tagging, word correction, lemmatization, and short text removal of the NLTK technology.
4. The App source code linking method based on user comments and developer wisdom according to claim 1, characterized in that: The specific implementation method of step 2 is as follows: Add a CLS token and a SEP token to the user comments in the data preprocessed in step 1. After the Embedding process, obtain the input of the pre-trained language model BERT. Invoke the pre-trained language model BERT, select the feature vector at the CLS position in its output, perform classification in the classification layer composed of the feed-forward neural network and the softmax function of the pre-trained language model BERT, and return the probability of each category to which the user comment belongs. Select the option with the largest probability value as the result of its intent classification.
5. A method for linking App source code based on user comments and developer wisdom according to claim 1, characterized in that: The specific implementation method of step 3 is as follows: Use the Sentence-BERT model to vectorize the user comments classified in the previous version in the data preprocessed in step 2 and the text sentences in the Issue data, and then calculate the similarity between the two through cosine similarity to extract valuable comments; As the similarity threshold between the user comments and the text in the Issue data gradually increases, the number of user comments generating link pairs with the text in the Issue data gradually decreases, and at the same time the number of valuable words gradually decreases; During the calculation process, compare the number of user comments that change with the decrease of the unit valuable word under different similarity thresholds, and dynamically select the optimal similarity threshold.
6. A method for linking App source code based on user comments and developer wisdom according to claim 1, characterized in that: The specific implementation method of step 4 is as follows: Use LDA to cluster the topics of the valuable comments in step 3 and the data classified by intent in the current version in the data preprocessed in step 2, and combine PyLDAvis visualization and topic relevance indicators to select the optimal number of clustering topics for LDA.
7. A method for linking App source code based on user comments and developer wisdom according to claim 1, characterized in that: The specific implementation method of step 5 is as follows: Semantically enrich the clustering topics of user comments using the text in the Issue data: Calculate the similarity between the clustering topics of user comments and the text in the Issue data using the asymmetric dice coefficient: Among them, is the set of words contained in the clustering topic of user comments a in, is the Issue text b is the set of words contained in, and the role of the min function is to compare the number of words in the two sets of words Set a similarity threshold, regard the links with similarity results greater than the similarity threshold as valid links, and combine the user comment clustering topics and the Issue texts linked to them in each valid link to form enriched user feedback information; Semantically enrich the source code components using the text in the Commit data: Extract the Commit information related to the source code, perform abstract syntax tree parsing on the source code, and combine the parsed source code components and the matching Commit information to form enriched source code components, where the abstract syntax tree parsing extracts package names, class names, method names, variable names, and comments.
8. A method for linking App source code based on user comments and developer wisdom according to claim 1, characterized in that: The specific implementation method of step 6 is as follows: Use the weighted asymmetric dice coefficient to calculate the similarity between the semantically enriched clustering topics and the source code components: where is the user comment clustering topic after semantic enrichment, is the source code component after semantic enrichment, is the word that appears in both and , df represents the word frequency, is the set of words included in, is the set of words included in. The role of the min function is to compare the number of words in the two sets of words. After completing the similarity calculation, the source code link is implemented through the potential source code recommendation algorithm.
Citation Information
Patent Citations
Method for automatically analyzing user comments in application store and recommending comments to developers
CN110827118A
Method and device for identifying App key functions based on user comments
CN111736804A