INFORMATION PROCESSING DEVICE, WARNING PRIORITY PREDICTION METHOD AND WARNING PRIORITY PREDICTION PROGRAM
A machine learning-based model predicts warning priorities across diverse projects by using non-dependent information, addressing the challenge of mixed warnings in static analysis tools and reducing costs by eliminating the need for project-specific models.
Patent Information
- Application Number
- DE102024138907
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-28
- Filing Date
- 2024-12-19
- Publication Date
- 2025-07-03
AI Technical Summary
Existing static analysis tools generate a mixture of true and false positive warnings, making it difficult for developers to accurately distinguish between them, and applying project-specific analysis devices increases development costs.
A machine learning-based prediction model that uses non-dependent information and warning labels to predict the priority of warnings across various product projects, incorporating features like static analysis tool type and settings, product-specific information, and source code metrics to classify warnings as true positives, false positives, and their correction needs.
Enables accurate prediction of warning priorities for different source codes, reducing the need for project-specific models and minimizing development costs by using a versatile model that classifies warnings with high accuracy.
Smart Images

Figure 00000013_0000 
Figure 00000014_0000 
Figure 00000015_0000
Abstract
Description
[0001] The present disclosure relates to an information processing apparatus, a warning priority prediction method, and a warning priority prediction program.
[0002] A static analysis tool was used to analyze a violation in a source code. Although the analysis result from a static analysis tool includes numerous warnings, these warnings are a mixture of a true positive warning and a false positive warning. A true positive warning indicates a location where a violation actually exists, while a false positive warning incorrectly indicates a location where no violation exists.
[0003] Therefore, a source code developer had to distinguish between a true positive warning and a false positive warning even from the numerous warnings included in the analysis result of a static analysis tool.
[0004] In response, Patent Document 1 describes an analysis device that prioritizes the static analysis result with respect to the source code being analyzed based on the user's judgment result on whether to modify the source code according to the static analysis result, information about source code metrics, and information about the program development project. The user's judgment result refers to the determination of true positive and false positive warnings. A true positive warning indicates a location where a violation truly exists, while a false positive warning incorrectly indicates a location where no violation exists. The user's judgment result indicating the true positive and false positive warnings is assigned to each warning as a warning label.
[0005] Patent Document 1: JP 2017-204090 A1
[0006] The analysis device described in Patent Document 1 prioritizes a static analysis result based on information specific to the product project, resulting in priority rankings that depend on the product project. Therefore, if the analysis device of Patent Document 1 is applied to the static analysis of source code from different product projects, the accuracy of the priority ranking of the static analysis result may decrease. In addition, constructing an analysis device for each product project may increase development costs and other expenses.
[0007] Therefore, the present disclosure aims to provide an information processing apparatus, a warning priority prediction method, and a warning priority prediction program that can accurately predict the priority of a warning from static analysis for various source codes used in various product projects.
[0008] According to one aspect of the present disclosure, an information processing device is provided that predicts a priority of a warning indicated by an analysis result of a source code by a static analysis tool. The information processing device includes: a learning unit configured to generate a prediction model, the prediction model being trained using machine learning with learning input data including the source code, warning information including a location of a warning in the source code indicated by the analysis result, product-specific information using the source code, and non-dependent information that does not depend on the product, and using a warning label associated with each warning as training data.wherein the analysis result of the static analysis tool for the source code serves as input and the priority of the warning serves as output; a receiving unit configured to receive the analysis result of the static analysis tool, which is a target of prediction by the prediction model; and a predicting unit configured to predict the priority of the warning indicated by the analysis result received by the receiving unit using the prediction model.
[0009] According to this configuration, a predictive machine learning model with warning labels associated with each warning as training data is used by the static analysis tool to determine the warning priority indicated by the source code analysis result. The warning label is assigned to each warning based on user judgment. The input data for learning the predictive model includes non-dependent information that is not dependent on the product. As a result, the predictive model is generated as a versatile machine learning model that is not specialized for specific source codes or products. Therefore, this configuration can accurately predict the warning priority indicated by the static analysis result for different source codes used in different product projects.
[0010] In the information processing device, the non-dependent information may include information regarding the type and settings of the static analysis tool used to analyze the source code.
[0011] The performance of static analysis varies depending on the type and settings of the static analysis tool used to analyze the source code. Therefore, by using information regarding the type and settings of the static analysis tool as input data to train the predictive model, the priority of the warning indicated by the static analysis result can be predicted with high accuracy. The type of static analysis tool includes different versions of the same type of static analysis tool. The settings of the static analysis tool can include, for example, programming language standards.
[0012] In the information processing device, the non-dependent information may include feature information of the static analysis tool according to the type of the static analysis tool. According to this configuration, by using the feature information of the static analysis tool as input data for learning the prediction model, a static analysis warning can be predicted with higher accuracy.
[0013] In the information processing device, the characteristic information may include the determinability of the static analysis and / or the benchmark evaluation result of the static analysis tool. According to this configuration, the priority of a warning indicated by the static analysis result can be predicted with high accuracy.
[0014] In the information processing device, the priority of the warning may be the probability of belonging to a specific warning flag. According to this configuration, the user can easily identify the location of the source code that needs to be corrected.
[0015] In the information processing device, the warning flags may be multi-valued, including at least one true positive indicating a location with an actual violation and one false positive indicating a location that was incorrectly warned of as having a violation. According to this configuration, the user can easily identify the location of the source code that needs to be corrected.
[0016] In the information processing device, the true positive can be divided into a warning that requires source code correction and a warning that does not require source code correction. According to this configuration, the user can easily identify the location of the source code that needs to be corrected.
[0017] In the information processing apparatus, the prediction unit may input the analysis result of the source code, the source code, product-specific information using the source code, and non-dependent information into the prediction model to predict the priority of the warning indicated by the analysis result.
[0018] In the above information processing apparatus, the prediction model may be generated by a machine learning model capable of ensemble learning.
[0019] According to one aspect of the present disclosure, a priority prediction method is provided that predicts a priority of a warning indicated by an analysis result of a source code by a static analysis tool. The priority prediction method includes: a first step of generating, by a learning unit, a prediction model trained using machine learning with learning input data including the source code, warning information including a location of a warning in the source code indicated by the analysis result, product-specific information using the source code, and non-dependent information that does not depend on the product, and using a warning label associated with each warning as training data, wherein the analysis result of the static analysis tool for the source code serves as input and the priority of the warning serves as output;a second step of receiving, by a receiving unit, the analysis result of the static analysis tool that is a target of prediction by the prediction model; and a third step of predicting, by a predicting unit, the priority of the warning indicated by the analysis result received by the receiving unit using the prediction model.
[0020] According to one aspect of the present disclosure, a warning priority prediction program is provided that causes a computer included in an information processing device to predict a priority of a warning indicated by an analysis result of a source code by a static analysis tool. The program causes the computer to function as: a learning unit configured to generate a prediction model trained using machine learning with learning input data including the source code, warning information including a location of a warning in the source code indicated by the analysis result, product-specific information using the source code, and non-dependent information that does not depend on the product, and using a warning label associated with each warning as training data.wherein the analysis result of the static analysis tool for the source code serves as input and the priority of the warning serves as output; a receiving unit configured to receive the analysis result of the static analysis tool, which is a target of prediction by the prediction model; and a predicting unit configured to predict the priority of the warning indicated by the analysis result received by the receiving unit using the prediction model.
[0021] According to the present disclosure, it is possible to accurately predict the priority of a warning from a static analysis for different source codes used in different product projects.
[0022] Objects, features, and advantages of the present disclosure will become more apparent from the following detailed description with reference to the accompanying drawings. In the drawings: Fig. 1 is a diagram illustrating the functional configuration of the warning priority prediction device according to an embodiment; Fig. 2 is a schematic diagram illustrating an overview of the alert classification model according to the embodiment; Fig. 3 is a schematic diagram illustrating the processing of generating an alert classification model using information from multiple product projects according to the embodiment; Fig. 4 is a schematic diagram illustrating the processing of prioritizing an alert using the alert classification model according to the embodiment; and Fig. 5 is a schematic diagram illustrating the processing of generating an alert classification model using information from multiple product projects according to another embodiment.
[0023] Embodiments of the present disclosure will be described with reference to the drawings. The embodiments described below are examples of how the present disclosure can be implemented and are not intended to limit the present disclosure to the specific configurations described below. In implementing the present disclosure, specific configurations suitable for the embodiments may be adopted as needed.
[0024] Fig. 1 is a diagram showing the functional configuration of the warning priority prediction device 10 according to the present embodiment. The warning priority prediction device 10 of this embodiment is an information processing device equipped with a CPU (Central Processing Unit) or other computing devices, and can be accessed via a web browser from a user terminal 12.
[0025] The warning priority prediction device 10 classifies warnings indicated by the analysis result (corresponding to a static analysis result) of a source code by a static analysis tool 20 using a prediction model generated by machine learning, and predicts the priority of the warning. The prediction model is referred to as an alert classification model.
[0026] The user terminal 12 is an information processing device, such as a laptop or desktop personal computer, and can access the alert priority prediction device 10. A user, such as a software developer, creates the source code for the software used in products using the user terminal 12.
[0027] A user sends the source code to the warning priority prediction device 10 from the user terminal 12 as needed according to the progress of the creation. The warning priority prediction device 10 analyzes the received source code using the static analysis tool 20 and outputs the analysis result.
[0028] The static analysis result includes a warning indicating the location of a violation in the source code. The warning contains a mixture of a true positive (TP), indicating an actual violation, and a false positive (FP), incorrectly indicating a violation when none exists. Therefore, it is necessary to determine whether each warning is a TP or an FP.
[0029] In response, the alert priority prediction device 10 of this embodiment classifies an alert indicated by the static analysis result into TP or FP using the alert classification model, thereby predicting the priority of the alert that a user should address. The alert priority prediction device 10 then sends the prediction result to the user terminal 12.
[0030] Furthermore, TPs can be further divided into a warning requiring source code correction (confirmed) and a warning that does not require source code correction (intentional). Therefore, the warning priority prediction device 10 can further classify TPs into a warning requiring correction and a warning that can be ignored, and predict the priority of the warning.
[0031] In other words, the prediction by the alert classification model of this embodiment includes a binary classification that classifies the alert indicated by the static analysis result into TP and FP and outputs the probability of belonging to each, and a ternary classification that classifies the alert into a warning requiring correction, a warning that can be ignored, and FP, and outputs the probability of belonging to each. Whether the prediction by the alert classification model is binary or ternary can be determined by the settings of the alert priority prediction device 10, or the alert priority prediction device 10 may have only one of these functions.
[0032] For example, in binary classification, the probability that an alert is a TP is predicted as a 90% TP probability and a 10% FP probability. In ternary classification, for a given alert, the probability that the alert is a corrective action alert is predicted as a 70% corrective action probability, the probability that the alert is an ignored alert is predicted as a 20% ignored probability, and the probability that the alert is a false positive (FP) is predicted as a 10% FP probability.
[0033] Thus, the alert priority prediction device 10 outputs the probability of belonging to a specific alert flag as the priority of the alert. The specific alert flags are multi-valued, including at least TP and FP. In other words, the alert with a high TP probability or a high correction need probability is considered a high-priority alert. This allows a user to easily identify the location of the source code that requires correction and prioritize the correction of the code according to the alert identified as a high-probability TP.
[0034] A warning label, such as TP, FP, corrective action warning, and ignored warning, is assigned to each warning. In the following description, the TP probability, FP probability, corrective action probability, and ignored probability predicted by the warning classification model are collectively referred to as a prediction probability.
[0035] As in Fig. 1, the warning priority prediction device 10 of this embodiment includes a static analysis tool 20, a source code metrics measurement tool 22, a database 24, a learning unit 26, a receiving unit 28, a prediction application programming interface (API) 30, and an output unit 32.
[0036] The static analysis tool 20 statically analyzes the source code received from the user terminal 12 and outputs the static analysis result. The static analysis result includes an alert ID assigned to each alert to identify the violation in the source code. The static analysis result is stored in the database 24 associated with the analyzed source code.
[0037] The source code metrics measurement tool 22 measures the metrics of the source code received from the user terminal 12. The source code metrics are stored in the database 24 associated with the measured source code.
[0038] The database 24 stores input data for learning (referred to as “learning input data”) and training data used to train the alert classification model, which is a machine learning model, and the classification result of an alert using the alert classification model.
[0039] The learning unit 26 generates the alert classification model through machine learning using the learning input data and the training data stored in the database 24. The learning unit 26 updates the alert classification model as needed when various data stored in the database 24 is added or updated.
[0040] The receiving unit 28 receives the static analysis result to be predicted by the alert classification model. The static analysis result is the analysis result of the source code by the static analysis tool 20, and the alert is unclassified. In addition to receiving the unclassified alert indicated by the static analysis result, the receiving unit 28 also receives various data necessary for determining the priority of the alert and derives these features.
[0041] The prediction API 30 uses the alert classification model trained (learned) by the learning unit 26 and determines the priority of the alert indicated by the static analysis result received by the receiving unit 28 using the alert classification model. In other words, the prediction API 30 functions as a classifier that outputs the predicted probability indicating the priority of the alert associated with the alert ID.
[0042] The output unit 32 sends the predicted probability for each alert, which is the prediction result of the alert classification model in the prediction API 30, to the user terminal 12.
[0043] The user terminal 12 displays the predicted probability for each alert issued by the alert classification model in a viewer. The viewer can sort alerts by TP probability. This allows the user to confirm the priority of the alert to be addressed and easily identify the alert that should be prioritized.
[0044] In addition, if the predicted probability received by the alert priority prediction device 10 is deemed appropriate, the user assigns a warning label to the alert ID according to the predicted probability and sends it back to the alert priority prediction device 10. Conversely, if the predicted probability received by the alert priority prediction device 10 is deemed inappropriate, the user assigns the warning label that the user considers appropriate to the alert ID and sends it back to the alert priority prediction device 10. The alert priority prediction device 10 stores the alert ID and a warning label received from the user terminal 12 in the database 24 associated with the static analysis result targeted by the alert classification model.Accordingly, the learning input data is accumulated to generate the alert classification model.
[0045] In this way, the alert priority prediction device 10 of this embodiment operates the alert classification model in conjunction with the static analysis tool 20. Therefore, it can collect labeled training data (training data) without requiring a user to perform an annotation task for training the alert classification model by machine learning.
[0046] While the warning priority prediction device 10 of this embodiment includes the learning unit 26, the functions of the learning unit 26 may be provided by another information processing device.
[0047] Fig. 2 is a schematic diagram showing an overview of the alert classification model according to this embodiment.
[0048] The alert classification model of this embodiment is a model in which, when an alert indicated by the static analysis result is input, a large number of weak learners are generated that predict an alert label based on various conditions of the input data. The prediction results of the alert label probability by these weak learners are aggregated into a single prediction value using methods such as majority voting or weighting to be output. Therefore, the alert classification model of this embodiment is generated by supervised learning using a machine learning model capable of ensemble learning, such as LightGBM.
[0049] The learning input data for generating the alert classification model of this embodiment through machine learning includes features based on source code, alert information, product project information, and non-dependent information (also referred to as independent information). The training data includes an alert label assigned to each alert indicated by the static analysis result based on the user's judgment.
[0050] The source code is the one analyzed by the static analysis tool 20. Instead of the source code itself, the learning input data may include the number of lines of code and the McCabe metric or cyclomatic complexity, which are included in metrics of the source code measured by the source code metrics measurement tool 22.
[0051] Here, the trends of the source code can vary depending on the product. Therefore, by using the source code or source code metrics as learning input data, the trends of different source codes are learned for each product project.
[0052] The warning information contains information about the occurrence location of a warning in the source code, which is specified by the static analysis result. The warning label, which is the training data, is associated with the information.
[0053] Product project information is product-specific information using the source code, such as the product name or model number. The product project information may include other information representing the characteristics of the product project in addition to the product name. By using the product project information as learning input data, the tendency of a warning label is learned for each product.
[0054] The non-dependent information includes information that does not depend on the product. By using the non-dependent information that does not depend on the product as learning input data, the alert classification model is generated as a versatile machine learning model that is not specialized for specific source code or products.
[0055] The non-dependent information in this embodiment includes information regarding the type and settings of the static analysis tool 20 used to analyze the source code. The type of the static analysis tool 20 includes, for example, the name, product number, and version of the static analysis tool 20. The settings of the static analysis tool 20 include, for example, the standard of the programming language used to create the source code. The programming language standard includes, for example, whether the same C language conforms to C99 or not. The information regarding the type and settings of the static analysis tool 20 is referred to as tool information.
[0056] The performance of static analysis varies depending on the type and settings of the static analysis tool 20 used to analyze the source code. Therefore, by using information regarding the type and settings of the static analysis tool 20 as learning input data, the warning label associated with the warning indicated by the static analysis result can be predicted with high accuracy regardless of the static analysis tool 20.
[0057] Additionally, the non-dependent information includes property information of the static analysis tool 20. The property information of the static analysis tool 20 includes the determinability of the static analysis and / or the benchmark evaluation result of the static analysis tool 20. The determinability of the static analysis refers to the determinability of the static analysis based on rules such as MISRA that correspond to the warning issued by the static analysis tool 20, indicating "determinable," "undeterminable," or "indeterminate" for each warning. The determinability and the benchmark evaluation result are based on publicly available information from the manufacturer of the static analysis tool 20, generally known information, or information based on past experience.
[0058] The property information of the static analysis tool 20 varies depending on the type and version of the static analysis tool 20 and is stored in the database 24 associated with the type and version of the static analysis tool 20. By using the property information of the static analysis tool 20 as learning input data for the alert classification model, a static analysis alert can be predicted with higher accuracy.
[0059] The learning unit 26 of this embodiment derives a feature based on the source code, warning information, product project information, and non-dependent information. An example of the derived feature is shown below. The derived features are not limited to the following, and it is not necessary to use all of the following features: the name (URL) of the source code base where the source code with the warning is maintained; the path to the source code file where the warning occurred; the name of the source code file where the warning occurred; the function name in the source code where the warning occurred; the line number in the source code where the warning occurred; the ID representing the type of alert; the severity of the warning; the name of the static analysis tool 20 that generated the warning; the version of the static analysis tool 20 that generated the warning; the determinability of static analysis based on rules such as MISRA that comply with the warning; the benchmark evaluation result of the static analysis tool 20 based on rules such as MISRA that correspond to the warning; the number of warnings in the source code function where the warning occurred; the number of warnings in the source code file where the warning occurred; the time elapsed since the warning occurred; the file-level source code metrics of the source code file where the warning occurred; the source code metrics at the function level of the source code file where the warning occurred; and the name of the product project where the warning occurred.
[0060] The learning unit 26 of this embodiment uses the above features as learning input data and generates the alert classification model through machine learning using the alert label assigned to each alert indicated by the static analysis result based on the user's judgment as training data. Such a combination of a feature and an alert label constitutes a single training data set. In other words, the more training data sets there are, the higher the accuracy of the alert classification model.
[0061] The prediction API 30 inputs the static analysis result of the source code, the source code, product project information, and non-dependent information into the alert classification model, and outputs the predicted probability of the alert marking, which indicates the priority of the alert indicated by the static analysis result.
[0062] Fig. 3 is a schematic diagram showing the processing of generating a warning classification model using information from a plurality of product projects 1 to N according to this embodiment. The processing of generating the warning classification model is executed by a program stored in a storage medium, such as a storage unit, provided in the warning priority prediction device 10.
[0063] When the program is executed, the procedure is executed according to the program.
[0064] As in Fig. 3, the feature that is the learning input data is derived by the learning unit 26 based on the source code metrics, the static analysis result, the product project information, the tool information, and the property information of the static analysis tool 20 for each product project 1 to N.
[0065] The property information of the static analysis tool 20 is stored in the database 24 as common information independent of the product project. The property information of the static analysis tool 20 corresponding to the name and version specified by the tool information is read from the database 24.
[0066] In addition, the warning label assigned to the warning indicated by the static analysis result of the source code for each product project 1 to N is used as training data.
[0067] The learning data sets, which are combinations of a feature and training data for each product project 1 to N, are stored in the database 24. The learning data sets are generated each time the source code for each product project 1 to N is modified or improved, stored in the database 24, and used to subsequently learn the alert classification model. The learning unit 26 generates the alert classification model through machine learning using a large number of learning data sets.
[0068] Fig. 4 is a schematic diagram showing the processing of prioritizing a warning using the warning classification model according to this embodiment. The processing of prioritizing a warning is executed by a program stored in a storage medium, such as a storage unit provided in the warning priority prediction device 10. When this program is executed, the process according to the program is executed.
[0069] The processing of prioritizing an alert is performed when the user sends the source code to the alert priority prediction device 10 via the user terminal 12 and issues an instruction to perform static analysis. According to this instruction, the static analysis tool 20 first performs static analysis on the source code and outputs the static analysis result. Additionally, the source code metrics measurement tool 22 measures the source code metrics of the source code.
[0070] The receiving unit 28 reads the product project information, tool information, and feature information of the static analysis tool 20 corresponding to the source code from the database 24. Based on the read information, the static analysis result, and the source code metrics, the receiving unit 28 derives the feature. The feature includes a warning indicated by the static analysis result with an unknown warning flag.
[0071] By inputting the feature to the alert classification model, the alert classification model outputs the predicted probability of the alert as the prediction result. The predicted probability of the alert is stored in the database 24 and sent to the user terminal 12. The user confirms the location of the source code requiring correction by checking the predicted probability of the alert sent to the user terminal 12 using a viewer.
[0072] As mentioned above, if the predicted probability is deemed appropriate, the user assigns a warning label to the warning ID according to the predicted probability and sends it back to the warning priority prediction device 10. If the predicted probability is deemed inappropriate, the user assigns a warning label to the warning ID based on the user's judgment and sends it back to the warning priority prediction device 10. The warning priority prediction device 10 stores the warning ID and the warning label received from the user terminal 12 in the database 24 associated with the static analysis result. The warning label thus stored in the database 24 is used as training data.
[0073] In Fig. 4, the static analysis result of the source code for product projects 1 to N is the prediction target of the warning classification model, but the static analysis result of the source code except product projects 1 to N can also be the prediction target of the warning classification model.
[0074] In addition, the alert priority prediction device 10 of this embodiment may update the alert classification model through machine learning using a newly stored learning data set in the database 24. The timing of the update may be, for example, when the performance deterioration of the alert classification model becomes apparent or when a predetermined number of new learning data sets have been accumulated.
[0075] For example, the performance degradation of the alert classification model may be a case when the frequency of an alert label assigned by the user, which differs from the predicted probability of the alert classification model, exceeds a predetermined value.
[0076] Furthermore, by updating the alert classification model, the predicted probability is determined by different versions of the alert classification model. Therefore, the viewer displayed on the user terminal 12 can filter the prediction results of the predicted probability for the alert by the version of the alert classification model.
[0077] As described above, the alert priority prediction device 10 of this embodiment uses non-product-independent information using the source code as learning input data for the alert classification model. This increases the versatility of the alert classification model, enabling the prediction of alert priorities for different source codes used in different product projects with a single alert classification model.
[0078] In other words, without using non-dependent information, it would be necessary to generate an alert classification model for each product project. On the other hand, by generating the alert classification model using non-dependent information as in this embodiment, it is not necessary to generate an alert classification model for each product project, thereby minimizing the cost of generating and operating the alert classification model. In addition, the number of training data sets used to generate the alert classification model increases, which, as expected, improves performance.
[0079] Although the present disclosure has been described using the above embodiments, the technical scope of the present disclosure is not limited to the range described in the above embodiments. Various changes or improvements may be made to the above embodiments without departing from the gist of the disclosure, and such changes or improvements are also included in the technical scope of the present disclosure.
[0080] In the above embodiments, the property information of the static analysis tool 20 is used to derive features, but the present disclosure is not limited thereto. As in Fig.5, the feature information of the static analysis tool 20 may not be used to derive features for each product project 1 to N. In this case, the alert classification model is trained using the features for each product project 1 to N and the feature information of the static analysis tool 20.
[0081] In the above embodiments, source code metrics are used to derive features, but the present disclosure is not limited thereto. Features may be derived using the source code itself instead of the source code metrics.
[0082] In the above embodiments, features are derived based on the learning input data, and the features are used to train the alert classification model, but the present disclosure is not limited thereto. The alert classification model may be trained using the learning input data itself instead of the features.
[0083] In the above embodiments, the warning label for which the predicted probability is obtained by the warning classification model is described as binary (TP and FP) or ternary (correction required warning, ignored warning, and FP), but the present disclosure is not limited thereto. The TP may be divided into three or more types, and the warning labels may be four or more values.
[0084] In the above embodiments, the non-dependent information is described as information related to the static analysis tool 20, but the present disclosure is not limited thereto. The non-dependent information may be other information that does not depend on the product used to develop the source code. For example, information related to the source code metrics measurement tool 22 (such as the name and version of the tool) may be used in addition to the information related to the static analysis tool 20. Controls and methods described in the present disclosure may be implemented by a special-purpose computer created by configuring a memory and a processor programmed to perform one or more particular functions embodied in computer programs.Alternatively, the controls and methods described in the present disclosure may be implemented by a special-purpose computer created by configuring a processor provided by one or more special-purpose hardware logic circuits. Alternatively, the controls and methods described in the present disclosure may be implemented by one or more special-purpose computers created by configuring a combination of memory and a processor programmed to perform one or more particular functions, and a processor provided by one or more hardware logic circuits. The computer programs may be stored in a tangible, non-transitory computer-readable medium as instructions that are executed by a computer. QUOTES CONTAINED IN THE DESCRIPTION
[0000] This list of documents submitted by the applicant was generated automatically and is included solely for the convenience of the reader. This list is not part of the German patent or utility model application. The DPMA assumes no liability for any errors or omissions. Cited patent literature
[0000] JP 2017-204090 A1
[0005]
Claims
[1] An information processing device (10) that predicts a priority of a warning indicated by an analysis result of a source code by a static analysis tool (20), the information processing device comprising: a learning unit (26) configured to generate a prediction model, the prediction model being trained using machine learning with learning input data including the source code, warning information including a location of a warning in the source code indicated by the analysis result, product-specific information using the source code, and non-dependent information that does not depend on the product, and using a warning label associated with each warning as training data, the analysis result of the static analysis tool for the source code serving as input and the priority of the warning as output; a receiving unit (28) configured to receive the analysis result of the static analysis tool, which is a target of the prediction by the prediction model; and a prediction unit (30) configured to predict the priority of the warning indicated by the analysis result received by the receiving unit using the prediction model. [2] The information processing apparatus according to claim 1, wherein the non-dependent information includes information regarding the type and settings of the static analysis tool used to analyze the source code. [3] The information processing apparatus according to claim 2, wherein the non-dependent information includes property information of the static analysis tool according to the type of the static analysis tool. [4] The information processing apparatus according to claim 3, wherein the characteristic information includes the determinability of the static analysis and / or the benchmark evaluation result of the static analysis tool. [5] An information processing apparatus according to claim 1 or claim 2, wherein the priority of the warning is the probability of belonging to a specific warning label. [6] The information processing apparatus according to claim 5, wherein the warning labels are multi-valued and include at least a true positive indicating a location with an actual violation and a false positive indicating a location incorrectly warned of as having a violation. [7] The information processing apparatus according to claim 6, wherein the true positive is divided into a warning requiring correction of the source code and a warning not requiring correction of the source code. [8] The information processing apparatus according to claim 1 or claim 2, wherein the prediction unit inputs the analysis result of the source code, the source code, product-specific information using the source code, and non-dependent information into the prediction model to predict the priority of the warning indicated by the analysis result. [9] An information processing apparatus according to claim 1 or claim 2, wherein the prediction model is generated by a machine learning model capable of ensemble learning. [10] A priority prediction method that predicts a priority of a warning indicated by an analysis result of a source code by a static analysis tool, the priority prediction method comprising: a first step of generating, by a learning unit, a prediction model trained using machine learning with learning input data including the source code, warning information including a location of a warning in the source code indicated by the analysis result, product-specific information using the source code, and non-dependent information that does not depend on the product, and using a warning label associated with each warning as training data, wherein the analysis result of the static analysis tool for the source code serves as input and the priority of the warning serves as output; a second step of receiving, by a receiving unit, the analysis result of the static analysis tool, which is a target of prediction by the prediction model; and a third step of predicting, by a predicting unit, the priority of the warning indicated by the analysis result received by the receiving unit using the predictive model. [11] A warning priority prediction program that causes a computer included in an information processing device to predict a priority of a warning indicated by an analysis result of a source code by a static analysis tool, functions as: a learning unit configured to generate a prediction model trained using machine learning with learning input data including the source code, warning information including a location of a warning in the source code indicated by the analysis result, product-specific information using the source code, and non-dependent information that does not depend on the product, and using a warning label associated with each warning as training data, wherein the analysis result of the static analysis tool for the source code serves as input and the priority of the warning serves as output; a receiving unit configured to receive the analysis result of the static analysis tool, which is a target of prediction by the prediction model; and a prediction unit configured to predict the priority of the warning indicated by the analysis result received by the receiving unit using the prediction model.
Citation Information
Patent Citations
Analysis device for static analysis result of source code, and analysis method
JP2017204090A