Information processing device, warning priority prediction method, and storage medium storing warning priority prediction program

By generating a warning classification model and using source code and non-dependent information for machine learning, the problem of low priority sorting of static analysis tools in different product projects is solved, and high-precision warning priority prediction is achieved, reducing development costs.

CN120234804APending Publication Date: 2025-07-01DENSO CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411879084.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-28
Filing Date
2024-12-19
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

In the prior art, static analysis tools have low priority sorting of source code analysis results for different product projects and are high in development costs, making it difficult to distinguish between true positive and false positive warnings with high accuracy.

Method used

The prediction model generated by machine learning is adopted, using source code, product inherent information and non-dependent information as input data, and generate warning classification models through machine learning to predict warning priority in static analysis results, and use prediction models to predict warning priority.

Benefits of technology

It realizes high-precision warning priority prediction of source code analysis results of different product projects, reduces development costs, and improves the accuracy and efficiency of warning classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234804A_ABST
    Figure CN120234804A_ABST
Patent Text Reader

Abstract

The invention provides an information processing apparatus, a warning priority prediction method, and a storage medium storing a warning priority prediction program. A learning unit (26) uses, as input data for learning, source code, warning information including a position at which a warning is generated in the source code indicated by a static analysis result, product-specific information using the source code, and non-dependent information including product-independent information. And a processing unit (28) that performs machine learning using a warning tag associated with each warning as teacher data, generates a warning classification model in which an input is a static analysis result for the source code and an output is a priority of the warning, and receives a static analysis result as a prediction target of the warning classification model. The prediction API (30) uses a warning classification model to predict the priority of the warning indicated by the static analysis result received by the reception unit (28).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an information processing apparatus, a warning priority prediction method, and a warning priority prediction program. Background Art

[0002] Conventionally, static analysis tools have been used to analyze violations in source code. The analysis results of static analysis tools contain many warnings, but the warnings mix true positive warnings and false positive warnings. A true positive is a warning indicating a location where a violation truly exists, and a false positive is a warning that wrongly warns a location where there is no violation.

[0003] Therefore, the developer of the source code needs to distinguish true positive warnings and false positive warnings from the many warnings included in the analysis results of the static analysis tool by themselves.

[0004] Therefore, Patent Document 1 describes an analysis apparatus that performs priority ranking of static analysis results related to the source code to be analyzed based on the judgment result of the user on whether to correct the source code corresponding to the static analysis result, information on source code metrics, and information on a program development project. In addition, the judgment result of the above user refers to the judgment result of true positive and false positive of the warning. A true positive is a location where a violation truly exists, and a false positive is a location where a warning is wrongly given to a location where there is no violation. The judgment result of the user indicating true positive and false positive is given to each warning as a warning label.

[0005] Patent Document 1: Japanese Unexamined Patent Application Publication No. 2017-204090

[0006] Since the analysis apparatus described in Patent Document 1 performs priority ranking of static analysis results based on information unique to a product item, the priority becomes a result dependent on the product item. Therefore, when applying the analysis apparatus of Patent Document 1 to static analysis of source codes of different product items, there is a possibility that the accuracy of priority ranking of static analysis results decreases. In addition, if an analysis apparatus is constructed for each product item, there is a possibility that development costs and the like increase. Summary of the Invention

[0007] Therefore, an object of the present disclosure is to provide an information processing apparatus, a warning priority prediction method, and a warning priority prediction program that can accurately predict the priority of warnings based on static analysis even for different source codes used for each of different product items.

[0008] The present disclosure adopts the following technical units to solve the above problems.

[0009] One aspect of the present disclosure is an information processing apparatus that predicts the priority of warnings shown in the analysis results of a static analysis tool for source code. The information processing apparatus includes: a learning unit that uses the source code, warning information including the generation positions of the warnings in the source code shown in the analysis results, product-specific information using the source code, and non-dependent information including information not dependent on the product as input data for learning, and performs machine learning using warning labels associated with each of the warnings as teacher data to generate a prediction model that takes as input the analysis results of the static analysis tool for the source code and outputs the priority of the warnings; a reception unit that receives the analysis results of the static analysis tool that are the prediction targets of the prediction model; and a prediction unit that uses the prediction model to predict the priority of the warnings shown in the analysis results received by the reception unit.

[0010] According to this configuration, the determination of the priority of warnings shown in the analysis results of a static analysis tool for source code uses a prediction model based on machine learning with warning labels associated with each warning as teacher data. The warning labels are labels attached to each warning through user judgment. Moreover, the input data for learning the prediction model uses non-dependent information not dependent on the product. As a result, the prediction model is not dedicated to a specific source code or product, but is generated as a general-purpose machine learning model. Therefore, this configuration can accurately predict the priority of warnings shown in the static analysis results even for different source codes used in each of different product projects.

[0011] In the above information processing apparatus, the non-dependent information may be information related to the type and settings of the static analysis tool used for the analysis of the source code.

[0012] The performance of static analysis by a static analysis tool used for the analysis of source code varies depending on its type and settings. Therefore, by using information related to the type and settings of the static analysis tool as input data for learning the prediction model, the priority of warnings shown in the static analysis results can be accurately predicted. In addition, different versions of the same type of static analysis tool are also included in the types of static analysis tools. Also, the settings of the static analysis tool include, for example, the standards of programming languages.

[0013] In the above information processing apparatus, the non-dependent information may include characteristic information of the static analysis tool corresponding to the type of the static analysis tool. According to this configuration, by using the characteristic information of the static analysis tool as input data for learning the prediction model, warnings based on static analysis can be predicted with higher accuracy.

[0014] In the above information processing apparatus, the characteristic information may also include at least one of the decidability of static analysis and the benchmark evaluation result of the static analysis tool. According to this configuration, it is possible to predict the priority of a warning shown by a static analysis result with high accuracy.

[0015] In the above information processing apparatus, the priority of the warning may also be the probability of belonging to a specific warning label. According to this configuration, the user can easily identify the position of the source code that needs to be corrected.

[0016] In the above information processing apparatus, the warning label may also be multi-valued and include at least a true positive indicating the position where a warning actually exists and a false positive that wrongly warns a position where there is no violation. According to this configuration, the user can easily identify the position of the source code that needs to be corrected.

[0017] In the above information processing apparatus, the true positive may also be divided into a warning that requires correction of the source code or a warning that does not require correction of the source code. According to this configuration, the user can easily identify the position of the source code that needs to be corrected.

[0018] In the above information processing apparatus, the prediction unit may also input the analysis result of the source code, the source code, information inherent to the product using the source code, and the non-dependent information into the prediction model to predict the priority of the warning shown by the analysis result.

[0019] In the above information processing apparatus, the prediction model may also be generated by a machine learning model capable of performing ensemble learning.

[0020] A priority prediction method according to an aspect of the present disclosure is a priority prediction method for predicting the priority of a warning shown by an analysis result of a static analysis tool for source code, and includes: a first step in which a learning unit uses the source code, warning information including the generation position of a warning in the source code shown by the analysis result, information inherent to a product using the source code, and non-dependent information including information not dependent on the product as learning input data, and performs machine learning using a warning label associated with each warning as teacher data to generate a prediction model that takes as input the analysis result of the source code by the static analysis tool and outputs the priority of the warning; a second step in which a reception unit receives the analysis result of the static analysis tool that is the prediction target of the prediction model; and a third step in which a prediction unit uses the prediction model to predict the priority of the warning shown by the analysis result received by the reception unit.

[0021] One aspect of the warning priority prediction program of the present disclosure causes a computer included in an information processing apparatus that shows information on the priority of warnings in the analysis result of a source code by a prediction static analysis tool to function as a learning unit, a reception unit, and a prediction unit. The learning unit uses the source code, warning information including the generation positions of the warnings in the source code shown in the analysis result, product-specific information using the source code, and non-dependent information including information independent of the product as learning input data, and performs machine learning using warning labels associated with each of the warnings as teacher data to generate a prediction model that takes the analysis result of the source code by the static analysis tool as input and outputs the priority of the warnings. The reception unit receives the analysis result of the static analysis tool that is the prediction target of the prediction model, and the prediction unit uses the prediction model to predict the priority of the warnings shown in the analysis result received by the reception unit.

[0022] According to the present disclosure, even for different source codes used for each of different product items, it is possible to accurately predict the priority of warnings based on static analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 It is a diagram showing the functional configuration of a warning priority prediction device according to an embodiment.

[0024] Figure 2 It is a schematic diagram showing an overview of a warning classification model according to an embodiment.

[0025] Figure 3 It is a schematic diagram showing a generation process of a warning classification model using information of multiple product items according to an embodiment.

[0026] Figure 4 It is a schematic diagram showing a warning priority determination process using a warning classification model according to an embodiment.

[0027] Figure 5 It is a schematic diagram showing a generation process of a warning classification model using information of multiple product items according to another embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0028] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. In addition, the embodiments described below show an example of a case where the present disclosure is implemented, and do not limit the present disclosure to the specific configurations described below. When implementing the present disclosure, specific configurations corresponding to the embodiments can be appropriately adopted.

[0029] Figure 1This is a diagram showing the functional configuration of the warning priority prediction device 10 of the present embodiment. The warning priority prediction device 10 of the present embodiment is an information processing device equipped with an arithmetic device such as a CPU (Central Processing Unit), and can be accessed via a web browser between the user terminal 12.

[0030] The warning priority prediction device 10 classifies the warnings shown in the analysis result (also referred to as the static analysis result) of the source code by the static analysis tool 20 according to the prediction model generated by machine learning, and predicts the priority of the warnings. Hereinafter, the prediction model will be referred to as the warning classification model.

[0031] The user terminal 12 is, for example, an information processing device such as a laptop or desktop personal computer, and can access the warning priority prediction device 10. Users such as software developers create the source code of the software used in the product through the user terminal 12.

[0032] The user appropriately sends the source code from the user terminal 12 to the warning priority prediction device 10 according to the progress of the creation. The warning priority prediction device 10 analyzes the received source code by the static analysis tool 20 and outputs its static analysis result.

[0033] The static analysis result includes warnings indicating the locations of violations generated in the source code. Among these warnings, there are true positives (also referred to as TPs) that warn of the locations where violations actually exist and false positives (also referred to as FPs) that erroneously warn of locations where there are no violations. Therefore, it is necessary to determine whether each warning is a TP or an FP.

[0034] Therefore, the warning priority prediction device 10 of the present embodiment predicts the priority order of the warnings that the user should address by classifying the warnings shown in the static analysis result into TPs or FPs using the warning classification model. Then, the warning priority prediction device 10 sends the prediction result to the user terminal 12.

[0035] In addition, TPs can be further subdivided into warnings that require source code correction (Confirmed), that is, warnings that require correction, or warnings that do not require source code correction (Intentional), that is, warnings that can be escaped. Therefore, the warning priority prediction device 10 can also further classify TPs into warnings that require correction and warnings that can be escaped to predict the priority order of the warnings.

[0036] That is, the prediction based on the warning classification model in the present embodiment has a binary classification prediction that classifies the warnings shown in the static analysis result into TP and FP binary values and outputs the probabilities belonging to the respective binary values, and a ternary classification prediction that classifies the warnings shown in the static analysis result into three values: warnings that require correction, warnings that can be detached, and FP, and outputs the probabilities belonging to the respective three values. Whether to make the prediction based on the warning classification model a binary classification or a ternary classification can be determined according to the setting of the warning priority prediction device 10, or the warning priority prediction device 10 may only have the function of either one.

[0037] In the binary classification, for example, the probability that a warning will be TP is predicted as a TP probability of 90%, and the probability that it will be FP is predicted as an FP probability of 10%. In addition, in the ternary classification, for example, the probability that a warning will be a warning that requires correction is predicted as a probability of requiring correction of 70%, the probability that it will be a detachable warning is predicted as a detachable probability of 20%, and the FP probability is 10%.

[0038] In this way, the warning priority prediction device 10 outputs the probability belonging to a specific warning label as the warning priority. Moreover, the specific warning label is a multi-value that includes at least TP and FP. That is, a warning with a high TP probability or a probability of requiring correction is a warning with a high priority. Thus, the user can easily identify the location of the source code that needs to be corrected and can preferentially correct the code corresponding to the warning determined to be a high-probability TP, etc.

[0039] In addition, warning labels such as TP, FP, warnings that require correction, and detachable warnings are labels attached to each warning. In the following description, the TP probability, FP probability, probability of requiring correction, and detachable probability predicted by the warning classification model are collectively referred to as prediction probabilities.

[0040] As Figure 1 shown, the warning priority prediction device 10 of the present embodiment, together with the static analysis tool 20, includes a source code metric measurement tool 22, a database 24, a learning unit 26, a reception unit 28, a prediction API (Application Programming Interface) 30, and an output unit 32.

[0041] As described above, the static analysis tool 20 performs static analysis on the source code received from the user terminal 12 and outputs the static analysis result. The static analysis result attaches a warning ID to each warning in order to identify each warning generated in the source code. In addition, the static analysis result is associated with the source code to be analyzed and stored in the database 24.

[0042] The source code metric measurement tool 22 measures the metrics of the source code received from the user terminal 12. The source code metrics are associated with the source code for which the measurement has been performed and stored in the database 24.

[0043] The database 24 stores the input data (also referred to as learning input data) and teacher data used for the learning (training) of the warning classification model which is a machine learning model, the classification results of the warnings using the warning classification model, and so on.

[0044] The learning unit 26 generates a warning classification model through machine learning using the learning input data and teacher data stored in the database 24. In addition, when various data stored in the database 24 are appended or updated, the learning unit 26 appropriately updates the warning classification model.

[0045] The reception unit 28 receives the static analysis result which is the prediction object of the warning classification model. This static analysis result is the analysis result of the source code by the static analysis tool 20, and the warnings are not classified. In addition, the reception unit 28 receives various data required for determining the priority of the warnings in addition to the unclassified warnings shown in the static analysis result, and derives their feature quantities.

[0046] The prediction API 30 deploys the warning classification model learned (trained) by the learning unit 26, and determines the priority of the warnings shown in the static analysis result received by the reception unit 28 using the warning classification model. That is, the prediction API 30 functions as a classifier for classifying the warnings by associating the prediction probability indicating the priority of the warnings shown in the static analysis result with the warning ID and outputting them.

[0047] The output unit 32 sends the prediction result based on the warning classification model in the prediction API 30, that is, the prediction probability for each warning, to the user terminal 12.

[0048] The user terminal 12 displays the prediction probability for each warning output by the warning classification model on the viewer. The viewer can sort the warnings according to the TP probability. Thus, the user can confirm the priority of the warnings that should be processed through the viewer and can easily identify the warnings that should be dealt with preferentially.

[0049] In addition, when the prediction probability received from the warning priority prediction device 10 is appropriate, the user attaches a warning label based on the prediction probability to the warning ID and sends it to the warning priority prediction device 10. On the other hand, when the prediction probability received from the warning priority prediction device 10 is inappropriate, the user attaches a warning label considered appropriate by the user himself / herself to the warning ID and sends it to the warning priority prediction device 10. The warning priority prediction device 10 associates the warning ID and the warning label received from the user terminal 12 with the static analysis result that is the prediction object of the warning classification model and stores them in the database 24. Thus, learning input data is accumulated for the warning classification model.

[0050] In this way, the warning priority prediction device 10 of the present embodiment collaborates with the static analysis tool 20 to apply the warning classification model. Therefore, it is possible to collect labeled training data (teacher data) without the need for a user or the like to perform the operation of attaching annotations for learning the warning classification model by machine learning.

[0051] In addition, although the warning priority prediction device 10 of the present embodiment is in a form having a learning unit 26, the function of the learning unit 26 may also be provided by other information processing devices.

[0052] Figure 2 It is a schematic diagram showing an outline of the warning classification model of the present embodiment.

[0053] The warning classification model of the present embodiment is a model that, when a warning shown in the static analysis result is input, generates a large number of weak learners that predict warning labels through branches based on various conditions of the input data, and aggregates the prediction results of the probabilities of these weak learners for the warning labels into one prediction value by methods such as majority voting or weighting and outputs it. Therefore, the warning classification model of the present embodiment is generated by supervised learning using a machine learning model capable of performing ensemble learning, such as LightGBM.

[0054] The learning input data for generating the warning classification model of the present embodiment by machine learning is feature quantities based on source code, warning information, product item information, and non-dependent information, and the teacher data is warning labels attached to each warning shown in the static analysis result according to the user's judgment.

[0055] The source code becomes the object of analysis by the static analysis tool 20. In addition, instead of the source code itself, the learning input data may use the number of lines of code, loop complexity, etc. included in the source code metrics measured by the source code metrics measurement tool 22.

[0056] Here, there are cases where the trends of source code vary for each product. Therefore, by using the source code or source code metrics as learning input data, it is possible to learn the trends of different source codes for product projects.

[0057] The warning information is information including the generation location of warnings in the source code shown by the static analysis results, and is associated with the warning labels as teacher data.

[0058] The product project information is information inherent to the product using the source code, such as the name and model of the product. In addition, the product project information may include not only the name of the product, etc., but also other information indicating the characteristics of the product project. By using the product project information as learning input data, it is possible to learn the trends of warning labels for each product.

[0059] The non-dependent information is information including information that does not depend on the product. By using the non-dependent information that does not depend on the product as learning input data, the warning classification model can be generated as a machine learning model that is not dedicated to a specific source code or product and has generality.

[0060] The non-dependent information of the present embodiment is information related to the type and settings of the static analysis tool 20 used for the analysis of the source code. The type of the static analysis tool 20 is, for example, the name, product number, and version of the static analysis tool 20. The settings of the static analysis tool 20 are, for example, the standard of the programming language used for the production of the source code. For the standard of the programming language, for example, even for the same C language, there is whether it is based on the C99 standard or the like. In addition, hereinafter, the information related to the type and settings of the static analysis tool 20 will be referred to as tool information.

[0061] Here, the performance of the static analysis performed by the static analysis tool 20 used for the analysis of the source code varies depending on its type and settings. Therefore, by using the information related to the type and settings of the static analysis tool 20 as learning input data, it is possible to accurately predict the warning labels associated with the warnings shown by the static analysis results regardless of the static analysis tool 20.

[0062] In addition, the non-dependent information also includes the characteristic information of the static analysis tool 20. The characteristic information of the static analysis tool 20 includes at least one of the decidability of the static analysis and the reference evaluation result of the static analysis tool 20. In addition, the decidability of the static analysis is the decidability of the static analysis of regulations such as MISRA corresponding to the warnings output by the static analysis tool 20, and shows "decidable", "undecidable", and "uncertain" for each warning. The decidability and the reference evaluation result are, for example, public information of the manufacturer of the static analysis tool 20, generally known information, information based on past empirical rules, etc.

[0063] The characteristic information of the static analysis tool 20 varies according to the type and version of the static analysis tool 20, and is associated with the type and version of the static analysis tool 20 and stored in the database 24. By using the characteristic information of the static analysis tool 20 as the learning input data of the warning classification model, it is possible to predict the warnings of the static analysis with higher accuracy.

[0064] Moreover, the learning unit 26 of the present embodiment derives feature amounts based on the source code, warning information, product item information, and non-dependency information. An example of the derived feature amounts is as follows. In addition, the derived feature amounts are not limited to the following, and it is not necessary to use all of the following feature amounts.

[0065] · The name (URL) of the source code library that manages the source code where the warning occurred

[0066] · The path to the source code file where the warning occurred

[0067] · The name of the source code file where the warning occurred

[0068] · The function name in the source code where the warning occurred

[0069] · The line number of the source code where the warning occurred

[0070] · The ID indicating the type of warning

[0071] · The severity of the warning

[0072] · The name of the static analysis tool 20 that caused the warning

[0073] · The version of the static analysis tool 20 that caused the warning

[0074] · The decidability of the static analysis of regulations such as MISRA corresponding to the warning

[0075] · The evaluation result of the static analysis tool 20 on the criteria of regulations such as MISRA corresponding to the warning

[0076] · The number of warnings generated in the function in the source code where the warning occurred

[0077] · The number of warnings generated in the source code file where the warning occurred

[0078] · The time elapsed since the warning occurred

[0079] · The source code metrics at the file level of the source code file where the warning occurred

[0080] · The source code metrics at the function level of the source code file where the warning occurred

[0081] · The name of the product item where the warning occurred

[0082] The learning unit 26 of this embodiment generates a warning classification model by performing machine learning using the above feature amounts as learning input data and using warning labels attached to each warning shown in the static analysis result used for deriving the feature amounts based on the user's judgment as teacher data. Such a combination of feature amounts and warning labels forms a training data set. That is, the larger the number of training data sets, the higher the accuracy of the warning classification model.

[0083] Moreover, the prediction API 30 inputs the static analysis result of the source code, the source code, product item information, and non-dependency information into the warning classification model, and outputs the prediction probability of the warning label indicating the priority of the warning shown in the static analysis result.

[0084] Figure 3 It is a schematic diagram showing the generation process of the warning classification model that uses information of a plurality of product items 1 to N in this embodiment. The generation process of the warning classification model is executed by a program stored in a storage medium such as a storage unit included in the warning priority prediction device 10. In addition, by executing this program, the method corresponding to the program is executed.

[0085] As Figure 3 shown, the learning unit 26 derives feature amounts as learning input data based on source code metrics, static analysis results, product item information, tool information, and characteristic information of the static analysis tool 20 for each of the product items 1 to N.

[0086] The characteristic information of the static analysis tool 20 is stored in the database 24 as common information regardless of the product item. Moreover, the characteristic information of the static analysis tool 20 corresponding to the name and version of the static analysis tool 20 shown in the tool information is read out from the database 24.

[0087] In addition, warning labels attached to the warnings shown in the static analysis results of each source code of the product items 1 to N are used as teacher data.

[0088] The combination of the feature amounts and the teacher data for each of the product items 1 to N, that is, the learning data set, is stored in the database 24. Whenever the source code of each of the product items 1 to N is corrected or improved, a learning data set is generated and stored in the database 24, and is used for subsequent learning of the warning classification model. Moreover, the learning unit 26 generates a warning classification model by performing machine learning using a large number of learning data sets.

[0089] Figure 4It is a schematic diagram showing the warning priority determination process using the warning classification model of the present embodiment. The warning priority determination process is executed by a program stored in a storage medium such as a storage unit included in the warning priority prediction device 10. In addition, by executing this program, a method corresponding to the program is executed.

[0090] When the user sends the source code to the warning priority prediction device 10 via the user terminal 12 and gives an execution instruction for static analysis, the warning priority determination process is performed. According to this execution instruction, first, the static analysis tool 20 executes the static analysis of the source code and outputs the static analysis result. In addition, the source code metric measurement tool 22 measures the source code metrics of the source code.

[0091] Then, the reception unit 28 reads out the product item information, tool information, and the characteristic information of the static analysis tool 20 corresponding to the source code from the database 24. The reception unit 28 derives feature quantities based on the read information, the static analysis result, and the source code metrics. Among these feature quantities, there are warnings whose static analysis results, that is, warning labels, are unknown.

[0092] By inputting the feature quantities into the warning classification model, the warning classification model outputs the prediction probability of the warning as the prediction result. The prediction probability of the warning is stored in the database 24 and is sent to the user terminal 12. The user confirms the position of the source code that needs to be corrected by checking the prediction probability of the warning sent to the user terminal 12 in the viewer.

[0093] In addition, as described above, when the prediction probability is appropriate, the user attaches a warning label based on the prediction probability to the warning ID and sends it to the warning priority prediction device 10. When the prediction probability is inappropriate, the user attaches a warning label considered appropriate by oneself to the warning ID and sends it to the warning priority prediction device 10. The warning priority prediction device 10 associates the warning ID and the warning label received from the user terminal 12 with the static analysis result and stores them in the database 24. In this way, the warning labels stored in the database 24 are used as teacher data.

[0094] In addition, in Figure 4 the prediction object of the warning classification model is the static analysis result of the source code of product items 1 to N, but the static analysis result that is the prediction object of the warning classification model can also be the static analysis result of the source code other than product items 1 to N.

[0095] In addition, the warning priority prediction device 10 of the present embodiment can also update the warning classification model by machine learning using the learning data set newly stored in the database 24. The timing of the update is, for example, the timing when the deterioration of the performance of the warning classification model appears, or the timing when a new learning data set of a specified amount or more has been accumulated. The deterioration of the performance of the warning classification model is, for example, the case where the frequency of the user attaching a warning label different from the prediction probability of the warning classification model is above a specified value.

[0096] In addition, by updating the warning classification model, it is possible to determine the prediction probability by different versions of the warning classification model. Therefore, in the viewer displayed on the user terminal 12, it is possible to filter the prediction results of the prediction probability of the warning according to the version of the warning classification model.

[0097] As described above, the warning priority prediction device 10 of the present embodiment uses non-dependent information that does not depend on the product using the source code as the learning input data for the warning classification model. Thereby, the versatility of the warning classification model is improved, and it is possible to predict the priority of warnings by one warning classification model even for different source codes used for each different product project.

[0098] That is, in the case of generating a warning classification model without using non-dependent information, it is necessary to generate a warning classification model for each product project. On the other hand, by generating a warning classification model using non-dependent information as in the present embodiment, it is not necessary to generate a warning classification model for each product project, and the generation and operation costs of the warning classification model can be minimized. In addition, since the number of learning data sets used for generating the warning classification model also increases, performance improvement can be further expected.

[0099] As described above, the present disclosure has been described using the above embodiment, but the technical scope of the present disclosure is not limited to the scope described in the above embodiment. Various changes or improvements can be made to the above embodiment without departing from the gist of the disclosure, and the modified or improved manner is also included in the technical scope of the present disclosure.

[0100] In the above embodiment, a method of deriving feature amounts using the characteristic information of the static analysis tool 20 has been described, but the present disclosure is not limited thereto. As Figure 5 shown, it is also possible not to use the characteristic information of the static analysis tool 20 for deriving the feature amounts of each product project 1 to N. In this case, the warning classification model is learned by the feature amounts of each product project 1 to N and the characteristic information of the static analysis tool 20.

[0101] In the above-described embodiment, the method of deriving feature quantities using source code metrics has been described, but the present disclosure is not limited thereto. Feature quantities may be derived using the source code itself instead of source code metrics.

[0102] In the above-described embodiment, the method of deriving feature quantities based on learning input data and using the feature quantities for learning a warning classification model has been described, but the present disclosure is not limited thereto. The warning classification model may be learned without using feature quantities and using the learning input data itself.

[0103] In the above-described embodiment, the method of making the warning labels for which prediction probabilities are obtained by the warning classification model be binary values of TP and FP, or be three values of warning requiring correction, warning escapable, and FP has been described, but the present disclosure is not limited thereto. TP may be further divided into three or more types, and the warning labels may be four or more values.

[0104] In the above-described embodiment, the method of making the non-dependent information be information related to the static analysis tool 20 has been described, but the present disclosure is not limited thereto. The non-dependent information may be other information as long as it is information that does not depend on the product used for the development of the source code. For example, in addition to the information related to the static analysis tool 20, information related to the source code metric measurement tool 22 (name and version of the tool) may also be used. The control unit and its method described in the present disclosure may be implemented by a dedicated computer provided by a processor and a memory configured to execute one or more functions embodied by a computer program. Alternatively, the control unit and its method described in the present disclosure may be implemented by a dedicated computer provided by a processor constituted by one or more dedicated hardware logic circuits. Alternatively, the control unit and its method described in the present disclosure may be implemented by one or more dedicated computers constituted by a combination of a processor programmed to execute one or more functions and a memory and a processor constituted by one or more hardware logic circuits. Further, the computer program may be stored as instructions executable by a computer in a non-transitory tangible recording medium readable by the computer.

Claims

1. An information processing device for predicting the priority of a warning indicated by a static analysis tool in an analysis result of source code, comprising: A learning unit, using the source code, warning information including the location where the warning in the source code shown by the analysis result is generated, information inherent to the product using the source code, and non-dependent information including information not dependent on the product as input data for learning, and using a warning label associated with each of the warnings as teacher data to perform machine learning, and generating a prediction model, the prediction model having input as the analysis result of the static analysis tool on the source code, and output as the priority of the warning; a receiving unit that receives the analysis result of the static analysis tool as a prediction target of the prediction model; and The prediction unit predicts the priority of the warning indicated by the analysis result received by the receiving unit using the prediction model.

2. The information processing device according to claim 1, wherein: The non-dependency information is information related to the type and setting of the static analysis tool used for analyzing the source code.

3. The information processing device according to claim 2, wherein: The non-dependency information includes characteristic information of the static analysis tool corresponding to the type of the static analysis tool.

4. The information processing device according to claim 3, wherein: The characteristic information includes at least one of the determinability of the static analysis and a benchmark evaluation result of the static analysis tool.

5. The information processing device according to claim 1 or 2, wherein: The priority of the above warning is the probability of belonging to a specific above warning label.

6. The information processing device according to claim 5, wherein: The above warning labels are multi-valued, including at least true positives that warn of locations where violations actually exist and false positives that mistakenly warn of locations where there are no violations.

7. The information processing device according to claim 6, wherein: The true positives are classified into warnings that require correction of the source code or warnings that do not require correction of the source code.

8. The information processing device according to claim 1 or 2, wherein: The prediction unit inputs the analysis result of the source code, the source code, information specific to a product using the source code, and the non-dependency information into the prediction model, and predicts a priority of a warning indicated by the analysis result.

9. The information processing device according to claim 1 or 2, wherein: The above prediction model is generated by a machine learning model capable of ensemble learning.

10. A priority prediction method is a priority prediction method for predicting the priority of a warning indicated by a static analysis tool in an analysis result of a source code, comprising: In the first step, the learning unit uses the source code, the warning information including the location where the warning in the source code shown by the analysis result is generated, the information inherent to the product using the source code, and the non-dependent information including the information not dependent on the product as input data for learning, and uses the warning label associated with each of the warnings as teacher data to perform machine learning, and generates a prediction model, the prediction model having the analysis result of the static analysis tool on the source code as input and the priority of the warning as output; In a second step, a receiving unit receives the analysis result of the static analysis tool as a prediction target of the prediction model; and In a third step, a prediction unit predicts a priority of the warning indicated by the analysis result received by the receiving unit using the prediction model.

11. A storage medium storing a warning priority prediction program, wherein: The warning priority prediction program is used to make a computer included in an information processing device that predicts the priority of a warning indicated by a static analysis tool in an analysis result of a source code function as a learning unit, a receiving unit, and a prediction unit. The learning unit uses the source code, warning information including the location where the warning in the source code shown by the analysis result is generated, information inherent to the product using the source code, and non-dependent information including information not dependent on the product as input data for learning, and uses warning labels associated with each of the warnings as teacher data to perform machine learning, and generates a prediction model, the prediction model having input as the analysis result of the source code by the static analysis tool, and output as the priority of the warning. The receiving unit receives the analysis result of the static analysis tool as a prediction target of the prediction model. The prediction unit predicts a warning priority level indicated by the analysis result received by the receiving unit using the prediction model.

Citation Information

Patent Citations

  • Analysis device for static analysis result of source code, and analysis method

    JP2017204090A