Information processing device, warning priority prediction method, and warning priority prediction program

The information processing apparatus uses machine learning to predict warning priorities in static analysis, addressing the challenge of mixed warnings across projects by classifying them accurately and reducing manual effort and costs.

JP2025104891APending Publication Date: 2025-07-10DENSO CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2023223052
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-28
Publication Date
2025-07-10

AI Technical Summary

Technical Problem

Existing static analysis tools produce a mix of true positive and false positive warnings, requiring manual differentiation by developers, which decreases accuracy when applied across different product projects and increases development costs.

Method used

An information processing apparatus using machine learning to predict warning priorities based on source code analysis, incorporating non-dependent information and product-specific data to generate a prediction model that accurately classifies warnings as true positives or false positives, regardless of the product project.

Benefits of technology

Enables accurate prediction of warning priorities across different source codes, reducing manual effort and development costs by automating the classification of true and false positives.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025104891000001_ABST
    Figure 2025104891000001_ABST
Patent Text Reader

Abstract

To provide an information processing device, a warning priority prediction method, and a warning priority prediction program that can accurately predict the priority of a warning from static analysis for different source codes used in different product projects.SOLUTION: In a warning priority prediction apparatus 10 that accesses a user terminal 12, a learning unit is configured to: generate a warning classification model, the warning classification model being trained using machine learning with learning input data including a source code, warning information including occurrence location of warning in the source code indicated by the static analysis result, product-specific information using the source code, and non-dependent information that does not depend on the product, and using a warning label associated with each warning as training data, with the analysis result of the static analysis on the source code as input and the priority of the warning as output. A reception unit 28 is configured to receive the analysis result of the static analysis, which is a target of prediction by the warning classification model. A prediction API 30 is configured to predict the priority of the warning indicated by the analysis result received, using the warning classification model.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing apparatus, a warning priority prediction method, and a warning priority prediction program.

Background Art

[0002] Conventionally, static analysis tools have been used to analyze violations of source code. Although the analysis results by the static analysis tools include a large number of warnings, the warnings include both true positive warnings and false positive warnings. A true positive is a warning indicating a location where there is a genuine violation, and a false positive is a warning that wrongly warns a location where there is no violation.

[0003] Therefore, the developer of the source code has to distinguish between true positive warnings and false positive warnings from among the large number of warnings included in the analysis results of the static analysis tool by themselves.

[0004] Thus, Patent Document 1 describes an analysis apparatus that ranks the static analysis results regarding the source code to be analyzed based on the determination result of whether to correct the source code corresponding to the static analysis result by the user, the information of the source code metrics, and the information of the program development project. Note that the determination result by the user is the determination result of true positive and false positive for the warning. A true positive is a location where there is a genuine violation, and a false positive is a location where a location without a violation is wrongly warned. The determination result of the user indicating this true positive and false positive is given as a warning label for each warning.

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0006] Since the analyzer described in Patent Document 1 prioritizes static analysis results based on information specific to a product project, the priority is a result that depends on the product project. Therefore, when applying the analyzer of Patent Document 1 to the static analysis of source codes of different product projects, the accuracy of prioritizing static analysis results may decrease. In addition, if an analyzer is constructed for each product project, development costs and the like may increase.

[0007] Therefore, an object of the present invention is to provide an information processing apparatus, a warning priority prediction method, and a warning priority prediction program that can accurately predict the priority of warnings by static analysis even for different source codes used for different product projects.

Means for Solving the Problems

[0008] The present invention employs the following technical means to solve the above problems. The claims and the reference numerals in parentheses described in this section are an example showing the correspondence relationship with the specific means described in the embodiments described later as one aspect, and do not limit the technical scope of the present invention.

[0009] An information processing apparatus according to an aspect of the present invention is an information processing apparatus (10) that predicts the priority of warnings indicated by the analysis results of source codes by a static analysis tool (20). As learning input data, the source code, warning information including the locations where warnings occur in the source code indicated by the analysis results, product-specific information using the source code, and non-dependent information including information that does not depend on the product are used. A learning unit (26) that performs machine learning using warning labels associated with each warning as teacher data to generate a prediction model in which the input is the analysis result of the static analysis tool for the source code and the output is the priority of the warning, a reception unit (28) that receives the analysis result of the static analysis tool to be predicted by the prediction model, and a prediction unit (30) that predicts the priority of the warning indicated by the analysis result received by the reception unit using the prediction model.

[0010] According to this configuration, a prediction model based on machine learning using warning labels associated with each warning as teacher data is used to determine the priority of warnings indicated by the analysis results of source code by a static analysis tool. The warning labels are assigned to each warning according to the user's judgment. Then, non-dependent information that does not depend on the product is used as input data for learning the prediction model. As a result, the prediction model is generated as a machine learning model with generality, rather than being specialized for specific source code or products. Therefore, this configuration can accurately predict the priority of warnings indicated by the static analysis results for different source codes used for different product projects.

[0011] In the information processing apparatus described above, the non-dependent information may be information regarding the type and settings of the static analysis tool used for the analysis of the source code.

[0012] The static analysis tools used for the analysis of source code have different static analysis performances depending on their types and settings. Therefore, by using information regarding the type and settings of the static analysis tool as input data for learning the prediction model, the priority of warnings indicated by the static analysis results can be accurately predicted. Note that the types of static analysis tools include different versions of the same type of static analysis tool. Also, the settings of the static analysis tool include, for example, the programming language specifications.

[0013] In the information processing apparatus described above, the non-dependent information may include characteristic information of the static analysis tool corresponding to the type of the static analysis tool. According to this configuration, by using the characteristic information of the static analysis tool as input data for learning the prediction model, warnings by the static analysis can be predicted with higher accuracy.

[0014] In the information processing apparatus described above, the characteristic information may include at least one of the determinability of the static analysis and the benchmark evaluation result of the static analysis tool. According to this configuration, the priority of warnings indicated by the static analysis results can be accurately predicted.

[0015] In the above information processing apparatus, the priority of the warning may be the probability of belonging to a specific warning label. According to this configuration, the user can easily recognize the location of the source code that needs to be corrected.

[0016] In the above information processing apparatus, the warning label may be a multi-value including at least a true positive that warns a location with a true violation and a false positive that wrongly warns a location without a violation. According to this configuration, the user can easily recognize the location of the source code that needs to be corrected.

[0017] In the above information processing apparatus, the true positive may be divided into a warning that requires correction of the source code or a warning that does not require correction of the source code. According to this configuration, the user can easily recognize the location of the source code that needs to be corrected.

[0018] In the above information processing apparatus, the prediction unit may input the analysis result of the source code, the source code, product-specific information using the source code, and the non-dependent information into the prediction model, and predict the priority of the warning indicated by the analysis result.

[0019] In the above information processing apparatus, the prediction model may be generated by a machine learning model capable of ensemble learning.

[0020] A priority prediction method according to one aspect of the present invention is a priority prediction method for predicting the priority of warnings indicated by the analysis result of source code by a static analysis tool. Using the source code as learning input data, warning information including the locations where warnings occur in the source code indicated by the analysis result, product-specific information using the source code, and non-dependent information including information not dependent on the product, and machine learning using warning labels associated with each warning as teacher data, a first step in which a learning unit generates a prediction model with the input being the analysis result of the static analysis tool for the source code and the output being the priority of the warning; a second step in which a reception unit receives the analysis result of the static analysis tool to be predicted by the prediction model; and a third step in which a prediction unit predicts the priority of the warning indicated by the analysis result received by the reception unit using the prediction model.

[0021] A warning priority prediction program according to one aspect of the present invention causes a computer included in an information processing apparatus that predicts the priority of warnings indicated by the analysis result of source code by a static analysis tool to function as a learning unit that generates a prediction model using the source code as learning input data, warning information including the locations where warnings occur in the source code indicated by the analysis result, product-specific information using the source code, and non-dependent information including information not dependent on the product, and machine learning using warning labels associated with each warning as teacher data, with the input being the analysis result of the static analysis tool for the source code and the output being the priority of the warning; a reception unit that receives the analysis result of the static analysis tool to be predicted by the prediction model; and a prediction unit that predicts the priority of the warning indicated by the analysis result received by the reception unit using the prediction model.

Advantages of the Invention

[0022] According to the present invention, it is possible to accurately predict the priority of warnings by static analysis even for different source codes used in different product projects.

Brief Description of the Drawings

[0023]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Mode for Carrying Out the Invention

[0024] Hereinafter, embodiments of the present invention will be described with reference to the drawings. Note that the embodiments described below are examples of cases where the present invention is implemented, and the present invention is not limited to the specific configurations described below. In implementing the present invention, specific configurations according to the embodiments may be appropriately adopted.

[0025] FIG. 1 is a diagram showing the functional configuration of the warning priority prediction device 10 of the present embodiment. The warning priority prediction device 10 of the present embodiment is an information processing device including an arithmetic device such as a CPU (Central Processing Unit), and can be accessed through a web browser between the user terminal 12.

[0026] The warning priority prediction device 10 classifies warnings indicated by the analysis result of the source code by the static analysis tool 20 (hereinafter referred to as "static analysis result") using a prediction model generated by machine learning, and predicts the priority of the warnings. Hereinafter, the prediction model is referred to as a warning classification model.

[0027] The user terminal 12 is an information processing device such as a laptop or desktop personal computer, and is capable of accessing the warning priority prediction device 10. A user such as a software developer creates the source code of the software to be used in the product using the user terminal 12.

[0028] The user appropriately transmits the source code from the user terminal 12 to the warning priority prediction device 10 according to the progress of the creation. The warning priority prediction device 10 analyzes the received source code using the static analysis tool 20 and outputs the static analysis result.

[0029] The static analysis result includes warnings indicating the positions of violations occurring in the source code. This warning contains a true positive (hereinafter also referred to as "TP") that warns of a truly violated location and a false positive (hereinafter also referred to as "FP") that erroneously warns of a non-violated location. Therefore, it is necessary to determine whether each warning is a TP or an FP.

[0030] Therefore, the warning priority prediction device 10 of the present embodiment predicts the priority order of the warnings that the user should address by classifying the warnings indicated by the static analysis result into TP or FP using a warning classification model. Then, the warning priority prediction device 10 transmits the prediction result to the user terminal 12.

[0031] Note that TP may be further subdivided into a warning that requires modification of the source code (Confirmed), i.e., a warning that requires correction, or a warning that does not require modification of the source code (Intentional), i.e., a warning that can be deviated from. Therefore, the warning priority prediction device 10 may further classify TP into a warning that requires correction and a warning that can be deviated from, and predict the priority order of the warnings.

[0032] That is, the prediction by the warning classification model of the present embodiment includes a binary classification prediction that classifies the warnings indicated by the static analysis results into two values, TP and FP, and outputs the probabilities belonging to each, and a ternary classification prediction that classifies the warnings indicated by the static analysis results into three values, warnings that need to be corrected, warnings that can deviate, and FP, and outputs the probabilities belonging to each. Whether the prediction by the warning classification model is a binary classification or a ternary classification may be determined by the setting for the warning priority prediction device 10, or the warning priority prediction device 10 may have only one of the functions.

[0033] In the binary classification, for example, the probability of being TP for a warning is predicted as a TP probability of 90%, and the probability of being FP is predicted as an FP probability of 10%. In the ternary classification, for example, the probability of being a warning that needs to be corrected for a warning is predicted as a probability of needing correction of 70%, the probability of being a warning that can deviate is predicted as a probability of being able to deviate of 20%, and the FP probability is 10%.

[0034] In this way, the warning priority prediction device 10 outputs the probability belonging to a specific warning label as the warning priority. And the specific warning label is a multi-value including at least TP and FP. That is, a warning with a high TP probability or a probability of needing correction is regarded as a warning with a high priority. Thereby, the user can easily recognize the location of the source code that needs to be corrected, and will preferentially correct the code corresponding to the warning that is considered to be TP with a high probability.

[0035] Note that warning labels such as TP, FP, warnings that need to be corrected, and warnings that can deviate are assigned for each warning. Also, in the following description, the TP probability, FP probability, probability of needing correction, and probability of being able to deviate predicted by the warning classification model are collectively referred to as prediction probabilities.

[0036] As shown in FIG. 1, the warning priority prediction device 10 of the present embodiment includes, together with the static analysis tool 20, a source code metrics measurement tool 22, a database 24, a learning unit 26, a reception unit 28, a prediction API (Application Programming Interface) 30, and an output unit 32.

[0037] As described above, the static analysis tool 20 statically analyzes the source code received from the user terminal 12 and outputs a static analysis result. The static analysis result assigns a warning ID to each warning in order to identify each warning that occurred in the source code. Further, the static analysis result is stored in the database 24 in association with the source code to be analyzed.

[0038] The source code metrics measurement tool 22 measures the metrics of the source code received from the user terminal 12. The source code metrics are stored in the database 24 in association with the measured source code.

[0039] The database 24 stores input data (hereinafter referred to as "learning input data") used for learning (training) a warning classification model that is a machine learning model, teacher data, classification results of warnings using the warning classification model, and the like.

[0040] The learning unit 26 generates a warning classification model by machine learning using the learning input data and teacher data stored in the database 24. Note that the learning unit 26 appropriately updates the warning classification model when various data stored in the database 24 is added or updated.

[0041] The reception unit 28 receives the static analysis result to be predicted by the warning classification model. This static analysis result is the analysis result of the source code by the static analysis tool 20, and the warning is unclassified. Note that the reception unit 28 receives various data necessary for determining the priority of the warning in addition to the unclassified warning indicated by the static analysis result, and derives these feature amounts.

[0042] The prediction API 30 deploys the warning classification model trained by the learning unit 26, and determines the priority of the warning indicated by the static analysis result received by the reception unit 28 using the warning classification model. That is, the prediction API 30 functions as a classifier that classifies the warning by associating and outputting a prediction probability indicating the priority of the warning indicated by the static analysis result with the warning ID.

[0043] The output unit 32 transmits to the user terminal 12 the prediction probability for each warning, which is the prediction result by the warning classification model in the prediction API 30.

[0044] The user terminal 12 displays the prediction probability for each warning output by the warning classification model on the viewer. The viewer can sort the warnings by the TP probability. As a result, the user can check the priority of the warnings to be processed by the viewer and can easily recognize the warnings that should be dealt with preferentially.

[0045] In addition, when the prediction probability received from the warning priority prediction device 10 is appropriate, the user assigns a warning label according to the prediction probability to the warning ID and transmits it to the warning priority prediction device 10. On the other hand, when the prediction probability received from the warning priority prediction device 10 is inappropriate, the user assigns a warning label that the user considers appropriate to the warning ID and transmits it to the warning priority prediction device 10. The warning priority prediction device 10 associates the warning ID and the warning label received from the user terminal 12 with the static analysis result that is the prediction target of the warning classification model and stores them in the database 24. Thereby, the learning input data for the warning classification model is accumulated.

[0046] In this way, the warning priority prediction device 10 of the present embodiment operates in cooperation with the static analysis tool 20 and the warning classification model. Therefore, labeled training data (teacher data) can be collected without the user or the like performing the work of attaching annotations for learning the warning classification model by machine learning.

[0047] Note that the warning priority prediction device 10 of the present embodiment has a configuration including a learning unit 26, but the functions of the learning unit 26 may be provided in other information processing devices.

[0048] FIG. 2 is a schematic diagram showing an overview of the warning classification model of the present embodiment.

[0049] When a warning indicated by the static analysis result is input, the warning classification model of this embodiment generates a large number of weak learners that predict warning labels through branches according to various conditions of the input data, and summarizes the prediction results of the probabilities of the warning labels by these weak learners into one prediction value by methods such as majority voting or weighting and outputs it. For this purpose, the warning classification model of this embodiment is generated by supervised learning using a machine learning model capable of ensemble learning, such as LightGBM.

[0050] The learning input data for generating the warning classification model of this embodiment by machine learning is feature quantities based on source code, warning information, product project information, and non-dependent information, and the teacher data is the warning labels assigned according to the user's judgment for each warning indicated by the static analysis result.

[0051] The source code is the one that has been the target of analysis by the static analysis tool 20. Note that, instead of the source code itself, the number of lines of code, cyclomatic complexity, etc. included in the source code metrics measured by the source code metrics measurement tool 22 may be used as the learning input data.

[0052] Here, the tendency of the source code may vary for each product. Therefore, by using the source code or source code metrics as the learning input data, the tendency of different source codes for each product project is learned.

[0053] The warning information is information including the location where the warning occurs in the source code indicated by the static analysis result, and the warning label, which is the teacher data, is associated with it.

[0054] The product project information is information specific to the product using the source code, such as the name and model number of the product. Note that the product project information may include not only the name of the product, etc., but also other information representing the characteristics of the product project. By using the product project information as the learning input data, the tendency of the warning labels for each product will be learned.

[0055] Non-dependent information is information that includes information not dependent on a product. By using non-dependent information not dependent on a product as learning input data, the warning classification model is generated as a machine learning model with generality, rather than being specific to a particular source code or product.

[0056] The non-dependent information of the present embodiment is information regarding the type and settings of the static analysis tool 20 used for the analysis of source code. The type of the static analysis tool 20 is, for example, the name, product number, and version of the static analysis tool 20. The settings of the static analysis tool 20 are, for example, the standard of the programming language used for creating the source code. The standard of the programming language is, for example, whether it conforms to C99 even if it is the same C language. Note that the information regarding the type and settings of the static analysis tool 20 is hereinafter referred to as tool information.

[0057] Here, the static analysis tool 20 used for the analysis of source code has different static analysis performances for each type and setting. Therefore, by using the information regarding the type and settings of the static analysis tool 20 as learning input data, it becomes possible to accurately predict the warning label associated with the warning indicated by the static analysis result regardless of the static analysis tool 20.

[0058] In addition, the non-dependent information also includes the characteristic information of the static analysis tool 20. The characteristic information of the static analysis tool 20 includes at least one of the determinability of static analysis and the benchmark evaluation result of the static analysis tool 20. Note that the determinability of static analysis is the determinability of static analysis with respect to a convention such as MISRA corresponding to the warning output by the static analysis tool 20, and indicates "determinable", "indeterminable", or "uncertain" for each warning. This determinability and benchmark evaluation result are, for example, public information by the manufacturer of the static analysis tool 20, generally known information, information based on past rules of thumb, etc.

[0059] The characteristic information of the static analysis tool 20 varies according to the type and version of the static analysis tool 20, and is stored in the database 24 in association with the type and version of the static analysis tool 20. By using the characteristic information of the static analysis tool 20 as the learning input data for the warning classification model, warnings by static analysis can be predicted with higher accuracy.

[0060] And the learning unit 26 of the present embodiment derives feature amounts based on the source code, warning information, product project information, and non-dependent information. An example of the derived feature amounts is shown below. Note that the derived feature amounts are not limited to the following, and it is not always necessary to use all of the following.

[0061] · The name (URL) of the source code base where the source code in which the warning occurred is managed · The path to the source code file in which the warning occurred · The name of the source code file in which the warning occurred · The function name in the source code in which the warning occurred · The line number of the source code in which the warning occurred · The ID representing the type of warning · The severity of the warning · The name of the static analysis tool 20 that caused the warning · The version of the static analysis tool 20 that caused the warning · The determinability of static analysis of a convention such as MISRA corresponding to the warning · The evaluation result of the static analysis tool 20 with respect to the benchmark of a convention such as MISRA corresponding to the warning · The number of warnings occurring in the function in the source code in which the warning occurred · The number of warnings occurring in the source code file in which the warning occurred · The time elapsed since the warning occurred · The file-level source code metrics of the source code file in which the warning occurred · The function-level source code metrics of the source code file in which the warning occurred · The name of the product project in which the warning occurred

[0062] The learning unit 26 of this embodiment uses the above feature amounts as learning input data, and generates a warning classification model by machine learning using, as teacher data, warning labels assigned according to the user's judgment for each warning indicated by the static analysis result used for deriving the feature amounts. Such a combination of feature amounts and warning labels is regarded as one training data set. That is, the accuracy of the warning classification model improves as the number of training data sets increases.

[0063] Then, the prediction API 30 inputs the static analysis result of the source code, the source code, the product project information, and the non-dependent information into the warning classification model, and outputs the prediction probability of the warning label indicating the priority of the warning indicated by the static analysis result.

[0064] FIG. 3 is a schematic diagram showing a generation process of a warning classification model using information of a plurality of product projects 1 to N of this embodiment. The generation process of the warning classification model is executed by a program stored in a storage medium such as a storage unit included in the warning priority prediction device 10. Note that when this program is executed, a method corresponding to the program is executed.

[0065] As shown in FIG. 3, the feature amounts, which are learning input data, are derived by the learning unit 26 based on the source code metrics, the static analysis result, the product project information, the tool information, and the characteristic information of the static analysis tool 20 for each of the product projects 1 to N.

[0066] The characteristic information of the static analysis tool 20 is stored in the database 24 as common information regardless of the product project. Then, the characteristic information of the static analysis tool 20 corresponding to the name and version of the static analysis tool 20 indicated by the tool information is read from the database 24.

[0067] Also, warning labels assigned to the warnings indicated by the static analysis results for the source code of each of the product projects 1 to N are used as teacher data.

[0068] A learning dataset, which is a combination of feature quantities and teacher data for each of product projects 1 to N, is stored in the database 24. The learning dataset is generated each time the source code for each of product projects 1 to N is modified or improved, stored in the database 24, and then used for learning of the subsequent warning classification model. Then, the learning unit 26 generates a warning classification model by machine learning using a number of learning datasets.

[0069] FIG. 4 is a schematic diagram showing the warning prioritization process using the warning classification model of the present embodiment. The warning prioritization process is executed by a program stored in a storage medium such as a storage unit included in the warning priority prediction device 10. When this program is executed, a method corresponding to the program is executed.

[0070] The warning prioritization process is performed when the user transmits the source code to the warning priority prediction device 10 via the user terminal 12 and issues an execution instruction for static analysis. In accordance with this execution instruction, first, the static analysis tool 20 executes static analysis of the source code and outputs the static analysis result. Also, the source code metrics measurement tool 22 measures the source code metrics of the source code.

[0071] Then, the reception unit 28 reads out the product project information, tool information, and characteristic information of the static analysis tool 20 corresponding to the source code from the database 24, and derives feature quantities based on the read information, the static analysis result, and the source code metrics. This feature quantity includes warnings with unknown warning labels in the static analysis result.

[0072] When the feature quantity is input to the warning classification model, the warning classification model outputs the prediction probability of the warning as the prediction result. The prediction probability of the warning is stored in the database 24 and transmitted to the user terminal 12. The user checks the location of the source code that needs to be corrected by viewing the prediction probability of the warning transmitted to the user terminal 12.

[0073] Also, as described above, when the prediction probability is appropriate, the user assigns a warning label according to the prediction probability to the warning ID and transmits it to the warning priority prediction device 10. When the prediction probability is not appropriate, the user assigns a warning label that is considered appropriate in confidence to the warning ID and transmits it to the warning priority prediction device 10. The warning priority prediction device 10 associates the warning ID and the warning label received from the user terminal 12 with the static analysis result and stores them in the database 24. The warning label stored in the database 24 in this way will be used as teacher data.

[0074] Note that in FIG. 4, the prediction target of the warning classification model is the static analysis result of the source code of product projects 1 to N. However, the static analysis result used as the prediction target of the warning classification model may be the static analysis result of the source code other than product projects 1 to N.

[0075] Also, the warning priority prediction device 10 of the present embodiment may update the warning classification model by machine learning using the learning data set newly stored in the database 24. The timing of the update is, for example, the timing when the performance degradation of the warning classification model appears, or the timing when a predetermined number or more of new learning data sets are accumulated. The performance degradation of the warning classification model is, for example, the case where the frequency of the warning label different from the prediction probability of the warning classification model being attached by the user becomes a predetermined value or more.

[0076] Also, when the warning classification model is updated, the prediction probability will be determined by different versions of the warning classification model. Therefore, in the viewer displayed on the user terminal 12, it is possible to filter the prediction result of the prediction probability for the warning by the version of the warning classification model.

[0077] As described above, the warning priority prediction device 10 of the present embodiment uses non-dependent information that does not depend on the product using the source code as learning input data for the warning classification model. As a result, the generality of the warning classification model is enhanced, and it becomes possible to predict the warning priority with a single warning classification model even for different source codes used for different product projects.

[0078] That is, when generating a warning classification model without using non-dependent information, it was necessary to generate a warning classification model for each product project. On the other hand, by generating a warning classification model using non-dependent information as in the present embodiment, it is not necessary to generate a warning classification model for each product project, and the generation and operation costs of the warning classification model can also be minimized. In addition, since the number of learning datasets used for generating the warning classification model also increases, more performance improvement can be expected.

[0079] As described above, the present invention has been described using the above embodiments, but the technical scope of the present invention is not limited to the scope described in the above embodiments. Various changes or improvements can be made to the above embodiments without departing from the gist of the invention, and the forms with such changes or improvements are also included in the technical scope of the present invention.

[0080] In the above embodiment, the form of using the characteristic information of the static analysis tool 20 for deriving the feature amount has been described, but the present invention is not limited to this. As shown in FIG. 5, it is not necessary to use the characteristic information of the static analysis tool 20 for deriving the feature amount for each of the product projects 1 to N. In this form, the warning classification model is learned using the feature amounts for each of the product projects 1 to N and the characteristic information of the static analysis tool 20.

[0081] In the above embodiment, the form of using source code metrics for deriving the feature amount has been described, but the present invention is not limited to this. The feature amount may be derived using the source code itself instead of the source code metrics.

[0082] In the above-described embodiment, the form in which feature amounts are derived based on learning input data and the feature amounts are used for learning the warning classification model has been described. However, the present invention is not limited to this. Instead of using feature amounts for learning the warning classification model, the warning classification model may be learned using the learning input data itself.

[0083] In the above-described embodiment, the form in which warning labels for obtaining prediction probabilities by the warning classification model are binary values of TP and FP, or ternary values of warnings to be corrected, warnings that can deviate, and FP has been described. However, the present invention is not limited to this. TP may be subdivided into three or more types, and the warning label may be a quaternary value or more.

[0084] In the above-described embodiment, the form in which non-dependent information is information regarding the static analysis tool 20 has been described. However, the present invention is not limited to this. The non-dependent information may be other information as long as it is information used for developing the source code and does not depend on the product. For example, information regarding the source code metrics measurement tool 22 (name and version of the tool) may be used in addition to the information regarding the static analysis tool 20.

Description of Reference Numerals

[0085] 10 ··· Warning priority prediction device, 20 ··· Static analysis tool, 24 ··· Database, 26 ··· Learning unit, 28 ··· Reception unit, 30 ··· Prediction API

Claims

1. An information processing apparatus (10) that predicts the priority of warnings indicated by the analysis result of source code by a static analysis tool (20), using, as learning input data, the source code, warning information including the locations where warnings occur in the source code indicated by the analysis result, product-specific information using the source code, and non-dependent information including information not dependent on the product, and machine learning using warning labels associated with each warning as teacher data, generating a prediction model that takes, as input, the analysis result of the static analysis tool for the source code and outputs the priority of the warning; a learning unit (26); a reception unit (28) that receives the analysis result of the static analysis tool to be predicted by the prediction model; a prediction unit (30) that predicts the priority of the warning indicated by the analysis result received by the reception unit using the prediction model; An information processing apparatus comprising:

2. The information processing apparatus according to claim 1, wherein the non-dependent information is information regarding the type and settings of the static analysis tool used for analyzing the source code.

3. The information processing apparatus according to claim 2, wherein the non-dependent information includes characteristic information of the static analysis tool according to the type of the static analysis tool.

4. The information processing apparatus according to claim 3, wherein the characteristic information includes at least one of the decidability of static analysis and the benchmark evaluation result of the static analysis tool.

5. The information processing apparatus according to claim 1 or claim 2, wherein the priority of the warning is the probability of belonging to a specific warning label.

6. The information processing apparatus according to claim 5, wherein the warning label is a multi-value including at least a true positive that warns of a location with a true violation and a false positive that erroneously warns of a location without a violation.

7. The information processing apparatus according to claim 6, wherein the true positive is divided into a warning that requires modification of the source code or a warning that does not require modification of the source code.

8. The information processing apparatus according to claim 1 or claim 2, wherein the prediction unit inputs the analysis result of the source code, the source code, product-specific information using the source code, and the non-dependent information into the prediction model and predicts the priority of the warning indicated by the analysis result.

9. The information processing apparatus according to claim 1 or claim 2, wherein the prediction model is generated by a machine learning model capable of ensemble learning.

10. A priority prediction method for predicting the priority of warnings indicated by the analysis result of source code by a static analysis tool, using, as learning input data, the source code, warning information including the locations where warnings occur in the source code indicated by the analysis result, product-specific information using the source code, and non-dependent information including information not dependent on the product, and machine learning using warning labels associated with each warning as teacher data, with the input being the analysis result of the static analysis tool for the source code and the output being the priority of the warning. A first step in which a learning unit generates a prediction model; A second step in which a reception unit receives the analysis result of the static analysis tool to be predicted by the prediction model; A third step in which a prediction unit predicts the priority of the warning indicated by the analysis result received by the reception unit using the prediction model; A priority prediction method having the above.

11. A computer included in an information processing apparatus that predicts the priority of warnings indicated by the analysis result of source code by a static analysis tool, using, as learning input data, the source code, warning information including the locations where warnings occur in the source code indicated by the analysis result, product-specific information using the source code, and non-dependent information including information not dependent on the product, and machine learning using warning labels associated with each warning as teacher data, with the input being the analysis result of the static analysis tool for the source code and the output being the priority of the warning. A learning unit that generates a prediction model; A reception unit that receives the analysis result of the static analysis tool to be predicted by the prediction model; A prediction unit that predicts the priority of the warning indicated by the analysis result received by the reception unit using the prediction model; A warning priority prediction program for causing the above to function.

Citation Information

Patent Citations

  • Code base risk analysis using static analysis and performance data

    JP2016004569A

  • Analysis device for static analysis result of source code, and analysis method

    JP2017204090A

  • False positive vulnerability detection using neural transformers

    US20230281317A1