A method, device and system for evaluating code design quality
Through the combination of static scanning results and artificial neural networks, the design quality of the code is automatically evaluated, and the problems of high labor cost of code evaluation and high false alarm rate in the existing technology are solved, achieving high accuracy and automated code review.
Patent Information
- Application Number
- CN201980093503.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-03-26
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2039-03-26
AI Technical Summary
The existing technology has problems such as high labor consumption and difficult to guarantee quality in code evaluation, and the warnings of static analysis tools are mostly false alarms, which require a lot of manpower to identify them.
By determining the probability of error-prone patterns in the code based on the code and inputting them into the artificial neural network, we predict whether the code violates predetermined design principles and its degree of quantification, thereby evaluating the design quality of the code.
It realizes full automation of the software code review process, improves evaluation accuracy, reduces false positives, and improves evaluation efficiency.
Smart Images

Figure CN113490920B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of software project development, and in particular to a method, device and system for evaluating code design quality. Background Art
[0002] Software plays an important role in many application environments such as modern electrical / electronic systems. Software quality plays a leading role in the overall functionality, reliability and quality of the entire system. Errors in software design can exist in many different forms, and the resulting system failures may endanger human life and safety, cost a lot of money to repair, cause customer dissatisfaction, and damage the company's reputation. Therefore, having appropriate software quality management capabilities is crucial to business success.
[0003] Code review is the process of systematically reviewing software source code, aiming to check the design quality of the code and find errors so that they can be resolved and improved. Effectively performing code review during the software development process can ensure that most software errors can be discovered and resolved early in the software development phase, thereby helping to improve the overall quality of the software and achieve rapid delivery of defect-free software.
[0004] So far, most code reviews are done manually, such as through informal walkthroughs, formal review meetings, or pair programming. These activities require a lot of manpower and also require the evaluators to be more senior or more experienced than ordinary developers. As a result, in practice, software development teams often ignore this process for a variety of reasons (e.g., time pressure, lack of suitable reviewers, etc.), or it is difficult to ensure the quality of code reviews.
[0005] It is also common to use static analysis tools to automate code evaluation. Static analysis tools can quickly scan source code based on predetermined quality inspection rules and identify patterns that are prone to software errors. They then alert developers in the form of warnings and provide suggestions on how to fix them.
[0006] However, violating the predetermined quality inspection rules does not necessarily lead to quality defects. Therefore, static analysis tools often generate a large number of warnings, but most of them are false positives that can be ignored, so it still takes a lot of manpower to identify the results to determine which of them are quality defects that really need to be fixed and which are just invalid warnings. Summary of the invention
[0007] The embodiments of the present invention provide a method, device and system for evaluating code design quality.
[0008] The technical solution of the embodiment of the present invention is as follows:
[0009] A method for evaluating the quality of code design, including:
[0010] Determining a probability of an error-prone pattern existing in the code based on a static scan result of the code;
[0011] Inputting the probability into an artificial neural network, and determining whether the code violates a predetermined design principle and a prediction result of the quantified degree of violation of the design principle based on the artificial neural network;
[0012] The design quality of the code is evaluated based on the prediction result.
[0013] It can be seen that the implementation mode of the present invention detects the existence of error-prone patterns in the code, predicts whether key design principles are violated in the software design process and the quantitative degree of violation of key design principles, and thereby evaluates the design quality of the code, thereby realizing full automation of the software code review process, overcoming the shortcomings of the prior art of directly evaluating the design quality based on quality inspection rules, which leads to too many error warnings, and has the advantage of high evaluation accuracy.
[0014] In one embodiment, the error-prone patterns include at least one of the following: shotgun modification; divergent changes; large-scale pre-design; scattered / redundant functionality; circular dependencies; wrong dependencies; complex classes; long methods; code duplication; long parameter lists; message chains; useless methods; and / or
[0015] The design principles include at least one of the following: separation of concerns principle; single responsibility principle; least knowledge principle; no duplication principle; up-front design minimization principle.
[0016] It can be seen that the embodiments of the present invention can further improve the evaluation accuracy through a large number of preset error-prone patterns and preset design principles.
[0017] In one embodiment, the method further comprises: receiving a modification record of the code;
[0018] The determining the probability of the existence of the error-prone pattern in the code based on the static scanning result of the code includes: determining the probability of the existence of the error-prone pattern in the code based on a composite logic conditional expression including the static scanning result and / or the modification record.
[0019] Therefore, the embodiments of the present invention can improve the accuracy of detecting error-prone patterns based on the composite logical conditional expression containing the modification record of the code.
[0020] In one embodiment, the determining the probability of the existence of the error-prone pattern in the code based on the compound logical conditional expression including the static scanning result and / or the modification record includes at least one of the following:
[0021] determining a probability of a shotgun modification being present based on a composite logical conditional expression including an afferent coupling metric, an efferent coupling metric, and a change method metric;
[0022] Determine the probability of existence of divergent changes based on a composite logic conditional expression including a version modification count metric, an instability metric, and an afferent coupling metric;
[0023] determining the probability of the existence of a large-scale pre-design based on a composite logical conditional expression including the number of lines of code, the number of lines of code changed, the number of classes, the number of classes changed, and the statistical means of the above metrics;
[0024] Determining the probability of the existence of dispersed / redundant functions based on a composite logical conditional expression including a structural similarity measure and a logical similarity measure;
[0025] Determine the probability of the existence of long methods based on a composite logical conditional expression containing a cyclomatic complexity measure;
[0026] Determine the probability of the existence of a complex class based on a composite logical conditional expression containing the maximum of the number of lines of code, the number of attributes, the number of methods, and the cyclomatic complexity of the methods of a given class;
[0027] Determine the probability of a long argument list based on a compound logical conditional expression involving the number of arguments;
[0028] Determines the probability of the existence of a message chain based on a compound logical conditional expression that includes the number of indirect calls.
[0029] It can be seen that the embodiments of the present invention respectively propose detection methods based on indexable metrics based on the causal attributes of the error-prone patterns, thereby realizing automatic detection of error-prone patterns.
[0030] In one embodiment, the predetermined threshold value in the complex logic conditional expression is adjustable.
[0031] Therefore, by adjusting the predetermined threshold value in the compound logic conditional expression, it can be applied to various application scenarios.
[0032] In one embodiment, the artificial neural network includes a connection between the error-prone pattern and the design principle; the prediction result of determining whether the code violates the predetermined design principle and the quantitative degree of violation of the design principle based on the artificial neural network includes:
[0033] Based on the connections in the artificial neural network and the probability of the existence of the error-prone mode, a prediction result of whether the code violates the design principle and the quantitative degree of violation of the design principle is predicted.
[0034] Therefore, the automatic generation of prediction results is achieved through artificial neural networks, ensuring the evaluation efficiency.
[0035] In one embodiment, the method further comprises:
[0036] The weights of the connections in the artificial neural network are adjusted based on a self-learning algorithm.
[0037] It can be seen that by adjusting the artificial neural network based on the self-learning algorithm, the artificial neural network becomes more and more accurate.
[0038] In one embodiment, the evaluating the design quality of the code based on the prediction result includes at least one of the following:
[0039] Evaluate the design quality of the code based on whether it violates the predetermined design principles, where not violating the predetermined design principles has a good design quality relative to violating the predetermined design principles;
[0040] The design quality of the code is evaluated based on the quantitative degree of violation of the design principle, wherein the quantitative degree of violation of the design principle is inversely proportional to the evaluation of the design quality.
[0041] In one embodiment, the connection between the error-prone pattern and the design principle includes at least one of the following:
[0042] The connection between shotgun modification and separation of concerns principle; the connection between shotgun modification and single responsibility principle; the connection between divergent change and separation of concerns principle; the connection between divergent change and single responsibility principle; the connection between large-scale pre-design and pre-design minimization principle; the connection between decentralized functions and the principle of no duplication; the connection between redundant functions and the principle of no duplication; the connection between circular dependency and separation of concerns principle; the connection between circular dependency and the principle of least knowledge; the connection between false dependency and separation of concerns principle; the connection between false dependency and the principle of least knowledge; the connection between complex class and single responsibility principle; the connection between long method and single responsibility principle; the connection between code duplication and the principle of no duplication: the connection between long parameter list and single responsibility principle; the connection between message chain and the principle of least knowledge; the connection between unused method and pre-design minimization principle. It can be seen that the embodiment of the present invention optimizes the connection form of artificial neural network by analyzing the corresponding relationship between error-prone patterns and design principles.
[0043] A device for evaluating code design quality, comprising:
[0044] A determination module configured to determine the probability of an error-prone pattern existing in the code based on a static scanning result of the code;
[0045] A prediction result determination module, configured to input the probability into an artificial neural network, and determine, based on the artificial neural network, a prediction result as to whether the code violates a predetermined design principle and the quantification degree of the violation of the design principle;
[0046] An evaluation module, configured to evaluate the design quality of the code based on the prediction result.
[0047] It can be seen that, through detecting the existence of error-prone patterns in the code, whether key design principles are violated during the software design process and the quantification degree of the violation of the key design principles, and thereby evaluating the design quality of the code, the embodiment of the present invention realizes the full automation of the software code review process, overcomes the shortcoming of excessive false alarms caused by directly evaluating the design quality based on quality inspection rules in the prior art, and has the advantage of high evaluation accuracy.
[0048] In one embodiment, the error-prone patterns include at least one of the following: shotgun surgery; divergent change; large-scale pre-design; scattered / redundant functions; cyclic dependencies; wrong dependencies; complex classes; long methods; code duplication; long parameter lists; message chains; useless methods; and / or
[0049] The design principles include at least one of the following: principle of separation of concerns; single responsibility principle; least knowledge principle; don't repeat yourself principle; minimization of pre-design principle.
[0050] It can be seen that, through a large number of preset error-prone patterns and preset design principles, the embodiment of the present invention can further improve the evaluation accuracy.
[0051] In one embodiment, the determination module is further configured to receive a modification record of the code, wherein determining the probability of the existence of error-prone patterns in the code based on the static scan result of the code includes: determining the probability of the existence of error-prone patterns in the code based on a composite logical conditional expression including the static scan result and / or the modification record.
[0052] Therefore, based on the composite logical conditional expression including the modification record of the code, the embodiment of the present invention can improve the accuracy of detecting error-prone patterns.
[0053] In one embodiment, determining the probability of the existence of error-prone patterns in the code based on the composite logical conditional expression including the static scan result and / or the modification record includes at least one of the following:
[0054] Determining the probability of the existence of shotgun surgery based on a composite logical conditional expression including an afferent coupling metric, an efferent coupling metric, and a changed method metric;
[0055] Determine the probability of existence of divergent changes based on a composite logic conditional expression including a version modification count metric, an instability metric, and an afferent coupling metric;
[0056] determining the probability of the existence of a large-scale pre-design based on a composite logical conditional expression including the number of lines of code, the number of lines of code changed, the number of classes, the number of classes changed, and the statistical means of the above metrics;
[0057] Determining the probability of the existence of dispersed / redundant functions based on a composite logical conditional expression including a structural similarity measure and a logical similarity measure;
[0058] Determine the probability of the existence of long methods based on a composite logical conditional expression containing a cyclomatic complexity measure;
[0059] Determine the probability of the existence of a complex class based on a composite logical conditional expression containing the maximum of the number of lines of code, the number of attributes, the number of methods, and the cyclomatic complexity of the methods of a given class;
[0060] Determine the probability of a long argument list based on a compound logical conditional expression involving the number of arguments;
[0061] Determines the probability of the existence of a message chain based on a compound logical conditional expression that includes the number of indirect calls.
[0062] It can be seen that the embodiments of the present invention respectively propose detection methods based on indexable metrics based on the causal attributes of the error-prone patterns, thereby realizing automatic detection of error-prone patterns.
[0063] In one embodiment, the artificial neural network includes connections between error-prone patterns and design principles; the prediction result determination module is configured to predict whether the code violates the design principle and the quantitative degree of violation of the design principle based on the connections in the artificial neural network and the probability of the existence of the error-prone pattern.
[0064] Therefore, the automatic generation of prediction results is achieved through artificial neural networks, ensuring the evaluation efficiency.
[0065] In one embodiment, the evaluation module is configured to evaluate the design quality of the code based on whether it violates the predetermined design principles, wherein not violating the predetermined design principles has a good design quality relative to violating the predetermined design principles; and to evaluate the design quality of the code based on the quantitative degree of violation of the design principles, wherein the quantitative degree of violation of the predetermined design principles is inversely proportional to the evaluation of the quality of the design.
[0066] In one embodiment, the connection between the error-prone pattern and the design principle includes at least one of the following:
[0067] The connection between shotgun modifications and the principle of separation of concerns; the connection between shotgun modifications and the single responsibility principle; the connection between divergent changes and the principle of separation of concerns; the connection between divergent changes and the single responsibility principle; the connection between large-scale up-front design and the principle of up-front design minimization; the connection between decentralized functionality and the principle of no duplication; the connection between redundant functionality and the principle of no duplication; the connection between circular dependencies and the principle of separation of concerns; the connection between circular dependencies and the principle of least knowledge; the connection between false dependencies and the principle of separation of concerns; the connection between false dependencies and the principle of least knowledge; the connection between complex classes and the single responsibility principle; the connection between long methods and the single responsibility principle; the connection between code duplication and the principle of no duplication; the connection between long parameter lists and the single responsibility principle; the connection between message chains and the principle of least knowledge; the connection between unused methods and the principle of up-front design minimization.
[0068] It can be seen that the embodiments of the present invention optimize the connection form of the artificial neural network by analyzing the corresponding relationship between error-prone patterns and design principles.
[0069] A system for evaluating code design quality, comprising:
[0070] A code repository, which is configured to store the code to be evaluated;
[0071] A static scanning tool, configured to statically scan the code to be evaluated;
[0072] An error-prone pattern detector configured to determine the probability of an error-prone pattern existing in the code based on a static scanning result output by a static scanning tool;
[0073] The artificial neural network is configured to determine whether the code violates a predetermined design principle and a prediction result of a quantitative degree of violation of the design principle based on the probability, wherein the prediction result is used to evaluate the design quality of the code.
[0074] Therefore, the implementation mode of the present invention detects the probability of the existence of error-prone patterns based on the static scanning results of the static scanning tool, and predicts whether the key design principles are violated in the software design process and the quantitative degree of violation of the key design principles based on the probability, and thereby evaluates the design quality of the code. It has the advantage of high evaluation accuracy and can effectively avoid false warnings.
[0075] In one embodiment, the code repository is further configured to store modification records of the code;
[0076] The static scanning tool is configured to detect the probability of the existence of an error-prone pattern in the code based on a compound logical conditional expression containing the static scanning result and / or the modification record.
[0077] Therefore, the embodiments of the present invention can improve the accuracy of detecting error-prone patterns based on the composite logical conditional expression containing the modification record of the code.
[0078] An apparatus for evaluating code design quality, comprising a processor and a memory;
[0079] The memory stores an application program executable by the processor, which is used to enable the processor to execute the method for evaluating code design quality as described in any one of the above items.
[0080] A computer-readable storage medium stores computer-readable instructions for executing the method for evaluating code design quality as described in any one of the above items. BRIEF DESCRIPTION OF THE DRAWINGS
[0081] Figure 1 The present invention is a flowchart of a method for evaluating code design quality according to an embodiment of the present invention.
[0082] Figure 2 4 is a structural diagram of an artificial neural network according to an embodiment of the present invention.
[0083] Figure 3 The structure diagram of the system for evaluating code design quality according to an embodiment of the present invention.
[0084] Figure 4 It is a structural diagram of a device for evaluating code design quality according to an embodiment of the present invention.
[0085] Figure 5 The present invention is a structural diagram of a device having a processor and a memory structure for evaluating code design quality according to an embodiment of the present invention.
[0086] The reference numerals are as follows:
[0087]
[0088] DETAILED DESCRIPTION
[0089] In order to make the technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and implementation methods. It should be understood that the specific implementation methods described herein are only used to illustrate the present invention and are not intended to limit the scope of protection of the present invention.
[0090] For the sake of brevity and intuitiveness in description, the scheme of the present invention is explained below by describing several representative implementations. A large number of details in the implementations are only used to help understand the scheme of the present invention. However, it is obvious that the technical scheme of the present invention may not be limited to these details when implemented. In order to avoid unnecessarily obscuring the scheme of the present invention, some implementations are not described in detail, but only a framework is given. Hereinafter, "including" means "including but not limited to", and "according to..." means "at least according to..., but not limited to only according to...". Due to the language habits of Chinese, when the number of a component is not specifically specified below, it means that the component can be one or more, or can be understood as at least one.
[0091] In the embodiment of the present invention, the static scanning results of the code and / or the modification records of the code (including the history of additions and subtractions / changes / revisions) can be used to estimate whether the key software design principles are followed and how well they are implemented during the development of the source code being reviewed. This is based on an obvious reason: if a key design principle is fully considered during the design process, the possibility of finding related error-prone patterns caused by violating the principle in the source code will be reduced accordingly. Therefore, by detecting whether there are related error-prone patterns in the source code, it is possible to estimate to what extent the software follows the key design principle during the design process, thereby evaluating the quality of the code design and determining areas where quality improvements can be made.
[0092] Figure 1 The invention relates to a method for evaluating code design quality according to an embodiment of the invention.
[0093] like Figure 1 As shown, the method includes:
[0094] Step 101: Determine the probability of an error-prone pattern existing in the code based on the static scanning result of the code.
[0095] Here, static scanning of code refers to a code analysis technology that scans program code through lexical analysis, syntax analysis, control flow, data flow analysis and other technologies without running the code to verify whether the code meets indicators such as standardization, security, reliability, and maintainability.
[0096] For example, static code scanning tools may include: Understand, Checkstyle, FindBugs, PMD, Fortify SCA, Checkmarx, CodeSecure, etc. Moreover, static scanning results may include: code metrics, quality defect warnings and related findings, etc. Among them, code metrics may include: maintainability index, cyclomatic complexity, inheritance depth, class coupling, number of lines of code, etc.
[0097] Here, error-prone patterns are often also referred to as bad code smells, which are diagnostic symptoms indicating that there may be quality problems in software design. Error-prone patterns can exist in source code units at different levels.
[0098] For example, error-prone patterns may include at least one of the following: Shortgun surgery; Divergent change; Big design up front (BDUF); Scattered functionality; redundant functionality; Cyclic dependency; Bad dependency; Complex class; Long method; Code duplication; Long parameter list; Message chains; Unused method, etc.
[0099] Among them: Divergent changes refer to a class that is always passively modified repeatedly for different reasons. Shotgun modifications are somewhat similar to divergent changes, which means that whenever a class needs a certain change, many small modifications must be made in many other different classes accordingly, that is, shotgun modifications. Large-scale pre-design refers to making a lot of designs in advance in the early stages of the project, especially when the requirements are not complete or clear. Dispersed functions and redundant functions refer to the same functions / high-level concerns being repeatedly implemented by multiple methods. Circular dependencies refer to two or more classes / modules / architecture components that are directly or indirectly dependent on each other. Wrong dependencies refer to classes / modules that need to use information from other classes / modules, which should not be closely related to them. Complex classes refer to classes that are too complex. Long methods refer to methods that have too many logical branches. Code duplication refers to the same code structure being repeated in different places in the software, mainly including: two functions in the same class contain the same expression; two sibling subclasses contain the same expression; two completely unrelated classes contain the same expression, etc. Long parameter lists refer to methods that require too many parameters. Message chains refer to classes / methods using methods or properties of another class that is not its direct friend. Dead methods are methods that are never used / called in the code.
[0100] For example, Table 1 is an exemplary illustration of error-prone patterns in C++ source code.
[0101]
[0102] Table 1
[0103] The above exemplary descriptions are typical examples of error-prone modes. Those skilled in the art will appreciate that such descriptions are merely exemplary and are not intended to limit the protection scope of the embodiments of the present invention.
[0104] In one embodiment, the method also includes: receiving modification records of the code; determining the probability of the existence of an error-prone pattern in the code based on the static scanning results of the code in step 101 includes: determining the probability of the existence of an error-prone pattern in the code based on a compound logical conditional expression, wherein the compound logical conditional expression includes both the static scanning results and the modification records, or includes the static scanning results but not the modification records, or includes the modification records but not the static scanning results.
[0105] Preferably, the probability of the presence of an error-prone pattern in the code can be determined based on a compound logical conditional expression including static scan results and modification records. Alternatively, the probability of the presence of an error-prone pattern in the code can be detected based on a compound logical conditional expression including only static scan results but not modification records.
[0106] Preferably, determining the probability of the existence of an error-prone pattern in the code based on a composite logical conditional expression including static scanning results and / or modification records includes at least one of the following:
[0107] (1) Determine the probability of the existence of a shotgun modification based on a composite logical conditional expression including an afferent coupling (Ca) metric, an efferent coupling (Ce) metric, and a changing method (CM) metric.
[0108] (2) Determine the probability of divergent changes based on a composite logical conditional expression including a revision number metric (Revision Number, RN), an instability metric (Instability, I) and an incoming coupling metric, where instability I = Ce / (Ca+Ce).
[0109] (3) Determine the probability of the existence of large-scale pre-design based on a composite logical conditional expression including the number of lines of code (LOC), the number of changed lines of code (LOCC), the number of classes (CN), the number of changed classes, and the statistical mean (AVERAGE) of the above metrics.
[0110] (4) Determine the probability of the existence of dispersed / redundant functions based on a composite logical conditional expression including a structural similarity (SS) metric and a logical similarity (LS) metric.
[0111] (5) Determine the probability of the existence of long methods based on a composite logical conditional expression including the cyclomatic complexity (CC) metric.
[0112] (6) Determine the probability of the existence of a complex class based on a composite logical conditional expression containing the number of code lines, the number of attributes (AN), the number of methods (AN), and the maximum method cyclomatic complexity (MCCmax) of the methods of a given class.
[0113] (7) Determine the probability of a long parameter list based on a compound logical conditional expression containing the number of parameters.
[0114] (8) Determine the probability of the existence of a message chain based on a composite logical conditional expression including the indirect calling number (ICN).
[0115] More preferably, the predetermined threshold value in any of the above-mentioned complex logic conditional expressions is adjustable.
[0116] A typical algorithm for determining error-prone patterns is described below.
[0117] For example: The compound logical conditional expression (Ca, TopValues(10%)) OR (CM, HigherThan(10))) AND (Ce, HigherThan(5)) is a typical example of a compound logical conditional expression used to determine shotgun modification. The compound logical conditional expression combines Ca, CM, and Ce. Among them: "OR" represents a logical OR relationship; "AND" represents a logical AND relationship.
[0118] Among them: The Ce metric counts the number of classes that have properties that a given class needs to access or methods that need to be called. The CM metric counts the number of methods that need to access properties of a given class or call methods of a given class. TopValues and HigherThan are parameterized filtering mechanisms with specific values (thresholds). TopValues selects those members whose specific indicator values fall within a specified maximum range among all given members. HigherThan selects all members whose specific indicator values are higher than a given threshold among all members. Therefore, the determination strategy based on the above compound logical conditional expression means: if the Ca metric value of a class falls within the highest 10% of all classes, or its CM metric value is higher than 10, and at the same time its Ce metric value is higher than 5, then it should be a suspicious factor causing shotgun modification.
[0119] In order to improve the flexibility and applicability of the determination algorithm, three risk levels can be introduced: low, medium and high. Moreover, different threshold values (thresholds) will be assigned to different risk levels.
[0120] For example, Table 2 is a typical example of a compound logic conditional expression that determines shotgun modification based on three risk levels:
[0121]
[0122] Table 2
[0123] The algorithm for determining shotgun modifications described in Table 2 is based on the metrics afferent coupling (Ca), efferent coupling (Ce), and change method (CM).
[0124] For example, Table 3 is a typical example of a compound logic conditional expression that determines divergent changes based on three risk levels:
[0125]
[0126] Table 3
[0127] In Table 3, the algorithm based on compound logical conditional expressions determines divergent changes based on the metric instability (1), afferent coupling (Ca), and the version modification number metric (RN). The version modification number metric (RN) is the total number of changes made to a given class and can be directly obtained from the change history.
[0128] Furthermore, the instability (I) was calculated by comparing the afferent and efferent dependencies represented by the afferent coupling Ca and the efferent coupling Ce.
[0129] Instability I = Ce / (Ca+Ce);
[0130] The rationale behind the algorithm in Table 3 is that if a class’s past revision history shows that it has changed frequently (measured by the number of revisions), or we predict that it is likely to change in the future (measured by instability), then it should be considered as suspected of divergent change.
[0131] For example, Table 4 is a typical example of determining a large-scale pre-designed composite logic conditional expression based on three risk levels:
[0132]
[0133] Table 4
[0134] Among them: the number of lines of code (LOC) metric counts the number of all lines of source code; the number of lines of code changed (LOCC) is the number of source code lines that have been changed (newly added, revised or deleted) in a given time period. The number of classes (CN) metric counts the total number of classes. The number of classes changed (CCN) metric counts the number of classes that have been revised in a given time period. The label AVERAGE is used to represent the average value of a given metric calculated based on a set of historical data. The basic principle of this algorithm is very clear, that is: if too many source code changes are made in a relatively short period of time, there is a high probability that a large-scale pre-design situation has occurred during the design process.
[0135] For C++ source code, redundant functionality (same or similar functionality implemented multiple times in different places) should be identified at the method / function level. The identification algorithm should be able to compare the similarity between different methods / functions.
[0136] Example: Table 5 is a typical example of a compound logic conditional expression for determining decentralized functions / redundant functions based on three risk levels:
[0137]
[0138] Table 5
[0139] In Table 5, if the code structure and logical control flow of two given methods are highly similar, then they are considered to be suspected of implementing similar functions. The algorithm determines the scattered functions / redundant functions based on the composite measure of structural similarity (SS) and logical similarity (LS).
[0140] The measurement methods of structural similarity and logical similarity are explained below.
[0141] Example: Table 6 is a typical example of a compound logic conditional expression for determining code structure similarity based on three risk levels:
[0142]
[0143] Table 6
[0144] In Table 6, the code structure similarity is determined based on parameter similarity (PS), return value similarity (RVS), and difference in lines of code (DLOC).
[0145] Among them: Parameter similarity (PS) measures the similarity between the input parameters of two given methods and is calculated as follows:
[0146] PS = STPN / PNmax;
[0147] Where PN represents the number of input parameters. PNmax means comparing the PN metrics of two given methods and selecting the maximum value. STPN represents the number of input parameters with the same data type.
[0148] For example, if the parameter lists for two given methods are:
[0149] Method A (float a, int b, int c); method B (float d, int e, long f, char g).
[0150] Then, STPN is 2 because both methods have at least one float input parameter and one int input parameter in common. PN for method A is 4 and PN for method B is 5, so the value of PNmax is 5. Therefore, the calculated result of the PS metric is 0.4 (2 / 5).
[0151] The metric Return Value Similarity (RVS) only takes the values 0 and 1. If two given methods have the same return value type, then the value of RVS is 1, otherwise the value of RVS is 0.
[0152] The difference in lines of code (DLOC) is the difference between the LOC of two given methods. It can only take positive values. So if method A has 10 lines of source code and method B has 15 lines of source code, then the DLOC between method A and method B is 5.
[0153] LOCmax means comparing the LOC of two given methods and choosing the maximum one.
[0154] Example: Table 7 is a typical example of a compound logic conditional expression for determining code logic similarity based on three risk levels, which is used to compare the similarity between the logic flows of two given methods.
[0155]
[0156] Table 7
[0157] In Table 7, the code logic similarity (LS) is compared based on the Cyclomatic Complexity (CC) metric, the Difference Cyclomatic Complexity (DCC) metric, and the Control Flow Similarity (CFS).
[0158] Cyclomatic Complexity (CC) is a well-known metric that measures the complexity of a method by counting the number of independent logic paths. Many static source code analysis tools provide the ability to calculate this metric.
[0159] The difference in cyclomatic complexity (DCC) is calculated based on the CC metric. It can take only positive values. Suppose the CC metric of method A is 20 and the CC metric of method B is 23, then the difference in cyclomatic complexity (DCC) between method A and method B is 3. CCmax means comparing the CC metrics of two given methods and selecting the maximum value.
[0160] Control flow similarity (CFS) measures the similarity between the control flow statements of two given methods. Take C++ language as an example. In C++ language, logical control statements include if, then, else, do, while, switch, case and for loops, etc. CFS is calculated based on the logical control blocks that appear in a given method.
[0161] Assume there are two given methods A and B, we put the control flow statements used in the methods into two sets, setA and setB:
[0162] setA={if, for, switch, if, if};
[0163] setB={if, switch, for, while, if};
[0164] Then the intersection between setA and setB represents the logical control blocks that appear in both methods: intersection = setA ∩ setB = {if, for, switch, if}
[0165] The metric CFS is calculated as follows:
[0166] CFS = (NE of the intersection) / MAX (NE of setA, NE of setB) = 4 / 5 = 0.8
[0167] Here: NE represents the number of elements in a set; MAX(NE of setA, NE of setB) represents selecting the maximum value between NE in setA and NE in setB.
[0168] The metric CFSS is used to measure the similarity between control flow sequences and is valid only when the same type of logic control blocks in one method are completely included in the other method. It takes two values: 0 and 1. If all logic control blocks in one method appear in the same order in the other method, CFSS will take the value 1; if any control blocks appear in a different order, CFSS should take the value 0.
[0169] For example, if the control blocks in the three given methods are as follows:
[0170] setA = {if, if, for, if}
[0171] setB = {if, for, if}
[0172] setC={switch, if, if, for}
[0173] You can see that setA completely contains setB. Each logic control block in setB appears in the same order in setA, so the CFSS of method A and method B is 1. Although setC includes all the elements in setB, they appear in a different order, so the CFSS of method B and method C is 0.
[0174] For the purpose of illustration, the following examples are provided for situations where the code forms are different but the logic is repeated.
[0175] Code 1:
[0176]
[0177] Code 2:
[0178]
[0179] The above code 1 and code 2 are different in form, but they are logically repeated. For human evaluators, it is easy to point out the redundant functions of the above code 1 and code 2, because the two methods implement very similar functions. However, it is difficult for most static analysis tools to detect such a situation where the code is different but the logic is similar. Usually these tools can only find simple duplication of code, that is, these repeated codes are generated by "copy-paste" of the same code block. The two methods here implement the same function, but the variable names used are different. Therefore, although the internal logic is very similar, due to some differences in the source code, the static analysis tool cannot determine that these are two logically repeated codes. For the embodiment of the present invention, the code structure similarity (SS) metric is first calculated. Since both methods take a float parameter and return a Boolean value, the values of the metrics PS and RVS are both 1. The metric LOC of the two methods is the same, so the metric DLOC=0. Then, the metric code structure similarity (SS) of the two given methods is classified as high (HIGH). For the logical similarity (LS) metric, since the two methods have the same cyclomatic complexity (CC) metric value, the metric DCC value is 0. The logic control blocks in both methods are the same ({if / else, if / else}) and appear in the same sequence, so the values of CFS and CFSS are both 1. Therefore, the logic similarity (LS) measure of the two given methods is also classified as high (HIGH).
[0180] After calculating the values of the metrics SS and LS, applying the algorithm in Table 5, the risk of these two methods having redundant functional issues will be classified as HIGH.
[0181] Example: Table 8 is a typical example of a complex logical conditional expression for determining a long method based on three risk levels.
[0182]
[0183] Table 8
[0184] As can be seen from Table 8, long methods can be determined based on cyclomatic complexity (CC).
[0185] Example: Table 9 is a typical example of a compound logical conditional expression for determining a complex class based on three risk levels.
[0186]
[0187] Table 9
[0188] In Table 9, the complex class is determined based on the number of source code lines (LOC), the number of attributes (AN), the number of methods (MN), and the maximum cyclomatic complexity (MCCmax) of the given class. Among them: the number of attributes (AN) calculates the number of attributes of a given class; the number of methods (MN) calculates the number of methods in a given class. The maximum cyclomatic complexity (MCCmax) of a given class is the maximum CC value of the methods in the given class.
[0189] The basic principle of Table 9 is that a class with too many lines of code and too many attributes and methods, or a class in which at least one method is very complex, will be identified as a suspect of a complex class.
[0190] Example: Table 10 is a typical example of a compound logical conditional expression that determines a long parameter list based on three risk levels.
[0191]
[0192] Table 10
[0193] In Table 10, the parameter number (PN) is used to measure the number of input parameters of a given method. The basic principle of the algorithm shown in Table 10 is that a method with more than 4 parameters should be regarded as having a long parameter list.
[0194] Example: Table 11 is a typical example of a compound logic conditional expression for determining a message chain based on three risk levels.
[0195]
[0196] Table 11
[0197] Message chains are determined based on the Indirect Call Number (ICN) in Table 11. The ICN metric counts the number of outgoing references to a given method that are not calls to its direct friends.
[0198] Here, calling direct friend means:
[0199] (1) A given method directly calls other methods in its class;
[0200] (2) A given method directly calls a public variable or method within its visible scope;
[0201] (3) A given method directly calls a method of the object in its input parameter;
[0202] (4) A given method calls a method of a local object created by it.
[0203] Apart from the above four cases, all other outgoing calls made by a given method will be classified as indirect calls and counted towards the ICN.
[0204] Those skilled in the art will appreciate that error-prone patterns such as code duplication, useless methods, circular dependencies, and incorrect dependencies can be directly identified by static code analysis tools, so the embodiments of the present invention will no longer describe their determination algorithms in detail.
[0205] Step 102: Input the probability determined in step 101 to an artificial neural network (ANN), and determine whether the code violates a predetermined design principle and the prediction result of the quantitative degree of violation of the design principle based on the ANN.
[0206] After determining the probability of the error-prone pattern, the embodiments of the present invention can evaluate the software design quality according to the prediction results of the artificial neural network, which can be achieved by mapping the probability of the existence of the error-prone pattern to the key design principles.
[0207] Good software design should aim to reduce the business risks associated with building technical solutions. It needs to be flexible enough to cope with changes in hardware and software technology, as well as changes in user needs and application scenarios. In the software industry, it is widely recognized that effectively following some key design principles can minimize R&D costs and maintenance workload, and improve software availability and scalability. Practice has proven that implementing and enforcing these design principles is essential to ensuring software quality.
[0208] For example, design principles may include at least one of the following: separation of concerns principle; single responsibility principle; least knowledge principle; no duplication principle; up-front design minimization principle, etc.
[0209] Table 12 is an exemplary description of the top five most well-known software design principles.
[0210]
[0211] Table 12
[0212] The above design principles are explained in more detail below.
[0213] The Separation Of Concerns (SOC) principle is one of the fundamental principles of object-oriented programming. If followed correctly, it results in loosely coupled and highly cohesive software. Therefore, error-prone patterns that reflect symptoms of tight coupling and incorrect use of component functionality, such as circular dependencies and incorrect dependencies, are clear signs that the SOC principle was not followed during the design process. In addition, SOC helps minimize the effort required to change the software. Therefore, it is also associated with error-prone patterns that make it difficult to implement changes (e.g., shotgun modifications, divergent changes), etc.
[0214] Single Responsibility Principle (SRP): If narrowed down to the class level, this principle can also be interpreted as: "There should not be more than one reason to change a class". Therefore, error-prone patterns that represent frequent code changes (e.g., shotgun modifications, divergent changes) may be related to violations of the SRP principle. In addition, methods / classes that take on too much responsibility are usually logically complex. Therefore, error-prone patterns that indicate inherent complexity (e.g., complex classes, long methods, and long parameter lists) are also related to this principle.
[0215] Principle Of Least Knowledge (LOD): This principle is also known as the Law of Demeter or LoD, which means "Don't talk to strangers". It states that a particular class should only talk to its "close friends" and not to "friends of friends". Therefore, error-prone patterns such as message chains indicate that this principle is not followed. This principle also opposes entanglement of a class with the details of other classes across different architectural levels. Therefore, error-prone patterns such as circular dependencies and false dependencies are also related to it.
[0216] Don't Repeat Yourself (DRY): This principle aims to avoid redundancy in source code, which can lead to logical inconsistencies and unnecessary maintenance efforts. Error-prone patterns, code duplication and redundant functionality are direct indicators of violations of the DRY principle.
[0217] Minimize Upfront Design: In general, this principle advocates that "large-scale" design is unnecessary and most designs should be carried out throughout the software development process. Therefore, the error-prone pattern BDUF is a clear violation of this principle. The Minimize Upfront Design principle is also known as YAGNI ("You Aren't Going to Need It"), which means that we should only do those designs that are strictly necessary to achieve our goals. Therefore, the discovery of useless methods also indicates that this principle may not be followed.
[0218] In the embodiments of the present invention, both the error-prone patterns and the design principles can be flexibly extended. If necessary, more error-prone patterns and design principles can be added to increase flexibility and applicability.
[0219] The above exemplary descriptions are typical examples of the design principles. Those skilled in the art will appreciate that such descriptions are merely exemplary and are not intended to limit the scope of protection of the embodiments of the present invention.
[0220] Table 13 is a correspondence table between error-prone patterns and design principles, where “x” indicates that the error-prone pattern violates the design principle.
[0221]
[0222]
[0223] Table 13
[0224] Embodiments of the present invention may use an artificial intelligence (AI) algorithm to simulate the above relationships in Table 13, and then evaluate compliance with these key design principles based on the occurrence of relevant error-prone patterns.
[0225] In one embodiment, the artificial neural network includes a connection between error-prone patterns and design principles; the prediction result of determining whether the code violates predetermined design principles and the quantified degree of violation of the design principles based on the artificial neural network includes: based on the connection in the artificial neural network and the probability of the existence of the error-prone pattern, predicting whether the code violates the design principles and the quantified degree of violation of the design principles.
[0226] Here, an artificial neural network is a computing model composed of a large number of nodes (or neurons) connected to each other. Each node represents a specific output function, called an activation function. The connection between each two nodes represents a weighted value for the signal passing through the connection, called a weight, which is equivalent to the memory of the artificial neural network. The output of the network varies depending on the connection mode of the network, the weight value and the activation function. Preferably, the method also includes: adjusting the weight of the connection in the artificial neural network based on a self-learning algorithm.
[0227] Preferably, the connection between error-prone patterns and design principles includes at least one of the following: the connection between shotgun modifications and the principle of separation of concerns; the connection between shotgun modifications and the single responsibility principle; the connection between divergent changes and the principle of separation of concerns; the connection between divergent changes and the single responsibility principle; the connection between large-scale up-front design and the principle of up-front design minimization; the connection between decentralized functions and the principle of no duplication; the connection between redundant functions and the principle of no duplication; the connection between circular dependencies and the principle of separation of concerns; the connection between circular dependencies and the principle of least knowledge; the connection between false dependencies and the principle of separation of concerns; the connection between false dependencies and the principle of least knowledge; the connection between complex classes and the single responsibility principle; the connection between long methods and the single responsibility principle; the connection between code duplication and the principle of no duplication; the connection between long parameter lists and the single responsibility principle; the connection between message chains and the principle of least knowledge; the connection between unused methods and the principle of up-front design minimization.
[0228] Among them, in order to infer the target (violation of design principles (such as SOC, SRP, LOD, DRY, YAGNI)) from the error-prone patterns (called extracted features in the algorithm below), artificial intelligence algorithms and artificial intelligence neural networks are applied to formulate the relationship between error-prone patterns and design principles.
[0229] Figure 2 4 is a structural diagram of an artificial neural network according to an embodiment of the present invention.
[0230] exist Figure 2 In the example, the artificial neural network includes an input layer 200, a hidden layer 202, and an output layer 204, wherein the output layer 204 includes a normalized exponential function (softmax) layer 400 and a maximum probability layer 600. The maximum probability layer 600 includes various design principles such as the pre-designed minimization principle 601, the single responsibility principle 602, and the separation of concerns principle 603. Figure 2 In the example, the neurons (also called units) represented as circles in the input layer 200 are error-prone patterns 201, and the input is linearly transformed using a weight matrix, and the input is nonlinearly transformed using an activation function, so as to increase the variability of the entire model and increase the representation of the knowledge model. The significance of the existence of artificial neural networks is based on the assumption that there is a hidden mathematical equation that can represent the relationship between the extracted features of the input and the output that violates the design principles, and the parameters in the mathematical equation are unknown. The artificial neural network can use a training data set to train the model, iteratively updating the internal parameters of a large number (for example, hundreds) of records in the data set until the trained parameters can make the equation conform to the relationship between the input and the output. At this point, these parameters embody the essence of the model, and the model can represent the connection between error-proneness and design principles.
[0231] Input layer 200: The input of the network is the features (probability of error-prone patterns) extracted from the code quality check. For example, if there is a piece of code that has a high risk of shotgun modification, then a number can be used to represent the risk level (e.g., high risk is 0.8, medium risk is 0.5, and low risk is 0.3). For example, if the model input is defined as 12 patterns, there are 12 numbers to represent the quality characteristics of this code. These numbers form a digital vector and will be used as the model input.
[0232] The output of the output layer 204 is the predicted category of violation of the design principle. For each code snippet, an attempt will be made to detect the probability of violating the principle, and the probability will be divided into three levels: low, medium and high risk. Therefore, for each of the various design principles such as the pre-design minimization principle 601, the single responsibility principle 602, the separation of concerns principle 603, etc., there will be three output units corresponding to the three categories, that is, the output of the hidden layer 202 is a 3-element vector, which is used as the input vector of the softmax layer 400 in the output layer 204. In the softmax layer 400, the elements in the vector are normalized to decimals in the range of 0 to 1 by the softmax function and summed to 1 in order to convert these elements into the probability of the category. The maximum probability layer 600 predicts the probability that the analyzed code violates a specific principle by selecting the category with the highest value as the final output.
[0233] against Figure 2 The training process of the artificial neural network shown in FIG. 1 includes forward propagation and back propagation. Forward propagation is a process based on the network of the current parameters (weights), using linear transformations of matrix multiplication, and nonlinear transformations through activation functions (sigmoid, tanh, relu) to enhance the model's degrees of freedom and generate prediction results (i.e., the probability of violating the principle). Combined with back propagation, the parameters used to calculate the prediction can be better updated to generate predictions that are closer to the true label (the principle violation does exist in this code, etc.). After a large number of training steps (for example, thousands of steps), the prediction is almost close to the true label, which means that the parameters in the model can form a relationship between the input (feature or metric) and the output (violation principle). In the back propagation step, the distance ("loss") is first calculated by the difference between the predicted value and the actual label, and the loss is generated by the cost function. "Loss" provides the network with a numerical measure of the prediction error. By applying the chain rule, the loss value can be propagated to each neuron in the network, and the parameters associated with the neuron are modified accordingly. By repeating this process many times (for example, a thousand times), the loss value becomes lower and lower, and the predicted value becomes closer and closer to the true label.
[0234] The training data of the artificial neural network includes error-prone patterns and manual evaluation of whether design principles are violated. The training data can be collected from different projects, which have completely different characteristics. Therefore, after training, the artificial neural network model can reflect the characteristics of projects in different industries, different development processes, and different situations. In other words, the trained model can be very flexible and scalable and can be applied to analyze different types of projects without pre-configuration according to the characteristics of the project, because the model can learn through the training data.
[0235] Another benefit of applying artificial neural networks is that the model is based on a self-learning algorithm. The model can update internal parameters to dynamically adjust its new data input. When the training model receives new input and propagates forward to predict violations of principles, the human reviewer can check the prediction results and provide a judgment. If the judgment is negative for the prediction, the system can obtain this round of input forward propagation and the final human judgment to form a new data sample point as a future training set and enable the model to train on previous errors and become more and more accurate after deployment. Prior to this step, the code review system successfully completed the steps of accessing user source code, code scanning for code quality analysis, determining error-prone patterns, and classifying error-prone patterns as design principle violations.
[0236] Step 103: Evaluate the design quality of the code based on the prediction result.
[0237] To evaluate software quality, a basic principle can be taken into account: if a design principle is correctly followed during the design process, its associated error-prone patterns should not be found in the source code. That is, if a certain error-prone pattern is identified, the principle is likely violated. Therefore, by observing the occurrence of error-prone patterns, we can measure the adherence to key design principles and then use the measurement to evaluate the design quality.
[0238] In one embodiment, evaluating the design quality of the code based on the prediction results includes at least one of the following: evaluating the design quality of the code based on whether it violates predetermined design principles, wherein not violating the predetermined design principles has good design quality relative to violating the predetermined design principles; evaluating the design quality of the code based on the quantitative degree of violation of the design principles, wherein the quantitative degree of violation of the design principles is inversely proportional to the evaluation of the quality of the design.
[0239] In one embodiment, in a typical software development environment, source code is stored in a central repository and all changes are controlled. The source code to be reviewed and its revision / change history information (which can be obtained from software configuration management tools such as ClearCase, Subversion, Git, etc.) will be used as the input for the code evaluation method of the present invention. When the user applies the code evaluation method of the present invention, the source code repository and the revision / change history are first accessed, and code metrics and quality-related findings are generated with the aid of a static code analysis tool. The error-prone pattern detector captures this information together with the revision / change information and calculates the likelihood of certain error-prone patterns existing in the source code. Subsequently, an AI model (such as an artificial neural network) reads these probabilities as input to predict the probability of violating design principles based on the pre-trained connection between error-prone patterns and design principles. Then, a design quality report can be generated based on the prediction results, including the predicted design principle violation information and error-prone patterns, and the user can accordingly improve the quality of their code.
[0240] Based on the above description, an embodiment of the present invention also proposes a system for evaluating code design quality.
[0241] Figure 3 It is a structural diagram of the system for evaluating code design quality according to an embodiment of the present invention.
[0242] As Figure 3 shown, the system for evaluating code design quality includes:
[0243] A code repository 300, configured to store the code 301 to be evaluated;
[0244] A static scanning tool 302, configured to statically scan the code 301 to be evaluated;
[0245] An error-prone pattern detector 303, configured to determine the probability of error-prone patterns existing in the code based on the static scanning results output by the static scanning tool 302;
[0246] An artificial neural network 304, configured to determine, based on the probability, a prediction result 305 of whether the code violates a predetermined design principle and the degree of quantification of the violation of the design principle, wherein the prediction result 305 is configured to evaluate the design quality of the code. For example, the artificial neural network 304 can automatically evaluate the design quality of the code based on the prediction result 305. Optionally, the artificial neural network 304 presents the prediction result 305, and the user manually evaluates the design quality of the code based on the display interface of the prediction result 305.
[0247] In one embodiment, the code repository 301 is further configured to store code modification records; the static scanning tool 302 is configured to determine the probability of the existence of error-prone patterns in the code based on a composite logical conditional expression containing static scanning results and / or modification records.
[0248] Based on the above description, an embodiment of the present invention further proposes a device for evaluating code design quality.
[0249] Figure 4 A structural diagram of an apparatus for evaluating code design quality according to an embodiment of the present invention.
[0250] like Figure 4 As shown, the device 400 for evaluating code design quality includes:
[0251] A determination module 401 is configured to determine the probability of an error-prone pattern existing in the code based on a static scanning result of the code;
[0252] A prediction result determination module 402 is configured to input the probability into the artificial neural network, and determine the prediction result of whether the code violates the predetermined design principle and the quantitative degree of violation of the design principle based on the artificial neural network;
[0253] The evaluation module 403 is configured to evaluate the design quality of the code based on the prediction result.
[0254] In one embodiment, error-prone patterns include at least one of the following: shotgun modifications; divergent changes; large-scale up-front design; scattered / redundant functionality; circular dependencies; incorrect dependencies; complex classes; long methods; code duplication; long parameter lists; message chains; useless methods.
[0255] In one embodiment, the design principles include at least one of the following: separation of concerns principle; single responsibility principle; least knowledge principle; no duplication principle; up-front design minimization principle.
[0256] In one embodiment, the determination module 401 is further configured to receive a modification record of the code, wherein determining the probability of the existence of an error-prone pattern in the code based on the static scanning result of the code includes: determining the probability of the existence of an error-prone pattern in the code based on a composite logical conditional expression including the static scanning result and the modification record.
[0257] In one embodiment, determining the probability of the presence of an error-prone pattern in the code based on a composite logical conditional expression including static scan results and / or modification records includes at least one of the following: determining the probability of the presence of a shotgun modification based on a composite logical conditional expression including an incoming coupling metric, an outgoing coupling metric, and a change method metric; determining the probability of the presence of a divergent change based on a composite logical conditional expression including a version modification count metric, an instability metric, and an incoming coupling metric; determining the probability of the presence of a large-scale pre-design based on a composite logical conditional expression including the number of code lines, the number of changed code lines, the number of classes, the number of modified classes, and the statistical means of the above metrics; determining the probability of the presence of a large-scale pre-design based on a composite logical conditional expression including a structure The probability of the existence of scattered functions is determined based on a composite logical conditional expression of similarity measurement and logical similarity measurement; the probability of the existence of redundant functions is determined based on a composite logical conditional expression containing structural similarity measurement and logical similarity measurement; the probability of the existence of long methods is determined based on a composite logical conditional expression containing cyclomatic complexity measurement; the probability of the existence of complex classes is determined based on a composite logical conditional expression containing the number of code lines, the number of attributes, the number of methods and the maximum cyclomatic complexity of methods of a given class; the probability of the existence of long parameter lists is determined based on a composite logical conditional expression containing the number of parameters; the probability of the existence of message chains is determined based on a composite logical conditional expression containing the number of indirect calls, and so on.
[0258] In one embodiment, the artificial neural network includes connections between error-prone patterns and design principles; the prediction result determination module 402 is configured to predict whether the code violates the design principle and the quantitative degree of violation of the design principle based on the connections in the artificial neural network and the probability of the existence of the error-prone pattern.
[0259] In one embodiment, the evaluation module 403 is configured to evaluate the design quality of the code based on whether it violates predetermined design principles, wherein not violating the predetermined design principles has good design quality relative to violating the predetermined design principles; and to evaluate the design quality of the code based on the quantitative degree of violation of the design principles, wherein the quantitative degree of violation of the predetermined design principles is inversely proportional to the evaluation of the design quality.
[0260] In one embodiment, the connection between error-prone patterns and design principles includes at least one of the following: the connection between shotgun modifications and the principle of separation of concerns; the connection between shotgun modifications and the single responsibility principle; the connection between divergent changes and the principle of separation of concerns; the connection between divergent changes and the single responsibility principle; the connection between large-scale up-front design and the principle of up-front design minimization; the connection between decentralized functions and the principle of no duplication; the connection between redundant functions and the principle of no duplication; the connection between circular dependencies and the principle of separation of concerns; the connection between circular dependencies and the principle of least knowledge; the connection between false dependencies and the principle of separation of concerns; the connection between false dependencies and the principle of least knowledge; the connection between complex classes and the single responsibility principle; the connection between long methods and the single responsibility principle; the connection between code duplication and the principle of no duplication; the connection between long parameter lists and the single responsibility principle; the connection between message chains and the principle of least knowledge; the connection between unused methods and the principle of up-front design minimization.
[0261] Figure 5 The present invention is a structural diagram of an apparatus for evaluating code design quality having a processor and a memory structure according to an embodiment of the present invention.
[0262] like Figure 5 As shown, the apparatus 500 for evaluating code design quality includes a processor 501 and a memory 502 .
[0263] The memory 502 stores an application program that can be executed by the processor 501 and is used to enable the processor 501 to execute the method for evaluating code design quality as described in the above item.
[0264] The memory 502 may be implemented as various storage media such as an electrically erasable programmable read-only memory (EEPROM), a flash memory (Flash memory), a programmable program read-only memory (PROM), etc. The processor 501 may be implemented as including one or more central processing units or one or more field programmable gate arrays, wherein the field programmable gate array integrates one or more central processing unit cores. Specifically, the central processing unit or the central processing unit core may be implemented as a CPU or an MCU.
[0265] It should be noted that not all steps and modules in the above processes and structure diagrams are necessary, and some steps or modules can be ignored according to actual needs. The execution order of each step is not fixed and can be adjusted as needed. The division of each module is only for the convenience of describing the functional division adopted. In actual implementation, a module can be implemented by multiple modules, and the functions of multiple modules can also be implemented by the same module. These modules can be located in the same device or in different devices.
[0266] The hardware modules in the various embodiments can be implemented mechanically or electronically. For example, a hardware module can include specially designed permanent circuits or logic devices (such as dedicated processors, such as FPGAs or ASICs) for performing specific operations. A hardware module can also include programmable logic devices or circuits (such as including general-purpose processors or other programmable processors) temporarily configured by software for performing specific operations. As for whether to specifically implement the hardware module mechanically, or using dedicated permanent circuits, or using temporarily configured circuits (such as configured by software), it can be determined based on cost and time considerations.
[0267] The present invention also provides a machine-readable storage medium storing instructions for causing a machine to execute the methods described herein. Specifically, a system or device equipped with a storage medium can be provided, on which software program code for implementing the functions of any one of the above-described embodiments is stored, and the computer (or CPU or MPU) of the system or device is caused to read and execute the program code stored in the storage medium. In addition, part or all of the actual operations can also be completed by an operating system or the like operating on the computer based on the instructions of the program code. The program code read from the storage medium can also be written to a memory provided in an expansion board inserted into the computer or to a memory provided in an expansion unit connected to the computer, and then based on the instructions of the program code, the CPU or the like installed on the expansion board or the expansion unit is caused to execute part and all of the actual operations, thereby implementing the functions of any one of the above-described embodiments.
[0268] Embodiments of the storage medium for providing program code include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROMs. Optionally, the program code can be downloaded from a server computer or the cloud via a communication network.
[0269] As described above, the above are only the preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
[0270] It should be noted that not all steps and modules in the above-mentioned processes and system structure diagrams are necessary, and some steps or modules can be ignored according to actual needs. The execution order of each step is not fixed and can be adjusted as needed. The system structure described in the above-mentioned embodiments can be a physical structure or a logical structure, that is, some modules may be implemented by the same physical entity, or some modules may be implemented by multiple physical entities, or some components in multiple independent devices may be implemented together.
[0271] The present invention is shown and described in detail above through the accompanying drawings and preferred embodiments. However, the present invention is not limited to these disclosed embodiments. Based on the above multiple embodiments, those skilled in the art can know that the code review methods in the above different embodiments can be combined to obtain more embodiments of the present invention, and these embodiments are also within the protection scope of the present invention.
Claims
1. A method for evaluating code design quality, It is characterized in that include: Determining the probability of the presence of an error-prone pattern in the code based on a static scanning result of the code (101), wherein the error-prone pattern is a diagnostic symptom indicating that there may be a quality problem in the software design; Inputting the probability into an artificial neural network, and determining whether the code violates a predetermined design principle and a prediction result of the quantified degree of violation of the design principle based on the artificial neural network (102); evaluating the design quality of the code based on the prediction result (103), The method further includes: receiving a modification record of the code, wherein determining the probability of the existence of an error-prone pattern in the code based on a static scanning result of the code includes: determining the probability of the existence of an error-prone pattern in the code based on a composite logical conditional expression including the static scanning result and the modification record, Wherein, the determining the probability of the existence of the error-prone pattern in the code based on the composite logical conditional expression including the static scanning result and the modification record includes at least one of the following: determining a probability of a shotgun modification being present based on a composite logical conditional expression including an afferent coupling metric, an efferent coupling metric, and a change method metric; Determine the probability of existence of divergent changes based on a composite logic conditional expression including a version modification count metric, an instability metric, and an afferent coupling metric; determining the probability of the existence of a large-scale pre-design based on a composite logical conditional expression including the number of lines of code, the number of lines of code changed, the number of classes, the number of classes changed, and the statistical means of the above metrics; Determining the probability of the existence of dispersed / redundant functions based on a composite logical conditional expression including a structural similarity measure and a logical similarity measure; Determine the probability of the existence of long methods based on a composite logical conditional expression containing a cyclomatic complexity measure; Determine the probability of the existence of a complex class based on a composite logical conditional expression containing the maximum of the number of lines of code, the number of attributes, the number of methods, and the cyclomatic complexity of the methods of a given class; Determine the probability of a long argument list based on a compound logical conditional expression involving the number of arguments; Determine the probability of a message chain existing based on a compound logical conditional expression including the number of indirect calls, The artificial neural network includes a connection between the error-prone pattern and the design principle, and the prediction result of determining whether the code violates the predetermined design principle and the quantitative degree of violation of the design principle based on the artificial neural network includes: Based on the connections in the artificial neural network and the probability of the existence of the error-prone mode, a prediction result of whether the code violates the design principle and the quantitative degree of violation of the design principle is predicted.
2. The method for evaluating code design quality according to claim 1, It is characterized in that The error-prone patterns include at least one of the following: shotgun modification; divergent changes; large-scale pre-design; scattered / redundant functions; circular dependencies; Wrong dependencies; complex classes; long methods; code duplication; long parameter lists; message chains; useless methods; and / or The design principles include at least one of the following: separation of concerns principle; single responsibility principle; The principle of least knowledge; the principle of not repeating oneself; Design minimally in advance.
3. The method for evaluating code design quality according to claim 1, It is characterized in that The predetermined threshold value in the complex logic conditional expression is adjustable.
4. The method for evaluating code design quality according to claim 1, It is characterized in that The method further includes: The weights of the connections in the artificial neural network are adjusted based on a self-learning algorithm.
5. The method for evaluating code design quality according to claim 1, It is characterized in that The evaluating the design quality of the code based on the prediction result comprises at least one of the following: Evaluate the design quality of the code based on whether it violates the predetermined design principles, where not violating the predetermined design principles has a good design quality relative to violating the predetermined design principles; The design quality of the code is evaluated based on the quantitative degree of violation of the design principle, wherein the quantitative degree of violation of the design principle is inversely proportional to the evaluation of the design quality.
6. The method for evaluating code design quality according to claim 1, It is characterized in that The connection between the error-prone patterns and design principles includes at least one of the following: The connection between shotgun changes and the separation of concerns principle; the connection between shotgun changes and the single responsibility principle; the connection between divergent changes and the separation of concerns principle; The connection between divergent change and the single responsibility principle; the connection between large-scale upfront design and the principle of minimal upfront design; The connection between decentralization of functionality and the principle of not duplicating; The connection between redundant functions and the don’t repeat yourself principle; The connection between circular dependencies and the separation of concerns principle; The connection between circular dependencies and the least knowledge principle; The connection between false dependencies and the separation of concerns principle; the connection between false dependency and the principle of least knowledge; The connection between complex classes and the single responsibility principle; the connection between long methods and the single responsibility principle; the connection between code duplication and the don't repeat principle; the connection between long parameter lists and the single responsibility principle; the connection between message chains and the least knowledge principle; The connection between unused parties and the principle of pre-design minimization.
7. A device (400) for evaluating code design quality, It is characterized in that include: A determination module (401) is configured to determine the probability of the presence of an error-prone pattern in the code based on a static scanning result of the code, wherein the error-prone pattern is a diagnostic symptom indicating that there may be a quality problem in the software design; A prediction result determination module (402) is configured to input the probability into an artificial neural network, and determine whether the code violates a predetermined design principle and a prediction result of a quantitative degree of violation of the design principle based on the artificial neural network; An evaluation module (403) is configured to evaluate the design quality of the code based on the prediction result. The determination module (401) is further configured to receive a modification record of the code, determine the probability of an error-prone pattern existing in the code based on a composite logic conditional expression including the static scanning result and the modification record, Wherein, the determining the probability of the existence of the error-prone pattern in the code based on the composite logical conditional expression including the static scanning result and the modification record includes at least one of the following: determining a probability of a shotgun modification being present based on a composite logical conditional expression including an afferent coupling metric, an efferent coupling metric, and a change method metric; Determine the probability of existence of divergent changes based on a composite logic conditional expression including a version modification count metric, an instability metric, and an afferent coupling metric; determining the probability of the existence of a large-scale pre-design based on a composite logical conditional expression including the number of lines of code, the number of lines of code changed, the number of classes, the number of classes changed, and the statistical means of the above metrics; Determining the probability of the existence of dispersed / redundant functions based on a composite logical conditional expression including a structural similarity measure and a logical similarity measure; Determine the probability of the existence of long methods based on a composite logical conditional expression containing a cyclomatic complexity measure; Determine the probability of the existence of a complex class based on a composite logical conditional expression containing the maximum of the number of lines of code, the number of attributes, the number of methods, and the cyclomatic complexity of the methods of a given class; Determine the probability of a long argument list based on a compound logical conditional expression involving the number of arguments; Determine the probability of a message chain existing based on a compound logical conditional expression including the number of indirect calls, Wherein, the artificial neural network comprises a connection between the error-prone pattern and the design principle, and the prediction result determination module (402) is further configured to: Based on the connections in the artificial neural network and the probability of the existence of the error-prone mode, a prediction result of whether the code violates the design principle and the quantitative degree of violation of the design principle is predicted.
8. The device (400) for evaluating code design quality according to claim 7, It is characterized in that The error-prone patterns include at least one of the following: shotgun modification; divergent changes; large-scale pre-design; scattered / redundant functions; circular dependencies; Wrong dependencies; complex classes; long methods; code duplication; long parameter lists; message chains; useless methods; and / or The design principles include at least one of the following: separation of concerns principle; single responsibility principle; The principle of least knowledge; the principle of not repeating oneself; Design minimally in advance.
9. The device (400) for evaluating code design quality according to claim 7, It is characterized in that The evaluation module (403) is configured to evaluate the design quality of the code based on whether it violates a predetermined design principle, wherein not violating the predetermined design principle has a good design quality relative to violating the predetermined design principle; The design quality of the code is evaluated based on the quantitative degree of violation of the design principle, wherein the quantitative degree of violation of the predetermined design principle is inversely proportional to the evaluation of the design quality.
10. The device (400) for evaluating code design quality according to claim 7, It is characterized in that The connection between the error-prone patterns and design principles includes at least one of the following: The connection between shotgun changes and the separation of concerns principle; the connection between shotgun changes and the single responsibility principle; the connection between divergent changes and the separation of concerns principle; The connection between divergent change and the single responsibility principle; the connection between large-scale upfront design and the principle of minimal upfront design; The connection between decentralization of functionality and the principle of not duplicating; The connection between redundant functions and the don’t repeat yourself principle; The connection between circular dependencies and the separation of concerns principle; The connection between circular dependencies and the least knowledge principle; The connection between false dependencies and the separation of concerns principle; the connection between false dependency and the principle of least knowledge; The connection between complex classes and the single responsibility principle; the connection between long methods and the single responsibility principle; the connection between code duplication and the don't repeat principle; the connection between long parameter lists and the single responsibility principle; the connection between message chains and the least knowledge principle; The connection between unused parties and the principle of pre-design minimization.
11. A system for evaluating the quality of code design, It is characterized in that include: A code repository (300) configured to store the code to be evaluated (301) and store modification records of the code (301); A static scanning tool (302), configured to statically scan the code to be evaluated (301); The error-prone pattern detector (303) is configured to determine the probability of the existence of an error-prone pattern in the code (301) based on a static scanning result output by a static scanning tool, wherein the error-prone pattern is a diagnostic symptom indicating that there may be a quality problem in software design, and the determination of the probability of the existence of the error-prone pattern in the code based on the static scanning result of the code comprises: determining the probability of the existence of the error-prone pattern in the code based on a composite logical conditional expression including the static scanning result and the modification record, Wherein, the determining the probability of the existence of the error-prone pattern in the code based on the composite logical conditional expression including the static scanning result and the modification record includes at least one of the following: determining a probability of a shotgun modification being present based on a composite logical conditional expression including an afferent coupling metric, an efferent coupling metric, and a change method metric; Determine the probability of existence of divergent changes based on a composite logic conditional expression including a version modification count metric, an instability metric, and an afferent coupling metric; determining the probability of the existence of a large-scale pre-design based on a composite logical conditional expression including the number of lines of code, the number of lines of code changed, the number of classes, the number of classes changed, and the statistical means of the above metrics; Determining the probability of the existence of dispersed / redundant functions based on a composite logical conditional expression including a structural similarity measure and a logical similarity measure; Determine the probability of the existence of long methods based on a composite logical conditional expression containing a cyclomatic complexity measure; Determine the probability of the existence of a complex class based on a composite logical conditional expression containing the maximum of the number of lines of code, the number of attributes, the number of methods, and the cyclomatic complexity of the methods of a given class; Determine the probability of a long argument list based on a compound logical conditional expression involving the number of arguments; Determine the probability of the existence of a message chain based on a composite logical conditional expression including the number of indirect calls; An artificial neural network (304) is configured to determine whether the code violates a predetermined design principle and a prediction result (305) of a quantitative degree of violation of the design principle based on the probability, wherein the prediction result (305) is used to evaluate the design quality of the code, The artificial neural network includes a connection between the error-prone pattern and the design principle, and the prediction result of determining whether the code violates the predetermined design principle and the quantitative degree of violation of the design principle based on the artificial neural network includes: Based on the connections in the artificial neural network and the probability of the existence of the error-prone mode, a prediction result of whether the code violates the design principle and the quantitative degree of violation of the design principle is predicted.
12. A device (500) for evaluating code design quality, It is characterized in that comprising a processor (501) and a memory (502); The memory (502) stores an application program that can be executed by the processor (501), and is used to enable the processor (501) to execute the method for evaluating code design quality according to any one of claims 1 to 6.
13. A computer-readable storage medium, It is characterized in that Computer-readable instructions are stored therein, and the computer-readable instructions are used to execute the method for evaluating code design quality according to any one of claims 1 to 6.
Citation Information
Patent Citations
Method for evaluating and predicting maintenance work load of open source software (OSS) based on code quality
CN104809066A
Similarity detection method of computer software source code
CN105426711A