Performance evaluation method and device for code large model and electronic equipment
By refining the general capability test data of the code big model into multiple sub-capacities and evaluating the score of each sub-capacity based on the feedback data, the problem that a single evaluation indicator cannot comprehensively measure the capabilities of the code big model is solved, and a more accurate and comprehensive performance evaluation is achieved.
Patent Information
- Application Number
- CN202510197461.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-27
AI Technical Summary
In the prior art, it is difficult to comprehensively measure the general capabilities of code large models, resulting in inaccurate and comprehensive performance evaluation.
By obtaining test data for the general ability of the code big model, parsing and refining it into subtest data of multiple sub-abilities, inputting it into the code big model, obtaining feedback data, and calculating the evaluation score of each sub-abilities based on the feedback data, and finally determining the evaluation result of the general ability.
It improves the accuracy and comprehensiveness of the performance evaluation of code large model, and can take into account the performance of various sub-capacities, avoiding the limitations of a single indicator.
Smart Images

Figure CN120216307A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of large model evaluation, for example, to a method and device for evaluating the performance of a code large model, and an electronic device. Background Art
[0002] Currently, in order to reduce the compilation difficulty of code, a code large model is usually used to assist developers in code compilation. The code large model can provide developers with various capabilities such as code generation, code annotation, and code interpretation. However, if there are defects in these capabilities of the code large model, it will result in low accuracy of the finally compiled code.
[0003] In the related art, usually a single evaluation metric (such as accuracy, recall, etc.) is relied on to evaluate multiple sub-capabilities in the general capabilities of the code large model. However, although a single evaluation metric can reflect the performance of the code large model in some aspects, it cannot comprehensively measure the general capabilities of the code large model. This leads to inaccurate and incomplete evaluation of the performance of the code large model. Therefore, how to improve the accuracy and comprehensiveness of the performance evaluation of the code large model has become an urgent technical problem to be solved. Summary of the Invention
[0004] To have a basic understanding of some aspects of the disclosed embodiments, a simple summary is given below. The summary is not a general review, nor is it intended to identify key / important elements or delineate the protection scope of these embodiments, but rather serves as a preface to the following detailed description.
[0005] Embodiments of the present disclosure provide a method and device for evaluating the performance of a code large model, and an electronic device, which can improve the accuracy and comprehensiveness of the performance evaluation of the code large model.
[0006] In some embodiments, the method for evaluating the performance of a code large model includes: obtaining first test data on the general capabilities of the code large model, and parsing the first test data to determine sub-test data of multiple sub-capabilities in the general capabilities; respectively inputting the sub-test data of each sub-capability into the code large model to obtain first feedback data of the code large model on each sub-capability; determining an evaluation score of each sub-capability of the code large model according to the first feedback data of each sub-capability; and determining an evaluation result of the general capabilities of the code large model according to the evaluation scores of each sub-capability.
[0007] Optionally, determining an evaluation score of each sub-capability of the code large model according to the first feedback data of each sub-capability includes: determining a current scoring criterion corresponding to the current sub-capability; scoring the first feedback data of the current sub-capability according to the current scoring criterion to obtain an evaluation score of the current sub-capability.
[0008] Optionally, the multiple sub - capabilities include: code understanding ability, code generation and completion ability, code conversion ability, unit test case generation ability, code diagnosis and optimization ability, and R & D Q&A ability.
[0009] Optionally, the performance evaluation method includes: obtaining second test data on the dedicated scenario capabilities of the code large - model; inputting the second test data into the code large - model to obtain second feedback data of the code large - model on multiple development scenarios; and determining the evaluation result of the dedicated scenario capabilities of the code large - model according to the second feedback data of the code large - model on multiple development scenarios.
[0010] Optionally, determining the evaluation result of the dedicated scenario capabilities of the code large - model according to the second feedback data of the code large - model on multiple development scenarios includes: determining the development ability score of the code large - model for each development scenario according to the second feedback data of the code large - model for each development scenario; and determining the evaluation result of the dedicated scenario capabilities of the code large - model according to the development ability scores of the code large - model for each development scenario.
[0011] Optionally, the multiple development scenarios include: website development scenario, desktop application development scenario, mobile application development scenario, database development scenario, big data development scenario, artificial intelligence development scenario, and embedded development scenario.
[0012] Optionally, the performance evaluation method includes: obtaining third test data on the application maturity of the code large - model; where the evaluation dimensions of the application maturity include data maturity, model maturity, and service maturity; inputting the third test data into the code large - model to obtain third feedback data; determining the maturity scores of the code large - model in data maturity, model maturity, and service maturity according to the third feedback data, and determining the evaluation result of the application maturity of the code large - model according to the maturity scores.
[0013] In some embodiments, a performance evaluation device for a code large - model includes: an acquisition module configured to acquire first test data on the general capabilities of the code large - model and parse the first test data to determine sub - test data of multiple sub - capabilities in the general capabilities; a test module configured to input the sub - test data of each sub - capability into the code large - model respectively to obtain first feedback data of the code large - model on each sub - capability; a calculation module configured to determine the evaluation score of each sub - capability of the code large - model according to the first feedback data of each sub - capability; and an evaluation module configured to determine the evaluation result of the general capabilities of the code large - model according to the evaluation scores of each sub - capability.
[0014] In some embodiments, a performance evaluation device for a code large - model includes a processor and a memory storing program instructions, and the processor is configured to be able to execute the performance evaluation method for a code large - model as described above.
[0015] In some embodiments, an electronic device includes: a device body; and a performance evaluation device for a code large model as described above, which is provided on the device body.
[0016] The performance evaluation method, device, and electronic device for a code large model provided by the embodiments of the present disclosure can achieve the following technical effects:
[0017] In the embodiments of the present disclosure, first, the first test data for the general capabilities of the code large model is parsed and refined into sub-test data for multiple sub-capabilities. Then, the sub-test data for each sub-capability is separately input into the code large model to obtain the first feedback data for each sub-capability of the code large model, and the evaluation score for each sub-capability of the code large model is calculated based on the first feedback data for each sub-capability. Finally, the evaluation result for the general capabilities of the code large model is determined based on the evaluation scores for each sub-capability. In this way, the evaluation of each sub-capability in the general capabilities is carried out independently, and the performance evaluation of the code large model is no longer limited to a single ability index. Moreover, by determining the evaluation result for the general capabilities of the code large model based on the evaluation scores for each sub-capability of the code large model, the performance of each sub-capability in the general capabilities of the code large model can be taken into account. Therefore, the embodiments of the present disclosure can improve the accuracy and comprehensiveness of the performance evaluation of the code large model.
[0018] The above general description and the following description are only exemplary and explanatory, and are not used to limit this application. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] One or more embodiments are exemplarily illustrated by the corresponding drawings. These exemplary illustrations and the drawings do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are shown as similar elements. The drawings do not constitute a scale limitation, and among them:
[0020] Figure 1 is a schematic diagram of an electronic device provided by an embodiment of the present disclosure;
[0021] Figure 2 is a schematic diagram of a performance evaluation method for a code large model provided by an embodiment of the present disclosure;
[0022] Figure 3 is a schematic diagram of another performance evaluation method for a code large model provided by an embodiment of the present disclosure;
[0023] Figure 4 is a schematic diagram of another performance evaluation method for a code large model provided by an embodiment of the present disclosure;
[0024] Figure 5It is a schematic diagram of a performance evaluation device for a code large model provided by an embodiment of the present disclosure;
[0025] Figure 6 It is a schematic diagram of another performance evaluation device for a code large model provided by an embodiment of the present disclosure. Detailed implementation manners
[0026] In order to more comprehensively understand the features and technical content of the embodiments of the present disclosure, the implementation of the embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. The accompanying drawings are only for reference and explanation, and are not intended to limit the embodiments of the present disclosure. In the following technical description, for the sake of explanation, numerous details are provided to give a thorough understanding of the disclosed embodiments. However, one or more embodiments may still be implemented without these details. In other instances, well-known structures and devices may be shown in a simplified manner to simplify the drawings.
[0027] In the description of the embodiments of the present disclosure, the terms "first", "second", etc. in the specification, claims and the above drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such data may be interchanged under appropriate circumstances, so as to implement the embodiments of the present disclosure described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion.
[0028] Unless otherwise specified, the term "a plurality of" features means two or more.
[0029] In the embodiments of the present disclosure, the character " / " features indicate that the front and rear objects are in an "or" relationship. For example, the A / B feature means: A or B.
[0030] The term "and / or" is a description of the association relationship of an object, and the feature indicates that there can be three relationships. For example, A and / or B, the feature means: A or B, or, the three relationships of A and B.
[0031] The term "corresponding" may refer to an association relationship or a binding relationship. A corresponding to B means that there is an association relationship or a binding relationship between A and B.
[0032] It should be noted that, without conflict, the embodiments in the embodiments of the present disclosure and the features in the embodiments may be combined with each other.
[0033] As Figure 1 shown, the electronic device 100 provided by the embodiments of the present disclosure includes a device body 110 and a performance evaluation device 500(600) for a code large model.
[0034] Specifically, the performance evaluation device 500(600) for a code large model is disposed on the device body 110.
[0035] Optionally, the performance evaluation device 600 for the code large model includes a processor. The processor can obtain and parse the first test data on the general capabilities of the code large model, and determine the sub-test data of multiple sub-capabilities in the general capabilities. The sub-test data of each sub-capability can be separately input into the code large model to obtain the first feedback data of the code large model on each sub-capability, and the evaluation score of each sub-capability of the code large model can be determined according to the first feedback data of each sub-capability. The evaluation result of the general capabilities of the code large model can be determined according to the evaluation scores of each sub-capability.
[0036] Combined with the above electronic device, an embodiment of the present disclosure provides a performance evaluation method for a code large model, as Figure 2 shown. The performance evaluation method includes:
[0037] S201, the processor obtains the first test data on the general capabilities of the code large model, and parses the first test data to determine the sub-test data of multiple sub-capabilities in the general capabilities.
[0038] Specifically, the first test data is a test data set pre-constructed by developers according to each sub-capability to be evaluated in the general capabilities of the code large model. Therefore, the sub-test data of multiple sub-capabilities in the general capabilities can be determined by parsing the first test data.
[0039] Specifically, by obtaining and parsing the first test data on the general capabilities of the code large model and further refining it into the sub-test data of multiple sub-capabilities, the subsequent evaluation of the code large model is no longer limited to a single or several main ability indicators, but covers all key sub-capabilities that the code large model may involve, such as code understanding, generation, repair, translation, etc. This fine-grained evaluation method can ensure the comprehensiveness of the performance evaluation of the code large model.
[0040] S202, the processor separately inputs the sub-test data of each sub-capability into the code large model to obtain the first feedback data of the code large model on each sub-capability.
[0041] Specifically, the sub-test data of each sub-capability is separately input into the code large model, and the first feedback data of the code large model on each sub-capability is collected. The independent and accurate evaluation of each sub-capability in the general capabilities can be carried out, so as to accurately identify the performance differences and potential problems of the code large model in different sub-capabilities.
[0042] S203, the processor determines the evaluation score of each sub-capability of the code large model according to the first feedback data of each sub-capability.
[0043] Specifically, according to the first feedback data of each sub - ability, it can be analyzed and determined how the code large - model performs when executing each sub - ability. According to the performance of the code large - model when executing each sub - ability, it can be evaluated whether each sub - ability of the code large - model can meet the usage requirements, and the evaluation score is a form of manifestation reflecting whether each sub - ability can meet the usage requirements. Therefore, the processor can determine the evaluation score of each sub - ability of the code large - model according to the first feedback data of each sub - ability.
[0044] S204, the processor determines the evaluation result of the general ability of the code large - model according to the evaluation score of each sub - ability.
[0045] Specifically, based on the range where the evaluation score of each sub - ability is located, it can be evaluated how each sub - ability in the general ability of the code large - model performs during the test. Therefore, according to the evaluation score of each sub - ability, the evaluation result of the general ability of the code large - model can be determined.
[0046] In the embodiments of the present disclosure, first, the first test data of the general ability of the code large - model is parsed and refined into sub - test data of multiple sub - abilities. Then, the sub - test data of each sub - ability is respectively input into the code large - model to obtain the first feedback data of the code large - model for each sub - ability, and the evaluation score of each sub - ability of the code large - model is calculated according to the first feedback data of each sub - ability. Finally, the evaluation result of the general ability of the code large - model is determined according to the evaluation score of each sub - ability. In this way, the evaluation of each sub - ability in the general ability is carried out independently, and the performance evaluation of the code large - model is no longer limited to a single ability index. And by the evaluation scores of each sub - ability of the code large - model, the evaluation result of the general ability of the code large - model is determined, which can take into account the performance of each sub - ability in the general ability of the code large - model. Therefore, the embodiments of the present disclosure can improve the accuracy and comprehensiveness of the performance evaluation of the code large - model.
[0047] In some embodiments, determining the evaluation score of each sub - ability of the code large - model according to the first feedback data of each sub - ability includes: determining the current scoring criterion corresponding to the current sub - ability; scoring the first feedback data of the current sub - ability according to the current scoring criterion to obtain the evaluation score of the current sub - ability.
[0048] It can be understood that the scoring criteria for different sub - abilities are different. Therefore, when evaluating each sub - ability, it is necessary to first determine the current scoring criterion corresponding to the current sub - ability.
[0049] Specifically, the current sub - ability can be scored by evaluating the degree to which the first feedback data of the current sub - ability meets each criterion in the current scoring criterion, so as to determine the evaluation score of the current ability.
[0050] Optionally, the multiple sub - capabilities include: code understanding ability, code generation and completion ability, code conversion ability, unit test case generation ability, code diagnosis and optimization ability, and R & D Q&A ability.
[0051] In some embodiments, the code understanding ability refers to the ability of the code large - model to deeply understand the code intention. Code understanding usually takes code as input and natural language as output, aiming to help developers quickly understand the code. The code understanding ability includes two sub - abilities: code interpretation and code annotation.
[0052] Optionally, the code interpretation ability refers to the ability to generate natural language to interpret the input code or code file.
[0053] Specifically, the scoring criteria corresponding to the code interpretation ability are shown in Table 1 below:
[0054] Table 1
[0055]
[0056] Specifically, for the code interpretation ability:
[0057] Input diversity refers to the number of types of code supported by the code large - model when performing code interpretation. The types of code include: line - level, block - level, function / class - level, single - file - level, and multi - file - level with logical relationships, etc.
[0058] Task diversity refers to the number of types of tasks supported by the code large - model when performing code interpretation. The types of tasks include: generating an explanation that meets the requirements according to the instruction based on the input code; generating a more reasonable explanation for the code through multiple - round conversations; generating a more reasonable explanation according to the context; generating a complete and reasonable explanation according to the code file; generating a design document according to the project - level code file, etc.
[0059] The completeness of supported languages refers to the number of types of programming languages and the number of types of interpretation languages supported by the code large - model when performing code interpretation. The types of programming languages include: Python, Java, JavaScript, C, C++, C#, TypeScript, and Golang, etc. The types of interpretation languages include: Chinese, English, and other natural languages.
[0060] Acceptability refers to the rating of the readability, integrity, security, and relevance to the source code of the code interpretation content output by the code large - model when performing code interpretation.
[0061] Exemplarily, the acceptability of the code interpretation content is evaluated according to the requirements of Table 2 below.
[0062] Table 2
[0063]
[0064] Specifically, a model for detecting the acceptable level of code interpretation content is pre - constructed. By inputting the code interpretation content into this model, the acceptability of the code interpretation content is determined.
[0065] Accuracy refers to: the average accuracy of the code interpretation content output by the code large - model during code interpretation compared to the standard interpretation content. Accuracy can be evaluated using metrics such as Exact Match (EM), text similarity, and keyword matching. Different weights are assigned to multiple metrics, the accuracy score is calculated, and the accuracy level to which the code interpretation content belongs is determined according to the range of the accuracy score.
[0066] Optionally, code annotation ability refers to the ability to add text annotations to the input code to explain the role, function, and implementation method of the code.
[0067] Specifically, the scoring criteria for code annotation ability are the same as those in the above - mentioned scoring criteria for code interpretation ability, and will not be elaborated here.
[0068] In some embodiments, code generation and completion ability refers to the ability to automatically generate or complete the required code according to the input of natural language or code. Its purpose is to help developers improve code development efficiency. Code generation and completion ability includes two sub - abilities: code generation and code completion.
[0069] Optionally, code generation ability refers to the ability to automatically generate the required code during the coding process according to the input of natural language or code.
[0070] Specifically, the scoring criteria for code generation ability are shown in Table 3 below:
[0071] Table 3
[0072]
[0073] Specifically, for code generation ability:
[0074] Task diversity refers to: the number of types of tasks supported by the code large - model during code generation. The types of tasks include: generating a code snippet with complete functionality according to a natural language instruction or annotation; generating a code snippet with complete functionality according to a function / class name; generating a code snippet with complete functionality according to context code; generating a code file with complete functionality according to a natural language instruction or document content; generating a code file with complete logic and functionality according to a natural language instruction or document content.
[0075] The completeness of supported languages refers to the number of programming languages supported by the code large model during code generation. The types of programming languages can include: Python, Java, JavaScript, C, C++, C#, TypeScript, Golang, etc.
[0076] Acceptability refers to the rating of the readability, integrity, security, and relevance to the source code of the code generated by the code large model during code generation. Exemplarily, the acceptability of the generated code is evaluated according to the requirements in Table 4 below.
[0077] Table 4
[0078]
[0079] Specifically, a model for detecting the acceptability level of the generated code is pre-built, and by inputting the generated code into this model, the acceptability of the generated code is determined.
[0080] Accuracy refers to the average accuracy of the code generated by the code large model compared to the standard code. Accuracy is evaluated using metrics such as Pass@k and CodeBLEU. Different weights are assigned to multiple metrics, and the accuracy score is calculated. The accuracy level of the generated code is determined according to the range in which the accuracy score falls. Among them, the Pass@k metric refers to the probability that the generated code passes the unit test. The CodeBLEU metric refers to the similarity between the generated code and the target code.
[0081] Optionally, code completion ability refers to automatically completing the missing code based on the input of the existing code when writing code.
[0082] Specifically, the scoring criteria for code completion ability are the same as those for the above-mentioned code generation ability, and will not be elaborated here.
[0083] In some embodiments, code conversion ability refers to rewriting and converting the source code while maintaining the original function to support different programming languages, different environments, architectures, etc. Both the input and output are code, and the purpose is to help developers improve the efficiency of quickly converting code.
[0084] Specifically, the scoring criteria for code conversion ability are shown in Table 5 below:
[0085] Table 5
[0086]
[0087] Specifically, for code conversion ability:
[0088] The conversion coverage refers to: the number of types supported by the code large model for code conversion between different programming languages, different frameworks, and different environments.
[0089] The task diversity refers to: the number of types of tasks supported by the code large model during code conversion. The types of tasks include: converting line-level, block-level, function / class-level, single-file-level, and multi-file-level code with logical relationships.
[0090] The acceptability refers to: the rating of the converted code output by the code large model in terms of readability, integrity, security, and relevance to the source code during code conversion. Exemplarily, the acceptability of the converted code is evaluated according to the requirements in Table 6 below.
[0091] Table 6
[0092]
[0093]
[0094] Specifically, a model for detecting the acceptability level of the converted code is pre-built, and by inputting the converted code into this model, the rating of the acceptability of the converted code is determined.
[0095] The accuracy refers to: the average accuracy of the converted code output by the code large model compared to the standard code (source code) during code conversion. The accuracy is evaluated using metrics such as Pass@k and CodeBLEU. Different weights are assigned to multiple metrics, the accuracy score is calculated, and the accuracy level of the converted code is determined according to the range in which the accuracy score falls.
[0096] In some embodiments, the unit test case generation ability refers to the ability to convert the input code into unit test cases. Through this function, it can help developers reduce the time spent writing unit test cases and improve the test coverage and quality and efficiency.
[0097] Specifically, the scoring criteria corresponding to the unit test case generation ability are shown in Table 7 below:
[0098] Table 7
[0099]
[0100]
[0101] Specifically, for the unit test case generation ability:
[0102] Task diversity refers to the number of types of tasks supported by the code large model when generating unit test cases. The types of tasks include: generating unit test cases for function-level code; generating unit test cases for class-level code; generating unit test cases for file-level code, etc.
[0103] The completeness of supported languages refers to the number of programming languages supported by the code large model when generating unit test cases. The types of programming languages can include: Python, Java, JavaScript, C, C++, C#, TypeScript, and Golang, etc.
[0104] The requirements for acceptability include: supporting mainstream unit test frameworks (such as JUnit, Pytest, etc.), naming that can reflect the test purpose, being readable and maintainable, conforming to the best practices and standards of unit testing, containing key comments, having reasonable execution speed and resource consumption, and being able to automatically stub external dependencies with reasonable stub usage.
[0105] The requirements for accuracy include: syntactic correctness (whether the generated unit test cases conform to syntactic correctness); compilation correctness (whether the generated unit test cases can be compiled successfully without compilation errors; whether there are compilation errors caused by code hallucinations or missing dependency information); coverage sufficiency (the coverage of the generated unit tests for the code under test, such as test scenario coverage, statement coverage, branch coverage, and method coverage, etc.); test verification strength (the test strength of the generated unit test cases, such as whether there are test assertions and the correctness of the assertions); error detection rate and false positive rate (the proportion of code errors detected by the generated unit test cases, and the proportion of false positives).
[0106] In some embodiments, the code diagnosis and optimization ability refers to the ability to discover problems, provide repair suggestions, and perform optimization by checking the source code in dimensions such as normality, accuracy, and security, with the aim of helping developers improve code quality. The code diagnosis and optimization ability includes three sub-abilities: code inspection, code repair, and code optimization.
[0107] Optionally, the code inspection ability refers to inspecting the input code or code file and outputting the problems detected.
[0108] Specifically, the scoring criteria corresponding to the code inspection ability are shown in Table 8 below:
[0109] Table 8
[0110]
[0111]
[0112] Specifically, for code inspection capabilities:
[0113] Input diversity refers to the number of types of input code supported by the code large model during code inspection. The types of code include: line level, block level, function / class level, single file level, and multi-file level with logical relationships, etc.
[0114] Task diversity refers to the number of types of tasks supported by the code large model during code inspection. The types of tasks include: code normality issue inspection, code smell inspection, code syntax error inspection, code logic error inspection, and code security vulnerability inspection, etc.
[0115] Completeness of supported languages refers to the number of programming languages supported by the code large model during code inspection. The types of programming languages include: Python, Java, JavaScript, C, C++, C#, TypeScript, and Golang, etc.
[0116] Acceptability refers to the rating of the code inspection results output by the code large model. Exemplarily, the acceptability of the code inspection results is evaluated according to the requirements in Table 9 below.
[0117] Table 9
[0118]
[0119]
[0120] Specifically, a model for detecting the acceptability level of code inspection results is pre-built, and by inputting the code inspection results into this model, the rating of the acceptability of the code inspection results is determined.
[0121] Accuracy refers to the accuracy level of the code inspection results comprehensively measured according to indicators such as the error detection rate and false positive rate. Different weights are assigned to the two indicators of the error detection rate and false positive rate, the accuracy score is calculated, and the accuracy level to which the code inspection results belong is determined according to the range in which the accuracy score is located.
[0122] Specifically, the error detection rate of the code inspection results is calculated according to the following expression:
[0123]
[0124] where, P W is the error detection rate, W1 is the number of correctly detected errors, and W is the total number of actual errors.
[0125] Specifically, the false positive rate of the code inspection results is calculated according to the following expression:
[0126]
[0127] Among them, P N is the false positive rate, N1 is the number of errors detected incorrectly, and N is the total number of errors detected.
[0128] Optionally, the code repair ability refers to the ability to repair defects in the code according to the code inspection results.
[0129] Specifically, the scoring criteria corresponding to the code repair ability are shown in Table 10 below:
[0130] Table 10
[0131]
[0132]
[0133] Specifically, for the code repair ability:
[0134] The input and output diversity refers to the number of types of code supported for input and output by the code large model during code repair. The types of code include: line level, block level, function / class level, single file level, and multi-file level with logical relationships, etc.
[0135] The task diversity refers to: the number of types of tasks supported by the code large model during code repair. The types of tasks include: syntax error repair, logical error repair, code smell repair, and security vulnerability repair, etc.
[0136] The completeness of supported languages refers to: the number of programming languages supported by the code large model during code repair. The types of programming languages include: Python, Java, JavaScript, C, C++, C#, TypeScript, and Golang, etc.
[0137] The acceptability refers to: the rating of the repaired code output by the code large model. Exemplarily, the acceptability of the repaired code is evaluated according to Table 11 below.
[0138] Table 11
[0139]
[0140]
[0141] Specifically, a model for detecting the acceptability level of the repaired code is pre-constructed, and by inputting the repaired code into this model, the rating of the acceptability of the repaired code is determined.
[0142] Accuracy refers to: comprehensively measuring the accuracy level of the repaired code based on indicators such as the error repair rate and the error introduction rate. Different weights are assigned to the two indicators of the error repair rate and the error introduction rate, the accuracy score is calculated, and the accuracy level to which the repaired code belongs is determined according to the range where the accuracy score is located.
[0143] Specifically, the error repair rate of the repaired code is calculated according to the following expression:
[0144]
[0145] where C W is the error repair rate, W1 is the number of errors correctly repaired, and W is the total number of errors.
[0146] Specifically, the error introduction rate of the repaired code is calculated according to the following expression:
[0147]
[0148] where B N is the error introduction rate, N1 is the number of new errors introduced during repair, and N is the total number of known errors and newly introduced errors.
[0149] Optionally, the code optimization ability refers to the ability to optimize the code according to natural language instructions and generate improved code.
[0150] Specifically, the scoring criteria corresponding to the code optimization ability are shown in Table 12 below:
[0151] Table 12
[0152]
[0153]
[0154] Specifically, for the code optimization ability:
[0155] The input and output diversity refers to: the number of types of code supported by the code large model during code optimization for the input code and the output code. The types of code include: line level, block level, function / class level, single file level, and multi-file level with logical relationships, etc.
[0156] The task diversity refers to: the number of types of tasks supported by the code large model during code optimization. The types of tasks include: normative optimization, performance optimization, and code refactoring.
[0157] The completeness of supported languages refers to the number of programming languages supported by the code large model during code optimization. The types of programming languages include: Python, Java, JavaScript, C, C++, C#, TypeScript, and Golang, etc.
[0158] Acceptability refers to the rating of the optimized code output by the code large model. Exemplarily, the acceptability of the optimized code is evaluated according to the requirements in Table 13 below.
[0159] Table 13
[0160]
[0161]
[0162] Specifically, a model for detecting the acceptability level of the optimized code is pre - constructed. By inputting the optimized code into this model, the rating of the acceptability of the optimized code is determined.
[0163] Accuracy refers to the accuracy level of the optimized code evaluated according to indicators such as CodeBLEU and refactoring accuracy. Different weights are assigned to the two indicators of CodeBLEU and refactoring accuracy, and the accuracy score is calculated. According to the range where the accuracy score is located, the accuracy level to which the optimized code belongs is determined.
[0164] In some embodiments, the R & D Q&A ability refers to the ability of the code large model to understand questions and generate answers in the face of questions raised by developers, with the aim of providing developers with the needs of answering questions and retrieving information.
[0165] Specifically, the scoring criteria for the code optimization ability are shown in Table 14 below:
[0166] Table 14
[0167]
[0168] Specifically, for the R & D Q&A ability:
[0169] P G Value refers to the completion rate of multi - round dialogue tasks when the code large model conducts R & D Q&A, which is calculated through the following expression:
[0170]
[0171] Where, P G is the completion rate of multi - round dialogue tasks, G1 is the number of multi - round dialogues that complete the task within the time, and G is the total number of multi - round dialogues.
[0172] The EM value refers to the EM value of the text generation result when the code large model conducts R & D Q&A, which is calculated through the following expression:
[0173]
[0174] Among them, c is the answer generated by the code large model during R & D Q&A, r is the correct answer of the manual standard, and Count exactmatch (c, r) is the number of sentences where the answer generated by the code large model and the correct answer are exactly matched, and Count(c) is the number of answers generated by the code large model during R & D Q&A.
[0175] Task diversity refers to the number of types of tasks supported by the code large model during R & D Q&A. The types of tasks include: answering questions related to common R & D fields; answering questions related to the learning and use of third-party libraries; providing consultation, troubleshooting, and solution methods for various questions.
[0176] Acceptability refers to the rating of the Q&A results output by the code large model. Exemplarily, the acceptability of the Q&A results is evaluated according to the requirements of Table 15 below.
[0177] Table 15
[0178]
[0179] The embodiments of the present disclosure provide another performance evaluation method for the code large model, as Figure 3 shown, this performance evaluation method includes:
[0180] S301, the processor obtains the first test data of the general capabilities of the code large model, and parses the first test data to determine the sub-test data of multiple sub-capabilities in the general capabilities.
[0181] S302, the processor inputs the sub-test data of each sub-capability into the code large model respectively, and obtains the first feedback data of the code large model for each sub-capability.
[0182] S303, the processor determines the evaluation score of each sub-capability of the code large model according to the first feedback data of each sub-capability.
[0183] S304, the processor determines the evaluation result of the general capabilities of the code large model according to the evaluation scores of each sub-capability.
[0184] S305, the processor obtains the second test data of the dedicated scenario capabilities of the code large model.
[0185] Specifically, the dedicated scenario capabilities of the code large model refer to the capabilities demonstrated by the code large model in various application development scenarios. The second test data is a test data set pre-constructed by developers according to multiple development scenarios to be evaluated in the dedicated scenario capabilities of the code large model.
[0186] Specifically, in addition to the general capabilities for code development, the code large model will also design dedicated scenario development capabilities to achieve the development of various information technologies (such as websites, desktop applications, mobile applications, databases, big data, artificial intelligence, and embedded systems, etc.). Therefore, after the evaluation of the general capabilities of the code large model is completed, it is necessary to obtain the second test data for the dedicated scenario capabilities of the code large model.
[0187] S306, the processor inputs the second test data into the code large model to obtain the second feedback data of the code large model for multiple development scenarios.
[0188] S307, the processor determines the evaluation result of the dedicated scenario capabilities of the code large model according to the second feedback data of the code large model for multiple development scenarios.
[0189] Specifically, by comparing the second feedback data of each development scenario with the standard feedback data, the development capabilities of the code large model for each development scenario can be evaluated. Therefore, according to the second feedback data of the code large model for multiple development scenarios, the evaluation result of the dedicated scenario capabilities of the code large model can be determined.
[0190] Optionally, determining the evaluation result of the dedicated scenario capabilities of the code large model according to the second feedback data of the code large model for multiple development scenarios includes: determining the development capability score of the code large model for each development scenario according to the second feedback data of the code large model for each development scenario; determining the evaluation result of the dedicated scenario capabilities of the code large model according to the development capability scores of the code large model for each development scenario.
[0191] Specifically, by comparing the second feedback data and the standard feedback data according to the scoring criteria corresponding to each development scenario, the deficiencies in the development capabilities of the code large model for each development scenario can be determined. Therefore, the development capability scores of the code large model for multiple development scenarios can be determined according to the second feedback data.
[0192] Specifically, based on the range where the development capability score of the code large model for each development scenario is located, the specific performance of the code large model when developing each scenario can be evaluated. Therefore, according to the development capability scores of the code large model for each development scenario, the evaluation result of the dedicated scenario capabilities of the code large model can be determined.
[0193] Optionally, multiple development scenarios include: web development scenario, desktop application development scenario, mobile application development scenario, database development scenario, big data development scenario, artificial intelligence development scenario, and embedded development scenario.
[0194] Specifically, for the development ability of the code large model in the web development scenario, the development ability score of the web development scenario can be determined from the following aspects: the types of tasks supported by the code large model when in the web development scenario, the web front-end development frameworks supported, the web back-end development frameworks supported, and the accuracy of the corresponding code generated, by comparing the second feedback data output by the code large model and the standard feedback data.
[0195] Specifically, for the development ability of the code large model in the desktop application development scenario, the development ability score of the desktop application development scenario can be determined from the following aspects: the types of tasks supported by the code large model when in the desktop application development scenario, the development frameworks supported, and the accuracy of the corresponding code generated, by comparing the second feedback data output by the code large model and the standard feedback data.
[0196] Specifically, for the development ability of the code large model in the mobile application development scenario, the development ability score of the desktop application development scenario can be determined from the following aspects: the types of tasks supported by the code large model when in the mobile application development scenario, the development frameworks supported, and the accuracy of the corresponding code generated, by comparing the second feedback data output by the code large model and the standard feedback data.
[0197] Specifically, for the development ability of the code large model in the database development scenario, the development ability score of the database development scenario can be determined from the following aspects: the types of tasks supported by the code large model when in the database development scenario, the types of databases supported, the complexity of the SQL statements supported, and the accuracy of the execution of the SQL statements generated, by comparing the second feedback data output by the code large model and the standard feedback data.
[0198] Specifically, for the development ability of the code large model in the big data development scenario, the development ability score of the big data development scenario can be determined from the following aspects: the types of tasks supported by the code large model when in the big data development scenario (including the types of data processing tasks supported and the types of data query tasks supported), the types of development languages supported, the processing frameworks supported, the complexity of the SQL statements supported, and the accuracy of the execution of the SQL statements generated, by comparing the second feedback data output by the code large model and the standard feedback data.
[0199] Specifically, for the development ability of the code large model in the artificial intelligence development scenario, the development ability score of the artificial intelligence development scenario can be determined from the following aspects: by comparing the second feedback data output by the code large model with the standard feedback data, the types of tasks supported by the code large model when in the artificial intelligence development scenario (including various tasks in the fields of machine learning, computer vision, and natural language processing), the supported processing frameworks and tools, and the accuracy of the generated corresponding code.
[0200] Specifically, for the development ability of the code large model in the embedded development scenario, the development ability score of the embedded development scenario can be determined from the following aspects: by comparing the second feedback data output by the code large model with the standard feedback data, the types of tasks supported by the code large model when in the embedded development scenario, the complexity of supporting embedded development, and the accuracy of the generated corresponding code.
[0201] In the embodiments of the present disclosure, after the general ability evaluation of the code large model is completed, the second test data of the dedicated scenario ability of the code large model is also obtained, and the dedicated scenario ability of the code large model is tested according to the second feedback data obtained by inputting the second test data into the code large model. In this way, the comprehensiveness of the performance evaluation of the code large model is improved.
[0202] The embodiments of the present disclosure provide another performance evaluation method for the code large model, as Figure 4 shown, the performance evaluation method includes:
[0203] S401, the processor obtains the first test data of the general ability of the code large model, and parses the first test data to determine the sub-test data of multiple sub-abilities in the general ability.
[0204] S402, the processor inputs the sub-test data of each sub-ability into the code large model respectively, and obtains the first feedback data of the code large model for each sub-ability.
[0205] S403, the processor determines the evaluation score of each sub-ability of the code large model according to the first feedback data of each sub-ability.
[0206] S404, the processor determines the evaluation result of the general ability of the code large model according to the evaluation scores of each sub-ability.
[0207] S405, the processor obtains the third test data of the application maturity of the code large model; wherein, the evaluation dimensions of the application maturity include data maturity, model maturity, and service maturity.
[0208] Specifically, the application maturity of the code large model refers to the performance of the code large model in actual applications. The third test data is a test data set pre-constructed by developers according to multiple dimensions for evaluating the application maturity of the code large model.
[0209] S406, the processor inputs the third test data into the code large model to obtain the third feedback data.
[0210] S407, the processor determines the maturity scores of the code large model in terms of data maturity, model maturity, and service maturity based on the third feedback data, and determines the evaluation result of the application maturity of the code large model according to the maturity scores.
[0211] Specifically, according to the third feedback data output by the code large model after the third test data is input into it, the performance of the code large model in the three dimensions of data maturity, model maturity, and service maturity can be evaluated. The maturity score is a form reflecting the performance of the code large model in data maturity, model maturity, and service maturity. Therefore, the maturity scores of the code large model in data maturity, model maturity, and service maturity are determined according to the third feedback data.
[0212] Specifically, for the data maturity of the code large model, it can be evaluated from the data classification and grading ability and data security protection ability of the code large model. Specifically, the data classification and grading ability is to evaluate whether the code large model has the ability to manage according to types and levels. The data security protection ability is to evaluate the security and compliance of the data used by the code large model.
[0213] Specifically, for the model maturity of the code large model, it can be evaluated from the accuracy performance and reasoning performance of the output information of the code large model. Specifically, the accuracy performance is evaluated from aspects such as the accuracy, generalization, robustness, performance of the evaluation task, and error analysis and optimization of the output content of the code large model according to the third feedback data. The reasoning performance is evaluated from aspects such as the single-machine reasoning performance, parallel reasoning performance, resource utilization, reasoning cost, and stability of the code large model according to the third feedback data.
[0214] Specifically, for the service maturity of the code large model, it can be evaluated from the traceability, risk assessment, and maintainability of the code large model. Specifically, the traceability is evaluated from aspects such as training traceability, deployment traceability, data traceability, model traceability, and service traceability according to the third feedback data. The risk assessment is evaluated from aspects such as external attack prevention, data protection, dialogue attack prevention, model output control, and continuous optimization according to the third feedback data. The maintainability is evaluated from aspects such as the deployment process, deployment environment, deployment mode, operation and maintenance ability, high-availability guarantee, scalability, compatibility, and feedback and optimization according to the third feedback data.
[0215] Specifically, based on the numerical range of the maturity score, the specific performance of the code large model in actual applications can be evaluated. Therefore, according to the maturity score, the evaluation result of the application maturity of the code large model can be determined.
[0216] In the embodiments of the present disclosure, after the general ability of the code large model is evaluated, the third test data of the application maturity of the code large model is also obtained, and the application maturity of the code large model is tested according to the third feedback data obtained by inputting the third test data into the code large model. In this way, the comprehensiveness of the performance evaluation of the code large model is improved.
[0217] Combined with Figure 5 As shown in
[0218] Combined with Figure 6 As shown in
[0219] In addition, when the logical instructions in the above-mentioned memory 602 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium.
[0220] The memory 602, being a computer-readable storage medium, can be used to store software programs and computer-executable programs, such as the program instructions / modules corresponding to the methods in the embodiments of the present disclosure. The processor 601 executes functional applications and data processing by running the program instructions / modules stored in the memory 602, that is, implements the performance evaluation method for the code large model in the above embodiments.
[0221] The memory 602 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the terminal device, etc. In addition, the memory 602 may include high-speed random access memory and may also include non-volatile memory.
[0222] The embodiments of the present disclosure provide a computing device, including: a device body; and the performance evaluation device for the code large model as described above, disposed on the device body.
[0223] The embodiments of the present disclosure provide a computer-readable storage medium storing computer-executable instructions, and the computer-executable instructions are configured to execute the above-mentioned security evaluation method for the large model.
[0224] The above description and the accompanying drawings fully illustrate the embodiments of the present disclosure, enabling those skilled in the art to practice them. Other embodiments may include structural, logical, electrical, process, and other changes. The embodiments merely represent possible variations. Unless explicitly required, individual components and functions are optional, and the order of operations may vary. Parts and features of some embodiments may be included in or replace parts and features of other embodiments. Moreover, the terms used in this application are only for describing the embodiments and do not limit the technical solutions described in this application. As used in the technical solutions described in this application, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to also include the plural forms. Similarly, as used in this application, the term "and / or" refers to any and all possible combinations including one or more of the associated listed items. Additionally, when used in this application, the term "comprise" and its variants "comprises" and / or "comprising" etc. mean the presence of the stated features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or groups thereof. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, or device comprising the element. Herein, each embodiment may focus on the differences from other embodiments, and the same or similar parts among the various embodiments may be referred to each other. For the methods, products, etc. disclosed in the embodiments, if they correspond to the method part disclosed in the embodiments, the relevant parts may refer to the description of the method part.
Claims
1. A performance evaluation method for a large code model, characterized in that: include: Acquire first test data of the general capability of the code big model, and parse the first test data to determine sub-test data of multiple sub-capabilities in the general capability; Inputting the sub-test data of each sub-capability into the code big model respectively, obtaining the first feedback data of the code big model for each sub-capability; Determine an evaluation score of each sub-capability of the code model according to the first feedback data of each sub-capability; According to the evaluation score of each sub-capability, the evaluation result of the general capability of the code model is determined.
2. The performance evaluation method according to claim 1, characterized in that: The evaluation score of each sub-capability of the code model is determined according to the first feedback data of each sub-capability, including: Determine the current scoring criteria corresponding to the current item sub-competency; The first feedback data of the current item sub-competency is scored according to the current scoring standard to obtain an evaluation score of the current item sub-competency.
3. The performance evaluation method according to claim 1, characterized in that: Multiple sub-abilities include: code comprehension ability, code generation and completion ability, code conversion ability, unit test case generation ability, code diagnosis and optimization ability, and R&D question-and-answer ability.
4. The performance evaluation method according to any one of claims 1 to 3, characterized in that: Also includes: Acquire second test data for the dedicated scenario capability of the code big model; Inputting the second test data into the code big model to obtain second feedback data of the code big model for multiple development scenarios; According to the second feedback data of the code big model for multiple development scenarios, the evaluation result of the dedicated scenario capability of the code big model is determined.
5. The performance evaluation method according to claim 4, characterized in that: According to the second feedback data of the code big model for various development scenarios, the evaluation results of the dedicated scenario capabilities of the code big model are determined, including: Determine the development capability score of the code big model for each development scenario according to the second feedback data of the code big model for each development scenario; According to the development capability score of the code big model for each development scenario, the evaluation result of the dedicated scenario capability of the code big model is determined.
6. The performance evaluation method according to claim 5, characterized in that: Various development scenarios include: website development scenarios, desktop application development scenarios, mobile application development scenarios, database development scenarios, big data development scenarios, artificial intelligence development scenarios and embedded development scenarios.
7. The performance evaluation method according to any one of claims 1 to 3, characterized in that: Also includes: Acquire third test data of application maturity of the code big model; wherein the evaluation dimensions of application maturity include data maturity, model maturity and service maturity; Inputting the third test data into the code macro model to obtain third feedback data; The maturity scores of the code big model in terms of data maturity, model maturity, and service maturity are determined based on the third feedback data, and the evaluation results of the application maturity of the code big model are determined based on the maturity scores.
8. A performance evaluation device for a large code model, characterized in that: include: An acquisition module is configured to acquire first test data of a general capability of a large code model, and parse the first test data to determine sub-test data of multiple sub-capabilities in the general capability; The testing module is configured to input the sub-test data of each sub-capability into the code big model respectively, and obtain the first feedback data of the code big model for each sub-capability; A calculation module, configured to determine an evaluation score of each sub-capability of the code model according to the first feedback data of each sub-capability; The evaluation module is configured to determine the evaluation result of the general capability of the code model according to the evaluation score of each sub-capability.
9. A performance evaluation device for a large code model, comprising a processor and a memory storing program instructions, characterized in that: The processor is configured to be able to execute the performance evaluation method for a large code model according to any one of claims 1 to 7.
10. An electronic device, characterized in that: include: Equipment body; The performance evaluation device for a large code model as described in claim 8 or 9 is arranged on the device body.