Quality evaluation method, electronic equipment and computer program product
By performing multi-dimensional performance characteristics analysis on the code information output by the target model, the comprehensive performance value and quality score are calculated, which solves the problem that traditional evaluation methods cannot comprehensively evaluate the quality of the model, and achieves a comprehensive and reliable evaluation of model performance, which is suitable for complex business scenarios.
Patent Information
- Application Number
- CN202510387480.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-29
AI Technical Summary
Traditional model quality evaluation methods mainly focus on the probability of the model generation results meeting the standards, and cannot fully consider the performance of the model in dimensions such as readability, operational efficiency, robustness and scalability, resulting in insufficient exploration of model potential and may mislead developers to ignore other aspects of code quality.
By obtaining multiple code information output from the target model, multi-dimensional performance characteristics analysis is carried out, including readability, operation efficiency, accuracy, security and maintainability, etc., the performance comprehensive value is calculated, and the average performance comprehensive value is used as the quality score of the target model, and the same or similar code information is eliminated to reduce the amount of data and calculation costs, which is suitable for the evaluation of complex business scenarios.
It realizes a comprehensive and accurate evaluation of model quality, reduces the interference of occasional situations on the evaluation results, improves the reliability and accuracy of the evaluation results, and is suitable for complex business scenarios where quality varies greatly in multiple dimensions.
Smart Images

Figure CN120386703A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of data processing, etc., and particularly relates to a quality evaluation method, an electronic device, and a computer program product. Background Art
[0002] In the current rapidly developing field of machine learning, comprehensively and accurately evaluating the quality of a model is a key link in promoting its progress. However, traditional evaluation methods (such as pass@K, etc.) have significant limitations. These methods mainly focus on whether the results generated by the model pass the test, and statistically calculate the passing probability of the results in a limited number of spot checks of the results. Obviously, this method can only reflect the performance of the model on a specific task and cannot comprehensively consider the performance of the model in other dimensions, such as code readability, running efficiency, robustness, and scalability, etc., and is not sufficient as an effective basis for optimizing and improving the model. Summary of the Invention
[0003] The present disclosure provides a quality evaluation method, an electronic device, and a computer program product.
[0004] According to one aspect of the present disclosure, there is provided a quality evaluation method, including: obtaining a plurality of code information output by a target model, where the code information is the processing result of the target model for a code task; analyzing various performance characteristics of the code information to determine a comprehensive performance value of the code information; and determining a quality score of the target model according to the comprehensive performance value of the code information.
[0005] In some embodiments, analyzing various performance characteristics of the code information to determine a comprehensive performance value of the code information includes: determining performance scores of each performance characteristic of the code information; and calculating the comprehensive performance value according to each performance score.
[0006] In some embodiments, determining the performance scores of each performance characteristic of the code information includes at least: performing readability analysis on the code information based on a code style analyzer to determine a readability score of the code information; determining a running efficiency score of the code information according to the running duration when running the code information and the resource consumption of the running system; comparing the running result of the code information with the expected result of the code task to determine a correctness score of the code information; and based on a security detection tool, identifying vulnerable parts of the code information and determining a security score of the code information according to the number and vulnerability attributes of the vulnerable parts.
[0007] In some embodiments, calculating the comprehensive performance value according to each of the performance scores includes: setting a first weight for each of the performance characteristics, where the first weight is used to measure the influence degree of the performance characteristic on the comprehensive performance of the code information; and performing a weighted sum of the performance scores of each of the performance characteristics based on the first weights of each of the performance characteristics, and using the result of the weighted sum as the comprehensive performance value.
[0008] In some embodiments, determining the quality score of the target model according to the comprehensive performance value of the code information includes: determining the top K comprehensive performance values in descending order of the comprehensive performance values of each of the code information, where K is any positive integer less than or equal to the total number of code information; and using the average value of the top K comprehensive performance values as the quality score of the target model.
[0009] In some embodiments, after obtaining a plurality of code information output by the target model, it further includes: removing the same or similar code information; and converting the unremoved code information into code information with a first format.
[0010] In some embodiments, after determining the quality score of the target model, it includes: setting a second weight for each of the code tasks processed by the target model, where the second weight is used to measure the difficulty of the code task; and performing a weighted sum of the quality scores corresponding to the target model when processing various code tasks according to the second weights of each of the code tasks to determine the comprehensive quality score of the target model.
[0011] In some embodiments, after determining the quality score of the target model, it includes: setting a second weight for each of the code tasks belonging to the same task type according to the task type of the code task, where the second weight is used to measure the difficulty of the code task; and performing a weighted sum of the quality scores corresponding to the target model when processing various code tasks of the target type according to the second weights of each of the code tasks to determine the quality score of the target type of the target model.
[0012] According to another aspect of the present disclosure, there is provided an electronic device, including: a memory that stores execution instructions; and a processor that executes the execution instructions stored in the memory, so that the processor executes the quality evaluation method of any of the above embodiments.
[0013] According to still another aspect of the present disclosure, there is provided a readable storage medium, in which execution instructions are stored, and when the execution instructions are executed by a processor, they are used to implement the quality evaluation method of any of the above embodiments.
[0014] According to another aspect of the present disclosure, there is provided a computer program product including a computer program which, when executed by a processor, implements the quality assessment method of any of the above embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The drawings illustrate exemplary embodiments of the present disclosure and, together with the description thereof, are used to explain the principles of the present disclosure. The drawings are included to provide a further understanding of the present disclosure and are incorporated in and constitute a part of this specification.
[0016] Figure 1 is an application scenario diagram of the quality assessment method according to an embodiment of the present disclosure.
[0017] Figure 2 is a flowchart of the quality assessment method according to an embodiment of the present disclosure.
[0018] Figure 3 is a schematic diagram of the process for determining the comprehensive performance value according to an embodiment of the present disclosure.
[0019] Figure 4 is a schematic diagram of the process for determining the quality score according to an embodiment of the present disclosure.
[0020] Figure 5 is a schematic diagram of the process for determining the comprehensive quality score according to an embodiment of the present disclosure.
[0021] Figure 6 is a schematic block diagram of the structure of the quality assessment device according to an embodiment of the present disclosure.
[0022] Figure 7 is a schematic block diagram of the structure of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] The present disclosure will be further described in detail below with reference to the drawings and examples. It can be understood that the specific examples described herein are only for explaining the relevant content and do not limit the present disclosure. Additionally, it should be noted that for the sake of convenience of description, only parts related to the present disclosure are shown in the drawings.
[0024] It should be noted that, without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other. The technical solutions of the present disclosure will be described in detail below with reference to the drawings and in combination with the embodiments.
[0025] In the current rapidly developing field of machine learning, comprehensively and accurately evaluating model quality is crucial for driving technological progress. However, traditional evaluation methods (such as pass@K, etc.) have significant limitations. These methods mainly judge model performance by statistically calculating the passing probability of the results generated by the model in a limited number of spot checks, that is, focusing on whether the model can pass a specific test. Although this method can to a certain extent reflect the performance of the model in some specific tasks, its consideration dimension is too single, ignoring non-functional requirements such as readability, and only focusing on the correctness of the results. This evaluation method not only limits the full exploration of the model's potential, but may also mislead developers to overly pursue the passing rate while ignoring other aspects of code quality.
[0026] For this reason, the present disclosure proposes a quality evaluation method.
[0027] Figure 1 It is a schematic diagram of the application scenario of the quality evaluation method according to an embodiment of the present disclosure. As Figure 1 shown, in this application scenario, it may include a server 100 and a terminal device 200. The server 100 and the terminal device 200 can be connected through a network or Bluetooth for data interaction. The server 100 can be a cloud server or a physical server, and the terminal device 200 can be an intelligent device such as a computer, a mobile phone, or a tablet. The server 100 is used to provide the basic data required to run the quality evaluation method, and the terminal device 200 executes the quality evaluation method of the present disclosure based on the basic data provided by the server 100.
[0028] Figure 2 It is a flowchart of the quality evaluation method according to an embodiment of the present disclosure. As Figure 2 shown, a quality evaluation method M200 is proposed. Through steps S210 to S230, multi-dimensional performance characteristics of the code information output by the target model are analyzed, and then the analysis results are processed. Finally, the obtained quality score can more comprehensively reflect the performance of the target model. And, compared with the binary evaluation method of the related technology, it is applicable to the model evaluation scenario of open tasks with non-unique answers.
[0029] The quality evaluation method M200 of the present disclosure can be run through the Figure 1 terminal device 200 therein, and the data required to run this method can be recorded on the server 100.
[0030] In step S210, multiple pieces of code information output by the target model are obtained.
[0031] The target model is a model that can generate code information according to code tasks, such as an LLM (Large Language Model). Before the target model is put into production, the quality assessment method of the present disclosure can be used to determine its quality.
[0032] Further, according to the application scenario of the target model, determine the code task, the expected code that the target model is expected to generate according to the code task, and the expected result generated by the expected code. The code task can be various types of code tasks in the application scenario of the target model. The code task can be in the form of natural language, including requirements for the code, requirements for the running result of the code, and data expected to be displayed by running the code, etc., which are not limited here. The expected code is the standard style of the code expected to be generated by the target model, including code style, etc. The expected result is the result expected to be generated by running the expected code, which can be the processing result of data, the web page expected to be displayed, etc.
[0033] The code information at least includes the code generated by the target model and is the processing result of the target model for the code task. The running result can be obtained by running the code information. For example, the code task can be to generate a library management page, the code information generated by the processing of the target model can be the code of the library management page, and the running result of the code information is the library management page. If the code task is of other types, the running result of the code can also be other, which will not be listed one by one here.
[0034] In some embodiments, after obtaining the code information, it is also necessary to preprocess multiple pieces of code information output by the target model, including: removing the same or similar code information; and converting the unremoved code information into code information with a first format.
[0035] Removing the same or similar code information can significantly reduce the amount of data to be processed, thereby reducing storage and computing costs; it can also speed up the subsequent processing steps and make the analysis more efficient. If there are a large number of identical code segments, it may cause the evaluation result to be biased towards these recurring patterns, thus affecting the accuracy and fairness of the overall evaluation. In addition, by removing these same or similar code information, it is also possible to ensure that the analysis includes as diverse code samples as possible, which helps to more comprehensively evaluate the performance of the model in different scenarios.
[0036] Convert the uneliminated code information into code information with a first format, making the code easier to read and understand. Standardizing the format helps in making direct comparisons between different codes as they follow the same structure and rules, reducing interference caused by format differences. Also, many code analysis tools require the input code to conform to specific format specifications, and the standardized code can be directly used in these tools without additional manual adjustment. Additionally, when the code follows a unified format, collaboration among team members becomes smoother, and the code is easier to maintain and update.
[0037] It should be noted that during the evaluation process, when inputting a code task to the target model, the target model should be controlled to output multiple code information for this task, such as fifty, without limitation here. Compared with using one or a few code information as the evaluation object, evaluating the performance of multiple code information can more comprehensively and realistically reflect the processing quality of the model for this task, avoiding interference from accidental situations to the evaluation results.
[0038] In step S220, analyze various performance characteristics of the code information to determine the comprehensive performance value of the code information.
[0039] Performance characteristics include: readability characteristics, running efficiency characteristics, correctness characteristics, security characteristics, maintainability characteristics, etc. In different application scenarios of the target model, different performance characteristics can be concerned according to needs, and one or more of the foregoing performance characteristics can be selected, without limitation here. Of course, the foregoing are only examples of performance characteristics, and there can be more performance characteristics, which are not listed one by one here and all fall within the protection scope of the present disclosure.
[0040] Different performance characteristics can be scored in different ways, such as using corresponding neural network models, without limitation here.
[0041] When performing multi-dimensional feature analysis on the code information, evaluate the performance score for each performance characteristic of the code information. Then, determine the comprehensive performance value of the code information according to each performance score. The comprehensive performance value can comprehensively reflect the comprehensive performance of the code information and is the result of considering the performance characteristics of each dimension. Compared with the related art where only the result after running the code is judged as correct or incorrect, through this method, a more comprehensive and accurate code performance situation can be understood.
[0042] It should be noted that both the performance score and the comprehensive performance value can be selected within a continuous numerical range, not limited to integers, and can more finely present the state of each code information, including subtle differences.
[0043] Perform readability analysis on the code information based on a code style analyzer to determine the readability score of the code information. Further, if the code information has good comments and a clear structure, a relatively high readability score can be obtained; if the code information basically conforms to the specifications, but there are some small problems and a small amount of unnecessary complex logic, a relatively medium readability score can be obtained; if the code information lacks necessary comments, has a chaotic structure, or violates the main coding specifications, a relatively low readability score should be given.
[0044] Determine the running efficiency score of the code information based on the running duration when running the code information and the resource consumption of the running system. Further, if the code information has a fast execution speed and low resource consumption, a relatively high running efficiency score can be obtained. If the performance of the code information is acceptable, but it may exhibit high resource consumption or long execution time in some scenarios, a relatively medium running efficiency score can be obtained. If the code information has obvious performance bottlenecks, long execution time, and high resource consumption, a relatively low running efficiency score can be obtained.
[0045] Compare the running result of the code information with the expected result of the code task to determine the correctness score of the code information. Further, if the code information can run correctly in all test cases and outputs a running result that is the same as or similar to the expected result, without any errors or exceptions, a relatively high correctness score can be obtained. If most of the test cases of the code information pass, but small problems may occur or further debugging is required under certain specific conditions, a relatively medium correctness score can be obtained. If the code information has obvious logical errors, cannot meet the expected functional requirements, and the running result is mostly or completely different from the expected result, a relatively low correctness score can be obtained.
[0046] Based on a security detection tool, identify the vulnerable parts of the code information, and determine the security score of the code information according to the number and attributes of the vulnerable parts. Further, if the code information undergoes strict security review, no serious security vulnerabilities are found, and it follows best security practices, a high security score can be obtained. If the code information has some minor security risks, but can be solved by simple repair measures, a medium security score can be obtained. If the code information contains serious security vulnerabilities and is vulnerable to attacks, a relatively low security score can be obtained.
[0047] Of course, the maintainability score of the code can also be determined, etc. For the code information generated by different code tasks and different application scenarios of the target model, the number and types of performance characteristics to be concerned about are not limited to the foregoing. Any performance characteristic and its performance score determination method fall within the protection scope of the present disclosure.
[0048] In step S230, the quality score of the target model is determined according to the comprehensive performance value of each code information.
[0049] The quality score can characterize the processing quality of the target model for processing code tasks. The quality score is obtained from the comprehensive performance value of the code information. Usually, it is the mean value of the comprehensive performance values of the code information, and the result calculated in this way is more representative. If only the comprehensive performance value of a single code information is used as the quality score of the target model, there may be problems of randomness and contingency, and it is not credible.
[0050] Of course, in order to reduce computing power waste, the top K comprehensive performance values of the code information can be selected in descending order of the comprehensive performance values of each code information for mean operation. Then, the mean value of the top K comprehensive performance values of the code information is used as the processing quality of the target model when processing this code task. Usually, the value of K can be selected according to the application scenario of the model and is not limited here. K should be any positive integer less than or equal to the total number of code information.
[0051] Usually, values such as 1, 5, 15, etc. can be selected for K. If K is 1, it means that the current evaluation requirement is to focus on the quality score when the target model can output the code with the best performance. There is no limit on the value of K.
[0052] In addition, the quality assessment method described in this disclosure, the quality score of the target model determined by averaging the comprehensive performance values of K code information, corresponds to paa@K in the related art, and this method can be named average@K.
[0053] By averaging the comprehensive performance values of the top K code information, the instability of the result caused by the fluctuation of the comprehensive performance value of a single code information is effectively reduced. The traditional pass@K method is prone to large variance in the evaluation result when there are slight differences in the generated code, which will affect the reliability of the model. average@K reduces this instability by aggregating the top K high-quality candidate solutions, making the evaluation result more reliable, and is particularly suitable for complex business scenarios where there are large differences in the generation quality in multiple dimensions.
[0054] In some embodiments, it further includes: setting a second weight for each code task processed by the target model, where the second weight is used to measure the difficulty of the code task; and weighted summing the quality scores corresponding to the target model when processing various code tasks according to the second weights of each code task to determine the comprehensive quality score of the target model.
[0055] In some embodiments, it further includes: after determining the quality score of the target model, including: setting a second weight for each code task belonging to the same task type according to the task type of the code task, where the second weight is used to measure the difficulty of the code task; and performing a weighted sum of the quality scores corresponding to the target model when processing various code tasks of the target type according to the second weights of the code tasks to determine the quality score of the target type of the target model.
[0056] The quality assessment method of the present disclosure analyzes the performance characteristics of each code information in multiple dimensions, so that the comprehensive performance value of the code information can comprehensively and truly reflect the quality of the code information. According to the descending order, the average value of the first K comprehensive performance values is used as the quality score of the target model, avoiding the interference of accidental situations on the quality assessment and making the result more credible.
[0057] The following provides a more detailed and comprehensive introduction to the specific implementation process of the quality assessment method of the present disclosure.
[0058] First, a code task is issued to the target model, and the target model can generate multiple code information according to the code task, and each code information can meet the requirements of the code task. The type of the code task should be the type that can be used in the application scenario of the target model, such as algorithm type, data processing type, interface type, etc., which are not limited here.
[0059] Among them, code tasks of the algorithm type usually focus on the specific logic or steps to solve problems, such as sorting algorithms, search algorithms, etc. The focus of such tasks is on the design and optimization of algorithms, rather than directly interacting with the user interface or external systems.
[0060] Code tasks of the data processing type mainly involve operations such as cleaning, transforming, and analyzing data. Although such tasks may generate some visualization results or reports, their main goal is still to process and analyze the data itself.
[0061] Code tasks of the interface type involve interactions between systems, such as calling APIs to obtain data, displaying information on the front-end page, etc. If it is desired that the output of the target model is a code page, this usually means that the task involves front-end display or interacting with certain back-end services to generate dynamic content, so it more conforms to the characteristics of the interface type.
[0062] Furthermore, preprocess the multiple code information output by the target model, including: removing the same or similar code information; and converting the unremoved code information into code information with a first format.
[0063] Removing identical or similar code information can significantly reduce the amount of data to be processed, thereby reducing storage and computing costs; it can also speed up subsequent processing steps and make the analysis more efficient. If there are a large number of identical code segments, it may cause the evaluation results to be biased towards these recurring patterns, thus affecting the accuracy and fairness of the overall evaluation. Additionally, by removing these identical or similar code information, it is possible to ensure that the analysis includes as diverse a set of code samples as possible, which helps to more comprehensively evaluate the performance of the model in different scenarios.
[0064] Convert the unremoved code information into code information with a first format, making the code easier to read and understand. Standardized formats facilitate direct comparison between different codes because they follow the same structure and rules, reducing interference caused by format differences. Moreover, many code analysis tools require the input code to conform to specific format specifications, and the standardized code can be directly used in these tools without additional manual adjustment. Additionally, when the code follows a unified format, collaboration among team members becomes smoother, and the code is easier to maintain and update.
[0065] Furthermore, after preprocessing the code information, the remaining code information needs to be analyzed one by one for multi-dimensional performance characteristics, and then the analysis results of each performance characteristic are aggregated to obtain the comprehensive performance value of the code information.
[0066] Figure 3 It is a schematic diagram of the process for determining the comprehensive performance value according to the embodiments of the present disclosure. As Figure 3 shown, it shows the process of determining the comprehensive performance value of the code information.
[0067] Suppose a certain code information has performance characteristics 1 to performance characteristic i, where i is any positive integer greater than or equal to 1. Among them, performance characteristics 1 to performance characteristic i can be readability characteristics, running efficiency characteristics, correctness characteristics, security characteristics, maintainability characteristics, etc., which are not limited here. After analyzing each performance characteristic, the corresponding performance scores of each performance characteristic can be obtained. For example, the performance score of performance characteristic 1 is f1, and the performance score of performance characteristic i is f1, and the performance scores are independent of each other.
[0068] Furthermore, according to the importance of each performance characteristic to the code information or the quality of the target model, a first weight is assigned to each performance characteristic. For example, the first weight of performance characteristic 1 is W1, and the first weight of performance characteristic i is Wi. The sum of the first weights of all performance characteristics should be 1, but the numerical values of the first weights of each performance characteristic are not limited and can be the same or different.
[0069] Further, multiply each first weight by the performance score of its corresponding performance characteristic, and sum up the results of multiplying all the first weights of the performance characteristics by their performance scores. Through the above-mentioned weighted summation calculation, the comprehensive performance value of the code information can be obtained.
[0070] That is, the comprehensive performance value of the code information can be: , where n is the serial number of the performance characteristic, and n is any value from 1 to i; Wn is the first weight of the nth performance characteristic; fn is the performance score of the nth performance characteristic.
[0071] The method for determining the performance score of each performance characteristic will be described below.
[0072] Based on the readability analysis of the code information by the code style analyzer, determine the readability score of the code information. Specifically, the code style analyzer can be a static analysis tool such as Pylint, which checks whether the code information follows predefined coding specifications, including indentation, naming rules, comment quality, etc. These tools can automatically detect non-compliant parts and give improvement suggestions. Calculate metrics such as the cyclomatic complexity and the number of lines of the code to evaluate the complexity of the code structure. Complex code information is usually more difficult to understand and maintain. Evaluate the quality and quantity of comments in the code to ensure that each function, class, and module has a clear explanation. Good documentation helps other developers quickly understand the code logic. Check the consistency between different parts of the code, such as variable naming rules, function signatures, etc. Code with high consistency is easier to understand and maintain.
[0073] Further, if the code information has good comments and a clear structure, a relatively high readability score can be obtained; if the code information basically conforms to the specifications but there are some small problems and a small amount of unnecessary complex logic, a relatively medium readability score can be obtained; if the code information lacks necessary comments, has a chaotic structure, or violates the main coding specifications, a relatively low readability score should be given.
[0074] Determine the running efficiency score of the code information according to the running duration and resource consumption of the running system when running the code information. Through benchmark testing tools such as Benchmark.js, measure the running duration and resource consumption of the code execution. The resource consumption includes CPU (Central Processing Unit), memory, I / O (input / output interface), etc. Different input scales can be set to simulate actual application scenarios. Evaluate whether the code effectively utilizes concurrent and asynchronous processing mechanisms to improve execution efficiency. For I / O-intensive tasks, asynchronous operations can significantly reduce waiting time. Check whether the code has memory leaks or excessive memory occupation, especially in long-running services.
[0075] Furthermore, if the code information has a fast execution speed and low resource consumption, a relatively high running efficiency score can be obtained. If the performance of the code information is acceptable, but it may exhibit high resource consumption or long execution time in certain scenarios, a relatively medium running efficiency score can be obtained. If the code information has obvious performance bottlenecks, long execution time, and high resource consumption, a relatively low running efficiency score can be obtained.
[0076] Compare the running result of the code information with the expected result of the code task to determine the correctness score of the code information. Write detailed unit test cases that cover all functional points and boundary conditions to ensure that the code can produce correct output in all situations. Pay special attention to the performance of the code under extreme inputs to ensure that it can gracefully handle exception situations and avoid crashes or unreasonable outputs.
[0077] Furthermore, if the code information can run correctly in all test cases and outputs a running result that is the same as or similar to the expected result without any errors or exceptions, a relatively high correctness score can be obtained. If most of the test cases of the code information pass, but small problems may occur or further debugging is required under certain specific conditions, a relatively medium correctness score can be obtained. If the code information has obvious logical errors, cannot meet the expected functional requirements, and the running result is mostly or completely different from the expected result, a relatively low correctness score can be obtained.
[0078] Based on security detection tools, identify the vulnerable parts of the code information and determine the security score of the code information according to the number and attributes of the vulnerable parts. The security detection tools can be static analysis tools such as SonarQube. Use the static analysis tool to scan the code information to find potential security vulnerabilities such as SQL injection, cross-site scripting attacks, buffer overflows, etc. The security detection tools can also be dynamic analysis tools such as OWASP ZAP. Use the dynamic analysis tool to simulate attack scenarios to detect the security of the code information during runtime, which can help discover problems that are only exposed during runtime. Check the third-party libraries and frameworks used in the project to ensure that they are up-to-date and have no known security vulnerabilities. It is also possible to use security detection tools to automatically monitor the security of dependencies. Evaluate the implementation of the code in terms of user privilege management and sensitive data protection to ensure compliance with the principle of least privilege and encrypted storage and transmission of sensitive data.
[0079] Further, if the code information passes a strict security review and no serious security vulnerabilities are found, and it follows best security practices, a high security score can be obtained. If there are some minor security risks in the code information but they can be resolved through simple repair measures, a medium security score can be obtained. If the code information contains serious security vulnerabilities and is vulnerable to attacks, a relatively low security score can be obtained.
[0080] Of course, the maintainability score of the code can also be determined, etc. For the code information generated by different code tasks and different application scenarios of the target model, the number and types of performance characteristics to be concerned about are not limited to the foregoing, and any performance characteristic and its performance score determination method fall within the protection scope of the present disclosure.
[0081] In practical applications, different types of code tasks have different focuses on the consideration of code quality. The present disclosure can flexibly adjust the evaluation criteria to adapt to different business needs. For example, in the evaluation of high-security code, low-latency code, or highly readable code, this method can be adjusted according to actual needs. This flexibility significantly improves the applicability of this method in complex business scenarios and helps the development team to conduct effective model evaluation and improvement according to actual needs.
[0082] Compared with traditional code evaluation methods, this method can not only focus on the syntactic correctness of the code, but also ensure the semantic consistency and rationality of the generated code by evaluating the logic and structure of the candidate code. This way makes the evaluation results more in line with the actual application requirements and effectively fills the gap that traditional evaluation tools cannot comprehensively evaluate the quality of code generation.
[0083] After obtaining the comprehensive performance value of each code information, the quality score of the target model for processing this code task should be determined based on these comprehensive performance values to measure the processing quality of the target model for this code task.
[0084] Figure 4 It is a schematic diagram of the quality score determination process according to an embodiment of the present disclosure. As Figure 4 shown, code information A to code information X are all the results obtained by the target model for processing a certain code task. Moreover, each code information has obtained the corresponding comprehensive performance value through the foregoing method. For example, the comprehensive performance value of code information A is FA, the comprehensive performance value of code information K is FK, and the comprehensive performance value of code information X is FX.
[0085] Further, arrange these comprehensive performance values in descending order. The obtained sequence is the code information A to code information X. In this sequence, the comprehensive performance value FA of code information A is the largest, and the comprehensive performance value FX of code information X is the smallest. It should be noted that for each code task, the number of code information that can be generated is not fixed, and after preprocessing the code information, the total amount of retained code information is usually less than the number of generated code information. Here, X is the total amount of retained code information after preprocessing, for example, 50.
[0086] Further, in the sequence of code information, extract the comprehensive performance values of the first K code information. The comprehensive performance values of the first K code information should be the top K highest values in this sequence. K should be a positive integer less than or equal to the total amount of code information. Using the comprehensive performance values of all code information as the basis for the quality score can reduce computing power waste and avoid contingency. K can usually be selected as values such as 1, 5, 15, etc. If K is 1, it means that the current evaluation requirement is to focus on the quality score when the target model can output the best performance code. There is no limit on the value of K.
[0087] Further, calculate the mean value of the comprehensive performance values of the first K code information. The obtained mean value is the quality score of the target model. That is, the quality score of the target model is: , where J is the serial number of the code information, and K is the number of code information that hits and is used to calculate the comprehensive performance value.
[0088] Further, after obtaining the quality score of the target model in processing the code task, the overall comprehensive quality score of the target model can also be determined by aggregating the quality scores of the target model for various code tasks.
[0089] By averaging the comprehensive performance values of the first K code information, the instability of the results caused by the fluctuations of the comprehensive performance values of individual code information is effectively reduced. The traditional pass@K method is prone to large variances in the evaluation results when there are slight differences in the generated codes, which will affect the reliability of the model. average@K reduces this instability by aggregating the first K high-quality candidate solutions, making the evaluation results more reliable, and is particularly suitable for complex business scenarios where there are large differences in the generation quality in multiple dimensions.
[0090] Figure 5 is a schematic diagram of the process for determining the comprehensive quality score according to the embodiments of the present disclosure.
[0091] As Figure 5As shown, assume that the target model has executed M code tasks such as code task AA to code task MM, and has obtained the quality scores for processing each code task. For example, the quality score for code task AA is FAA, and the quality score for code task MM is FMM. Further, a second weight is configured for each code task, and the second weight is associated with the difficulty of the code task. Further, multiply the quality score of each code task by its weight, and add the results of multiplying the quality scores of each code task by their weights. The result obtained through the weighted operation can be used as the comprehensive quality score of the target model. The number of M is not limited.
[0092] The comprehensive quality score of the target model can be: , where Q is the serial number of the code task.
[0093] The second weight is used to measure the difficulty of the code task. For example, the difficulty of the code task "generate a library management page" is lower than that of "generate a library management page including name, book code, and having a book screening function", so the second weight of the former is lower than that of the latter. However, the sum of the second weights of all code tasks should be 1.
[0094] In some implementation sets, the processing quality of the target model for each type of code task can also be determined according to the task type of each code task. The task type can be an interface type, a data processing type, an algorithm type, etc., which is not limited here.
[0095] Specifically, according to the task type of the code task, a second weight is set for each code task belonging to the same task type, and the second weight is used to measure the difficulty of the code task; and according to the second weight of each code task, the quality scores corresponding to the target model when processing various code tasks of the target type are weighted and summed to determine the quality score of the target type of the target model.
[0096] In practical applications, different types of code tasks have different emphases on the consideration of code quality. The present disclosure can flexibly adjust the evaluation criteria to adapt to different business requirements. For example, in the evaluation of high-security code, low-latency code, or highly readable code, this method can be adjusted according to actual needs. This flexibility significantly improves the applicability of this method in complex business scenarios, helping the development team to conduct effective model evaluation and improvement according to actual needs.
[0097] The quality evaluation method of the present disclosure analyzes the performance characteristics of each code information in multiple dimensions, so that the performance comprehensive value of the code information can comprehensively and truly reflect the quality of the code information. In descending order, the mean of the first K performance comprehensive values is used as the quality score of the target model, avoiding the interference of accidental situations on the quality evaluation and making the result more credible.
[0098] Figure 6 It is a schematic block diagram of the structure of a quality evaluation device according to an embodiment of the present disclosure. Figure 6 It shows each module of the quality evaluation device 600, including: An acquisition module 610, configured to acquire a plurality of code information output by a target model, where the code information is a processing result of the target model for a code task; an analysis module 620, configured to analyze various performance characteristics of the code information to determine a comprehensive performance value of the code information; a quality score determination module 630, configured to according to the comprehensive performance value of the code information.
[0099] The quality evaluation device 600 of the present disclosure may be in the form of computer software, and each module of the quality evaluation device 600 may be in the form of computer software modules.
[0100] Each module of the quality evaluation device 600 of the present disclosure is set to implement each step of the quality evaluation method, and its execution principle and steps can be referred to the foregoing, and will not be elaborated herein.
[0101] Figure 7 It is a schematic block diagram of the structure of an electronic device according to an embodiment of the present disclosure. As Figure 7 shown, the present disclosure also provides an electronic device 1000, including: a processor 1200 and a memory 1300, where the memory 1300 stores execution instructions; the processor 1200 executes the execution instructions stored in the memory 1300, so that the processor 1200 executes the quality evaluation method.
[0102] The hardware structure of the electronic device 1000 can be implemented by using a bus architecture. The bus architecture may include any number of interconnected buses and bridges, depending on the specific application of the hardware and the overall design constraints. The bus 1100 connects various circuits including one or more processors 1200, a memory 1300, and / or hardware modules together. The bus 1100 can also connect various other circuits 1400 such as peripheral devices, voltage regulators, power management circuits, external antennas, etc.
[0103] The bus 1100 may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Component (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, only one connection line is used in this figure, but it does not mean that there is only one bus or one type of bus.
[0104] The present disclosure also provides a readable storage medium storing a computer program, which is used to implement the above method when executed by a processor. A "readable storage medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. More specific examples of the readable storage medium include the following: an electrical connection part (electronic device) having one or more wirings, a portable computer diskette case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable read-only memory (CDROM), etc.
[0105] The present disclosure also provides a computer program product. The method of the present disclosure can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed, the processes or functions of the present disclosure are executed in whole or in part.
[0106] The computer program or instructions can be stored in a readable storage medium, or transmitted from one readable storage medium to another. For example, the computer program or instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless manner. The readable storage medium can be any available medium that can be accessed, or a data storage device such as a server or data center integrating one or more available media. The available medium can be a magnetic medium, such as a floppy disk, a hard disk, or a magnetic tape; it can also be an optical medium, such as a digital video disc; or it can be a semiconductor medium, such as a solid state drive. The computer-readable storage medium can be a volatile or non-volatile storage medium, or can include both volatile and non-volatile types of storage media.
[0107] Those skilled in the art should understand that the embodiments of the present disclosure can be provided as a method, an electronic device, a readable storage medium, or a computer program product. Therefore, the present disclosure can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present disclosure can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0108] The present disclosure is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the present disclosure. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce a means for implementing the functions specified in one or more of the flows Figure 1 one or more flows and / or blocks Figure 1 or a means for implementing the functions specified in one or more of the blocks.
[0109] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction means that implements the functions specified in one or more of the flows Figure 1 one or more flows and / or blocks Figure 1 or a means for implementing the functions specified in one or more of the blocks.
[0110] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the flows Figure 1 one or more flows and / or blocks Figure 1 or a means for implementing the functions specified in one or more of the blocks.
[0111] In the description of this specification, the description with reference to terms such as "one embodiment / way", "some embodiments / ways", "example", "specific example", or "some examples", etc. means that the specific features, structures, or characteristics described in connection with the embodiment / way or example are included in at least one embodiment / way or example of the present disclosure. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment / way or example. Moreover, the specific features, structures, or characteristics described can be combined in a suitable manner in any one or more embodiments / ways or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments / ways or examples described in this specification and the features of different embodiments / ways or examples.
[0112] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In the description of the present disclosure, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0113] Those skilled in the art should understand that the above embodiments are merely for clearly illustrating the present disclosure and are not intended to limit the scope of the present disclosure. For those skilled in the art, other changes or modifications can be made based on the above disclosure, and these changes or modifications are still within the scope of the present disclosure.
Claims
1. A quality assessment method, characterized in that, Including: Obtain multiple pieces of code information output by the target model, where the code information is the processing result of the target model for a code task; Analyze various performance characteristics of the code information to determine the comprehensive performance value of the code information; Determine the quality score of the target model according to the comprehensive performance value of the code information.
2. The quality assessment method according to claim 1, wherein Analyze various performance characteristics of the code information to determine the comprehensive performance value of the code information, including: Determine the performance scores of each performance characteristic of the code information; and Calculate the comprehensive performance value according to each of the performance scores.
3. The quality assessment method according to claim 2, wherein Determine the performance scores of each performance characteristic of the code information, including at least: Based on a code style analyzer, perform readability analysis on the code information to determine the readability score of the code information; Determine the running efficiency score of the code information according to the running duration when running the code information and the resource consumption of the running system; Compare the running result of the code information with the expected result of the code task to determine the correctness score of the code information; and Based on a security detection tool, identify the vulnerable parts of the code information, and determine the security score of the code information according to the number and vulnerability attributes of the vulnerable parts.
4. The quality assessment method according to claim 2, characterized in that Calculate the comprehensive performance value according to each of the performance scores, including: Set a first weight for each of the performance characteristics, where the first weight is used to measure the influence degree of the performance characteristic on the comprehensive performance of the code information; and Based on the first weights of each of the performance characteristics, perform weighted summation on the performance scores of each of the performance characteristics, and use the result of the weighted summation as the comprehensive performance value.
5. The quality assessment method according to claim 1, wherein Determine the quality score of the target model according to the comprehensive performance value of the code information, including: Determine the top K comprehensive performance values in descending order according to the comprehensive performance values of each of the code information, where K is any positive integer less than or equal to the total number of code information; and Use the average value of the top K comprehensive performance values as the quality score of the target model.
6. The quality assessment method according to claim 1, wherein After obtaining multiple pieces of code information output by the target model, it further includes: Eliminate the same or similar code information; and Convert the uneliminated code information into code information with a first format.
7. The quality assessment method according to claim 1, characterized in that After determining the quality score of the target model, it includes: Set a second weight for each of the code tasks processed by the target model, where the second weight is used to measure the difficulty of the code task; and According to the second weights of each of the code tasks, perform weighted summation on the quality scores corresponding to the target model when processing various code tasks to determine the comprehensive quality score of the target model.
8. The quality assessment method according to claim 1 or 7, characterized in that, After determining the quality score of the target model, it includes: According to the task type of the code task, set a second weight for each of the code tasks belonging to the same task type, where the second weight is used to measure the difficulty of the code task; and According to the second weights of each of the code tasks, perform weighted summation on the quality scores corresponding to the target model when processing various code tasks of the target type to determine the quality score of the target type of the target model.
9. An electronic device, characterized in that, Including: A memory that stores execution instructions; and a processor that executes the execution instructions stored in the memory, so that the processor executes the quality assessment method according to any one of claims 1 to 8.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the quality assessment method according to any one of claims 1 to 8.