Quality evaluation method, electronic equipment and computer program product

By conducting content analysis and running interface analysis of the automatic generated code, combining grammatical structure and visual effects, and calculating comprehensive quality scores, the problem of difficulty in comprehensively evaluating code quality in the existing technology is solved, and efficient and accurate code quality evaluation is achieved.

CN120197614APending Publication Date: 2025-06-24KE COM (BEIJING) TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510259122.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The prior art is difficult to fully evaluate the quality of automatic generation code, especially when dealing with complex logic and multimodal outputs, the evaluation is inefficient and the results are not accurate and reliable.

Method used

By conducting content analysis and running interface analysis on the code, the first and second mass scores of the code are determined respectively, and the comprehensive mass score is calculated based on the grammatical structure and visual effects.

Benefits of technology

A comprehensive evaluation of the functionality, structural similarity and visual effects of the code is achieved, improving evaluation efficiency and accuracy, and ensuring that the generated code meets high standards in terms of functionality and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197614A_ABST
    Figure CN120197614A_ABST
Patent Text Reader

Abstract

The invention provides a quality evaluation method, electronic equipment and a computer program product. The quality evaluation method comprises the steps of performing content analysis on a code, and determining a first quality score of the code; comparing the similarity between the operation interface of the code and the reference interface, and determining a second quality score of the code; and determining a comprehensive mass score of the code according to the first mass score and the second mass score.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of data processing and the like, and particularly relates to a quality assessment method, an electronic device, and a computer program product. Background Art

[0002] With the wide application of code generation models in software development, ensuring the quality of automatically generated code has become crucial. High-quality code not only needs to be functionally correct, but also should have good readability, reasonable structure, and consistency to meet the actual business needs and reduce subsequent maintenance costs. However, related technologies mainly rely on simple compilation checks and manual reviews, lacking automated and multi-dimensional evaluation methods, and it is difficult to comprehensively evaluate the functionality, structural similarity, and visual effects of code, especially when dealing with complex logic and multi-modal outputs, resulting in low evaluation efficiency and inaccurate and unreliable results. Summary of the Invention

[0003] The present disclosure provides a quality assessment method, an electronic device, and a computer program product.

[0004] According to one aspect of the present disclosure, a quality assessment method is provided, including: performing content analysis on the code to determine a first quality score of the code; comparing the similarity between the running interface of the code and a reference interface to determine a second quality score of the code; and determining a comprehensive quality score of the code based on the first quality score and the second quality score.

[0005] In some embodiments, performing content analysis on the code to determine the first quality score of the code includes: determining the text content similarity between the code and a reference code, where the reference code is the code of the reference interface; determining the syntactic content similarity between the syntactic structure of the code and the syntactic structure of the reference code; and using the mean between the text content similarity and the syntactic content similarity as the first quality score.

[0006] In some embodiments, determining the text content similarity between the code and a reference code includes: determining the word frequency of each word segment in the code; determining the inverse document frequency of the word segment according to the total number of code samples in a first code set that contain the word segment, where the inverse document frequency is negatively correlated with the total number of code samples containing the word segment, and the first code set is a code sample set related to the application scenario of the code; multiplying the word frequency by the inverse document frequency to determine the importance score of the word segment; arranging the importance scores of the word segments in the order of their appearance in the code to form a first vector; and obtaining the text content similarity according to the first cosine value between the first vector and a second vector of the reference code, where the first cosine value is negatively correlated with the first quality score.

[0007] In some embodiments, determining the syntactic content similarity between the syntactic structure of the code and that of the reference code includes: forming a first syntax tree according to the syntactic structure of the code, where the first syntax tree includes multiple nodes representing the syntactic units of the code, and according to the attributes of the corresponding syntactic units, the nodes have levels; determining the number of identical nodes with the same levels and representing the same syntactic units between the first syntax tree and a second syntax tree of the reference code; and determining the syntactic content similarity according to the number of identical nodes.

[0008] In some embodiments, comparing the similarity between the running interface of the code and a reference interface to determine the second quality score of the code includes: using a vectorization model to convert the running interface of the code into a third vector, where the third vector can represent the interface content and interface features of the running interface; and calculating the second cosine value between the third vector and a fourth vector of the reference interface to obtain the second quality score, where the second cosine value is negatively correlated with the second quality score.

[0009] In some embodiments, determining the comprehensive quality score of the code according to the first quality score and the second quality score includes: normalizing both the first quality score and the second quality score into a target score range to obtain a first normalized quality score and a second normalized quality score; and calculating the mean between the first normalized quality score and the second normalized quality score and using the mean as the comprehensive quality score.

[0010] In some embodiments, standardizing both the first mass fraction and the second mass fraction into a target score range to obtain a first standardized mass fraction and a second standardized mass fraction includes: determining a first maximum value and a first minimum value of a scoring range corresponding to the first mass fraction, and determining a second maximum value and a second minimum value of a scoring range corresponding to the second mass fraction; dividing the difference between the first mass fraction and the first minimum value by the difference between the first maximum value and the first minimum value to obtain the first standardized mass fraction, where the first standardized mass fraction is within the target score range, and the target score range is from 0 to 1; and dividing the difference between the second mass fraction and the second minimum value by the difference between the second maximum value and the second minimum value to obtain the second standardized mass fraction, where the second standardized mass fraction is within the target score range.

[0011] In some embodiments, before determining the first mass fraction of the code, it includes: installing dependencies of the code in a blank project to create a running environment suitable for the code.

[0012] In some embodiments, after creating a running environment suitable for the code, it includes: running the code in the running environment to determine the running state of the code; and when the running state is abnormal, removing the code.

[0013] According to another aspect of the present disclosure, there is provided an electronic device, including: a memory that stores execution instructions; and a processor that executes the execution instructions stored in the memory, such that the processor executes the quality assessment method of any embodiment of the present disclosure.

[0014] According to still another aspect of the present disclosure, there is provided a readable storage medium in which execution instructions are stored, and when the execution instructions are executed by a processor, they are used to implement the quality assessment method of any embodiment of the present disclosure.

[0015] According to yet another aspect of the present disclosure, there is provided a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the quality assessment method of any embodiment of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The drawings illustrate exemplary embodiments of the present disclosure and, together with the description thereof, are used to explain the principles of the present disclosure. These drawings are included to provide a further understanding of the present disclosure and are included in this specification and form a part of this specification.

[0017] Figure 1 is an application scenario diagram of the quality assessment method according to an embodiment of the present disclosure.

[0018] Figure 2 It is a flowchart of a quality assessment method according to an embodiment of the present disclosure.

[0019] Figure 3 It is a schematic diagram of the process for determining the comprehensive quality score according to an embodiment of the present disclosure.

[0020] Figure 4 It is a schematic block diagram of the structure of a quality assessment device according to an embodiment of the present disclosure.

[0021] Figure 5 It is a schematic block diagram of the structure of an electronic device according to an embodiment of the present disclosure. Specific embodiments

[0022] The following further elaborates on the present disclosure in conjunction with the accompanying drawings and examples. It can be understood that the specific examples described herein are only used to explain the relevant content and do not limit the present disclosure. Additionally, it should be noted that for ease of description, only parts related to the present disclosure are shown in the drawings.

[0023] It should be noted that, without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other. The following will detail the technical solutions of the present disclosure with reference to the accompanying drawings and embodiments.

[0024] With the wide application of code generation models in software development, ensuring the quality of automatically generated code has become crucial. High-quality code not only needs to be functionally correct but also should have good readability, reasonable structure, and consistency to meet actual business requirements and reduce subsequent maintenance costs. However, current technologies mainly rely on simple compilation checks and manual reviews. These methods can only verify whether the code can be successfully compiled and run, but cannot evaluate the internal quality of the code, such as logical structure and style consistency. Although manual review can complement the deficiencies of compilation checks, due to its reliance on individual experience and subjective judgment, it is difficult to ensure consistency and objectivity, and is time-consuming and laborious. Especially when dealing with complex logic and multi-modal outputs (such as user interface design), existing methods are inadequate, resulting in low evaluation efficiency and inaccurate and unreliable results.

[0025] Therefore, the present disclosure proposes a quality assessment method.

[0026] Figure 1 It is a schematic diagram of the application scenario of the quality assessment method according to an embodiment of the present disclosure. As Figure 1As shown in the figure, in this application scenario, it may include a server 100 and a terminal device 200. The server 100 and the terminal device 200 can be connected through a network or Bluetooth to perform data interaction. The server 100 can be a cloud server or a physical server, and the terminal device 200 can be an intelligent device such as a computer, a mobile phone, or a tablet. The server 100 is used to provide the basic data required for the operation quality assessment method, and the terminal device 200 executes the quality assessment method of the present disclosure based on the basic data provided by the server 100.

[0027] Figure 2 is a flowchart of the quality assessment method according to an embodiment of the present disclosure. As Figure 2 shown, a quality assessment method M200 is proposed. Through steps S210 to S230, content analysis and running interface analysis are performed on the code, and based on the analysis results of the two dimensions, the comprehensive quality score of the code is determined. Obviously, the solution of the present disclosure comprehensively considers the functions, structures, and performances of the code, filling the gap in the field of code generation quality assessment.

[0028] The quality assessment method M200 of the present disclosure can be run through the Figure 1 terminal device 200 in the figure, and the data required to run this method can be recorded on the server 100.

[0029] In step S210, content analysis is performed on the code to determine the first quality score of the code.

[0030] The code is the object to be quality-assessed, and a running interface can be obtained through running. If the code depends on an external database or the like during running, the code should contain information about the dependencies, such as the database address, API key, etc.

[0031] Content analysis essentially includes text content analysis and syntax content analysis of the code. It can be understood that when describing a page, the code generates semantics and logic through the stacking of multiple word segments, and the word segments can be keywords, variable names, operators, etc. We use one or more word segments that can express syntactic meanings as a syntactic unit, and different syntactic units have corresponding attributes according to the syntactic meanings they express. These syntactic units with different attributes constitute the syntactic structure of the code, and different codes have different syntactic habits, so they usually correspond to different syntactic structures.

[0032] When performing text content analysis on the code, the text content similarity between the code and the reference code can be determined by comparing them. In this way, the accuracy of the code in terms of expression form can be evaluated to ensure that it not only has correct functions but also good readability and consistency.

[0033] When performing syntactic content analysis on the code, more attention is paid to the logical structure and language specifications of the code. Compare the number of codes at the same level between the code and the reference code, and determine the operation steps from the code to the reference code, such as deletion, replacement, etc. Through syntactic content analysis, the logical clarity and structural rationality of the code can be evaluated.

[0034] It should be noted that the code is the result generated by the code generation model according to the code task. Before the code generation model is put into production, the quality of the code it generates should be inspected, and then the quality of the code generation model can be determined. The reference code is the expected result that the code generation model is expected to output, and it is the standard answer that meets the application requirements of the application scenario of the code generation model.

[0035] The present disclosure evaluates the quality of the code from two dimensions of text content and syntactic content by comparing the code with the reference code. Compared with the related art that only evaluates whether the code can be compiled and run, it provides data support for the optimization and upgrade of the code generation model, and further enables the code generation model to generate code that is easy to maintain and has clear logic.

[0036] The first quality score is the result of content analysis of the code and is a quantification of the content dimension of the code. The first quality score should be within the corresponding scoring range, such as 0 to 100, and the first quality score can be any value within this scoring range. The specific values of the scoring range and the first quality score are not limited here. It should be noted that the first quality score can be an integer or a decimal, and it should be able to fully reflect the quality status of the code at the content level. Compared with the binary evaluation criteria (such as "pass" or "fail"), the way of expressing the code quality numerically can fully reflect the true quality of the code.

[0037] By evaluating the quality of the code from two dimensions of text and syntactic logic, compared with the traditional technical means that only evaluate whether the code can be compiled and run, a more comprehensive and detailed evaluation system is provided. This method can not only provide data support for the optimization and upgrade of the code generation model, but also help the development team generate code that is more convenient to maintain and has clear logic.

[0038] In step S220, compare the similarity between the running interface of the code and the reference interface, and determine the second quality score of the code.

[0039] The running interface is the result of the code running. In the present disclosure, the code describes the layout of the display interface, the logical functions of each component, etc. By running the code in the target running environment, the running interface of the code can be obtained.

[0040] The reference interface is the running result of the reference code, and the reference code is the code generated by expecting the code generation model to execute the code task, which is the standard code answer that conforms to the code task. During the acceptance process of the code generation model, by comparing the similarity between the running interface and the reference interface, the quality of the code output by the code generation model can be inspected, and then the quality of the code generation model can be evaluated. Only when the code generation model can output an interface that is the same as or similar to the reference interface can it be characterized that the execution result of the model for the code task meets the requirements. Otherwise, even if the syntax logic and text content of the code match the reference code, it cannot be put into production.

[0041] The second quality score is a score for the similarity between the running interface and the reference interface. The greater the similarity between the two, the higher the score; conversely, the lower the score. The score for the similarity between the running interface and the reference interface should fall within the scoring range of the second quality score, and the scoring range of the second quality score can be different from that of the first quality score. For example, the scoring range of the first quality score can be from 0 to 100, and the scoring range of the second quality score is from -1 to 1.

[0042] It should be noted that the second quality score can be any value within its scoring range, which can be an integer or a decimal, so as to accurately evaluate the difference degree of the second quality scores of each running interface. Compared with the binary evaluation mechanism, it is more accurate and effective.

[0043] In step S230, according to the first quality score and the second quality score, the comprehensive quality score of the code is determined.

[0044] Since the evaluation criteria of the first quality score and the second quality score are different, their scoring ranges are different, and it is difficult to obtain the comprehensive quality score through simple mean calculation. Therefore, before calculating the mean of the two, the first quality score and the second quality score should be standardized so that they are respectively standardized to the same target score range. This provides support for the evaluation of multi-modal (i.e., text modality and image modality).

[0045] After the first quality score and the second quality score are respectively standardized, the first standardized quality score and the second standardized quality score are obtained. Furthermore, calculate the mean of the first standardized quality score and the second standardized quality score to obtain the comprehensive quality score that can represent the comprehensive quality of the code.

[0046] The quality evaluation method of the present disclosure analyzes the code content and the code running interface in a multi-modal and multi-dimensional manner, and uses "score" as the result of the quality evaluation. Compared with the binary evaluation method of the related technology, it can obtain more real, reliable and accurate quality evaluation results.

[0047] Figure 3 It is a schematic diagram of the comprehensive quality score determination process according to an embodiment of the present disclosure. The following will be combined with Figure 3 to provide a more comprehensive and complete description of the determination process of the comprehensive quality score of the present disclosure.

[0048] In step 301, install the dependencies of the code in a blank project to create a running environment for the code.

[0049] Specifically, before evaluating the quality of the code, a brand-new and clean blank project will be created for each code to be evaluated. This blank project is specifically prepared for the code to ensure its independence and isolation. In this way, it can be ensured that each code runs in its own environment, avoiding interference between different codes. By creating an isolated running environment, the accuracy and reliability of the code running results can be ensured, preventing misjudgment caused by external factors (such as other codes or configurations).

[0050] Furthermore, there are usually some external dependencies when the code runs, such as an external database, etc. The dependency files containing these dependencies will be recorded in the code. By analyzing the dependency files in the code files, the third-party libraries and tools required for the code to run can be identified. Common dependency declaration files include requirements.txt in Python, package.json in Node.js, etc. These files list all the dependencies required for the code to run. By parsing these files, the system can automatically identify all necessary dependencies without manual intervention, improving efficiency and accuracy.

[0051] After installing the dependencies, the system will further configure the running environment of the project to ensure that the code can run in the correct environment. To further improve isolation, the system may create and activate a virtual environment. The virtual environment can isolate the dependencies of the project and avoid conflicts with the global environment or other projects. If the code requires specific environment variables (such as API keys, database connection strings, etc.), these variables can also be automatically configured to ensure that the code can run normally.

[0052] In step 302, run the code in the running environment to determine the running state of the code.

[0053] Run the code in the configured running environment to check its basic functionality. The main purpose of this step is to verify whether the code can be successfully compiled and run in the current environment without involving complex logical analysis.

[0054] Specifically, for compiled languages (such as Java, C++), the code can be attempted to be compiled. During the compilation process, any syntax errors or type mismatches and other issues will be detected. For interpreted languages (such as Python, JavaScript), the code can be directly run to check for runtime exceptions (such as uncaught exceptions, null pointer references, etc.). Further, all output information during the running process (including standard output and error output) is recorded for subsequent analysis and debugging.

[0055] In step 303, it is determined whether the running state is abnormal.

[0056] If the code cannot be successfully compiled, the evaluation can be immediately stopped and the code can be marked as in an abnormal state and the code is unavailable. Common compilation errors include: syntax errors (such as missing semicolons, mismatched parentheses), type mismatches (such as variable types not matching the expectations), duplicate definitions (such as duplicate function or variable names), etc., which will not be listed one by one here. If the code compiles normally but has runtime exceptions, it should also be marked as in an abnormal state. Common runtime exceptions include: uncaught exceptions (such as division by zero error, array out-of-bounds), resource loading failures (such as file not found, database connection failure), performance issues (such as infinite loop, memory overflow), etc., which will not be listed one by one here.

[0057] If any errors or exceptions are found during the trial run, the code should be immediately marked as unavailable and scored 0 points. This approach ensures that only code in a normal state that can run successfully will enter the subsequent detailed evaluation stage.

[0058] The code marked as unavailable will be excluded from further quality evaluation because they cannot meet the most basic functional requirements. Further, a detailed error report can be provided, indicating the specific error location and reason to help developers quickly locate and fix the problem.

[0059] Through the above steps, the basic functionality of the code is ensured, that is, the code can be successfully compiled and run in the specified environment. This is the first hurdle in the evaluation process and also a prerequisite for subsequent evaluations. Basic functionality is the first step in code quality evaluation. Only by passing this step can deeper analysis (such as code similarity, structural rationality, etc.) be carried out. By automatically checking the basic functionality, the evaluation efficiency can be greatly improved and the need for manual intervention can be reduced.

[0060] In step 304, for the code with a normal running state, content analysis of the code is carried out to determine the first quality score of the code.

[0061] Specifically, determine the text content similarity between the code and the reference code, where the reference code is the code of the reference interface; determine the syntactic content similarity between the syntactic structure of the code and the syntactic structure of the reference code; and use the mean between the text content similarity and the syntactic content similarity as the first quality score.

[0062] In some embodiments, when determining the text content similarity, the following steps are mainly adopted: determine the term frequency of each token in the code; according to the occurrence frequency of the token in the first code set, determine the inverse document frequency of the token, where the inverse document frequency is negatively correlated with the occurrence frequency, and the first code set is a code set related to the application scenario of the code; multiply the term frequency by the inverse document frequency to determine the importance score of the token; arrange the importance scores of each token in the order of their appearance in the code to form a first vector; and obtain the text content similarity according to the first cosine value between the first vector and the second vector of the reference code, where the first cosine value is negatively correlated with the first quality score.

[0063] The first code set is a code set related to the application scenario of the code. It usually contains multiple code samples, and these samples are all written for the same or similar application scenarios. For example, if the code is used to generate a user interface, then the first code set may contain multiple code samples of different versions of the user interface. Ensure that the code samples in the first code set have the same business logic and functional requirements, so as to more accurately evaluate the quality of the generated code.

[0064] Term Frequency (TF) refers to the frequency of a certain token (such as a keyword, variable name, function name, etc.) appearing in a code snippet. The term frequency TF(t) is the number of times the token t appears in the code / the total number of tokens in the code. A high term frequency means that the token is relatively important in the code, but the term frequency alone cannot distinguish between common and keyword tokens.

[0065] The inverse document frequency measures the general importance of a token in the entire first code set. The inverse document frequency IDF(t) = log (the total number of code samples in the first code set / the number of code samples containing the token t).

[0066] The inverse document frequency is negatively correlated with the number of code samples containing the token t. If a token appears in many code samples, its inverse document frequency is low; if it appears in only a few code samples, its inverse document frequency is high. By using the inverse document frequency, the importance of common tokens is reduced and the weight of rare tokens is increased.

[0067] Multiply the word frequency by the inverse document frequency to obtain the importance score of the word segmentation: TF-IDF(t) = TF(t) × IDF(t). By comprehensively considering the importance of the word segmentation in a single code snippet and its universality in the entire first code set, the importance of the word segmentation can be measured more accurately.

[0068] Further, arrange the TF-IDF values of each word segmentation in the order of their appearance in the code to form a first vector, which represents the content characteristics of the code snippet. This step converts the code into a numerical vector, facilitating subsequent mathematical operations and comparisons.

[0069] Perform the same steps on the reference code as on the code to obtain a second vector of the reference code. The process of obtaining the second vector will not be elaborated here.

[0070] Measure the similarity degree of the two code snippets at the content level by calculating the first cosine value between the first vector and the second vector of the reference code. The value range of the first cosine value is between 0 and 1. The closer the value is to 1, the more similar the two code snippets are; the value close to 0 indicates that the two code snippets are quite different.

[0071] Further, determine the text content similarity according to the first cosine value. The first cosine value is negatively correlated with the first quality score.

[0072] Although the higher the cosine similarity, the better, in some cases, a higher cosine similarity may not necessarily mean high-quality code. Therefore, the scoring mechanism can be adjusted according to specific situations, so that there is a certain negative correlation between the first cosine value and the first quality score. Specifically: if the generated code is too similar to the reference code, there may be a suspicion of plagiarism, and at this time, the first quality score may be appropriately reduced. On the contrary, if the generated code is innovative while maintaining functional correctness, the first quality score may be increased.

[0073] In some embodiments, when determining the syntax content similarity, the following steps are mainly adopted: form a first syntax tree according to the syntax structure of the code. The first syntax tree includes multiple nodes representing the syntax units of the code. According to the attributes of the corresponding syntax units, the nodes have levels; determine the number of identical nodes with the same levels and representing the same syntax units between the first syntax tree and the second syntax tree of the reference code; and determine the syntax content similarity according to the number of identical nodes.

[0074] First, parse the generated code into the first Abstract Syntax Tree (AST). An AST is a tree-like structure that can capture the logical structure and control flow information of the code. Each node in the AST represents a syntactic unit, such as a function definition, variable declaration, conditional statement, etc. These nodes are hierarchically arranged according to the syntactic structure of the code. The nodes have a hierarchical relationship, where the parent node represents a higher-level syntactic structure and the child node represents more specific details. For example, a function definition node may contain a parameter list node and a function body node.

[0075] Specifically, use the corresponding compiler or parsing tool (such as Python's ast module, JavaScript's Esprima library, etc.) to convert the code snippet into an AST. Identify and classify each node to ensure that each node can correctly reflect its corresponding syntactic unit. For example, a function definition node should contain information such as the function name, parameter list, and function body.

[0076] Similarly, parse the reference code into the second syntax tree. This process is the same as the AST parsing of the code, ensuring that the two syntax trees have the same structure and node types.

[0077] Furthermore, use a structure matching algorithm to compare the first syntax tree of the code and the second syntax tree of the reference code, and find the same nodes with the same hierarchy and representing the same syntactic unit. Compare the hierarchical positions of the nodes in the two syntax trees to ensure that they are at the same level. Check whether the syntactic units represented by the nodes are the same. For example, two function definition nodes should not only be at the same level but also have the same function name and parameter list.

[0078] Specifically, start from the root node and gradually traverse the two syntax trees downward to find nodes with the same attributes. For each successfully matched node, continue to recursively match its child nodes until all levels have been compared. Record the number of successfully matched nodes and their hierarchical information.

[0079] Furthermore, count the number of nodes with the same hierarchy and representing the same syntactic unit between the first syntax tree and the second syntax tree. This number reflects the similarity degree of the syntactic structures of the two code segments.

[0080] Calculate the final syntactic content similarity score based on the number and weight of the same nodes. This score reflects the similarity between the generated code and the reference code at the syntactic and logical levels.

[0081] The syntactic content similarity score can be the ratio of the number of same nodes to the total number of nodes; it can also be the ratio of the sum of the weights of the same nodes to the sum of the weights of all nodes.

[0082] In step 305, compare the similarity between the running interface and the reference interface to determine the second quality score of the code.

[0083] Specifically, use a vectorization model to convert the running interface of the code into a third vector, which can represent the interface content and features of the running interface; and calculate the second cosine value between the third vector and the fourth vector of the reference interface to obtain the second quality score, where the second cosine value is negatively correlated with the second quality score.

[0084] First, run the code in the configured running environment to obtain the running interface and take a screenshot of this running interface. This screenshot is the image basis for evaluating the running interface. For example, after the code runs, it can display web pages, mobile application interfaces, or data visualization charts, etc.

[0085] The vectorization model (Contrastive Language-Image Pre-training, CLIP) converts the screenshot of the running interface into a high-dimensional vector, that is, the third vector. This process includes the following steps: preprocess the generated interface screenshot, such as operations like resizing, cropping, and normalization to meet the input requirements of the CLIP model. Through the image encoder part of the CLIP model, extract the key features in the image and convert them into a fixed-length high-dimensional vector. This vector can represent the content and features of the running interface, such as layout, color, element position, etc. The vector contains information at different levels, from the global layout to local details, ensuring a comprehensive reflection of the overall structure and visual effect of the interface.

[0086] Similarly, the system also inputs the screenshot of the reference interface into the CLIP model to generate the corresponding fourth vector. Ensure that the reference interface screenshot and the running interface screenshot go through the same preprocessing steps for comparability during comparison.

[0087] Measure the similarity between the two images by calculating the cosine value between the third vector and the fourth vector. The value range of the cosine value is between 0 and 1. The closer the value is to 1, the more similar the two images are; the value close to 0 indicates that the two images are quite different.

[0088] Based on the calculated second cosine value, the second quality score of the code can be obtained. This score reflects the consistency of the code with the reference implementation in terms of visual effects.

[0089] Although a higher cosine value usually means better visual consistency, in some cases, an overly high similarity may indicate that the generated code lacks innovation or is suspected of plagiarism. Therefore, the scoring mechanism can be adjusted according to specific situations so that there is a certain negative correlation between the second cosine value and the second quality score.

[0090] In step 306, the comprehensive quality score of the code is determined according to the first quality score and the second quality score.

[0091] Specifically, both the first quality score and the second quality score are normalized to the target score range to obtain the first normalized quality score and the second normalized quality score; and the mean value between the first normalized quality score and the second normalized quality score is calculated, and the mean value is used as the comprehensive quality score.

[0092] More specifically, the first maximum value and the first minimum value of the scoring range corresponding to the first quality score are determined, and the second maximum value and the second minimum value of the scoring range corresponding to the second quality score are determined; the difference between the first quality score and the first minimum value is divided by the difference between the first maximum value and the first minimum value to obtain the first normalized quality score, and the first normalized quality score is within the target score range, and the target score range is 0 to 1; and the difference between the second quality score and the second minimum value is divided by the difference between the second maximum value and the second minimum value to obtain the second normalized quality score, and the second normalized quality score is within the target score range.

[0093] For example, if the scoring range of the first quality score is 0 to 100, then the first maximum value is 100 and the first minimum value is 0. If the first quality score is 87, then the first normalized quality score is (87 - 0) / (100 - 0), that is, 0.87. If the scoring range of the second quality score is -1 to 1, then the second maximum value is 1 and the second minimum value is -1. If the second quality score is 0.5, then the second normalized quality score is [0.5 - (-1)] / [1 - (-1)], that is, 0.75. Obviously, regardless of whether the scoring ranges are the same, after the normalization process, both the first normalized quality score and the second normalized quality score are within the target score range, that is, 0 to 1.

[0094] Furthermore, by calculating the mean value of the first normalized quality score and the second normalized quality score, the comprehensive quality score of the code can be determined.

[0095] It should be noted that the normalization process of the quality score and the second quality score is not limited to the foregoing, and any method that can convert the two to the same target score range falls within the protection scope of the present disclosure.

[0096] The present disclosure adopts a modular design to ensure that each step (such as dependency installation, code running, similarity calculation, etc.) runs as an independent module, thereby providing high flexibility and maintainability. This design not only allows each module to be independently replaced and upgraded, but also allows customization according to specific requirements. For example: Different similarity calculation models can be easily replaced according to specific evaluation requirements (such as switching from TF-IDF cosine similarity to similarity calculation of BERT embedding vectors), or more dimensions of scoring mechanisms can be added (such as introducing code complexity analysis or performance metrics). This flexibility enables the system to adapt to different types of code generation tasks and provide accurate evaluation results whether focusing on text similarity or logical consistency.

[0097] This method also supports automated dependency resolution and installation functions, which can identify and process dependency declarations in code files and automatically download and configure the required libraries and tools. This feature not only improves the efficiency of code evaluation but also reduces the need for manual intervention, especially suitable for scenarios where a large number of generated codes need to be quickly evaluated.

[0098] In addition, the good scalability of the system enables it to handle code evaluation of multiple programming languages. By implementing corresponding dependency resolution and abstract syntax tree (AST) generation modules for different languages, new programming language support can be seamlessly integrated. Specifically: By developing specialized parsers and AST generators for each programming language, the system can flexibly handle the syntax and structural characteristics of different languages. For example, for Python code, the ast module can be used to parse the code; while for JavaScript code, the Esprima library can be used. This modular design enables the system to easily expand to support multiple programming languages such as C++, Java, Ruby, etc.

[0099] This scalability not only enhances the applicable scope of this disclosure but also enables this disclosure to adapt to various complex business scenarios. Whether in front-end user interface code generation, back-end service logic implementation, or in fields such as data visualization and machine learning model development, a reliable scoring mechanism can be provided to ensure that the generated code is not only functionally correct but also meets high standards in terms of user experience and interface consistency. Therefore, this disclosure has broad application prospects, can significantly improve the quality and efficiency of code generation, and promote the development of automated code generation technology.

[0100] The quality assessment method of the present disclosure combines the syntax structure of the code (AST) and the image similarity of the rendering results, breaking through the limitations of a single modality. It not only evaluates the running correctness of the code but also delves into the logical structure and visual output consistency, comprehensively considering the function, structure, and performance of the code. The use of syntax tree similarity calculation avoids misjudgment problems in simple text similarity evaluation and improves the evaluation accuracy. The automated dependency management and full-process verification functions enhance the evaluation efficiency and reduce the manual error rate. The multi-dimensional scoring system that combines TF-IDF cosine similarity, AST similarity, and image similarity makes the evaluation more comprehensive and is particularly suitable for code evaluation that generates visual elements. Finally, this solution is not only applicable to general code generation quality evaluation but also provides a reliable scoring mechanism for complex business scenarios, ensuring that the generated code meets high standards in terms of both function and user experience, significantly improving the measurability and interpretability of code generation quality.

[0101] Figure 4 It is a structural schematic block diagram of a quality assessment device according to an embodiment of the present disclosure.

[0102] As Figure 4 shown, the present disclosure provides a quality assessment device 400, including: a first analysis module 410 for performing content analysis on the code to determine a first quality score of the code; a second analysis module 420 for comparing the similarity between the running interface of the code and a reference interface to determine a second quality score of the code; and a comprehensive quality score determination module 430 for determining the comprehensive quality score of the code according to the first quality score and the second quality score.

[0103] The quality assessment device 400 of the present disclosure may be in the form of computer software, and each module of the quality assessment device 400 may be in the form of a computer software module.

[0104] Each module of the quality assessment device 400 of the present disclosure is set to implement each step of the quality assessment method. Its execution principle and steps can be referred to the foregoing, and will not be elaborated here.

[0105] Figure 5 It is a structural schematic block diagram of an electronic device according to an embodiment of the present disclosure. As Figure 5 shown, the present disclosure further provides an electronic device 1000, including: a processor 1200 and a memory 1300, where the memory 1300 stores execution instructions; the processor 1200 executes the execution instructions stored in the memory 1300, so that the processor 1200 executes the quality assessment method.

[0106] The hardware structure of the electronic device 1000 can be implemented using a bus architecture. The bus architecture can include any number of interconnected buses and bridges, depending on the specific application of the hardware and the overall design constraints. The bus 1100 connects various circuits including one or more processors 1200, a memory 1300, and / or hardware modules together. The bus 1100 can also connect various other circuits 1400 such as peripheral devices, voltage regulators, power management circuits, external antennas, etc.

[0107] The bus 1100 can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Component (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, only one connecting line is used in this figure, but it does not mean that there is only one bus or one type of bus.

[0108] The present disclosure also provides a readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, it is used to implement the above method. The "readable storage medium" can be any device that can contain, store, communicate, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. More specific examples of the readable storage medium include the following: an electrical connection part with one or more wirings (electronic device), a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable read-only memory (CDROM), etc.

[0109] The present disclosure also provides a computer program product. The method of the present disclosure can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed, the processes or functions of the present disclosure are executed in whole or in part.

[0110] Computer programs or instructions can be stored in a readable storage medium or transmitted from one readable storage medium to another. For example, the computer programs or instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless manner. The readable storage medium can be any available medium that can be accessed or a data storage device such as a server or data center integrating one or more available media. The available medium can be a magnetic medium, such as a floppy disk, hard disk, or magnetic tape; it can also be an optical medium, such as a digital video disc; or it can be a semiconductor medium, such as a solid-state drive. The computer-readable storage medium can be a volatile or non-volatile storage medium, or can include both volatile and non-volatile types of storage media.

[0111] Those skilled in the art should understand that the embodiments of the present disclosure can be provided as a method, an electronic device, a readable storage medium, or a computer program product. Therefore, the present disclosure can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present disclosure can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

[0112] The present disclosure is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the present disclosure. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or multiple flows and / or blocks

[0113] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in Figure 1 one or more of the flows Figure 1 or multiple flows and / or blocks

[0114] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 steps of the functions specified in one block or multiple blocks.

[0115] In the description of this specification, the description with reference to the terms "one embodiment / way", "some embodiments / ways", "example", "specific example", or "some examples", etc. means that the specific features, structures, or characteristics described in connection with the embodiment / way or example are included in at least one embodiment / way or example of the present disclosure. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment / way or example. Moreover, the specific features, structures, or characteristics described can be combined in a suitable manner in any one or more embodiments / ways or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments / ways or examples described in this specification and the features of different embodiments / ways or examples.

[0116] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" can explicitly or implicitly include at least one of these features. In the description of the present disclosure, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0117] Those skilled in the art should understand that the above embodiments are only for clearly explaining the present disclosure and are not intended to limit the scope of the present disclosure. For those skilled in the art, other changes or modifications can be made on the basis of the above disclosure, and these changes or modifications are still within the scope of the present disclosure.

Claims

1. A quality assessment method, characterized in that: include: performing content analysis on the code to determine a first quality score for the code; Comparing the similarity between the running interface of the code and the reference interface to determine a second quality score of the code; as well as A comprehensive quality score of the code is determined according to the first quality score and the second quality score.

2. The quality assessment method according to claim 1, characterized in that: Performing content analysis on the code to determine a first quality score of the code includes: Determining a text content similarity between the code and a reference code, the reference code being a code of the reference interface; determining a grammatical content similarity between a grammatical structure of the code and a grammatical structure of the reference code; and The average of the text content similarity and the grammatical content similarity is taken as the first quality score.

3. The quality assessment method according to claim 2, characterized in that: Determining the text content similarity between the code and a reference code, including: Determining the frequency of each word in the code; Determine an inverse text frequency of the word segmentation according to the total number of code samples containing the word segmentation in a first code set, wherein the inverse text frequency is negatively correlated with the total number of code samples containing the word segmentation, and the first code set is a set of code samples related to an application scenario of the code; Multiplying the word frequency by the inverse text frequency to determine the importance score of the word segment; Arranging the importance scores of the participles according to the order in which the participles appear in the code to form a first vector; and The text content similarity is obtained according to a first cosine value between the first vector and the second vector of the reference code, and the first cosine value is negatively correlated with the first quality score.

4. The quality assessment method according to claim 2, characterized in that: Determining the grammatical content similarity between the grammatical structure of the code and the grammatical structure of the reference code includes: According to the grammatical structure of the code, a first grammatical tree is formed, wherein the first grammatical tree includes a plurality of nodes representing grammatical units of the code, and the nodes have a hierarchy according to the attributes of the corresponding grammatical units; Determining the number of identical nodes having the same level and representing the same syntax unit between the first syntax tree and the second syntax tree of the reference code; and The grammatical content similarity is determined according to the number of identical nodes.

5. The quality assessment method according to claim 1, characterized in that: Comparing the similarity between the running interface of the code and the reference interface to determine the second quality score of the code includes: Converting the code running interface into a third vector using a vectorization model, wherein the fourth vector can represent the interface content and interface features of the running interface; and A second cosine value between the third vector and a fourth vector of the reference interface is calculated to obtain the second mass score, wherein the second cosine value is negatively correlated with the second mass score.

6. The quality assessment method according to claim 1, characterized in that: Determining a comprehensive quality score of the code according to the first quality score and the second quality score includes: Normalizing both the first quality score and the second quality score to a target score range to obtain a first normalized quality score and a second normalized quality score; and An average value between the first standardized quality score and the second standardized quality score is calculated, and the average value is used as the comprehensive quality score.

7. The quality assessment method according to claim 6, characterized in that: Standardizing the first quality score and the second quality score to a target score range to obtain a first standardized quality score and a second standardized quality score, comprising: Determine a first maximum value and a first minimum value of a score range corresponding to the first quality score, and determine a second maximum value and a second minimum value of a score range corresponding to the second quality score; Dividing the difference between the first quality score and the first minimum value by the difference between the first maximum value and the first minimum value to obtain the first standardized quality score, wherein the first standardized quality score is within the target score range, and the target score range is 0 to 1; and The second standardized quality score is obtained by dividing the difference between the second quality score and the second minimum value by the difference between the second maximum value and the second minimum value, and the second standardized quality score is within the target score range.

8. The quality assessment method according to claim 1, characterized in that: Prior to determining a first quality score of the code, comprising: Install the dependencies of the code in a blank project to create a running environment suitable for the code.

9. An electronic device, characterized in that: include: A memory storing execution instructions; as well as A processor, wherein the processor executes the execution instructions stored in the memory, so that the processor executes the quality assessment method according to any one of claims 1 to 8.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the quality assessment method according to any one of claims 1 to 8 is implemented.

Citation Information

Cited By

  • New code generation method and device based on historical code library, storage medium and electronic equipment

    CN121387255A