Evaluation method and computing device
By obtaining the number of new codes, total generated codes and total adopted codes, combining context information and syntax segmentation, building prompt words, and using the encoding record database to evaluate the language model, the problem of accurately evaluating the code generation ability of private domain projects is solved, and the code generation efficiency and quality are improved.
Patent Information
- Application Number
- CN202510330402.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-22
AI Technical Summary
How to accurately evaluate the ability of language models to generate private domain project code to improve code generation efficiency and quality in selecting the appropriate language model.
By obtaining the number of new codes, total generated codes and total adopted codes, using indicators such as completion rate and adoption rate, the evaluation value of the language model is calculated, combined with context information and grammatical segmentation of code clusters, the prompt words are constructed, the generated and adopted codes are recorded, and the code records database is used for evaluation.
It realizes the accurate evaluation of the code generation ability of the language model in private domain projects, improves the accuracy of selecting suitable models and the quality of generating codes, and enhances the applicability and accuracy of the language model.
Smart Images

Figure CN120353673A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of servers, and in particular, to an evaluation method and a computing device. Background Art
[0002] Due to its ability to improve development efficiency and reduce repetitive work, language models are increasingly widely used in the field of programming. Currently, language models can be applied to generate source codes for private projects. Private project codes refer to the source codes of software projects developed by enterprises or individuals for specific needs or products.
[0003] Accurately evaluating the ability of a language model to generate private project codes is crucial for developers to select appropriate language models to improve the generation efficiency and quality of codes. Therefore, how to accurately evaluate the ability of a language model in generating private project codes has become an urgent technical problem to be solved. Summary of the Invention
[0004] Embodiments of this application provide an evaluation method and a computing device for accurately evaluating the ability of a language model to generate private project codes.
[0005] In a first aspect, embodiments of this application provide an evaluation method applied to a computing device. The computing device can be a server or other electronic devices with computing functions, and embodiments of this application do not specifically limit it. When the computing device is a server, the computing device can be a blade server, a rack server, or a server with other structures, and embodiments of this application do not specifically limit it.
[0006] Specifically, the computing device obtains the number of newly added codes, the total number of generated codes, and the total number of adopted codes. Among them, the number of newly added codes is the codes that are newly added and meet the requirements of the object to be developed within a preset time interval. The total number of generated codes is the total number of codes generated after processing the development information of the object to be developed n times based on the language model within a preset time interval. The total number of adopted codes is the number of codes adopted among the total number of generated codes, and n is an integer greater than or equal to 1. Then, the computing device uses the number of newly added codes, the total number of generated codes, and / or the total number of adopted codes to obtain a first evaluation value of the language model; the first evaluation value is used to evaluate the code generation ability of the language model.
[0007] It can be understood that when the object to be developed is a private project, the computing device uses the number of newly added codes that can directly reflect the actual increased code volume in the private project, or the total number of generated codes that directly reflects the total amount of codes generated by the language model for the private project, or the total number of adopted codes that directly reflects the total number of generated codes adopted, which are data directly reflecting the actual development situation of the language model in the private project, to obtain the first evaluation value. Thus, the first evaluation value can be used to accurately evaluate the code generation ability of the language model in the private project.
[0008] In one implementation, the computing device can obtain the completion rate of the language model based on the number of newly added codes and the total number of adopted codes; the completion rate of the language model has a positive correlation with the number of newly added codes and a negative correlation with the total number of adopted codes; using the completion rate of the language model, obtain the first evaluation value of the language model. That is, the computing device obtains the first evaluation value through the completion rate. At this time, the first evaluation value indicates the contribution degree of the generated code, so it can accurately evaluate the code generation ability of the language model.
[0009] In another implementation, the computing device can determine the adoption rate of the language model based on the total number of adopted codes and the total number of generated codes; the adoption rate of the language model has a positive correlation with the total number of adopted codes and a negative correlation with the total number of generated codes; according to the adoption rate of the language model, obtain the first evaluation value of the language model. That is, the computing device obtains the first evaluation value through the adoption rate. At this time, the first evaluation value indicates the quality of the generated code, so it can accurately evaluate the code generation ability of the language model.
[0010] In yet another implementation, the computing device can obtain the completion rate of the language model based on the number of newly added codes and the total number of adopted codes; the completion rate of the language model has a positive correlation with the number of newly added codes and a negative correlation with the total number of adopted codes; determine the adoption rate of the language model based on the total number of adopted codes and the total number of generated codes; the adoption rate of the language model has a positive correlation with the total number of adopted codes and a negative correlation with the total number of generated codes; obtain the first evaluation value of the language model according to the adoption rate of the language model and the completion rate of the language model. That is, the computing device can use the adoption rate and the completion rate to comprehensively evaluate the code generation ability of the language model and further improve the evaluation accuracy.
[0011] In still another implementation, the computing device can obtain the first code file at the start time of the preset time interval and the second code file at the end time of the preset time interval; the first code file is used to store the codes that meet the requirements of the object to be developed before the start time, and the second code file is used to store the codes that meet the requirements of the object to be developed before the end time; obtain the number of newly added codes according to the first code file and the second code file. Since the code file is the most direct and specific result manifestation in the software development process, the computing device can accurately obtain the number of newly added codes by comparing the code files at different times.
[0012] In yet another implementation, the total number of generated codes includes n times of code generation, and each time of code generation is obtained in the following way: obtain the context information in the code file, where the code file is used to store the codes that meet the requirements of the object to be developed at the current moment; construct a prompt based on the context information and the development information of the object to be developed; input the prompt into the language model to obtain the generated code. Thus, the computing device can construct a prompt based on the context information, ensuring that the code generated by the language model better meets the actual requirements of the private domain project, thereby further improving the evaluation accuracy.
[0013] Furthermore, the computing device can also perform syntax splitting on the code file to obtain multiple groups of code clusters; the codes in each group of code clusters among the multiple groups of code clusters have the same syntax; construct a prompt based on the context information and the development information of the object to be developed, in combination with the multiple groups of code clusters and / or the location information of the code file; the location information of the code file indicates the code information included in the corresponding location of the code file. That is, the computing device can further utilize the location information and syntax splitting of the code file to guide the code generated by the language model to meet the actual requirements, thereby further improving the evaluation accuracy.
[0014] In yet another implementation, the computing device can obtain the number of deleted codes and the number of requests; the number of deleted codes is the number of codes deleted in the code file within a preset time interval; the number of requests is the number of times the language model is accessed within a preset time interval; if any one of the number of requests, the number of deleted codes, and the number of newly added codes is not 0, determine the coding time as the preset time interval; determine the second evaluation value of the language model according to the coding time and the number of newly added codes; the second evaluation value of the language model is used to evaluate the coding efficiency of the language model; obtain the comprehensive evaluation value of the language model according to the second evaluation value of the language model and the first evaluation value of the language model, and the comprehensive evaluation value is used to evaluate the coding ability of the language model. That is, the computing device can comprehensively evaluate the coding ability of the language model by using the coding efficiency and the code generation score, making the evaluation accuracy of evaluating the generation ability of the language model higher.
[0015] In yet another implementation, the computing device can write the coding record data of the language model into the coding record database to evaluate the coding ability of the language model by using the coding record data in the coding record database; where the coding record data includes the coding time, the newly added codes corresponding to the number of newly added codes, the generated codes corresponding to the total number of generated codes, the adopted codes corresponding to the total number of adopted codes, the first evaluation value, the number of requests, the adoption rate, and / or the completion rate. Thus, the computing device can write the coding record data into the coding record database, facilitating the manager to select a suitable target language model.
[0016] In a second aspect, an embodiment of the present application provides an evaluation device, which is applied to a computing device. The computing device can be a server or other electronic devices with computing functions, and the embodiments of the present application do not specifically limit it. When the computing device is a server, the computing device can be a blade server, a rack server, or a server with other structures, and the embodiments of the present application do not specifically limit it. The device includes:
[0017] The first acquisition unit is used to acquire the number of newly added codes, the total number of generated codes, and the total number of adopted codes for the computing device.
[0018] Among them, the number of newly added codes is the codes that are newly added within a preset time interval and meet the requirements of the object to be developed. The total number of generated codes is the total number of codes generated after processing the development information of the object to be developed n times based on the language model within a preset time interval. The total number of adopted codes is the number of codes adopted among the total number of generated codes, and n is an integer greater than or equal to 1;
[0019] The second acquisition unit is used to obtain a first evaluation value of the language model by using the number of newly added codes, the total number of generated codes, and / or the total number of adopted codes; the first evaluation value is used to evaluate the code generation ability of the language model.
[0020] In a third aspect, an embodiment of the present application provides a computing device, including:
[0021] A memory for storing programs;
[0022] A processor for executing the programs stored in the memory. When the programs stored in the memory are executed, the processor is used to execute the method according to any one of the first aspect.
[0023] In a fourth aspect, the present application provides a computer storage medium for storing a computer program. When the computer program is executed, it is used to implement the method provided by any one of the embodiments in the first aspect of the present application.
[0024] In a fifth aspect, the present application provides a computer program product containing instructions. When it runs on at least one computing device, it enables at least one computing device to implement the method provided by any one of the embodiments in the first aspect of the present application.
[0025] Any of the above-provided evaluation methods, corresponding computing devices, computer-readable storage media, computer program products, etc. are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 It is a schematic diagram of an application scenario provided by an embodiment of the present application;
[0027] Figure 2 It is a flowchart of an evaluation method provided by an embodiment of the present application;
[0028] Figure 3 It is another interaction diagram of an evaluation method provided by an embodiment of the present application;
[0029] Figure 4 It is another interaction diagram of an evaluation method provided by an embodiment of the present application;
[0030] Figure 5 It is a schematic diagram of recording coded records into the adopted code database and the coded record database provided by an embodiment of the present application;
[0031] Figure 6 It is a schematic diagram of the structure of an evaluation device provided by an embodiment of the present application;
[0032] Figure 7 It is a schematic diagram of the structure of a server provided by an embodiment of the present application. Detailed implementation manners
[0033] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0034] First, the technical terms related to the embodiments of the present application are introduced.
[0035] Code: including private domain project code and non-private domain project code.
[0036] Among them, the private domain project code usually refers to the source code of a software project developed by an enterprise or an individual for specific needs or products. For example, the private domain project code can be the source code of a software project corresponding to a private domain e-commerce platform, that is, an e-commerce platform that an enterprise hopes to develop through private domain traffic for e-commerce sales. The private domain project code may involve core business logics, algorithms, and database designs, etc. Therefore, the private domain project code usually needs to be kept confidential.
[0037] Language model: Also known as a Large Language Model (LLM), it is an artificial intelligence model that uses deep learning algorithms. Through training with a large amount of data, it learns the patterns and structures of language, enabling it to understand and generate natural language text. In the embodiments of this application, the language model can be trained with a large amount of code data to learn the syntax, structure, and logic of different programming languages, thereby obtaining the ability to generate code. The language model can be, for example, DeepSeek, or it can be Wenyan Yixin or Tongyi Qianwen, etc. The embodiments of this application do not specifically limit it.
[0038] Currently, language models can significantly improve code development efficiency, reduce repetitive work, and help developers focus on more creative tasks, etc., so they are increasingly widely used in the programming field. However, there are many current language models, and how to select a language model with stronger code generation ability and higher coding efficiency has become an urgent problem to be solved. Therefore, it is necessary to accurately evaluate the code generation ability of language models, especially the ability to generate code for private domain projects.
[0039] The embodiments of this application provide an evaluation method, which accurately evaluates the code generation ability of language models in private domain projects by using the number of newly added codes that directly reflect the actual increased code volume in the object to be developed, or the total number of generated codes that directly reflect the total code volume generated by the language model, or the total number of adopted generated codes that directly reflect the total number of adopted generated codes.
[0040] To better illustrate the evaluation method provided by the embodiments of this application, the application scenarios of the embodiments of this application are introduced below.
[0041] Exemplarily, attached Figure 1 is a schematic diagram of an application scenario provided by the embodiments of this application. It specifically relates to a computing device 101 and a developer device 102. The developer device 102 and the computing device 101 can be interconnected through a network, such as through a communication network including at least one switch.
[0042] In the actual application scenario, the developer device 102 can also be called a developer workstation, which is oriented to developers and is the front-end interface for developers to interact with the computing device 101. In the embodiments of this application, the developer device 102 can be an electronic device deployed with a code generation plugin, such as a mobile phone, a laptop, a wearable electronic device (such as a smart watch), a tablet computer, a server, a computer, etc. The developer can start the code generation plugin through the developer device 102 to initiate the code generation process.
[0043] Meanwhile, developers can also input the development information of the object to be developed through the developer device 102. Among them, the development information is used to implement specific development functions. For example, if the object to be developed is "coupons of a certain e-commerce system", the development information corresponding to "coupons of the e-commerce system" includes, but is not limited to, coupon types, discount amounts or discount ratios, applicable products or product categories, etc. Specifically, the developer device 102 displays an operation interface for generating code, and developers can input the development information of the object to be developed through this operation interface to achieve interaction with the device 102 to be developed.
[0044] After the developer device 102 (specifically, the code generation plugin) receives the development information of the object to be developed input by the developer, it sends the development information to the computing device 101. Based on the received development information, the computing device 101 requests the language model to generate code corresponding to the development information. It should be noted that in actual applications, multiple language models are deployed on the computing device 101. The operation interface of the developer device 101 also includes a language model selection area, which includes multiple language models deployed on the computing device 101. Developers select the target language model in the language model selection area and send the selected result to the computing device 101. The computing device 101 can process the development information by calling the target language model based on the received result to generate code that meets the development information. The computing device 101 can return the generated code to the corresponding developer device 102 to be built into the corresponding code file of the developer device 101.
[0045] The computing device 101 refers to an electronic device with computing functions. For example, the computing device 101 can be a server or other electronic devices, and the embodiments of the present application do not specifically limit it. When the computing device 101 is a server, the computing device 101 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), as well as large databases and artificial intelligence platforms. When the above server is a server cluster or distributed system composed of multiple physical servers, the multiple physical servers can form a blockchain, and each physical server is a node on the blockchain. The physical type of the server area can be a rack server, a high-density server, a GPU server, a tower server, or a blade server, an all-in-one cabinet server, etc., and the present application does not specifically limit it.
[0046] In one example, the computing device 101 includes a model deployment module 101-1 and a model evaluation module 101-2. In the embodiments of the present application, the model deployment module 101-1 deploys at least one language model. After the code generation plugin in the developer device 101 obtains the development information, it sends the development information to the language model that can be called by the model deployment module 101-1 to generate code corresponding to the development requirements. Further, the operation interface of the code generation plugin can display multiple deployed language models, such as displaying Language Model A, Language Model B, and Language Model C. The developer can specify a language model through the operation interface, such as specifying Language Model C. The code generation plugin can call Language Model C from the computing device 101 based on the development information to generate code corresponding to the development information.
[0047] In an actual application scenario, different language models may have different code generation capabilities. To help developers select the best language model to perform the code generation operation, the computing device 101 further includes a model evaluation module 101-2. The model evaluation module 101-2 is used to evaluate the code generated by the language model. In one example, the model evaluation module 101-2 can send the evaluation result to the developer device 102 for display. Thus, the developer device 102 can select the required language model to perform code generation based on the evaluation result.
[0048] It should be noted that the above application method of the evaluation result is only a schematic representation. In actual use, it can also be applied in other aspects, such as evaluating the coding efficiency of developers, or evaluating the difficulty of developing private project codes, etc. The embodiments of the present application do not specifically limit this.
[0049] In the embodiments of the present application, the model evaluation module 101-2 is specifically configured to obtain the number of newly added codes, the total number of generated codes, and the total number of adopted codes within a preset time interval. Among them, the number of newly added codes is the codes that are newly added and meet the requirements of the object to be developed within the preset time interval. The total number of generated codes is the total number of codes generated after processing the development information of the object to be developed n times based on the language model within the preset time interval. The total number of adopted codes is the number of codes adopted among the total number of generated codes, and n is a natural number. Then, using the number of newly added codes, the total number of generated codes, and / or the total number of adopted codes, obtain the first evaluation value of the language model; the first evaluation value is used to evaluate the code generation ability of the language model. Thus, by directly reflecting the actual development situation of the language model in the private project data, obtain the first evaluation value. Thus, the first evaluation value can be used to accurately evaluate the code generation ability of the language model in the private project.
[0050] It should be noted that in the embodiments of the present application, the computing device 101 and the developer device 102 belong to different devices as an example for illustrative purposes. In actual application scenarios, the developer device 102 may also be located in the same computing device as the computing device 101. At this time, the developer device 102 and the computing device 101 are respectively the developer module and the computing module of the same computing device, and the embodiments of the present application do not specifically limit this.
[0051] It should be noted that in the embodiments of the present application, the computing device 101 is connected to one developer device 102. In actual use scenarios, the computing device can be connected to multiple developer devices at the same time. After multiple developer devices start the code generation plugin, the computing device 101 can manage the code generation of multiple developer devices at the same time.
[0052] It should be noted that the above application scenarios and the hardware structure of the computing device are only schematic representations. In actual use, those skilled in the art can also make adjustments according to needs. For example, the computing device 101 communicates with multiple developer devices at the same time to obtain development information sent by multiple developer devices, provides corresponding language models for multiple developer devices respectively, and performs code generation operations. And the computing device 101 obtains the number of newly added codes, the total number of generated codes and / or the total number of adopted codes generated by multiple developer devices for the corresponding objects to be developed, and conducts comprehensive analysis to obtain the evaluation result of the language model. The embodiments of the present application do not specifically limit this.
[0053] The following will describe in detail the evaluation method provided by the embodiments of the present application with reference to the accompanying drawings. To enable those skilled in the art to better understand the evaluation method provided by the embodiments of the present application, the following will use the Figure 1 application scenario as an example for description.
[0054] Embodiment 1
[0055] Attached Figure 2 is a flowchart of an evaluation method provided by an embodiment of the present application. The method includes the following content:
[0056] Within a preset time interval, the developer inputs the development information corresponding to the object to be developed n times through the developer device 102, which is processed by the same language model in the computing device 101. Among them, each development information can be the same or different.
[0057] S210. The computing device 101 obtains the number of newly added codes, the total number of generated codes, and the total number of adopted codes within the preset time interval.
[0058] Specifically, the computing device 101 obtains the number of newly added codes, the total generated codes, and the total adopted codes for the object to be developed within a preset time interval. The object to be developed is a software project that requires source code development. In the embodiments of the present application, the object to be developed specifically refers to a private domain project.
[0059] The number of newly added codes refers to the number of codes newly obtained within a preset time interval that meet the requirements of the object to be developed. The number of newly added codes can reflect the actual increase in the amount of code of the object to be developed within the preset time interval. Therefore, the number of newly added codes can be used to evaluate the coding efficiency of the language model. For example, the object to be developed is user login and registration, and the preset time interval is 1 hour (h). Within this 1h, the developer device 102 obtains the codes for the user login and registration interface, including the codes corresponding to the text input box, the codes corresponding to the button click event, etc. These codes are all newly obtained and can implement the functional requirements of user login and registration. The total number of lines of these codes is the number of newly added codes corresponding to this preset time interval.
[0060] In the embodiments of the present application, after the developer device 102 obtains the codes that meet the requirements of the object to be developed, it will store the codes in a code file. Among them, the code file is the most direct and specific result manifestation in the software development process. Therefore, the computing device 101 can directly and accurately obtain the number of newly added codes that reflect the requirements of the object to be developed through the code file.
[0061] Specifically, the computing device 101 counts the first code file D1 corresponding to the start time of the preset time interval and the second code file D2 corresponding to the end time. Among them, the number of newly added lines of code in D2 compared to D1 is the number of newly added codes. Exemplarily, if D1 includes 10 lines of code, and the codes from the first line to the tenth line are a1 to a10 in sequence. D2 includes 11 lines of code, and the codes from the first line to the eleventh line are b1 to b11 in sequence. Among them, the code contents of a1 to a9 are the same as those of b1 to b9, and the code contents of b10 to b11 do not exist in D1. At this time, the number of newly added codes is 2, and the corresponding newly added codes are b10 to b11.
[0062] In the embodiments of the present application, for the convenience of description, the computing device marks the newly added codes as a.
[0063] The total number of generated codes refers to the total number of codes generated after processing the development information of the object to be developed by the language model n times within a preset time interval. The total number of generated codes can directly reflect the total amount of private domain project codes generated by the language model within the preset time interval. For example, the object to be developed is still user login and registration, and the preset time interval is 1 hour. Within this 1 hour, the computing device first uses the language model to process the information in the text input box to generate the text input box code, simply referred to as the first generated code f1, and then uses the language model to process the button click event information to generate the button click event code, simply referred to as the second generated code f2. Then the total number of generated codes = the number of codes of f1 + the number of codes of f2. Further, for example, the number of codes of f1 is 110 lines, and the number of codes of f2 is 86 lines, then the total number of generated codes is 196 lines.
[0064] It should be noted that n is not a fixed number. n is determined by the number of requests made by the code generation plugin to the language model for code generation within the preset time interval. n can be 0, or it can be an integer greater than or equal to 1. Exemplarily, if within a certain preset time interval, the code generation plugin does not request the language model, that is, the number of requests is 0, then n is 0. If the number of requests made by the code generation plugin to the language model is 3 times, then n is 3.
[0065] It can be understood that not all the codes in the generated codes may be available. There may be situations where some codes are available, some codes may not be available, or all codes may not be available. In the embodiments of the present application, the available codes in the generated codes are called adopted codes. In one application, the codes that are built into the code file and have not changed after waiting for a preset duration are called adopted codes. Wherein, the preset duration is a duration set by those skilled in the art according to needs. For example, the preset duration is 30s. Exemplarily, if the generated codes are c1 to c10, and after being written into the code file, the codes that have not changed after waiting for 30s are c1 to c9, then the adopted codes are c1 to c9. It should be noted that c1 represents the first line of code in the code file, and c10 represents the 10th line of code in the code file.
[0066] S220. The computing device 101 obtains the first evaluation value of the language model by using the number of newly added codes, the total number of generated codes, and / or the total number of adopted codes within the preset time interval.
[0067] Among them, the first evaluation value is used to evaluate the code generation ability of the language model. The code generation ability of the language model is related to the first evaluation value. For example, it can be a positive correlation relationship, that is, the higher the first evaluation value, the stronger the code generation ability of the language model. Another example is that it can be a negative correlation relationship, the lower the first evaluation value, the lower the code generation ability of the language model. For ease of understanding, the following will take the example where the code generation ability of the language model is related to the first evaluation value for illustration.
[0068] In an embodiment of the present application, the computing device 101 can use any one of the number of newly added codes, the total number of generated codes, and the total number of adopted codes within a preset time interval (this number is referred to as the target number) to obtain a first evaluation value. The first evaluation value has a positive correlation with the target number, and the more the target number, the higher the first evaluation value.
[0069] However, considering that only one number is considered during evaluation, there will be a problem of low evaluation accuracy. For this reason, the embodiment of the present application can obtain the first evaluation value in the following manner:
[0070] Example 1: The computing device 101 obtains the completion rate of the language model according to the number of newly added codes and the total number of adopted codes. Among them, the completion rate has a positive correlation with the number of newly added codes and a negative correlation with the total number of adopted codes. For example, the completion rate = the total number of adopted codes / the number of newly added codes. For instance, if the number of newly added codes is 100 lines and the total number of adopted codes is 10 lines, then the completion rate is 0.1. It can be understood that the completion rate directly reflects the contribution degree of the generated codes among the newly added codes.
[0071] Next, the computing device 101 obtains the first evaluation value of the computing device according to the completion rate of the language model. Among them, the first evaluation value has a positive correlation with the completion rate, and the higher the completion rate, the higher the first evaluation value. For example, the first evaluation value = the completion rate.
[0072] Example 2: The computing device 101 determines the adoption rate of the language model according to the total number of adopted codes and the total number of generated codes. Among them, the adoption rate has a positive correlation with the number of adopted codes and a negative correlation with the total number of generated codes. For example, the adoption rate = the total number of adopted codes / the total number of generated codes. For example, if the total number of adopted codes is 10 lines and the total number of generated codes is 50 lines, then the adoption rate is 0.2. It can be understood that the adoption rate can determine the quality of the codes generated by the language model, and the higher the adoption rate, the stronger the code generation ability of the language model.
[0073] Next, the computing device 101 determines the first evaluation value according to the adoption rate. Among them, the first evaluation value has a positive correlation with the adoption rate, and the higher the adoption rate, the higher the first evaluation value. For example, the first evaluation value = the completion rate.
[0074] In the above Example 1 and Example 2, the computing device 101 determines the first evaluation value by using the contribution degree or the quality of the generated codes. Compared with obtaining the first evaluation value only from one dimension, the obtained first evaluation value can more accurately reflect the code generation ability of the language model.
[0075] Example 3: The computing device 101 uses the number of newly added codes, the total number of generated codes, and the total number of adopted codes within a preset time interval to obtain the first evaluation value.
[0076] Specifically, the computing device 101 first obtains the completion rate in the manner of Example 1, and then obtains the adoption rate in the manner of Example 2. The computing device obtains the first evaluation value based on the completion rate and the adoption rate. Specifically, the computing device 101 calculates the harmonic mean of the completion rate and the adoption rate, and takes the harmonic mean of the completion rate and the adoption rate as the first evaluation value. Exemplarily, if the completion rate is l1 and the adoption rate is l2, then the first average value Thus, when the completion rate and the adoption rate are close, the first average value is the average of the two. When the completion rate and the adoption rate differ greatly, the first average value is the smaller of the two. The harmonic mean can comprehensively reflect the overall performance of the completion rate and the adoption rate, avoid the excessive influence of extreme values of a single indicator on the evaluation value, and further improve the accuracy of the obtained first evaluation value.
[0077] The accuracy of the first evaluation value described in the embodiments of the present application refers to the accuracy of reflecting the ability of the language model to generate code. The higher the accuracy, the more it can reflect the ability of the language model to generate code.
[0078] The embodiments of the present application may also obtain the first evaluation value in other ways, and the embodiments of the present application do not specifically limit it.
[0079] In summary, the embodiments of the present application use the number of newly added codes, the total number of generated codes, or the total number of adopted codes of the private domain project within a preset time interval to calculate the first evaluation value and evaluate the code generation ability of the large model for the private domain project. Since the number of newly added codes can directly reflect the actual increased code volume in the private domain project, the total number of generated codes directly reflects the total amount of codes generated by the large model for the private domain project, and the total number of adopted codes directly reflects the number of generated codes that are adopted, that is, the number of newly added codes, the total number of generated codes, and the total number of adopted codes of the private domain project can directly reflect the actual development situation of the large model in the private domain project. Therefore, the first evaluation value calculated using the number of newly added codes, the total number of generated codes, or the total number of adopted codes of the private domain project can accurately reflect the code generation ability of the large model in the private domain project.
[0080] Embodiment 2
[0081] In the embodiments of the present application, after the computing device 101 receives the development information sent by the code generation plugin, it can construct a prompt word based on the development information and the context information obtained from the monitored code file, and input the prompt word into the language model to accurately generate the code of the private domain project, thereby accurately obtaining the number of newly added codes, the total number of generated codes, and the total number of adopted codes within a preset time interval. Thus, the accuracy of the first evaluation value obtained based on the above-mentioned number of newly added codes, the total number of generated codes, and the total number of adopted codes is higher.
[0082] Appendix Figure 3Another evaluation method interaction diagram provided by the embodiments of this application. The code file in this method is specifically a private domain project code file, that is, the private domain project code file stores the code corresponding to a certain private domain project. This method includes S311 to S316, and S210 to S220. Among them, S311 to S316 are the private domain project code generation process, and S210 to S220 are the language model evaluation process. This method includes the following content:
[0083] S311. After the developer device 101 receives the developer input information, it sends the input information to the computing device 101.
[0084] Among them, the input information is the development requirement information of the private domain project.
[0085] S312. The computing device 101 obtains the context information of the private domain project code according to the private domain project code file.
[0086] The context information refers to the information other than the private domain project code itself in the code file, including but not limited to file path, code language, code generation mode, dependency relationship of the code file, etc. The context information can help the language model better understand the structure and dependency relationship of the code. Therefore, adopting the context information of the private domain project code can enable the language model to obtain more accurate code.
[0087] In the embodiments of this application, the computing device 101 can monitor the private domain project code file corresponding to the code generation plugin, and when receiving the input information, directly obtain the context information from the currently monitored private domain project code file.
[0088] S313. The computing device 101 constructs a prompt word according to the context information of the private domain project code and the input information.
[0089] The prompt word is used to guide or stimulate the input text for the language model to perform a specific task, so as to help the language model understand the type of task the user wants to execute or the required output format. In the embodiments of this application, the computing device 101 constructs a prompt word according to the context information of the private domain project code and the input information.
[0090] Exemplarily, if the context information of the private domain project code includes: file path, code language, code generation mode, dependency relationship, etc. The input information is a functional requirement description, then the prompt word constructed by the computing device 101 includes: file path, code language, code generation mode, dependency relationship, and functional requirement description.
[0091] It can be understood that using the context information of the private domain project code and the input information to construct the prompt word helps the language model better understand the background and purpose of the code, and thus helps improve the accuracy of the code generated by the language model.
[0092] Furthermore, the computing device 101 can split the monitored private project code file by syntax to obtain multiple groups of code clusters. Among them, the code in each group of code clusters uses the same syntax. For example, 3 groups of code clusters are obtained, namely code cluster A, code cluster B, and code cluster C. The code in code cluster A are all class codes, the code in code cluster B are all function codes, and the code in code cluster C are all variable codes.
[0093] The computing device 101 can construct a prompt word based on the context information, input information, and multiple groups of code clusters. At this time, the prompt word can clearly identify and distinguish code languages, such as class codes, function codes, and variable codes, etc., so it helps to further improve the generation accuracy of the language model.
[0094] In another example, the computing device 101 can also record the location information of the private project code file. The location information indicates the information included in the corresponding location of the code file. For example, the location information includes the file name and the function module name, etc. The computing device 101 uses the location information, context information, input information, and multiple groups of code clusters of the private project code file to construct a prompt word. The prompt word generated in this way has the characteristics of context awareness, clear syntax structure, location information, etc. Therefore, it helps to assist the language model in generating high-quality code. In addition, it also helps to improve development efficiency, reduce errors, etc.
[0095] The embodiments of the present application can also construct prompt words in other ways, and the embodiments of the present application do not specifically limit this.
[0096] S314. The computing device 101 inputs the prompt word into the language model for processing to obtain the generated code.
[0097] After the computing device 101 obtains the generated code, it embeds the code file into the code file. Among the code files, the computing device 101 marks the generated code as f, marks the generated code type as h, marks the code file name as ml, marks the function module described in the file as m2, and marks the content of the generated code as e.
[0098] Among them, the possible values of h include but are not limited to functions, classes, constants, and comments, etc.
[0099] S315. The computing device 101 obtains the adopted code based on the generated code.
[0100] The computing device 101 records the time when the generated code enters the code file. When the entry time reaches the preset time, the unchanged code in the generated code is recorded as the adopted code. For the convenience of subsequent processing, the computing device 101 further marks the adopted code in the code file as g.
[0101] It should be noted that in actual use, the codes in the generated code may all be unavailable. For the codes that are all unavailable, the adopted code at this time is 0, which is also called the non-existent adopted code. In this case, the operation of S316 is no longer executed.
[0102] S316. The computing device 101 calculates the similarity between the generated code and the adopted code. If the similarity is greater than or equal to the preset similarity threshold, the adopted code and the generated code are recorded in the adopted code database.
[0103] In the embodiments of the present application, the adopted code database is a system for storing and managing code generation, adoption, and related data, which is used to help developers track the code generation process, evaluate the performance of the language model, and provide data support for the iterative training of the language model. Further, for the convenience of tracking and management, the adopted code database also includes but is not limited to the following: generated code, adopted code, recording time, developer, coding time, model name, code type, code file, file ownership, and generated code content, etc.
[0104] Among them, the recording time refers to the timestamp when the code generation occurs. The developer refers to the identifier of the developer who executes the code generation request. The coding time refers to the time from the start of coding to the end. The model name refers to the name of the large model used to generate the code. The code type is also called the classification of the code, such as function, class, method, variable, constant, etc. The code file refers to the file name to which the code belongs. The file ownership refers to the project functional module to which the code file belongs. The generated code content refers to the specific process of the language model generating the code, including the request data and the return data.
[0105] Among them, the computing device 101 calculates the similarity between the generated code and the adopted code. If the similarity is greater than or equal to the preset similarity threshold, the adopted code and the generated code are recorded in the adopted code database. For example, the preset similarity threshold is 0.9. If the computing device 101 calculates that the similarity between the generated code and the adopted code is greater than 0.9, the adopted code and the generated code are recorded in the adopted code database, otherwise they are not recorded in the adopted code database.
[0106] Exemplary illustration: If the generated code is def login(username,password):if username == "admin" and password == "admin123":return "Login successful"else:return "Username or password error". The adopted code is def authenticate(user,passcode):if user == "admin" and passcode == "admin123":return "Authentication passed"else:return "Authentication failed".
[0107] Then, it can be calculated by cosine similarity. Specifically, the vocabulary of the generated code: def, login, username, password, admin, admin123, return, login successful, username or password error. The vocabulary of the adopted code: def, authenticate, user, passcode, admin, admin123, return, authentication passed, authentication failed. Generate code vectors (1 indicates the existence of the vocabulary, 0 indicates non-existence): [1, 1, 0, 1, 0, 1, 0, 1, 1, 1, 1, 0, 1, 0]. Adopt code vectors: [1, 0, 1, 0, 1, 0, 1, 1, 1, 1, 0, 1, 0, 1]. Then the similarity is 0.5.
[0108] The code generation plugin records the request records with a similarity greater than the preset threshold into the adopted code database. For example, the preset threshold is 0.9. If the similarity is greater than 0.9, the request record is recorded into the adopted code database. If the similarity is not greater than 0.9, the request record does not need to be included in the adopted code database.
[0109] Thus, by calculating the similarity between the adopted code and the generated code by the computing device 101, ensuring the training and iteration of the language model using the data in the adopted code database can improve the accuracy and applicability of the language model when generating code.
[0110] It should be noted that developers will continuously input the development information of the private domain project through the developer device 102, and the computing device 101 will continuously obtain the generated code and the adopted code.
[0111] S210. The computing device 101 obtains the number of newly added codes, the total number of generated codes, and the total number of adopted codes within a preset time interval.
[0112] It should be noted that S210 and S316 can also be executed synchronously, or S316 can be executed first and then S210, or S210 can be executed first and then S316. The embodiments of the present application do not specifically limit this.
[0113] S220. The computing device 101 uses the number of newly added codes, the total number of generated codes, and / or the total number of adopted codes within a preset time interval to obtain the first evaluation value of the language model.
[0114] It should be noted that steps S311 - S316 and steps S210 - S220 are parallel steps and do not affect each other.
[0115] In summary, after receiving the input information, the computing device 101 can construct a prompt word based on the input information and the context information obtained from the code file, input the prompt word into the language model, accurately generate the private domain project code, and thereby accurately obtain the number of newly added codes, the total number of generated codes, and the total number of adopted codes within a preset time interval. Furthermore, the first evaluation value obtained based on the above-mentioned number of newly added codes, the total number of generated codes, and the total number of adopted codes has a higher accuracy. In addition, by including the generated codes or adopted codes with a higher similarity in the adopted code database, the language model can be trained using the adopted code database, thereby improving the accuracy and applicability of the language model when generating codes.
[0116] Embodiment 3
[0117] Furthermore, the computing device 101 can record the encoding information in the encoding record database uniformly. Thus, the computing device 101 can obtain the encoding information from the encoding record database to determine the code generation ability of the language model.
[0118] Appendix Figure 4 is another interaction diagram of the evaluation method provided by the embodiments of the present application. Based on the appendix Figure 3 , S210 is further refined into S410 to S413. The method also specifically refines S220 into S414 and S415. And S416 is added after S415. The method specifically includes the following content:
[0119] S410. The computing device 101 counts the private domain project code files once every preset time interval to obtain the number of newly added codes and the number of deleted codes.
[0120] Among them, the number of deleted codes is the number of codes deleted within the preset time interval in the code file, that is, the number of codes reduced in the private domain project code file counted this time compared with the private domain project code file counted last time.
[0121] For example, the private domain code file counted last time includes 10 lines of code, namely a1 to a10 codes. The private domain code file counted this time includes 11 lines of code, namely a1 to a9, b10 to b11 codes. The newly added codes are b10 to b11, and the deleted code is a10. The number of newly added codes is 2, and the number of deleted codes is 1. It can be understood that through the number of newly added codes and the number of deleted codes, the ability of the model to generate codes to meet user needs can be evaluated.
[0122] Furthermore, the code generation plugin marks the newly added codes as a and the deleted codes as b for subsequent statistics.
[0123] S411. The computing device 101 records the number of requests corresponding to the preset time interval.
[0124] Among them, the number of requests is the number of times the language model is accessed within a preset time interval.
[0125] S412. If the computing device 101 determines that any one of the number of newly added codes, the number of deleted codes, and the number of requests is not 0, it records the coding time as the preset time interval.
[0126] If the number of newly added codes, the number of deleted codes, and the number of requests are all 0, then the computing device 101 records the coding time as 0.
[0127] For the convenience of subsequent recording, the coding time is recorded as c.
[0128] S413. The computing device 101 counts the total number of generated codes and the total number of adopted codes within each preset time interval.
[0129] S414. The computing device 101 determines the adoption rate based on the total number of generated codes and the total number of adopted codes within the same preset time interval; and determines the completion rate based on the total number of adopted codes and the number of newly added codes within the same preset time interval.
[0130] S415. The computing device 101 calculates the first evaluation value of the language model based on the adoption rate and the completion rate obtained within the same preset time interval.
[0131] S416. The computing device 101 determines whether the coding time is 0. If it is not 0, it executes S417.
[0132] If the coding time is 0, S417 is no longer executed.
[0133] S417. The computing device 101 records the coding time, the total generated codes, the total adopted codes, the number of newly added codes, the completion rate, the adoption rate, and the first evaluation value into the coding record database.
[0134] The coding record database is used to track coding activities during the development process. This database can help analyze and evaluate development efficiency, code quality, and the role of the large model in the coding process. The following are the key data included in the coding record database: recording time, developer, coding time, model name, number of newly added codes, number of deleted codes, number of requests, total generated codes, total adopted codes, completion rate, adoption rate, and the first evaluation value.
[0135] Exemplary illustration, such as Figure 5The figure shows a schematic diagram of encoding records being incorporated into an adopted code database and an encoding record database provided by an embodiment of the present application. Specifically, computing device 101 counts the encodings and incorporates the encodings into the adopted code database from the following dimensions: generated code, adopted code, recording time, developer, encoding time, model name, code type, code file, file ownership, and generated code content. Computing device 101 counts the encodings and incorporates the encodings into the encoding record database from the following dimensions: recording time, developer, encoding time, model name, number of newly added codes, number of deleted codes, number of requests, total number of generated codes, total number of adopted codes, completion rate, adoption rate, and first evaluation value.
[0136] It should be noted that, during the specific incorporation process, every time computing device 101 generates a code, and after the similarity between the generated code and the corresponding adopted code meets or exceeds a preset similarity threshold, it is incorporated into the adopted code database. And computing device 101 incorporates the encodings into the encoding record database at every preset time interval.
[0137] It should be noted that the encoding record database and / or the adopted code database can be visually displayed. Thus, developers can evaluate the language model based on the data in the encoding record database, or select appropriate data such as generated codes and adopted codes from the adopted code database for fine-tuning and iterative training of the language model, so as to improve the accuracy and applicability of the model when generating codes.
[0138] In one example, computing device 101 can obtain the first evaluation value from the encoding record database, and by comparing the first evaluation values of different language models at different preset time intervals, select a target language model that meets the requirements of the object to be developed. For example, if the encoding record database includes that the first evaluation value of model A at time interval 1 is 0.85 and at time interval 2 is 0.88, and the first evaluation value of model B at time interval 1 is 0.75 and at time interval 2 is 0.78. Since the first evaluation value of model A at time interval 1 and the first average value at time interval 2 are both better than those of model B, model A is the target language model.
[0139] In another example, computing device 101 can obtain the number of requests and the number of newly added codes at the same preset time interval from the encoding record database, and determine the second evaluation value of the language model, where the second evaluation value can evaluate the encoding efficiency of the language model. Computing device 101 obtains the comprehensive evaluation value of the language model based on the second evaluation value and the first evaluation value at the same moment. For example, the harmonic mean of the first evaluation value and the second evaluation value can be the comprehensive evaluation value. Thus, by evaluating the language model through the comprehensive evaluation value, the evaluation result is more accurate.
[0140] Furthermore, it should be noted that in actual use, the evaluation method provided in the embodiments of the present application can be applied not only to the private domain project code provided in the embodiments of the present application, but also to the entire process of software development such as testing and design.
[0141] In addition, the embodiments of the present application provide an evaluation device, which is applied to a computing device. The computing device can be a server or other electronic devices with computing functions, and the embodiments of the present application do not specifically limit it. When the computing device is a server, the computing device can be a blade server, a rack server, or a server with other structures, and the embodiments of the present application do not specifically limit it.
[0142] Appendix Figure 6 FIG. 6 is a schematic structural diagram of an evaluation device provided in an embodiment of the present application. The device 600 includes:
[0143] A first acquisition unit 601, configured to acquire the number of newly added codes, the total number of generated codes, and the total number of adopted codes for the computing device.
[0144] Among them, the number of newly added codes is the codes that are newly added within a preset time interval and meet the requirements of the object to be developed. The total number of generated codes is the total number of codes generated after processing the development information of the object to be developed n times based on the language model within a preset time interval. The total number of adopted codes is the number of codes adopted among the total number of generated codes, and n is an integer greater than or equal to 1;
[0145] A second acquisition unit 602, configured to obtain a first evaluation value of the language model by using the number of newly added codes, the total number of generated codes, and / or the total number of adopted codes; the first evaluation value is used to evaluate the code generation ability of the language model.
[0146] Optionally, the second acquisition unit 602 is specifically configured to: obtain the completion rate of the language model according to the number of newly added codes and the total number of adopted codes; the completion rate of the language model is positively correlated with the number of newly added codes and negatively correlated with the total number of adopted codes; use the completion rate of the language model to obtain the first evaluation value of the language model.
[0147] Optionally, the second acquisition unit 602 is further configured to: determine the adoption rate of the language model according to the total number of adopted codes and the total number of generated codes; the adoption rate of the language model is positively correlated with the total number of adopted codes and negatively correlated with the total number of generated codes; obtain the first evaluation value of the language model according to the adoption rate of the language model.
[0148] Optionally, the second obtaining unit 602 is further configured to: obtain the completion rate of the language model according to the number of newly added codes and the total number of adopted codes; the completion rate of the language model is positively correlated with the number of newly added codes and negatively correlated with the total number of adopted codes; determine the adoption rate of the language model according to the total number of adopted codes and the total number of generated codes; the adoption rate of the language model is positively correlated with the total number of adopted codes and negatively correlated with the total number of generated codes; obtain the first evaluation value of the language model according to the adoption rate of the language model and the completion rate of the language model. That is, the computing device can use the adoption rate and the completion rate to comprehensively evaluate the code generation ability of the language model, and further improve the evaluation accuracy.
[0149] Optionally, the first obtaining unit 601 is specifically configured to: the computing device can obtain a first code file at the start time of a preset time interval and a second code file at the end time of the preset time interval; the first code file is used to store codes that meet the requirements of the object to be developed before the start time, and the second code file is used to store codes that meet the requirements of the object to be developed before the end time; obtain the number of newly added codes according to the first code file and the second code file. Since the code file is the most direct and specific result in the software development process, the computing device can accurately obtain the number of newly added codes by comparing the code files at different times.
[0150] In another implementation, the total number of generated codes includes n times of generated codes, and each time of generated codes is obtained through the following method: obtain the context information in the code file, and the code file is used to store codes that meet the requirements of the object to be developed at the current time; construct a prompt word based on the context information and the development information of the object to be developed; input the prompt word into the language model to obtain the generated code.
[0151] Furthermore, the computing device can also segment the code file by syntax to obtain multiple groups of code clusters; each group of code clusters in the multiple groups of code clusters has the same syntax; construct a prompt word based on the context information and the development information of the object to be developed, in combination with the multiple groups of code clusters and / or the location information of the code file; the location information of the code file indicates the code information included in the corresponding location of the code file
[0152] Optionally, the first obtaining unit 601 is further configured to: obtain the number of deleted codes and the number of requests; the number of deleted codes is the number of codes deleted in the code file within a preset time interval; the number of requests is the number of times of accessing the language model within a preset time interval; if any one of the number of requests, the number of deleted codes, and the number of newly added codes is not 0, determine the coding time as the preset time interval; determine the second evaluation value of the language model according to the coding time and the number of newly added codes; the second evaluation value of the language model is used to evaluate the coding efficiency of the language model; obtain the comprehensive evaluation value of the language model according to the second evaluation value of the language model and the first evaluation value of the language model, and the comprehensive evaluation value is used to evaluate the coding ability of the language model.
[0153] Optionally, the apparatus further includes a writing unit configured to write the encoded record data of the language model into an encoded record database, so as to evaluate the encoding ability of the language model by using the encoded record data in the encoded record database; wherein, the encoded record data includes encoding time, newly added codes corresponding to the number of newly added codes, generated codes corresponding to the total number of generated codes, adopted codes corresponding to the total number of adopted codes, a first evaluation value, the number of requests, adoption rate, and / or completion rate.
[0154] Furthermore, an embodiment of the present application also provides a server.
[0155] As Figure 7 shown, an embodiment of the present application also provides a server 700. The application scenario of the server 700 is not specifically limited. For example, the server 700 is introduced by taking a server as an example, and the type of the server is not specifically limited either. For example, the server can be a rack server or an edge server. The server can be located in a data center or in other areas, which is not specifically limited in the embodiments of the present application.
[0156] The server 700 includes a processor 701 and a memory 703. The memory 703 is electrically connected to the processor 701 respectively; the memory 703 is used to store program instructions for the evaluation methods involved in the above embodiments; the processor 701 is used to call the corresponding parts of the above program instructions, so that the server can execute the evaluation methods involved in the above embodiments.
[0157] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using a software program, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the processes or functions in the embodiments of the present application are generated in whole or in part.
[0158] An embodiment of the present application also provides a computer program product containing instructions. The computer program product can be software or a program product containing instructions that can run on a computing device or be stored in any available medium. When the computer program product runs on a computing device, the computing device is caused to execute the above evaluation method.
[0159] The embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center that includes one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive), etc. The computer-readable storage medium includes instructions that direct the computing device to execute the above evaluation method.
[0160] The descriptions of the processes or structures corresponding to the above respective drawings each have their own focuses. For parts not detailed in a certain process or structure, reference can be made to the relevant descriptions of other processes or structures.
[0161] As described above, the above are only the specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. An evaluation method, characterized in that, The method includes: Obtaining the number of newly added codes, the total number of generated codes, and the total number of adopted codes within a preset time interval; The number of newly added codes is the number of codes newly obtained and meeting the requirements of the object to be developed within the preset time interval. The total number of generated codes is the total number of codes generated after processing the development information of the object to be developed n times based on a language model within the preset time interval, where n is an integer greater than or equal to 0. The total number of adopted codes is the number of codes adopted among the total generated codes; Using the number of newly added codes, the total number of generated codes, and / or the total number of adopted codes to obtain a first evaluation value of the language model; the first evaluation value is used to evaluate the code generation ability of the language model.
2. The method according to claim 1, characterized in that, The step of using the number of newly added codes, the total number of generated codes, and / or the total number of adopted codes to obtain the first evaluation value of the language model includes: Obtaining the completion rate of the language model according to the number of newly added codes and the total number of adopted codes; the completion rate of the language model is positively correlated with the number of newly added codes and negatively correlated with the total number of adopted codes; Using the completion rate of the language model to obtain the first evaluation value of the language model.
3. The method according to claim 1, wherein The step of using the number of newly added codes, the total number of generated codes, and / or the total number of adopted codes to obtain the first evaluation value of the language model includes: Determining the adoption rate of the language model according to the total number of adopted codes and the total number of generated codes; the adoption rate of the language model is positively correlated with the total number of adopted codes and negatively correlated with the total number of generated codes; Obtaining the first evaluation value of the language model according to the adoption rate of the language model.
4. The method according to claim 1, wherein The step of using the number of newly added codes, the total number of generated codes, and / or the total number of adopted codes to obtain the first evaluation value of the language model includes: Obtaining the completion rate of the language model according to the number of newly added codes and the total number of adopted codes; the completion rate of the language model is positively correlated with the number of newly added codes and negatively correlated with the total number of adopted codes; Determining the adoption rate of the language model according to the total number of adopted codes and the total number of generated codes; the adoption rate of the language model is positively correlated with the total number of adopted codes and negatively correlated with the total number of generated codes; Obtaining the first evaluation value of the language model according to the adoption rate of the language model and the completion rate of the language model.
5. The method according to claim 1, wherein Obtaining the number of newly added codes includes: Obtaining a first code file at the start time of the preset time interval and a second code file at the end time of the preset time interval; the first code file is used to store the codes that meet the requirements of the object to be developed before the start time, and the second code file is used to store the codes that meet the requirements of the object to be developed before the end time; Obtaining the number of newly added codes according to the first code file and the second code file.
6. The method according to claim 1, characterized in that, The total number of generated codes includes n times of code generation, and each time of code generation is obtained through the following method: Obtain context information in the code file, where the code file is used to store the code that meets the requirements of the object to be developed at the current moment; Construct a prompt based on the context information and the development information of the object to be developed; Input the prompt into the language model to obtain generated code.
7. The method according to claim 6, characterized in that, The method further includes: Segment the code file by syntax to obtain multiple groups of code clusters; the code in each group of code clusters in the multiple groups of code clusters has the same syntax; The constructing a prompt based on the context information and the development information of the object to be developed includes: Based on the context information and the development information of the object to be developed, combine the multiple groups of code clusters and / or the location information of the code file to construct the prompt; the location information of the code file indicates the code information included in the corresponding location of the code file.
8. The method according to claim 1, wherein The method further includes: Obtain the number of deleted codes and the number of requests; the number of deleted codes is the number of codes deleted in the code file within the preset time interval; the number of requests is the number of times the language model is accessed within the preset time interval; If any one of the number of requests, the number of deleted codes, and the number of new codes is not 0, determine the coding time as the preset time interval; Determine a second evaluation value of the language model according to the coding time and the number of new codes; the second evaluation value of the language model is used to evaluate the coding efficiency of the language model; Obtain a comprehensive evaluation value of the language model according to the second evaluation value and the first evaluation value of the language model, where the comprehensive evaluation value is used to evaluate the coding ability of the language model.
9. The method according to any one of claims 1-8, characterized in that, The method further includes: Write the coding record data of the language model into the coding record database to evaluate the coding ability of the language model by using the coding record data in the coding record database; Wherein, the coding record data includes coding time, new codes corresponding to the number of new codes, generated codes corresponding to the total number of generated codes, adopted codes corresponding to the total number of adopted codes, the first evaluation value, the number of requests, the adoption rate, and / or the completion rate.
10. A computing device, characterized in that, Includes a memory and a processor; The memory is used to store programs; The processor is used to execute the program stored in the memory. When the program stored in the memory is executed, the processor is used to execute the evaluation method according to any one of claims 1-9.