Method, device, equipment, medium and product for performance evaluation of language model
Through the performance evaluation of multi-source truth data, multiple predictive answers and truth answers of the language model are obtained, and the one-sidedness and unreliability of evaluation results caused by single source truth data is solved, and more accurate and effective model performance evaluation and optimization guidance are achieved.
Patent Information
- Application Number
- CN202411918595.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2044-12-24
AI Technical Summary
When evaluating the performance of language models, the prior art relies on single-source truth-value data, which leads to the one-sidedness and unreliability of the evaluation results, making it difficult to provide effective guidance for further optimization and improvement of the model.
By obtaining multiple predicted answers to the question and multiple true answers, with different answer forms and sources, the performance evaluation of multi-source true data is carried out to ensure the quality and coverage of true data.
The coverage and stability of the model during performance evaluation are improved, more accurate and effective guidance is provided, and a strong guarantee for the further optimization and improvement of the language model.
Smart Images

Figure CN119358686B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure generally relate to the field of computers, and in particular, relate to methods and apparatuses, devices, media, and products for performance evaluation of language models. Background Art
[0002] In recent years, language models have developed rapidly and have become an influential and highly-watched technology. Large Language Model (LLM), as a representative of language models, is widely used for its powerful natural language understanding and generalization capabilities. One of the core functions of language models is to significantly simplify the previously complex and cumbersome knowledge acquisition process. Users can overcome the high knowledge barriers of the past and achieve functions equivalent to those difficult-to-use professional tools, thereby reducing the entry threshold and learning costs of knowledge.
[0003] The evaluation of language model capabilities and performance undoubtedly plays an important role. This is directly related to the performance of the model in practical applications and is a key means of measuring whether it can meet processing requirements and achieve expected results. Through a scientific and comprehensive evaluation process, various problems that may cause degradation in the language model can be identified in a timely and accurate manner, and based on this, optimization strategies can be formulated and implemented in a targeted manner, thereby achieving the improvement and enhancement of model accuracy and robustness, so that the language model can meet the continuously growing processing needs and ensure that the language model can run stably and efficiently in various application scenarios. Summary of the invention
[0004] An embodiment of the present disclosure provides a solution for performance evaluation of a language model.
[0005] In a first aspect of the present disclosure, a method for performance evaluation of a language model is provided, the method comprising obtaining, for each question in a question set, multiple predicted answers to the question through a language model, the multiple predicted answers each having a different answer form, the answer forms including a first type of structured query language (SQL) statement, a visual uniform resource locator (URL), a second type of SQL statement, an SQL execution result, a recall field, and a URL execution result. The method also includes obtaining multiple true value answers to the question, the multiple true value answers each corresponding to a different answer source, the answer sources including a first SQL source, a visual URL source, a second SQL source, an SQL execution result source, a recall field source, and a URL execution result source. The method also includes obtaining multiple comparison results based on a comparison between the multiple predicted answers and the multiple true value answers, the multiple comparison results indicating the differences between the multiple predicted answers and the multiple true value answers. The method also includes determining multiple comparison scores corresponding to the multiple comparison results, and determining a performance score of the language model based on the multiple comparison scores, the performance score indicating the question-answering performance of the language model.
[0006] In a second aspect of the present disclosure, a device for performance evaluation of a language model is provided, the device comprising a predicted answer acquisition module, configured to acquire multiple predicted answers to each question in a question set through a language model, the multiple predicted answers each having a different answer form, the answer forms including a first type of structured query language SQL statement, a visual uniform resource locator URL, a second type of SQL statement, an SQL execution result, a recall field, and a URL execution result. The device also comprises a true value answer acquisition module, configured to acquire multiple true value answers for the question, the multiple true value answers each corresponding to a different answer source, the answer sources including a first SQL source, a visual URL source, a second SQL source, an SQL execution result source, a recall field source, and a URL execution result source. The device also comprises a comparison module, configured to acquire multiple comparison results based on a comparison between multiple predicted answers and multiple true value answers, the multiple comparison results indicating the differences between multiple predicted answers and multiple true value answers. The device also comprises a scoring module, configured to determine multiple comparison scores corresponding to the multiple comparison results, and determine a performance score of the language model based on the multiple comparison scores, the performance score indicating the question-answering performance of the language model.
[0007] According to a third aspect of the present disclosure, an electronic device is provided. The computing device includes a processor and a memory, wherein instructions are stored in the memory, and when the instructions are executed by the processor, the processor executes a method or process according to an embodiment of the present disclosure.
[0008] According to a fourth aspect of the present disclosure, a machine-readable storage medium is provided, wherein machine-executable instructions are stored on the machine-readable storage medium, and when the machine-executable instructions are executed by a processor, the processor is enabled to execute a method or process according to an embodiment of the present disclosure.
[0009] In a fifth aspect of the present disclosure, a computer program product is provided. The computer program product is tangibly stored on a non-transitory computer-readable storage medium and includes a computer program, which, when executed by a processor of a computer, causes the processor to perform a method or process according to an embodiment of the present disclosure.
[0010] Please note that the summary is provided to introduce a series of concepts in a simplified form, which will be further described in the detailed description below. The summary is not intended to identify key features or essential features of the disclosure, nor is it intended to limit the scope of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The above and other purposes, features and advantages of the present disclosure will become more clearly understood by describing the embodiments of the present disclosure in more detail in conjunction with the accompanying drawings, in which:
[0012] Figure 1 is a diagram schematically illustrating an example environment in which methods and / or processes according to embodiments of the present disclosure may be implemented;
[0013] Figure 2 is a flowchart schematically illustrating a method for performance evaluation of a language model according to an embodiment of the present disclosure;
[0014] FIG3 schematically illustrates an example of a workflow for performance evaluation of a language model according to an embodiment of the present disclosure;
[0015] Figure 4 is a diagram schematically illustrating an example process of forward data enhancement according to an embodiment of the present disclosure;
[0016] Figure 5 is a diagram schematically illustrating an example process of reverse data enhancement according to an embodiment of the present disclosure;
[0017] Figure 6 is a diagram schematically illustrating an apparatus for performance evaluation of a language model according to an embodiment of the present disclosure;
[0018] Figure 7 is a schematic block diagram of an example device that may be used to implement embodiments according to the present disclosure.
[0019] Throughout the drawings, same or similar reference numbers generally refer to same or similar elements. DETAILED DESCRIPTION
[0020] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.
[0021] In the description of the embodiments of the present disclosure, the term "including" and its variations should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc. may refer to different or the same objects, unless explicitly indicated to be different.
[0022] It is understandable that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, scope of use, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0023] For example, in response to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested to be performed will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, application, server, or storage medium that performs the operation of the technical solution of the present disclosure according to the prompt message.
[0024] As an optional but non-limiting implementation, in response to receiving an active request from the user, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0025] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that meet the relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0026] As mentioned above, it is necessary to conduct performance evaluation on language models. In the performance evaluation of language models such as natural language understanding and text-to-structured query language conversion (Text2SQL), accuracy and robustness are crucial. In order to comprehensively and accurately verify the accuracy and robustness of language models in various application scenarios, it is expected to use high-quality annotated datasets to evaluate model performance.
[0027] The performance of language models such as LLM often shows a high degree of scenario dependence, which means that the model may have performance advantages in a certain application scenario, but its ability deteriorates in other application scenarios (for example, answer bias, irrelevant or wrong output generation, etc.), which is obviously contrary to expectations. In order to fully verify the capabilities of language models in all aspects, it is necessary to ensure the coverage of performance evaluation.
[0028] Ground Truth (GT), also known as benchmark truth, is the true truth or a relative standard that is considered close to such a truth. At present, some evaluation schemes usually rely on the truth (GroundTruth) from a single source. This truth scheme based on a single source has some significant limitations. First, data from a single source may be biased and cannot fully reflect the performance of the language model in different scenarios in the performance evaluation of the language model. Secondly, the annotation quality of the truth data from a single source may be inconsistent, resulting in reduced reliability of the evaluation results. Finally, the truth scheme based on a single source is difficult to capture complex natural language phenomena, which limits the comprehensive evaluation of the model's capabilities.
[0029] In the performance evaluation of language models, truth solutions based on a single source or type can usually only cover specific scenarios or conditions. Answers from a single source may be biased or deviated, resulting in evaluation results that are not universal. In addition, the quality of truth data from a single source may not be high, and there may be problems such as noise and incorrect labeling, which will directly affect the accuracy of the evaluation results. It should also be noted that if data from a single source is relied upon, if there is a problem with the data source (for example, data loss, incorrect labeling, etc.), the entire evaluation process will be seriously affected.
[0030] The evaluation results of a single type of GT data may not accurately reflect the generalization ability of the language model, that is, the performance of the model in unseen data or scenarios. Due to the one-sidedness and unreliability of the evaluation results, it is difficult to provide effective guidance for the further optimization and improvement of the model. In addition, in the actual scenario of annotation, there are situations where different users and teams may be accustomed to using different tools and methods to annotate data. For example, some annotators may be accustomed to using structured query language (SQL) for data query and annotation, while other annotators may prefer to use business intelligence (BI) tools for data analysis and report generation, and so on. The annotated true value data may come from multiple sources and types.
[0031] At least in order to solve at least some of the above and other potential problems, an embodiment of the present disclosure proposes a solution for performance evaluation of a language model. A solution for performance evaluation of a language model is provided, which includes obtaining multiple predicted answers to the question through a language model for each question in a question set, and each of the multiple predicted answers has a different answer form, and these answer forms include a first type of structured query language SQL statement, a visual uniform resource locator URL, a second type of SQL statement, an SQL execution result, a recall field, and a URL execution result. The solution also includes obtaining multiple true value answers for the question, and each of the multiple true value answers corresponds to a different answer source, and these answer sources include a first SQL source, a visual URL source, a second SQL source, an SQL execution result source, a recall field source, and a URL execution result source. The solution also includes obtaining multiple comparison results based on the comparison between the multiple predicted answers and the multiple true value answers, and the multiple comparison results indicate the differences between the multiple predicted answers and the multiple true value answers. The solution also includes determining multiple comparison scores corresponding to the multiple comparison results, and determining the performance score of the language model based on the multiple comparison scores, and the performance score indicates the question-answering performance of the language model.
[0032] According to the performance evaluation scheme for language models of the embodiments of the present disclosure, consideration of multi-source true value data is introduced in the performance evaluation of language models. The performance evaluation strategy based on multi-source true value data can ensure the quality of true value data and improve the coverage and stability of the model during performance evaluation, thereby providing strong guarantees for the generalization and understanding capabilities of the model. The performance evaluation results obtained in this way can provide accurate and effective guidance for further optimization and improvement of the language model, and help improve and enhance the performance of the model.
[0033] Reference below Figures 1 to 7It should be understood that these exemplary embodiments are provided only to enable those skilled in the art to better understand and implement the embodiments of the present disclosure, and are not intended to limit the scope of the present disclosure in any way.
[0034] Figure 1 1 is a schematic diagram schematically illustrating an example environment 100 in which methods and / or processes according to embodiments of the present disclosure may be implemented. Figure 1 As shown in , the example environment 100 includes a knowledge base 110, a computing system 120 including computing nodes (eg, a locally arranged computing node 121 and a computing node 122 arranged in the cloud), and a performance evaluation result 130. Figure 1 The implementation environment of the embodiment of the present disclosure is described by taking a hybrid computing system as an example. It should be understood that this is only an exemplary and non-restrictive description, and other different types of computing systems are also feasible, such as distributed or centralized computing systems. The appropriate computing environment configuration can be selected according to actual use needs.
[0035] exist Figure 1 In the figure, only limited components and exemplary connection relationships are shown. It should be understood that this is for the purpose of ease of explanation and ease of illustration and is not intended to limit the scope of the present disclosure, and other different components may also exist. For example, a display component and an input component, etc. In an exemplary and non-limiting manner, the performance evaluation result 130 of the language model output by the computing system 120 and the scores of each sub-item can be displayed on the display component to indicate the evaluation result for the language model, and the labeled question-answer pairs can be entered into the knowledge base through the input component.
[0036] According to an embodiment of the present disclosure, the knowledge base 110 can store a question set. The question set includes multiple questions, and these questions cannot have problems such as unclear, subjective, ambiguous, unanswerable or wrongly labeled, because such degradation problems will cause distortion of the evaluation results. For example, if the language model gives a "wrong" answer to an extremely ambiguous question, then what should be questioned is the question itself, not the reasoning ability of the model.
[0037] In addition, the knowledge base 110 may store true value answers. According to an embodiment of the present disclosure, true value answers from multiple sources may be stored in the knowledge base 110. The knowledge base 110 may include multiple sub-knowledge bases, each of which corresponds to an answer source and is configured to store true value answers from the answer source. The computing system 120 may obtain the corresponding true value answers from the knowledge base 110 for performing performance evaluation for the language model according to an embodiment of the present disclosure.
[0038] As described above, at least one computing node included in the computing system 120 can perform processing and operations corresponding to the multi-source performance evaluation according to the embodiment of the present disclosure. In an exemplary and non-restrictive manner, the computing node 121 can be arranged locally at the user, and the computing node 122 can be arranged in the cloud. The cloud can indicate a service model built based on cloud or distributed technology. In this mode, computing resources, storage resources, etc. are coupled together through a network to form a schedulable and scalable resource collection. At least a portion of these resources can be dynamically accessed or allocated to complete tasks without the need to locally own or manage these physical resources.
[0039] A computing node refers to a computing resource or computing power, and may be a device with computing power. For example, a computing node may be provided with a processor and a memory, etc., or may be equipped with a dedicated accelerator (such as a graphics processing unit (GPU)). In addition, a computing node may store and maintain data. In addition, the model (e.g., a large language model) involved in the performance evaluation scheme for a language model according to an embodiment of the present disclosure may be a model deployed on at least one of these computing nodes, and of course an external model may also be called.
[0040] like Figure 1 As shown in the example, the local computing node 121 and the computing node 122 in the cloud can be interconnected via a network for communication. Figure 1 Taking the architecture in the example, computing nodes 121 and 122 can communicate via a network to achieve, for example, data synchronization or sharing between nodes. Multiple computing nodes in the example environment 100 can process computing tasks in parallel. In addition, multiple computing nodes in the example environment 100 can have redundancy and fault tolerance mechanisms to ensure reliable execution of computing tasks. When a computing node fails, the system can automatically migrate the job to other nodes that are working normally to ensure the continuity and availability of the task.
[0041] Examples of computing nodes may include supercomputers, personal computers, laptop computers, vehicle-mounted computing devices, mobile devices (such as smart phones, tablet computers, etc.), wearable electronic devices, multimedia devices, personal digital assistants (PDAs), smart home devices, consumer electronic products, or a combination of any one or more of the above devices, etc. It should be understood that the computing nodes described herein are merely exemplary and non-restrictive, and for example, other different types of computing nodes may also be used.
[0042] Combined with the above Figure 1 Describes an example environment in which the methods and / or processes according to embodiments of the present disclosure may be implemented. Figure 2A method 200 for performance evaluation of a language model according to an embodiment of the present disclosure is described. Through this method 200, the coverage of performance evaluation can be improved and model tuning can be facilitated.
[0043] Figure 2 2 is a flowchart schematically illustrating a method 200 for performance evaluation of a language model according to an embodiment of the present disclosure. At 210, for each question in a question set, a plurality of predicted answers to the question are obtained through a language model, and each of the plurality of predicted answers has a different answer form, and the answer forms include a first type of structured query language (SQL) statement, a visual uniform resource locator (URL), a second type of SQL statement, an SQL execution result, a recall field, and a URL execution result.
[0044] According to an embodiment of the present disclosure, a language model to be evaluated can be input to obtain a prediction output corresponding to the question. For the same question, multiple predicted answers each having a different answer form can be obtained through the language model, that is, a question can correspond to multiple predicted answers, and the answer forms of the multiple predicted answers are different. The predicted answer acquisition process will be iteratively executed for each question in the question set until all the predicted answers corresponding to each question are acquired. In an application scenario such as Text2SQL, the answer form of the answer corresponding to the question in text form includes a first type of structured query language SQL statement (e.g., ANSI-SQL), a second type of SQL statement (e.g., CH-SQL), and SQL execution results. These answer forms also include visualization URLs associated with BI tools, recall fields, and URL execution results.
[0045] At 220, multiple truth answers are obtained for the question, each of which corresponds to a different answer source, including a first SQL source, a visualization URL source, a second SQL source, a SQL execution result source, a recall field source, and a URL execution result source. According to an embodiment of the present disclosure, in response to a corresponding question being input into a language model to be evaluated, multiple truth answers corresponding to the question are obtained, each of which corresponds to a different answer source. In other words, one question may correspond to multiple truth answers, each of which comes from a corresponding answer source. For example, a first truth answer whose answer form is a first type of SQL statement comes from a first SQL source. The truth answer acquisition process will be iteratively performed for each question in the question set until all the truth answers corresponding to each question are acquired.
[0046] By way of example, the truth answers annotated with SQL may come from a first SQL source, a second SQL source, and a SQL execution result source. The first SQL source may be configured to provide annotated SQL statements of a first type, the second SQL source may be configured to provide annotated SQL statements of a second type, and the SQL execution result source may be configured to provide annotated SQL execution results. In addition, the truth answers annotated with a BI tool may come from a visualization URL source, a recall field source, and a URL execution result source. The visualization URL source may be configured to provide annotated visualization URL sources, the recall field source may be configured to provide annotated recall fields, and the URL execution result source may be configured to provide annotated URL execution results.
[0047] At 230, based on the comparison between the multiple predicted answers and the multiple true value answers, multiple comparison results are obtained, and the multiple comparison results indicate the differences between the multiple predicted answers and the multiple true value answers. For each question in the question set, multiple predicted answers in multiple answer forms are obtained, and multiple true value answers from different answer sources are obtained, and these true value answers can correspond to the obtained predicted answers respectively. According to an embodiment of the present disclosure, by comparing the obtained multiple predicted answers with the corresponding multiple true value answers, multiple comparison results are obtained to know the differences between these predicted answers and the corresponding true value answers.
[0048] At 240, multiple comparison scores corresponding to the multiple comparison results are determined, and a performance score of the language model is determined based on the multiple comparison scores, the performance score indicating the question-answering performance of the language model. After obtaining multiple comparison results between multiple predicted answers and corresponding multiple true value answers, multiple such multiple comparison results can be converted into multiple comparison scores, which can indicate the model performance of the language model in various sub-items, such as the model's capabilities in terms of corresponding answer forms, answer sources or types, etc. Such multiple comparison scores will be further synthesized into a performance score of the language model. In the following, the above-mentioned various processes according to the embodiments of the present disclosure will be further described in detail.
[0049] According to the method 200 for performance evaluation of a language model of an embodiment of the present disclosure, consideration of multi-source true value data is introduced in the performance evaluation of a language model. The performance evaluation strategy based on multi-source true value data can ensure the quality of the true value data, and can improve the coverage and stability of the model during performance evaluation, thereby providing a strong guarantee for the generalization and understanding capabilities of the model. The performance evaluation results obtained in this way can provide accurate and effective guidance for further optimization and improvement of the language model, and help improve and enhance the performance of the model.
[0050] Figure 3Aschematically illustrates the correspondence between the answer types or forms based on SQL according to an embodiment of the present disclosure, and Figure 3B The corresponding relationship between the answer forms based on the BI tool according to the embodiment of the present disclosure is schematically illustrated. It should be understood that the answer forms described here are exemplary, and the embodiments of the present disclosure are not limited to these answer forms, and may also include other different answer forms.
[0051] According to an embodiment of the present disclosure, multiple predicted answers to a question may include a first predicted answer in the form of a first type of SQL statement, and the first type of SQL statement may be an ANSI-SQL statement. SQL is a standard language for managing and operating relational databases. By writing SQL query statements, data that meets specific conditions can be extracted from the database as standard answers (i.e., true value answers). According to the questions in the question set used for evaluation, the answers can be annotated with SQL statements, and the language model can also output the corresponding predicted answers in the form of SQL statements based on the input question.
[0052] According to an embodiment of the present disclosure, a third predicted answer and a fourth predicted answer for a question are determined based on a first predicted answer whose answer form is a first type of SQL statement (e.g., ANSI-SQL statement), the answer form of the third predicted answer is a second type of SQL statement (e.g., CH-SQL statement), and the answer form of the fourth predicted answer is an SQL execution result. Figure 3A As shown, whether for the predicted answer or the true answer, taking ANSI-SQL as the starting point, CH-SQL and SQL execution results can be additionally provided. In other words, given ANSI-SQL, both CH-SQL and SQL execution results can be known. It should be understood that ANSI-SQL can be used for annotation and the true value annotation of CH-SQL and SQL can be obtained based on this, or CH-SQL and SQL execution results can be used for annotation directly. In addition, the language model can output the predicted answer of ANSI-SQL and the predicted answer of CH-SQL and SQL execution results can be obtained based on this, or the predicted answer of CH-SQL and SQL execution results can be directly output.
[0053] According to an embodiment of the present disclosure, multiple predicted answers to a question may include a second predicted answer in the form of a visualization URL. There is a class of BI tools for data analysis and visualization. According to the questions in the question set for evaluation, the visualization URL generated by such BI tools can be used as a true value annotation, that is, a standard answer is provided, and the language model can output the predicted answer of the corresponding visualization URL according to the input question.
[0054] According to an embodiment of the present disclosure, a fifth predicted answer and a sixth predicted answer for a question are determined based on the second predicted answer (e.g., a visualization URL), the answer form of the fifth predicted answer is a recall field, and the answer form of the sixth predicted answer is a URL execution result. The recall field may indicate various attributes of data recalled from a database or knowledge base. Figure 3B As shown, whether for the predicted answer or the true value answer, the recall field and the URL execution result can be additionally provided with the visualization URL as the starting point. In other words, given the visualization URL, both the recall field and the URL execution result can be known. It should be understood that the visualization URL can be used for annotation and the true value annotation of the recall field and the URL execution result can be obtained based on this, or the recall field and the URL execution result can be directly used for annotation. In addition, the language model can output the predicted answer of the visualization URL and obtain the predicted answer of the recall field and the URL execution result based on this, or the predicted answer of the recall field and the URL execution result can be directly output.
[0055] Figure 4 The exemplary process 400 of model performance evaluation based on multi-source truth values according to an embodiment of the present disclosure is schematically illustrated. In order to increase the diversity and completeness of truth value data to overcome data bias and provide effective guidance for the generalization ability of language models, the embodiment of the present disclosure adopts a performance evaluation strategy based on multi-source truth values. Figure 4 As shown, multi-source truth value 410 may include truth value annotations from answer source 1 411, answer source 2 412, answer source 3 413, ..., answer source N 41N, etc., where N is a positive integer.
[0056] The application scenario of Text2SQL is used here for illustrative but non-limiting explanation. These answer sources may include a first SQL source configured to provide annotated SQL statements of a first type (e.g., ANSI-SQL statements), a visualization URL source configured to provide annotated visualization URLs, a second SQL source configured to provide annotated SQL statements of a second type (e.g., CH-SQL statements), a SQL execution result source configured to provide annotated SQL execution results, a recall field source configured to provide annotated recall fields, and a URL execution result source configured to provide annotated URL execution results, and the like.
[0057] In some embodiments, formatting of homologous truth answers from the same source can be performed using a unified formatting strategy to obtain formatted homologous truth answers with a consistent format, and the formatted homologous truth answers can be organized in a corresponding sub-knowledge base in the knowledge base. In this way, it is convenient to access truth answers of the same source or the same type. Similar formatting can also be performed on the predicted answers output by the language model.
[0058] like Figure 4 As shown in , for a question in the question set, multiple predicted answers can be obtained via the language model to be evaluated, namely predicted answer 1 421, predicted answer 2 422, predicted answer 3 423, ..., predicted answer N 42N, etc., and these predicted answers each have a different answer form. For example, the answer form of predicted answer 1 421 is an ANSI-SQL statement, the answer form of predicted answer 2 422 is a visualization URL, and so on. In response to the question being input to the language model to be evaluated, multiple true value answers corresponding to the question can be obtained from multiple answer sources. For example, the ANSI-SQL true value answer corresponding to predicted answer 1 421 can be obtained from answer source 1 411 (i.e., the first SQL source) configured to provide ANSI-SQL true value answers, and the visualization URL true value answer corresponding to predicted answer 1 422 can be obtained from answer source 2 411 (i.e., the visualization URL source) configured to provide visualization URLs, and so on. It should be understood that this process will be iteratively executed until each question in the predetermined question and the multiple answer forms corresponding to each question are traversed.
[0059] According to an embodiment of the present disclosure, the corresponding true value answer is compared with the predicted answer to obtain a comparison result indicating the difference between them, such as Figure 4 The comparison results 1 431, 2 432, 3 433, N 43N, etc. are shown in FIG. For example, the predicted answer 1 421 can be compared with the corresponding ANSI-SQL true value answer from the answer source 1 411 to obtain the comparison result 1 431, the predicted answer 2 422 can be compared with the corresponding visualization URL true value answer from the answer source 2 412 to obtain the comparison result 2 432, and so on.
[0060] In some embodiments, by calculating the difference measure between the predicted answer of each answer form in multiple answer forms and the corresponding true value answer, the corresponding comparison result for each answer form can be obtained, wherein the difference measure can include absolute error, relative error, etc. As described above, for each question in the question set, multiple predicted answers in multiple answer forms and multiple true value answers from multiple answer sources corresponding thereto will be obtained. For all questions in the question set, the comparison between the predicted answers of each question obtained and the corresponding true value answers can be organized in different answer forms. For example, in the ANSI-SQL answer form, multiple comparisons between the ANSI-SQL predicted answers of each question and the corresponding ANSI-SQL true value answers can be organized together to obtain multiple comparison results in the ANSI-SQL answer form. In other words, multiple comparison results between the predicted answers of each question in the same answer form and the corresponding true value answers will be, for example, averaged or representative values such as the maximum or minimum values will be taken to indicate the item differences of the answer form.
[0061] In some embodiments, the comparison of SQL can be based on the code abstract syntax tree (AST). Similarity calculation is an effective method. AST is a tree structure that represents the grammatical structure of the code. By converting the SQL query to AST, the syntax and structure of the two SQL queries can be compared more accurately, rather than just a superficial string match.
[0062] Figure 5 The exemplary scoring process 500 of the sub-item scores and the fusion score organized in the form of answers according to an embodiment of the present disclosure is schematically illustrated. It should be understood that for the sake of convenience of illustration and easy understanding, a limited number of answer forms are described herein, but the embodiments of the present invention are not limited to these illustrative answer types, and can also be applied to (more or fewer) other different answer types.
[0063] For multiple questions of the question set 510, the comparison between the predicted answers obtained via the language model to be evaluated and the corresponding true answers from the answer source are organized in the form of answers. For example, based on the comparison between the predicted ANSI-SQL 521 and the true ANSI-SQL 531, a comparison result 1 541 for ANSI-SQL can be obtained, based on the comparison between the predicted CH-SQL 522 and the true CH-SQL 532, a comparison result 2 542 for CH-SQ can be obtained, and based on the comparison between the predicted SQL execution result 523 and the true SQL execution result 533, a comparison result 3 543 for the SQL execution result can be obtained. A corresponding comparison score for each answer form can be calculated based on the corresponding comparison result for each answer form, such as Figure 4As shown in , comparison score 1 is 551, comparison score 2 is 552, and comparison score 3 is 553.
[0064] According to an embodiment of the present disclosure, a corresponding weight may be assigned to each answer form. These weights may be normalized, for example, by a softmax function to ensure that the sum of the multiple weights assigned to the multiple answer forms is 1. Figure 5 As shown in , the weight of the ANSI-SQL answer type is W1, the weight of the CH-SQL answer type is W2, and the weight of the SQL execution result is W3, where the sum of W1, W2, and W3 is 1. The performance score 560 of the language model is calculated by weighted averaging the corresponding comparison scores based on the corresponding weights. In addition to the weighted average, a voting mechanism or other data fusion algorithms may also be used.
[0065] In some embodiments, after obtaining the comparison score of each item (i.e., comparison score 1 551, comparison score 2 552, comparison score 3 553), in response to determining that at least one of the first comparison score for the first type of SQL statement (i.e., ANSI-SQL) (i.e., comparison score 1 551), the second comparison score for the second type of SQL statement (i.e., CH-SQL) (i.e., comparison score 2 552), or the third comparison score for the SQL execution result (i.e., SQL execution result) (i.e., comparison score 3 553) satisfies the corresponding score threshold, the performance score 560 of the language model is determined to be a first performance score (e.g., 1), and the first performance score indicates that the answer output by the language model is accurate. If the comparison scores 1 551, 2 552, and 3 553 are equal, the performance score of the language model is determined to be 1. If none of the above two answers meet the corresponding score threshold, the performance score 560 of the language model is determined as a second performance score (e.g., -1), and the second performance score indicates that the answer output by the language model is wrong. If it is found that there is a null value or a zero value in the true value data, or there is no corresponding true value data, a warning score (e.g., 0) is given to indicate manual review.
[0066] According to an embodiment of the present disclosure, multiple comparison scores (i.e., comparison score 1 551, comparison score 2 552, comparison score 3 553) and a performance score 560 of a language model can be provided to a user in a visual manner. An evaluation report can be generated, including various indicators of model performance, difference analysis results, and visualization charts, to help the development team quickly understand the performance of the model and the direction of improvement. The prediction results of the model are evaluated using the integrated or fused performance score 560. Common evaluation indicators include accuracy, recall, and F1 score.
[0067] Figure 6Schematically illustrates an apparatus 600 for evaluating the performance of a language model according to an embodiment of the present disclosure. The apparatus 600 may include multiple units or modules for performing the steps or actions in the method or process discussed above. Figure 6 As shown in , the device 600 includes: a predicted answer acquisition module 610, which is configured to obtain multiple predicted answers to the question through a language model for each question in the question set, and the multiple predicted answers each have a different answer form, and the answer form includes a first type of structured query language SQL statement, a visual uniform resource locator URL, a second type of SQL statement, an SQL execution result, a recall field and a URL execution result; a true value answer acquisition module 620, which is configured to be configured to obtain multiple true value answers for the question, and the multiple true value answers each correspond to a different answer source, and the answer source includes a first SQL source, a visual URL source, a second SQL source, an SQL execution result source, a recall field source and a URL execution result source; a comparison module 630, which is configured to obtain multiple comparison results based on the comparison between the multiple predicted answers and the multiple true value answers, and the multiple comparison results indicate the differences between the multiple predicted answers and the multiple true value answers; and a scoring module 640, which is configured to determine a plurality of comparison scores corresponding to the plurality of comparison results, and determine a performance score of the language model based on the plurality of comparison scores, and the performance score indicates the question-answering performance of the language model.
[0068] In some embodiments, the multiple predicted answers to the question include a first predicted answer whose answer form is the first type of SQL statement and a second predicted answer whose answer form is the visualized URL, and the device 600 also includes: a first answer determination module, configured to determine a third predicted answer and a fourth predicted answer for the question based on the first predicted answer, the answer form of the third predicted answer being the second type of SQL statement, and the answer form of the fourth predicted answer being the SQL execution result; and a second answer determination module, configured to determine a fifth predicted answer and a sixth predicted answer for the question based on the second predicted answer, the answer form of the fifth predicted answer being the recall field, and the answer form of the sixth predicted answer being the URL execution result.
[0069] In some embodiments, the multiple true answers to the question include: a first true answer in the form of an SQL statement of the first type, the first true answer corresponding to the first predicted answer and coming from the first SQL source, the first SQL source being configured to provide annotated SQL statements of the first type; a second true answer in the form of an visualization URL, the second true answer corresponding to the second predicted answer and coming from the visualization URL source, the visualization URL source being configured to provide annotated visualization URLs; a third true answer in the form of an SQL statement of the second type, the third true answer corresponding to the third predicted answer and coming from the second SQL source, the second SQL source being configured To provide a labeled second type of SQL statement; the answer form is a fourth true value answer of the SQL execution result, the fourth true value answer corresponds to the fourth predicted answer and comes from the SQL execution result source, the SQL execution result source is configured to provide a labeled SQL execution result; the answer form is a fifth true value answer of the recall field, the fifth true value answer corresponds to the fifth predicted answer and comes from the recall field source, the recall field source is configured to provide a labeled recall field; and the answer form is a sixth true value answer of the URL execution result, the sixth true value answer corresponds to the sixth predicted answer and comes from the URL execution result source, the URL execution result source is configured to provide a labeled URL execution result.
[0070] In some embodiments, the multiple truth-value answers are stored in a knowledge base, and the device 600 further includes: a formatting module configured to obtain formatted homologous truth-value answers having a consistent format by performing formatting processing on homologous truth-value answers from the same source with a unified formatting strategy; and a storage module configured to organize the formatted homologous truth-value answers in corresponding sub-knowledge bases in the knowledge base.
[0071] In some embodiments, the comparison module 630 is further configured to obtain a corresponding comparison result for each answer form by calculating a difference metric between the predicted answer of each answer form in multiple answer forms and the corresponding true answer, wherein the difference metric includes an absolute error and a relative error.
[0072] In some embodiments, the scoring module 640 is further configured to: calculate a corresponding comparison score for each answer form based on the corresponding comparison result for each answer form; assign a corresponding weight to each answer form, and the sum of multiple weights assigned to the multiple answer forms is 1; and calculate the performance score of the language model by weighted averaging the corresponding comparison scores based on the corresponding weights.
[0073] The scoring module 640 is further configured to: calculate a corresponding comparison score for each answer form based on a corresponding comparison result for each answer form; in response to determining that at least one of the first comparison score for the first type of SQL statement, the second comparison score for the second type of SQL statement, or the third comparison score for the SQL execution result satisfies a corresponding score threshold, determine the performance score of the language model as a first performance score, the first performance score indicating that the answer output by the language model is accurate; and in response to determining that the first comparison score, the second comparison score, or the third comparison score all satisfy the corresponding score threshold, determine the performance score of the language model as a second performance score, the second performance score indicating that the answer output by the language model is incorrect.
[0074] In some embodiments, the apparatus 600 further includes a providing module configured to provide the plurality of comparison scores and the performance score of the language model to a user in a visual manner.
[0075] Figure 7 FIG. 1 is a block diagram of an electronic device 700 according to some embodiments of the present disclosure. The device 700 may be a device or apparatus described in an embodiment of the present disclosure. Figure 7 As shown, the device 700 includes a central processing unit (CPU) and / or a graphics processing unit (GPU) 701, which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) 702 or loaded from a storage unit 708 to a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the device 700 can also be stored. The CPU / GPU 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704. Although not shown in FIG. Figure 7 As shown in FIG. 7 , device 700 may further include a coprocessor.
[0076] A number of components in the device 700 are connected to the I / O interface 705, including: an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a disk, an optical disk, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the device 700 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0077] The various methods or processes described above may be performed by the CPU / GPU 701. For example, in some embodiments, the methods may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 708. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the CPU / GPU 701, one or more steps or actions in the methods or processes described above may be performed.
[0078] In some embodiments, the methods and processes described above may be implemented as a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present disclosure.
[0079] Computer readable storage medium can be a tangible device that can hold and store instructions used by an instruction execution device. Computer readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (non-exhaustive list) of computer readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, for example, a punch card or a convex structure in a groove on which instructions are stored, and any suitable combination thereof. The computer readable storage medium used here is not interpreted as a transient signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated by a waveguide or other transmission medium (for example, a light pulse by an optical fiber cable), or an electrical signal transmitted by a wire.
[0080] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in the computer-readable storage medium in each computing / processing device.
[0081] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages, and conventional procedural programming languages. Computer-readable program instructions may be executed entirely on a user's computer, partially on a user's computer, as a separate software package, partially on a user's computer, partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., using an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may be customized by utilizing the state information of the computer-readable program instructions, and the electronic circuit may execute the computer-readable program instructions, thereby implementing various aspects of the present disclosure.
[0082] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device that implements the functions / actions specified in one or more boxes in the flowchart and / or block diagram is generated. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing device, and / or other equipment to work in a specific manner, so that the computer-readable medium storing the instructions includes a manufactured product, which includes instructions for implementing various aspects of the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0083] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operating steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0084] The flow chart and block diagram in the accompanying drawings show the possible architecture, function and operation of the equipment, method and computer program product according to multiple embodiments of the present disclosure. In this regard, each frame in the flow chart or block diagram can represent a part of a module, program segment or instruction, and a part of the module, program segment or instruction includes one or more executable instructions for realizing the specified logical function. In some alternative implementations, the function marked in the frame can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous frames can actually be executed substantially in parallel, and they can also be executed in the opposite order sometimes, depending on the functions involved. It should also be noted that each frame in the block diagram and / or flow chart, and the combination of frames in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.
[0085] Various embodiments of the present disclosure have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the various embodiments, practical applications, or technical improvements to the technology in the market, or to enable other persons of ordinary skill in the art to understand the various embodiments disclosed herein.
Claims
1. A method for evaluating the performance of a language model, comprising: For each question in the question set, a plurality of predicted answers to the question are obtained through a language model, wherein each of the plurality of predicted answers has a different answer form, and the answer form includes a first type of structured query language SQL statement, a visual uniform resource locator URL, a second type of SQL statement, an SQL execution result, a recall field, and a URL execution result; Obtain multiple true value answers to the question, each of the multiple true value answers corresponds to a different answer source, the answer source including a first SQL source, a visualization URL source, a second SQL source, a SQL execution result source, a recall field source, and a URL execution result source; Based on the comparison between the plurality of predicted answers and the plurality of true value answers, obtaining a plurality of comparison results, the plurality of comparison results indicating the differences between the plurality of predicted answers and the plurality of true value answers; as well as determining a plurality of comparison scores corresponding to the plurality of comparison results, and determining a performance score of the language model based on the plurality of comparison scores, the performance score indicating question answering performance of the language model, The multiple predicted answers to the question include a first predicted answer in the form of an answer of the first type of SQL statement and a second predicted answer in the form of an answer of the visualization URL, and the method further includes: determining a third predicted answer and a fourth predicted answer for the question based on the first predicted answer, wherein the answer form of the third predicted answer is the second type of SQL statement, and the answer form of the fourth predicted answer is the SQL execution result; as well as A fifth predicted answer and a sixth predicted answer for the question are determined based on the second predicted answer, the answer form of the fifth predicted answer being the recall field, and the answer form of the sixth predicted answer being the URL execution result.
2. The method of claim 1, wherein the plurality of true value answers to the question include: The answer form is a first true answer to the SQL statement of the first type, the first true answer corresponds to the first predicted answer and comes from the first SQL source, the first SQL source is configured to provide annotated SQL statements of the first type; an answer in the form of a second true answer to the visualization URL, the second true answer corresponding to the second predicted answer and coming from the visualization URL source, the visualization URL source being configured to provide annotated visualization URLs; The answer form is a third true answer to the SQL statement of the second type, the third true answer corresponds to the third predicted answer and comes from the second SQL source, the second SQL source is configured to provide annotated SQL statements of the second type; The answer form is a fourth true value answer of the SQL execution result, the fourth true value answer corresponds to the fourth predicted answer and comes from the SQL execution result source, and the SQL execution result source is configured to provide annotated SQL execution results; an answer in the form of a fifth truth answer to the recall field, the fifth truth answer corresponding to the fifth predicted answer and coming from the recall field source, the recall field source being configured to provide annotated recall fields; as well as The answer form is a sixth true value answer of the URL execution result, the sixth true value answer corresponds to the sixth predicted answer and comes from the URL execution result source, and the URL execution result source is configured to provide annotated URL execution results.
3. The method according to claim 2, wherein the plurality of true value answers are stored in a knowledge base, the method further comprising: Obtaining formatted homologous truth-value answers having a consistent format by performing formatting processing on homologous truth-value answers from the same source using a unified formatting strategy; as well as The formatted homologous truth answers are organized into corresponding sub-knowledge bases in the knowledge base.
4. The method according to claim 1, wherein obtaining the plurality of comparison results comprises: The corresponding comparison result for each answer form is obtained by calculating the difference metric between the predicted answer of each answer form in multiple answer forms and the corresponding true answer, wherein the difference metric includes an absolute error and a relative error.
5. The method of claim 4, wherein determining the performance score of the language model comprises: Calculating a corresponding comparison score for each answer form based on the corresponding comparison results for each answer form; Assigning a corresponding weight to each answer form, wherein the sum of the multiple weights assigned to the multiple answer forms is 1; as well as The performance score of the language model is calculated by weighted averaging the corresponding comparison scores based on the corresponding weights.
6. The method of claim 4, wherein determining the performance score of the language model comprises: Calculating a corresponding comparison score for each answer form based on the corresponding comparison results for each answer form; In response to determining that at least one of a first comparison score for the first type of SQL statement, a second comparison score for the second type of SQL statement, or a third comparison score for a SQL execution result satisfies a corresponding score threshold, determining the performance score of the language model as a first performance score indicating that an answer output by the language model is accurate; as well as In response to determining that none of the first comparison score, the second comparison score, or the third comparison score satisfies a corresponding score threshold, determining the performance score of the language model as a second performance score indicating that an answer output by the language model is incorrect.
7. The method according to claim 1, further comprising: The multiple comparison scores and the performance score of the language model are provided to a user in a visual manner.
8. A device for evaluating the performance of a language model, comprising: A predicted answer acquisition module is configured to acquire, for each question in the question set, a plurality of predicted answers to the question through a language model, wherein each of the plurality of predicted answers has a different answer form, and the answer form includes a first type of structured query language SQL statement, a visual uniform resource locator URL, a second type of SQL statement, an SQL execution result, a recall field, and a URL execution result; A true value answer acquisition module is configured to acquire multiple true value answers to the question, each of which corresponds to a different answer source, and the answer source includes a first SQL source, a visualization URL source, a second SQL source, an SQL execution result source, a recall field source, and a URL execution result source; A comparison module, configured to obtain a plurality of comparison results based on a comparison between the plurality of predicted answers and the plurality of true value answers, wherein the plurality of comparison results indicate differences between the plurality of predicted answers and the plurality of true value answers; as well as a scoring module configured to determine a plurality of comparison scores corresponding to the plurality of comparison results, and determine a performance score of the language model based on the plurality of comparison scores, the performance score indicating the question-answering performance of the language model, The multiple predicted answers to the question include a first predicted answer in the form of an SQL statement of the first type and a second predicted answer in the form of an answer of the visualized URL, and the device further includes: a first answer determination module configured to determine a third predicted answer and a fourth predicted answer for the question based on the first predicted answer, wherein the answer form of the third predicted answer is the second type of SQL statement, and the answer form of the fourth predicted answer is the SQL execution result; as well as The second answer determination module is configured to determine a fifth predicted answer and a sixth predicted answer for the question based on the second predicted answer, wherein the answer form of the fifth predicted answer is the recall field, and the answer form of the sixth predicted answer is the URL execution result.
9. An electronic device, comprising: processor; as well as A memory coupled to the processor, wherein instructions are stored in the memory, and when the instructions are executed by the processor, the processor is caused to perform the method according to any one of claims 1 to 7.
10. A machine-readable storage medium having machine-executable instructions stored thereon, wherein the machine-executable instructions, when executed by a processor, cause the processor to perform the method according to any one of claims 1 to 7.
11. A computer program product comprising a computer program which, when executed by a processor of a computer, causes the processor to perform the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Assessment method and system of large language model, electronic equipment and storage medium
CN119088914A