Large model evaluation method, device, equipment, system and program product
By extracting instructions, comparing answer information and evaluating the response quality of the big model question and answer data, and conducting a comprehensive evaluation of multiple scores, the subjectivity and inconsistency of the big model evaluation in the existing technology is solved, and a more accurate and automated evaluation process is achieved.
Patent Information
- Application Number
- CN202510042948.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art has subjectivity and contingency in evaluating the quality of large model responses, and it is impossible to conduct quantitative evaluation of quantifiable indicators, resulting in inconsistent evaluation in multiple rounds of dialogue scenarios, affecting the accuracy of the evaluation.
By obtaining the conversation data obtained through the target big model Q&A, the instructions are extracted and scored based on the following relationship between the question and answer conversations, the answer information is extracted and compared, the reply quality evaluation is performed, and a comprehensive evaluation is carried out in combination with multiple scores to achieve multi-dimensional evaluation of the big model.
Automation of large-scale model evaluation and multi-dimensional evaluation are realized, reducing manual participation, reducing subjective bias, and improving the accuracy of evaluation.
Smart Images

Figure CN120106210A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to large model evaluation methods, devices, equipment, systems and program products. Background Art
[0002] Currently, large models are increasingly used. In the daily use of various large language models, people often need to have multiple rounds of conversations with the models to achieve their goals. How to evaluate the quality of the large model's responses has become a difficult problem.
[0003] Generally, you can have multiple rounds of conversations with the large model through human-computer interaction, and score each round of conversations according to pre-set scoring criteria.
[0004] However, the manual evaluation process is subjective or accidental, and it is impossible to use quantifiable indicators to quantitatively evaluate multiple rounds of dialogue. In scenarios with large amounts of dialogue data, inconsistent evaluations may occur, affecting the accuracy of large model evaluations. Summary of the invention
[0005] In view of this, the present application provides a large model evaluation method to improve the accuracy of large model evaluation.
[0006] According to a first aspect of an embodiment of the present application, a method for evaluating a large model is provided, which can be applied to a system or program including a large model evaluation function in a terminal device, and specifically includes:
[0007] Acquire dialogue data obtained through question-answering of a target large model, wherein the dialogue data includes multiple rounds of question-answering dialogues;
[0008] Extracting instructions from the question-and-answer dialogues in sequence based on the follow-up relationship between the question-and-answer dialogues to obtain dialogue instructions, and scoring the dialogue instructions to obtain a first score;
[0009] Extracting answer information corresponding to the question-and-answer dialogue, and comparing the answer information with previous answer information to obtain a second score, wherein the previous answer information is determined based on answer information of a previous round of the question-and-answer dialogue;
[0010] Performing a response quality evaluation on the answer information to obtain a third score;
[0011] A target score is obtained by combining the first score, the second score and the third score, so as to evaluate the target macro model by using the target score.
[0012] Optionally, in some possible implementations, extracting instructions from the question-and-answer dialogues in sequence based on the follow-up relationship between the question-and-answer dialogues to obtain dialogue instructions, and scoring the dialogue instructions to obtain a first score, includes:
[0013] Identifying a conversation identifier in the conversation data to determine a conversation boundary;
[0014] Dividing the conversation data according to the conversation boundary and sorting them according to a preset method to obtain a text conversation sequence;
[0015] Determining the instruction follow-up type corresponding to the question-and-answer dialogue according to the follow-up relationship between the question-and-answer dialogues in the text dialogue sequence;
[0016] Extracting instructions from the question-and-answer dialogue in sequence based on the instruction follow-up type to obtain the dialogue instruction;
[0017] The dialogue instruction is scored using a comment model to obtain a first score.
[0018] Optionally, in some possible implementations, extracting instructions from the question-and-answer dialogue in sequence based on the instruction follow-up type to obtain the dialogue instruction includes:
[0019] In the case where it is determined that the instruction following type is a multi-round following type, determining the preceding answer information and the prompt word information corresponding to the question-and-answer dialogue based on the multi-round following type;
[0020] Combining the prompt word information and the previous answer information to obtain the current dialogue content;
[0021] The current conversation content is subjected to instruction extraction according to preset rules to obtain the conversation instruction.
[0022] Optionally, in some possible implementations, extracting instructions from the question-and-answer dialogue in sequence based on the instruction follow-up type to obtain the dialogue instruction includes:
[0023] In the case where it is determined that the instruction following type includes multiple types, determining combination information between the multiple types;
[0024] configuring the dialogue content corresponding to the question-and-answer dialogue based on the combination information to obtain current dialogue content;
[0025] The current conversation content is subjected to instruction extraction according to preset rules to obtain the conversation instruction.
[0026] Optionally, in some possible implementations, the training process of the comment model includes:
[0027] Obtain training corpus based on multi-round question-answering dialogue configuration;
[0028] Extracting instructions from the training corpus to obtain training instructions, and determining the instruction type corresponding to each training instruction;
[0029] Evaluate the training instructions of different instruction types based on preset dimensions to obtain training scores;
[0030] Combining the training score with the training corpus to obtain training data;
[0031] The review model is trained based on the training data.
[0032] Optionally, in some possible implementations, extracting answer information corresponding to the question-and-answer dialogue, and comparing the answer information with previous answer information to obtain a second score includes:
[0033] Extracting key elements from the conversation data in text order;
[0034] The key elements are collected to configure a key information library;
[0035] Extracting answer information corresponding to the question-and-answer dialogue, and matching the preceding answer information in the key information database according to the answer information to obtain a text information chain;
[0036] The answer information is compared with the text information link for conflicts to obtain the second score.
[0037] Optionally, in some possible implementations, performing a conflict comparison between the answer information and the text information link to obtain the second score includes:
[0038] Performing part-of-speech tagging on the answer information and the text information chain to obtain a part-of-speech tagging result, and determining first comparison information according to the part-of-speech tagging result;
[0039] Determine semantic vectors in the answer information and the text information chain, and determine second comparison information according to differences indicated by the semantic vectors;
[0040] Determine logic information corresponding to the answer information and the text information chain, and determine third comparison information according to the conflict degree indicated by the logic information;
[0041] The second score is obtained by combining the first comparison information, the second comparison information and the third comparison information.
[0042] Optionally, in some possible implementations, performing a reply quality evaluation on the answer information to obtain a third score includes:
[0043] determining key information in the answer information;
[0044] Answering the question-and-answer dialogue using a trusted resource to obtain a trusted response;
[0045] Comparing the key information with the confidence reply to obtain fourth comparison information;
[0046] Answering the question-answer dialogue using a large confidence model to obtain a response sequence;
[0047] Determining fifth comparison information according to the order of the answer information in the reply sequence;
[0048] The fourth comparison information and the fifth comparison information are combined to obtain the third score.
[0049] Optionally, in some possible implementations, combining the first score, the second score, and the third score to obtain a target score, so as to evaluate the target macro model by using the target score, includes:
[0050] configuring a hierarchical model by using the first score, the second score, and the third score as criterion layers;
[0051] Output weight information through the hierarchical structure model;
[0052] linearly weighting the first score, the second score, and the third score based on the weight information to obtain the target score;
[0053] The target macro model is evaluated by the target score.
[0054] Optionally, in some possible implementations, the method further includes:
[0055] Divide the conversation data into training set and validation set;
[0056] Calculating a validation score based on the validation set;
[0057] Performing confidence evaluation on the verification set to obtain a confidence score;
[0058] Determining a correlation coefficient by comparing the verification score and the confidence score, and adjusting the weight information based on the correlation coefficient;
[0059] A target score is determined based on the adjusted weight information.
[0060] According to a second aspect of an embodiment of the present application, a large model evaluation device is provided, comprising:
[0061] An acquisition unit, configured to acquire dialogue data obtained through question-answering of a target large model, wherein the dialogue data includes multiple rounds of question-answering dialogues;
[0062] a determining unit, configured to extract instructions from the question-and-answer dialogues in sequence based on the follow-up relationship between the question-and-answer dialogues to obtain dialogue instructions, and score the dialogue instructions to obtain a first score;
[0063] The determination unit is further configured to extract answer information corresponding to the question-and-answer dialogue, and compare the answer information with previous answer information to obtain a second score, wherein the previous answer information is determined based on answer information of a previous round of the question-and-answer dialogue;
[0064] The determining unit is further configured to perform a reply quality evaluation on the answer information to obtain a third score;
[0065] An evaluation unit is used to combine the first score, the second score and the third score to obtain a target score, so as to evaluate the target macro model through the target score.
[0066] Optionally, in some possible implementations, the determination unit is used to identify dialogue identifiers in the dialogue data to determine dialogue boundaries when extracting instructions from the question-and-answer dialogues in sequence based on the follow-up relationship between the question-and-answer dialogues to obtain dialogue instructions and scoring the dialogue instructions to obtain a first score; divide the dialogue data according to the dialogue boundaries, and sort them in a preset manner to obtain a text dialogue sequence; determine the instruction follow-up type corresponding to the question-and-answer dialogue according to the follow-up relationship between the question-and-answer dialogues in the text dialogue sequence; extract instructions from the question-and-answer dialogues in sequence based on the instruction follow-up type to obtain the dialogue instructions; and score the dialogue instructions according to a comment model to obtain a first score.
[0067] Optionally, in some possible implementations, the determination unit is used to extract instructions from the question-and-answer dialogue in sequence based on the instruction following type to obtain the dialogue instruction. When it is determined that the instruction following type is a multi-round following type, the determination unit determines the preceding answer information and prompt word information corresponding to the question-and-answer dialogue based on the multi-round following type; combines the prompt word information and the preceding answer information to obtain the current dialogue content; and extracts instructions from the current dialogue content according to preset rules to obtain the dialogue instruction.
[0068] Optionally, in some possible implementations, the determination unit is used to extract instructions from the question-and-answer dialogue in sequence based on the instruction following type to obtain the dialogue instruction, and when it is determined that the instruction following type includes multiple types, determine combination information between the multiple types; configure the dialogue content corresponding to the question-and-answer dialogue based on the combination information to obtain the current dialogue content; and extract instructions from the current dialogue content according to preset rules to obtain the dialogue instruction.
[0069] Optionally, in some possible implementations, the determination unit is used to obtain training corpus configured based on multiple rounds of question-and-answer dialogues; perform instruction extraction on the training corpus to obtain training instructions, and determine the instruction type corresponding to each training instruction; evaluate training instructions of different instruction types based on preset dimensions to obtain training scores; combine the training scores and the training corpus to obtain training data; and train the comment model based on the training data.
[0070] Optionally, in some possible implementations, the determination unit is used to extract key elements in the dialogue data in text order when extracting answer information corresponding to the question and answer dialogue and comparing the answer information with previous answer information to obtain a second score; collect the key elements to configure a key information library; extract the answer information corresponding to the question and answer dialogue, and match the previous answer information in the key information library based on the answer information to obtain a text information chain; and perform a conflict comparison between the answer information and the text information chain to obtain the second score.
[0071] Optionally, in some possible implementations, the determination unit is used to, when performing a conflict comparison between the answer information and the text information chain to obtain the second score, perform part-of-speech tagging on the answer information and the text information chain to obtain a part-of-speech tagging result, and determine first comparison information based on the part-of-speech tagging result; determine semantic vectors in the answer information and the text information chain, and determine second comparison information based on the differences indicated by the semantic vectors; determine logical information corresponding to the answer information and the text information chain, and determine third comparison information based on the degree of conflict indicated by the logical information; and obtain the second score by combining the first comparison information, the second comparison information, and the third comparison information.
[0072] Optionally, in some possible implementations, the determination unit is used to determine the key information in the answer information when performing a reply quality evaluation on the answer information to obtain a third score; answer the question and answer dialogue through trusted resources to obtain a trusted reply; compare the key information with the trusted reply to obtain fourth comparison information; answer the question and answer dialogue through a trusted big model to obtain a reply sequence; determine fifth comparison information based on the order of the answer information in the reply sequence; and combine the fourth comparison information and the fifth comparison information to obtain the third score.
[0073] Optionally, in some possible implementations, the evaluation unit is used to use the first score, the second score and the third score as criterion layers to configure a hierarchical model when obtaining a target score by combining the first score, the second score and the third score to evaluate the target large model through the target score; output weight information through the hierarchical model; linearly weight the first score, the second score and the third score based on the weight information to obtain the target score; and evaluate the target large model through the target score.
[0074] Optionally, in some possible implementations, the evaluation unit is further used to divide the conversation data into a training set and a verification set; calculate a verification score based on the verification set; perform confidence evaluation on the verification set to obtain a confidence score; determine a correlation coefficient by comparing the verification score and the confidence score, and adjust the weight information based on the correlation coefficient; and determine a target score based on the adjusted weight information.
[0075] According to a third aspect of an embodiment of the present application, there is provided a large model evaluation device, including an input-output component and a processor;
[0076] The input and output components are used to obtain conversation data;
[0077] The processor is used to evaluate the big model corresponding to the dialogue data obtained by the input-output component by executing the big model evaluation method as described in the first aspect or any implementation of the first aspect.
[0078] According to a fourth aspect of an embodiment of the present application, there is provided a large model evaluation system, including an interactive client and a server;
[0079] The interactive client is used to obtain the conversation data and send the conversation data to the server, and to display the large model evaluation result output by the server;
[0080] The server is used to evaluate the big model corresponding to the dialogue data obtained by the interactive client by executing the big model evaluation method described in the first aspect or any implementation of the first aspect.
[0081] According to a fifth aspect of an embodiment of the present application, a computer program product is provided, including: a computer program, which, when executed by a processor, implements the large model evaluation method described in the first aspect or any implementation manner of the first aspect.
[0082] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:
[0083] By obtaining the dialogue data obtained through the question-and-answer of the target large model, the dialogue data includes multiple rounds of question-and-answer dialogues; then extracting the instructions of the question-and-answer dialogue in turn based on the follow-up relationship between the question-and-answer dialogues to obtain dialogue instructions, and scoring the dialogue instructions to obtain a first score; and extracting the answer information corresponding to the question-and-answer dialogue, and comparing the answer information with the previous answer information to obtain a second score, the previous answer information is determined based on the answer information of the previous round of the question-and-answer dialogue; and evaluating the reply quality of the answer information to obtain a third score; and then combining the first score, the second score and the third score to obtain the target score, so as to evaluate the target large model through the target score. Thus, a multi-dimensional evaluation process is realized. Due to the multi-dimensional indicator configuration based on the characteristics of multi-round dialogues, automatic evaluation is realized, which can greatly reduce the degree of manual participation, reduce the deviation caused by personal subjective views, and improve the accuracy of large model evaluation. BRIEF DESCRIPTION OF THE DRAWINGS
[0084] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0085] Figure 1 A diagram of the network architecture that runs the large model evaluation system;
[0086] Figure 2 A process architecture diagram of a large model evaluation provided in an embodiment of the present application;
[0087] Figure 3 A flowchart of a large model evaluation method provided in an embodiment of the present application;
[0088] Figure 4 A schematic diagram of an interface for scoring a review model provided in an embodiment of the present application;
[0089] Figure 5 A schematic diagram of a scenario of a large model evaluation method provided in an embodiment of the present application;
[0090] Figure 6 A schematic diagram of the structure of a large model evaluation device provided in an embodiment of the present application;
[0091] Figure 7 A schematic diagram of the structure of a terminal device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0092] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0093] It should be understood that the large model evaluation method provided in the present application can be applied to a system or program including a large model evaluation function in a terminal device, such as an intelligent assistant application. Specifically, the large model evaluation system can be run on Figure 1 In the network architecture shown in Figure 1 As shown in FIG. 1 , it is a network architecture diagram of the operation of the large model evaluation system. As can be seen from the figure, the large model evaluation system can provide a large model evaluation process with multiple information sources, that is, the conversation data is determined by the acquisition operation on the terminal side, and then sent to the server to evaluate the performance of the large model corresponding to the conversation data; it can be understood that Figure 1 A variety of terminal devices are shown in FIG. 1 . The terminal devices may be computer devices. In actual scenarios, more or fewer types of terminal devices may participate in the large model evaluation process. The specific number and type depend on the actual scenario and are not limited here. In addition, Figure 1 One server is shown in the figure, but in actual scenarios, multiple servers may be involved, especially in multi-disciplinary scenarios, and the specific number of servers depends on the actual scenario.
[0094] In this embodiment, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal can be a smart phone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, etc., but is not limited to this. The terminal and the server can be directly or indirectly connected through wired or wireless communication, and the terminal and the server can be connected to form a blockchain network, which is not limited in this application.
[0095] It can be understood that the above-mentioned large model evaluation system can run on a personal mobile terminal, for example: as an application such as a smart assistant, it can also run on a server, and it can also run on a third-party device to provide a large model evaluation to obtain the large model evaluation processing results of the information source; the specific large model evaluation system can be run in the above-mentioned device in the form of a program, or it can be run as a system component in the above-mentioned device, and it can also be used as a cloud service program. The specific operation mode depends on the actual scenario and is not limited here.
[0096] Currently, large models are increasingly used. In the daily use of various large language models, people often need to have multiple rounds of conversations with the models to achieve their goals. How to evaluate the quality of the large model's responses has become a difficult problem.
[0097] Generally, you can have multiple rounds of conversations with the large model through human-computer interaction, and score each round of conversations according to pre-set scoring criteria.
[0098] However, the manual evaluation process is subjective or accidental, and it is impossible to use quantifiable indicators to quantitatively evaluate multiple rounds of dialogue. In scenarios with large amounts of dialogue data, inconsistent evaluations may occur, affecting the accuracy of large model evaluations.
[0099] In order to solve the above problems, this application proposes a large model evaluation method, which is applied to Figure 2 In the process framework of large model evaluation shown in Figure 2 As shown, it is a process architecture diagram of a large model evaluation provided in an embodiment of the present application, which shows that the user determines the conversation data through the interactive operation between the terminal and the large model, so that the server performs a multi-dimensional evaluation process of follow-up relationship, content inheritance and reply quality based on the conversation data.
[0100] It can be understood that the large model evaluation method provided in the present application can be written as a program as a processing logic in a hardware system, or as a large model evaluation device, which implements the above processing logic in an integrated or external manner. As an implementation method, the large model evaluation device obtains the dialogue data obtained through the question and answer of the target large model, and the dialogue data includes multiple rounds of question and answer dialogues; then, based on the follow-up relationship between the question and answer dialogues, the question and answer dialogues are sequentially extracted to obtain dialogue instructions, and the dialogue instructions are scored to obtain a first score; and the answer information corresponding to the question and answer dialogue is extracted, and the answer information is compared with the previous answer information to obtain a second score, and the previous answer information is determined based on the answer information of the previous round of the question and answer dialogue; and the answer information is evaluated for the quality of the reply to obtain a third score; and then the first score, the second score and the third score are combined to obtain the target score, so as to evaluate the target large model through the target score. Thus, a multi-dimensional evaluation process is realized. Since multi-dimensional indicators are configured according to the characteristics of multi-round dialogues, automatic evaluation is realized, which can greatly reduce the degree of manual participation, reduce the deviation caused by personal subjective views, and improve the accuracy of large model evaluation.
[0101] Combined with the above process architecture, the following will introduce the large model evaluation method in this application. Please refer to Figure 3 , Figure 3 A flowchart of a large model evaluation method provided in an embodiment of the present application, the embodiment of the present application at least includes the following steps:
[0102] 301. Acquire dialogue data obtained through question-answering of a target large model, where the dialogue data includes multiple rounds of question-answering dialogues.
[0103] In this embodiment, the target large model is a large language model, which is pre-trained on massive text data through deep learning technology to achieve efficient understanding and generation of natural language. Correspondingly, the dialogue data is the data generated by the question-and-answer interaction process between the user and the large language model, and because the dialogue has temporal or logical continuity, the dialogue data includes multiple rounds of question-and-answer dialogues.
[0104] It is understandable that the input process of the question-and-answer dialogue can be carried out through text input, voice input, or other somatosensory interaction methods, and the specific form of interaction depends on the actual scenario.
[0105] 302. Extracting instructions from the question-answer dialogues in sequence based on the follow-up relationship between the question-answer dialogues to obtain dialogue instructions, and scoring the dialogue instructions to obtain a first score.
[0106] In this embodiment, the follow-up relationship between question-and-answer dialogues is the content or logical association relationship between different rounds of question-and-answer dialogues; and the dialogue instructions are the characteristic parameters in the question-and-answer dialogue, which are used to directly reflect the content in the question-and-answer dialogue, and the dialogue instructions can be used as constraints of the large language model to trigger the large language model to answer questions.
[0107] It is understandable that since the dialogue data includes multiple rounds of dialogue, the large language model may forget the instruction requirements of several rounds of dialogue, or ignore the constraints of the system prompt; or the instructions of the current round may not be fully executed, and there may be partial or incorrect execution of multiple intentions or multi-level instructions. Therefore, this embodiment evaluates different dialogues from the perspective of the overall dialogue through the fine-grained division of the follow-up relationship in the instruction dimension.
[0108] Specifically, since the constraints of the dialogue instructions can be divided into single round and multi-round, and the constraints will be reflected in the system prompt or query input by the user. Therefore, the process of instruction extraction can be carried out for different types of dialogue instructions, that is, firstly identify the dialogue identifier in the dialogue data to determine the dialogue boundary; then divide the dialogue data according to the dialogue boundary, and sort it according to the preset method to obtain the text dialogue sequence; and determine the instruction follow-up type corresponding to the question-answer dialogue according to the follow-up relationship between the question-answer dialogues in the text dialogue sequence; then extract the instructions of the question-answer dialogue in turn based on the instruction follow-up type to obtain the dialogue instruction; and then score the dialogue instruction through the comment model to obtain the first score.
[0109] Among them, the conversation identifier is an identifier configured by means of timestamp, conversation sequence identifier, etc., which is used to determine the boundaries of each round of conversation, that is, to assign a unique identifier to each round of conversation, to facilitate the subsequent overall processing and correlation analysis of multiple rounds of conversations, and then to combine each round of conversation into a text conversation sequence in chronological order or logical order.
[0110] Specifically, different instruction following types may include a single-round following type and a multi-round following type; wherein the definition of the single-round following type includes:
[0111] a. Structure: For example, the answer needs to be expanded in the format of total score and total, or output in a specific markdown format such as a table, etc.;
[0112] b. Genre: such as official documents, promotional materials, Xiaohongshu copywriting, etc.
[0113] c. Perspective: for example, first-person perspective, official perspective, objective evaluation, industry perspective, etc.;
[0114] d. Style: For example, the answer needs to be personified, humorous, formal, two-dimensional, etc.;
[0115] e. Language: For example, the answer needs to be given in English, French, Japanese, classical Chinese, Sichuan dialect, etc.;
[0116] The dimensions defined here are just an example and other dimensions can be expanded.
[0117] In addition, the definition of multi-round follow-up types includes:
[0118] a. Personality requirements: At the beginning of multiple rounds or through system prompts, the model is required to perform tasks step by step according to the established workflow in the next set of multiple rounds of interactions. And it usually has a more distinct personality perspective and language style;
[0119] b. Long conversation dependency: In the following conversation, you need to follow the set instructions. For example, "In the following conversation, you need to translate every sentence I say."
[0120] The dimensions defined here are just an example and other dimensions can be expanded.
[0121] It can be seen that through the above definitions of different instruction follow-up types, the extraction process of multi-round follow-up types needs to consider the situation of the previous answer. That is, when the instruction follow-up type is determined to be a multi-round follow-up type, the previous answer information and prompt word information corresponding to the question-and-answer dialogue are determined based on the multi-round follow-up type; then the prompt word information and the previous answer information are merged to obtain the current dialogue content; and the current dialogue content is extracted according to the preset rules to obtain the dialogue instruction. For example, to judge Q3 of the multi-round dialogue, it is necessary to merge the system prompt (if any), Q1, Q2, and Q3.
[0122] In a possible scenario, since a multi-round dialogue often has multiple rounds, the instruction type of the entire dialogue often includes one or more of the above instruction types, and there are some combinations. Therefore, when it is determined that the instruction follow-up type includes multiple types, the combination information between the multiple types can be determined; then the dialogue content corresponding to the question-and-answer dialogue is configured based on the combination information to obtain the current dialogue content; and the instruction extraction of the current dialogue content is performed according to the preset rules to obtain the dialogue instruction. Among them, the combination information includes but is not limited to the following combination forms: chain (instruction A needs to be executed first, and then instruction B), judgment (if the condition is met, then instruction A is executed, otherwise instruction B is executed), etc.
[0123] It can be understood that for the configuration of preset rules in the above-mentioned instruction extraction process, grammar rules and keyword rules can be formulated to extract instructions, or through machine learning or deep learning methods (training a sequence labeling model (such as BiLSTM+CRF model) or an entity extraction model, etc.), and post-processing, that is, dealing with reference resolution problems and semantic regularization, and can be matched and verified with predefined instruction templates or knowledge bases to check whether the extracted instructions are complete and accurate.
[0124] After the above-mentioned different types of instructions are extracted, the skill-based comment model scoring process can be carried out, that is, the instruction type is judged through the classification model or the large model prompt project, and the extracted instruction requirements are classified according to the set dimensions. The subsequent comment model will evaluate and assign scores based on these sub-dimensions.
[0125] Specifically, for the training process of the comment model, you can first obtain training corpus based on a multi-round question-answering dialogue configuration; then perform instruction extraction on the training corpus to obtain training instructions, and determine the instruction type corresponding to each training instruction; and evaluate the training instructions of different instruction types based on preset dimensions to obtain training scores, where the preset dimensions may include structural requirements, text requirements, style requirements and other requirements for describing the text; then combine the training scores and the training corpus to obtain training data; and train the comment model based on the training data.
[0126] Therefore, for the training of the comment model, we can collect a small amount of real users and multi-round conversation data (training corpus) of a large model, and then obtain the content and type of instructions through instruction extraction. Then, through manual labeling, evaluate them according to the instruction type to obtain the comments and scores under each sub-dimension. Then, we can further obtain the total score of the instruction in a weighted manner (different types of instructions can be scored differently according to the degree of difficulty). In this way, we can obtain the training data of the comment model, thereby ensuring that the output results of the comment model can be better aligned with human judgment.
[0127] In a possible scenario, after the user enters multiple rounds of dialogue, the command extraction and comment model scoring are performed for each round in turn, and the answer command follow-up score (first score) and question analysis for that round are output, where the output content can be used Figure 4 The method shown is shown. Figure 4 A schematic diagram of an interface for scoring a comment model provided in an embodiment of the present application; the figure shows that answer instructions are displayed on a display interface along with scores and question analysis, and the reasons for such scoring are given, thereby increasing the interpretability of the evaluation platform.
[0128] 303. Extract answer information corresponding to the question-and-answer dialogue, and compare the answer information with previous answer information to obtain a second score, where the previous answer information is determined based on answer information of a previous round of the question-and-answer dialogue.
[0129] In this embodiment, the process of comparing the answer information with the previous answer information is the consideration of the dialogue data after multiple rounds of inheritance. That is, because the user completes a set of multiple rounds of interaction, and the tone and length of the answer of the large language model after several rounds of dialogue have changed significantly compared with the previous rounds, there is a loss or tampering of some historical information. For example, in the task of completing the novel generation, the "main character" set in the previous interaction has some information of the character tampered in the subsequent dialogue, resulting in problems such as context mismatch or logical conflict. Therefore, this embodiment determines by matching the previous answer information, and then evaluates the content inheritance dimension.
[0130] Specifically, for the evaluation process of the content inheritance dimension, the key elements in the dialogue data can be first extracted in text order; then the key elements are collected to configure a key information library; and the answer information corresponding to the question-and-answer dialogue is extracted, and the preceding answer information is matched in the key information library according to the answer information to obtain a text information chain; and then the answer information is compared with the text information chain for conflicts to obtain a second score.
[0131] It is understandable that the above process of comparing the answer information with the text information chain for conflicts, i.e., for content inheritance, is mainly to determine whether the current round of replies will have any information or logic conflicts with the previous ones. Unlike instruction following, this requires merging the previous answers and extracting key information, and then comparing whether the current round of answers conflict with the previous answers.
[0132] In a possible scenario, for the evaluation process of the above content inheritance dimension, taking the novel generation task as an example, the evaluation process includes:
[0133] A. Information collection steps: Obtain novel texts from multiple rounds of interactions in real time, and organize and store them in chronological order or interaction round order (text order).
[0134] B. Key information extraction step: Use natural language processing algorithms to process the collected text, extract key elements such as nouns (protagonists), core verbs (main plots, etc.), qualifiers (story background, etc.), and build a key information database.
[0135] C. Content inheritance step: When processing the current round of information, query the key information database, match and integrate the previous key information related to the current information, and build a complete information chain.
[0136] D. Conflict detection (conflict comparison) step: After completing content inheritance, compare and analyze the results generated in the current round with the inherited information chain, and perform conflict detection from multiple levels such as syntax, semantics, and logic. Points are assigned according to the degree of conflict.
[0137] Specifically, the conflict comparison process can be performed based on multiple levels such as grammar, semantics, and logic. At the grammatical level, the answer information and the text information chain can be tagged with parts of speech to obtain the results of part-of-speech tagging, and the first comparison information can be determined based on the results of part-of-speech tagging; that is, part-of-speech tagging and dependency analysis are performed, specifically using natural language processing models in deep learning, such as pre-trained language models based on the Transformer architecture (such as variants of BERT, GPT, etc.), to tag the text with parts of speech, and determine the part-of-speech category of each word (noun, verb, adjective, etc.). Then, through dependency analysis, find out the grammatical dependency relationship between words (subject-predicate relationship, verb-object relationship, etc.). When detecting conflicts, compare the part-of-speech tagging results and dependency relationships of the current round text with the corresponding part in the inherited information chain.
[0138] At the semantic level, the semantic vectors in the answer information and the text information chain can be determined, and the second comparison information can be determined according to the difference indicated by the semantic vector; that is, the comparison is performed through word vector and semantic similarity calculation, specifically with the help of word vector representation methods in deep learning, such as word vector representation obtained by Word2Vec, FastText, etc. Calculate the semantic similarity between the key words in the current round text and the corresponding words in the inherited information chain. For words that express the same concept or similar concepts, their semantic similarity should be within a reasonable range. For example, when describing a specific place in a novel, the previous text has always used the word "castle", but the current round frequently uses "modern apartment" to refer to it. The semantic similarity between the two is extremely low through word vector calculation, indicating that there may be a conflict in semantics. At the same time, semantic vector representation can also be performed on sentences and even paragraph levels to measure the semantic differences between them as a whole, and judge whether there is a semantic conflict and assign points based on the degree of difference.
[0139] At the logical level, we can determine the logical information corresponding to the answer information and the text information chain, and determine the third comparison information based on the degree of conflict indicated by the logical information; that is, to identify logical relationships, specifically using classification algorithms in machine learning (such as support vector machines, decision trees, etc., after training with a large amount of labeled logical relationship text data) to identify logical relationships in the text, such as causal relationships, transitional relationships, parallel relationships, etc. Check whether the content generated in the current round is logically consistent with the inherited information chain. For example, the previous text explained the causal logic that the team fell into crisis because of the betrayal of a certain character, but the current round stated that the team successfully overcame the crisis for no reason, which violated the causal logical connection established previously. It can be judged that there is a logical conflict, and then points are assigned based on factors such as the importance of the conflict.
[0140] After performing the comparisons at the above multiple levels, the first comparison information, the second comparison information and the third comparison information can be combined to obtain a second score, thereby improving the comprehensiveness of the conflict comparison and improving the accuracy of the content inheritance dimension scoring.
[0141] 304. Evaluate the quality of the answer information to obtain a third score.
[0142] In this embodiment, in order to evaluate the quality of answers in the dialogue data, the reply quality is used as an important criterion for in-depth analysis. For example, the generation of repeated meaningless words and sentences, the generated answers are redundant and have no useful information, the generated text content is not beautiful enough, the generated format is not good, and the degree of anthropomorphism is not enough, which directly affects the user's interactive experience.
[0143] Specifically, the definition of the evaluation dimension of reply quality can be measured by two dimensions: accuracy and winning rate. Among them, accuracy is used to indicate that the reply can fully answer the question, and the information in the reply must be true and reliable. For example, there should be no errors when introducing the dates of historical events, names of people and other details; when providing technical parameters, product specifications and other data, they must also be accurate. The reply should be logical. The reasoning process from premise to conclusion should be reasonable to avoid self-contradictory or unreasonable situations. For example, you cannot first say that a product is suitable for all people, and then say that the product will cause severe allergic reactions to specific people. The grammar is correct and the structure is clear. In addition, the winning rate is used to indicate that the reply should be better than the replies of multiple competing products. For example, when generating an essay, the generated article needs to have good writing style, and direct judgment is more subjective. You can choose the generation results of the mainstream large model in the industry and get the ranking through the AB test method.
[0144] Therefore, for the response quality evaluation process, the key information in the answer information can be first determined; then the question-answering dialogue can be answered through trusted resources (such as trusted network resources) to obtain trusted responses; then the key information can be compared with the trusted responses to obtain fourth comparison information (accuracy); and the question-answering dialogue can be answered through a trusted macro model (such as a trusted general macro model) to obtain a response sequence; then the fifth comparison information (winning rate) can be determined based on the order of the answer information in the response sequence; and the fourth comparison information and the fifth comparison information can be combined to obtain a third score.
[0145] It is understandable that the accuracy detection process above uses the RAG method, that is, extracting key information (knowledge entities, time, place, game results, etc.) from the answer, and searching selected websites with high credibility. After obtaining the standard answer, the accuracy of the information in the answer is analyzed through large model prompt engineering judgment or semantic similarity comparison, and a corresponding score is given.
[0146] As for the detection of the winning rate, we send the conversation history to the mainstream big model in the industry through rejection sampling to obtain multiple generated results. We also use the trained reinforcement learning reward model (confidence big model) to score and sort the multiple answers, and get the ranking and score of the responses to be evaluated, thereby expanding the scope of evaluation for the target big model and making the evaluation of the response quality more representative.
[0147] 305. Combining the first score, the second score, and the third score to obtain a target score, so as to evaluate the target macro model through the target score.
[0148] In this embodiment, by combining the above-mentioned instruction following, content inheritance and reply quality evaluation process, the following can be obtained: Figure 5 The execution scenario shown, Figure 5 A scenario diagram of a large model evaluation method provided for an embodiment of the present application; the figure shows that firstly, the dialogue data is extracted by the slot extraction model to obtain dialogue instructions, and then the dialogue instructions are classified by type through the classification model, and a multi-dimensional evaluation process is performed through the comment scoring model (comment model) to obtain the first score; and based on the comment model, the second score is obtained by comparing the content inheritance of the answer information with the previous answer information. Furthermore, the winning rate is evaluated by combining rejection sampling with the reward model of reinforcement learning, and the accuracy is evaluated in combination with RAG technology to obtain the third score, and finally the above scores are comprehensively evaluated to obtain the target score, thereby realizing an automated evaluation process based on artificial intelligence methods.
[0149] Combined with the scores of the above three dimensions, a comprehensive evaluation is conducted. The above steps can be automated and performed simultaneously. Compared with traditional methods, they are more granular, more accurate, more convenient and faster. Through the above evaluation scenarios, a multi-dimensional, automated, and highly reliable multi-round dialogue evaluation platform is configured to help quantify the multi-round capabilities of large models. Specifically, a multi-dimensional evaluation of the multi-round dialogue evaluation is conducted, and detailed definitions are given from three aspects: command following, multi-round inheritance, and reply quality. It can be evaluated and scored from multiple dimensions, and a more comprehensive score can be obtained through the construction of an evaluation index system; and in the evaluation of each dimension, the above model is configured for automated evaluation, which can greatly reduce the degree of manual participation and reduce the deviation caused by personal subjective views. At the same time, it is more convenient and quick, and releases human resources.
[0150] Furthermore, the comprehensive evaluation process of the above-mentioned multiple scores can be a weighted evaluation process, that is, firstly taking the first score, the second score and the third score as the criterion layer to configure the hierarchical model; then outputting the weight information through the hierarchical model; and linearly weighting the first score, the second score and the third score based on the weight information to obtain the target score; and then evaluating the target large model through the target score.
[0151] It is understandable that before weighting each score, the score can be standardized. Assume that the instruction following score is recorded as X1, the context inheritance score is recorded as X2, and the answer quality score is recorded as X3. First calculate the sample mean and sample standard deviation of each dimension. Z-standardized score Zi = (Xi-sample mean of the corresponding dimension) / sample standard deviation of the corresponding dimension (i = 1, 2, 3), so that each dimension score will be converted to a standard normal distribution with a mean of 0 and a standard deviation of 1.
[0152] In addition, the configuration of the above-mentioned hierarchical model is used for weight calculation, that is, the analytic hierarchy process (AHP) is adopted to construct a hierarchical model, with comprehensive evaluation as the target layer, instruction following, context inheritance, and answer quality as the criterion layer. By establishing a judgment matrix, the relative importance of each criterion is compared and the weight vector is calculated.
[0153] Furthermore, for the above weighted process, that is, using the linear weighted method, assuming that the standardized instruction following, context inheritance and answer quality scores are X1', X2', X3', respectively, and the weights are w1, w2, w3, respectively, the comprehensive score S = w1×X1'+w2×X2'+w3×X3'.
[0154] In one possible scenario, in order to improve the accuracy of the weight information, cross-validation can also be performed, that is, the conversation data is divided into a training set and a validation set; then the validation score is calculated based on the validation set; and a confidence assessment is performed on the validation set to obtain a confidence score; then the correlation coefficient is determined by comparing the validation score and the confidence score, and the weight information is adjusted based on the correlation coefficient; and then the target score is determined based on the adjusted weight information.
[0155] It can be seen that through the cross-validation method, the data set is divided into a training set and a validation set, a comprehensive scoring system is constructed in the training set, and then the comprehensive score is calculated in the validation set, compared with the actual evaluation results (such as the manually labeled comprehensive quality level), and the correlation coefficient (such as the Pearson correlation coefficient) is calculated to evaluate the correlation between the comprehensive score and the actual evaluation results; thereby adjusting the weights and gradually iterating the scoring system.
[0156] In combination with the above embodiments, it can be known that by obtaining the dialogue data obtained through the question-and-answer of the target large model, the dialogue data includes multiple rounds of question-and-answer dialogues; then based on the follow-up relationship between the question-and-answer dialogues, the question-and-answer dialogues are sequentially extracted to obtain dialogue instructions, and the dialogue instructions are scored to obtain a first score; and the answer information corresponding to the question-and-answer dialogue is extracted, and the answer information is compared with the previous answer information to obtain a second score, and the previous answer information is determined based on the answer information of the previous round of the question-and-answer dialogue; and the answer information is evaluated for reply quality to obtain a third score; and then the first score, the second score and the third score are combined to obtain the target score, so as to evaluate the target large model through the target score. Thus, a multi-dimensional evaluation process is realized. Since multi-dimensional indicators are configured according to the characteristics of multi-round dialogues, automated evaluation is realized, which can greatly reduce the degree of manual participation, reduce the deviation caused by personal subjective views, and improve the accuracy of large model evaluation.
[0157] In order to better implement the above solution of the embodiment of the present application, the following also provides related devices for implementing the above solution. Figure 6 , Figure 6 This is a schematic diagram of the structure of a large model evaluation device provided in an embodiment of the present application. The evaluation device 600 includes:
[0158] An acquisition unit 601 is used to acquire dialogue data obtained through question-answering of a target large model, wherein the dialogue data includes multiple rounds of question-answering dialogues;
[0159] A determining unit 602 is used to extract instructions from the question-and-answer dialogues in sequence based on the follow-up relationship between the question-and-answer dialogues to obtain dialogue instructions, and score the dialogue instructions to obtain a first score;
[0160] The determining unit 602 is further configured to extract answer information corresponding to the question-and-answer dialogue, and compare the answer information with previous answer information to obtain a second score, wherein the previous answer information is determined based on answer information of a previous round of the question-and-answer dialogue;
[0161] The determining unit 602 is further configured to evaluate the quality of the answer information to obtain a third score;
[0162] The evaluation unit 603 is used to obtain a target score by combining the first score, the second score and the third score, so as to evaluate the target large model through the target score.
[0163] Optionally, in some possible implementations, the determination unit 602 is used to identify dialogue identifiers in the dialogue data to determine dialogue boundaries when extracting instructions from the question-and-answer dialogues in sequence based on the follow-up relationship between the question-and-answer dialogues to obtain dialogue instructions and scoring the dialogue instructions to obtain a first score; divide the dialogue data according to the dialogue boundaries, and sort them in a preset manner to obtain a text dialogue sequence; determine the instruction follow-up type corresponding to the question-and-answer dialogue according to the follow-up relationship between the question-and-answer dialogues in the text dialogue sequence; extract instructions from the question-and-answer dialogues in sequence based on the instruction follow-up type to obtain the dialogue instructions; and score the dialogue instructions through a comment model to obtain a first score.
[0164] Optionally, in some possible implementations, the determination unit 602 is used to extract instructions from the question-and-answer dialogue in sequence based on the instruction following type to obtain the dialogue instruction. When it is determined that the instruction following type is a multi-round following type, determine the preceding answer information and prompt word information corresponding to the question-and-answer dialogue based on the multi-round following type; merge the prompt word information and the preceding answer information to obtain the current dialogue content; and extract instructions from the current dialogue content according to preset rules to obtain the dialogue instruction.
[0165] Optionally, in some possible implementations, the determination unit 602 is used to extract instructions from the question-and-answer dialogue in sequence based on the instruction following type to obtain the dialogue instruction, and when it is determined that the instruction following type includes multiple types, determine combination information between the multiple types; configure the dialogue content corresponding to the question-and-answer dialogue based on the combination information to obtain the current dialogue content; and extract instructions from the current dialogue content according to preset rules to obtain the dialogue instruction.
[0166] Optionally, in some possible implementations, the determination unit 602 is used to obtain training corpus configured based on a multi-round question-and-answer dialogue; perform instruction extraction on the training corpus to obtain training instructions, and determine the instruction type corresponding to each training instruction; evaluate the training instructions of different instruction types based on preset dimensions to obtain training scores; combine the training scores and the training corpus to obtain training data; and train the comment model based on the training data.
[0167] Optionally, in some possible implementations, the determination unit 602 is used to extract key elements in the dialogue data in text order when extracting answer information corresponding to the question and answer dialogue and comparing the answer information with previous answer information to obtain a second score; collect the key elements to configure a key information library; extract the answer information corresponding to the question and answer dialogue, and match the previous answer information in the key information library according to the answer information to obtain a text information chain; and perform a conflict comparison between the answer information and the text information chain to obtain the second score.
[0168] Optionally, in some possible implementations, the determination unit 602 is used to perform part-of-speech tagging on the answer information and the text information chain to obtain part-of-speech tagging results when performing a conflict comparison between the answer information and the text information chain to obtain the second score, and determine first comparison information based on the part-of-speech tagging results; determine semantic vectors in the answer information and the text information chain, and determine second comparison information based on the differences indicated by the semantic vectors; determine logical information corresponding to the answer information and the text information chain, and determine third comparison information based on the degree of conflict indicated by the logical information; and combine the first comparison information, the second comparison information, and the third comparison information to obtain the second score.
[0169] Optionally, in some possible implementations, the determination unit 602 is used to determine the key information in the answer information when performing a reply quality evaluation on the answer information to obtain a third score; answer the question and answer dialogue through trusted resources to obtain a trusted reply; compare the key information with the trusted reply to obtain fourth comparison information; answer the question and answer dialogue through a trusted big model to obtain a reply sequence; determine fifth comparison information based on the order of the answer information in the reply sequence; and combine the fourth comparison information and the fifth comparison information to obtain the third score.
[0170] Optionally, in some possible implementations, the evaluation unit 603 is used to use the first score, the second score and the third score as criterion layers to configure a hierarchical model when obtaining a target score by combining the first score, the second score and the third score to evaluate the target large model through the target score; output weight information through the hierarchical model; linearly weight the first score, the second score and the third score based on the weight information to obtain the target score; and evaluate the target large model through the target score.
[0171] Optionally, in some possible implementations, the evaluation unit 603 is further used to divide the conversation data into a training set and a verification set; calculate a verification score based on the verification set; perform confidence evaluation on the verification set to obtain a confidence score; determine a correlation coefficient by comparing the verification score and the confidence score, and adjust the weight information based on the correlation coefficient; and determine a target score based on the adjusted weight information.
[0172] By obtaining the dialogue data obtained through the question-and-answer of the target large model, the dialogue data includes multiple rounds of question-and-answer dialogues; then extracting the instructions of the question-and-answer dialogue in turn based on the follow-up relationship between the question-and-answer dialogues to obtain dialogue instructions, and scoring the dialogue instructions to obtain a first score; and extracting the answer information corresponding to the question-and-answer dialogue, and comparing the answer information with the previous answer information to obtain a second score, the previous answer information is determined based on the answer information of the previous round of the question-and-answer dialogue; and evaluating the reply quality of the answer information to obtain a third score; and then combining the first score, the second score and the third score to obtain the target score, so as to evaluate the target large model through the target score. Thus, a multi-dimensional evaluation process is realized. Due to the multi-dimensional indicator configuration based on the characteristics of multi-round dialogues, automatic evaluation is realized, which can greatly reduce the degree of manual participation, reduce the deviation caused by personal subjective views, and improve the accuracy of large model evaluation.
[0173] The large model evaluation device provided in this embodiment belongs to the same application concept as the method provided in the above embodiments of this application, and can execute the large model evaluation method provided in any of the above embodiments of this application, and has the corresponding functional modules and beneficial effects of executing the large model evaluation method. For technical details not fully described in this embodiment, please refer to the specific processing content of the large model evaluation method provided in the above embodiments of this application, and will not be repeated here.
[0174] The functions implemented by the above acquisition unit 601, determination unit 602 and evaluation unit 603 may be implemented by the same or different processors respectively, which is not limited in the embodiment of the present application.
[0175] It should be understood that each unit in the above device can be implemented in the form of a processor calling software. For example, the device includes a processor, the processor is connected to a memory, and instructions are stored in the memory. The processor calls the instructions stored in the memory to implement any of the above methods or realize the functions of each unit of the device, wherein the processor can be a general-purpose processor, such as a CPU or a microprocessor, etc., and the memory can be a memory in the device or a memory outside the device. Alternatively, the unit in the device can be implemented in the form of a hardware circuit, and the functions of some or all units can be realized by designing the hardware circuit. The hardware circuit can be understood as one or more processors; for example, in one implementation, the hardware circuit is an ASIC, and the functions of some or all units above are realized by designing the logical relationship of the components in the circuit; for another example, in another implementation, the hardware circuit can be implemented by PLD, taking FPGA as an example, which can include a large number of logic gate circuits, and the connection relationship between the logic gate circuits is configured by a configuration file, so as to realize the functions of some or all units above. All units of the above device can be implemented in the form of a processor calling software, or in the form of a hardware circuit, or in the form of a processor calling software, and the remaining part can be implemented in the form of a hardware circuit.
[0176] In an embodiment of the present application, a processor is a circuit with the ability to process signals. In one implementation, the processor may be a circuit with the ability to read and run instructions, such as a CPU, a microprocessor, a GPU, or a DSP; in another implementation, the processor may implement certain functions through the logical relationship of a hardware circuit, and the logical relationship of the hardware circuit is fixed or reconfigurable, such as a hardware circuit implemented by an ASIC or PLD, such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document to implement the hardware circuit configuration can be understood as the process of the processor loading instructions to implement the functions of some or all of the above units. In addition, it can also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as an NPU, TPU, DPU, etc.
[0177] It can be seen that each unit in the above device can be one or more processors (or processing circuits) configured to implement the above method, such as: CPU, GPU, NPU, TPU, DPU, microprocessor, DSP, ASIC, FPGA, or a combination of at least two of these processor forms.
[0178] In addition, all or part of the units in the above device can be integrated together, or can be implemented independently. In one implementation, these units are integrated together and implemented in the form of a SOC. The SOC may include at least one processor for implementing any of the above methods or implementing the functions of each unit of the device. The type of the at least one processor may be different, for example, including a CPU and an FPGA, a CPU and an artificial intelligence processor, a CPU and a GPU, etc.
[0179] The embodiment of the present application also provides a large model evaluation device; the large model evaluation device includes a processor and an input-output component. The input-output component is used to obtain conversation data;
[0180] The processor is used to evaluate the big model corresponding to the dialogue data obtained by the input-output component by executing any one of the big model evaluation methods in the above embodiments.
[0181] The above-mentioned interface circuit can be any interface circuit that can realize the data communication function, for example, it can be a USB interface circuit, a Type-C interface circuit, a serial port circuit, a PCIE circuit, etc.
[0182] Optionally, the embodiment of the present application further provides a system, wherein the interactive client is used to obtain the conversation data and send the conversation data to the server, and to display the large model evaluation result output by the server;
[0183] The server is used to evaluate the big model corresponding to the conversation data obtained by the interactive client by executing the big model evaluation method described in any one of the above embodiments.
[0184] Another embodiment of the present application also provides a large model evaluation device, see Figure 7 As shown, the device includes:
[0185] Memory 700 and processor 710;
[0186] The memory 700 is connected to the processor 710 and is used to store programs;
[0187] The processor 710 is used to implement the large model evaluation method disclosed in any of the above embodiments by running the program stored in the memory 700.
[0188] Specifically, the above-mentioned large model evaluation device may further include: a bus, a communication interface 720 , an input device 730 and an output device 740 .
[0189] The processor 710, the memory 700, the communication interface 720, the input device 730 and the output device 740 are connected to each other via a bus.
[0190] A bus may include a pathway that transfers information between components of a computer system.
[0191] Processor 710 may be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the scheme of the present invention. It may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0192] The processor 710 may include a main processor, and may also include a baseband chip, a modem, and the like.
[0193] The memory 700 stores a program for executing the technical solution of the present invention, and may also store an operating system and other key services. Specifically, the program may include a program code, and the program code includes a computer operation instruction. More specifically, the memory 700 may include a read-only memory (ROM), other types of static storage devices that can store static information and instructions, a random access memory (RAM), other types of dynamic storage devices that can store information and instructions, a disk storage, a flash, and the like.
[0194] The input device 730 may include a device for receiving data and information input by a user, such as a keyboard, a mouse, a camera, a scanner, a light pen, a voice input device, a touch screen, a pedometer, or a gravity sensor.
[0195] Output device 740 may include devices that allow information to be output to a user, such as a display screen, printer, speaker, etc.
[0196] The communication interface 720 may include any transceiver or the like to communicate with other devices or communication networks, such as Ethernet, Radio Access Network (RAN), Wireless Local Area Network (WLAN), etc.
[0197] The processor 710 executes the program stored in the memory 700 and calls other devices, which can be used to implement each step of any large model evaluation method provided in the above embodiments of the present application.
[0198] In addition to the above-mentioned methods and devices, an embodiment of the present application may also be a computer program product, which includes computer program instructions, which, when executed by a processor, enable the processor to execute the steps of the large model evaluation method according to various embodiments of the present application described in any of the above embodiments of this specification.
[0199] The computer program product may be written in any combination of one or more programming languages to write program codes for performing the operations of the embodiments of the present application, including object-oriented programming languages, such as Java, C++, etc., and conventional procedural programming languages, such as "C" language or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as an independent software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0200] In addition, the embodiment of the present application may also be a storage medium on which a computer program is stored. The computer program is executed by a processor to execute the steps of the large model evaluation method according to various embodiments of the present application described in any of the above embodiments of this specification, and specifically the following steps may be implemented:
[0201] Obtaining dialogue data obtained through question-answering of the target large model, where the dialogue data includes multiple rounds of question-answering dialogues;
[0202] Based on the follow-up relationship between the question-answer dialogues, the question-answer dialogues are sequentially extracted to obtain dialogue instructions, and the dialogue instructions are scored to obtain a first score;
[0203] Extracting answer information corresponding to the question-and-answer dialogue, and comparing the answer information with previous answer information to obtain a second score, where the previous answer information is determined based on answer information of a previous round of the question-and-answer dialogue;
[0204] Evaluate the quality of the answer information to obtain a third score;
[0205] The first score, the second score, and the third score are combined to obtain a target score, so as to evaluate the target macro model through the target score.
[0206] For the aforementioned method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the order of the actions described, because according to the present application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.
[0207] It should be noted that each embodiment in this specification is described in a progressive manner, and each embodiment focuses on the differences from other embodiments, and the same or similar parts between the embodiments can be referred to each other. For the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0208] The steps in the methods of each embodiment of the present application can be adjusted in sequence, combined and deleted according to actual needs, and the technical features recorded in each embodiment can be replaced or combined.
[0209] The modules and sub-modules in the devices and terminals of the various embodiments of the present application can be combined, divided and deleted according to actual needs.
[0210] In the several embodiments provided in the present application, it should be understood that the disclosed terminals, devices and methods can be implemented in other ways. For example, the terminal embodiments described above are only schematic, for example, the division of modules or submodules is only a logical function division, and there may be other division methods in actual implementation, for example, multiple submodules or modules can be combined or integrated into another module, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or modules, which can be electrical, mechanical or other forms.
[0211] The modules or submodules described as separate components may or may not be physically separated, and the components of the modules or submodules may or may not be physical modules or submodules, that is, they may be located in one place, or they may be distributed on multiple network modules or submodules. Some or all of the modules or submodules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0212] In addition, each functional module or submodule in each embodiment of the present application may be integrated into one processing module, or each module or submodule may exist physically separately, or two or more modules or submodules may be integrated into one module. The above-mentioned integrated modules or submodules may be implemented in the form of hardware or in the form of software functional modules or submodules.
[0213] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0214] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly by hardware, software units executed by a processor, or a combination of the two. The software units may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0215] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0216] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A large model evaluation method, characterized in that: include: Acquire dialogue data obtained through question-answering of a target large model, wherein the dialogue data includes multiple rounds of question-answering dialogues; Extracting instructions from the question-and-answer dialogues in sequence based on the follow-up relationship between the question-and-answer dialogues to obtain dialogue instructions, and scoring the dialogue instructions to obtain a first score; Extracting answer information corresponding to the question-and-answer dialogue, and comparing the answer information with previous answer information to obtain a second score, wherein the previous answer information is determined based on answer information of a previous round of the question-and-answer dialogue; Performing a response quality evaluation on the answer information to obtain a third score; A target score is obtained by combining the first score, the second score and the third score, so as to evaluate the target macro model by using the target score.
2. The method according to claim 1, characterized in that The extracting instructions from the question-and-answer dialogues in sequence based on the follow-up relationship between the question-and-answer dialogues to obtain dialogue instructions, and scoring the dialogue instructions to obtain a first score, includes: Identifying a conversation identifier in the conversation data to determine a conversation boundary; Dividing the conversation data according to the conversation boundary and sorting them according to a preset method to obtain a text conversation sequence; Determining the instruction follow-up type corresponding to the question-and-answer dialogue according to the follow-up relationship between the question-and-answer dialogues in the text dialogue sequence; Extracting instructions from the question-and-answer dialogue in sequence based on the instruction follow-up type to obtain the dialogue instruction; The dialogue instruction is scored using a comment model to obtain a first score.
3. The method according to claim 2, characterized in that The extracting instructions from the question-and-answer dialogue in sequence based on the instruction follow-up type to obtain the dialogue instruction includes: In the case where it is determined that the instruction following type is a multi-round following type, determining the preceding answer information and the prompt word information corresponding to the question-and-answer dialogue based on the multi-round following type; Combining the prompt word information and the previous answer information to obtain the current dialogue content; The current conversation content is subjected to instruction extraction according to preset rules to obtain the conversation instruction.
4. The method according to claim 2, characterized in that: The extracting instructions from the question-and-answer dialogue in sequence based on the instruction follow-up type to obtain the dialogue instruction includes: In the case where it is determined that the instruction following type includes multiple types, determining combination information between the multiple types; configuring the dialogue content corresponding to the question-and-answer dialogue based on the combination information to obtain current dialogue content; The current conversation content is subjected to instruction extraction according to preset rules to obtain the conversation instruction.
5. The method according to claim 2, characterized in that: The training process of the review model includes: Obtain training corpus based on multi-round question-answering dialogue configuration; Extracting instructions from the training corpus to obtain training instructions, and determining the instruction type corresponding to each training instruction; Evaluate the training instructions of different instruction types based on preset dimensions to obtain training scores; Combining the training score with the training corpus to obtain training data; The review model is trained based on the training data.
6. The method according to claim 1, characterized in that The extracting answer information corresponding to the question-answer dialogue, and comparing the answer information with previous answer information to obtain a second score, includes: Extracting key elements from the conversation data in text order; The key elements are collected to configure a key information library; Extracting answer information corresponding to the question-and-answer dialogue, and matching the preceding answer information in the key information database according to the answer information to obtain a text information chain; The answer information is compared with the text information link for conflicts to obtain the second score.
7. The method according to claim 6, characterized in that The step of performing a conflict comparison between the answer information and the text information link to obtain the second score includes: Performing part-of-speech tagging on the answer information and the text information chain to obtain a part-of-speech tagging result, and determining first comparison information according to the part-of-speech tagging result; Determine semantic vectors in the answer information and the text information chain, and determine second comparison information according to differences indicated by the semantic vectors; Determine logic information corresponding to the answer information and the text information chain, and determine third comparison information according to the conflict degree indicated by the logic information; The second score is obtained by combining the first comparison information, the second comparison information and the third comparison information.
8. The method according to claim 1, characterized in that The step of evaluating the quality of the answer information to obtain a third score includes: determining key information in the answer information; Answering the question-and-answer dialogue using a trusted resource to obtain a trusted response; Comparing the key information with the confidence reply to obtain fourth comparison information; Answering the question-answer dialogue using a large confidence model to obtain a response sequence; Determining fifth comparison information according to the order of the answer information in the reply sequence; The fourth comparison information and the fifth comparison information are combined to obtain the third score.
9. The method according to claim 1, characterized in that: The step of combining the first score, the second score, and the third score to obtain a target score, so as to evaluate the target macro model by using the target score, comprises: configuring a hierarchical model by using the first score, the second score, and the third score as criterion layers; Output weight information through the hierarchical structure model; linearly weighting the first score, the second score, and the third score based on the weight information to obtain the target score; The target macro model is evaluated by the target score.
10. The method according to claim 9, characterized in that The method further comprises: Divide the conversation data into training set and validation set; Calculating a validation score based on the validation set; Performing confidence evaluation on the verification set to obtain a confidence score; Determining a correlation coefficient by comparing the verification score and the confidence score, and adjusting the weight information based on the correlation coefficient; A target score is determined based on the adjusted weight information.
11. A large model evaluation device, characterized in that: include: An acquisition unit, configured to acquire dialogue data obtained through question-answering of a target large model, wherein the dialogue data includes multiple rounds of question-answering dialogues; a determining unit, configured to extract instructions from the question-and-answer dialogues in sequence based on the follow-up relationship between the question-and-answer dialogues to obtain dialogue instructions, and score the dialogue instructions to obtain a first score; The determination unit is further configured to extract answer information corresponding to the question-and-answer dialogue, and compare the answer information with previous answer information to obtain a second score, wherein the previous answer information is determined based on answer information of a previous round of the question-and-answer dialogue; The determining unit is further configured to perform a reply quality evaluation on the answer information to obtain a third score; An evaluation unit is used to combine the first score, the second score and the third score to obtain a target score, so as to evaluate the target macro model through the target score.
12. A large model evaluation device, comprising an input and output component and a processor; The input and output components are used to obtain conversation data; The processor is used to evaluate the big model corresponding to the dialogue data obtained by the input-output component by executing the big model evaluation method as described in any one of claims 1 to 10.
13. A large model evaluation system, including an interactive client and a server; The interactive client is used to obtain the conversation data and send the conversation data to the server, and to display the large model evaluation result output by the server; The server is used to evaluate the big model corresponding to the dialogue data obtained by the interactive client by executing the big model evaluation method described in any one of claims 1 to 10.
14. A computer program product, characterized in that include: A computer program, wherein when the computer program is executed by a processor, the large model evaluation method according to any one of claims 1 to 10 is implemented.
Citation Information
Cited By
Large model safety performance evaluation method based on analytic hierarchy process and Shapley value fusion
CN120337234A
Medical large model evaluation system
CN120408119A
Large model and agent evaluation method and device based on multi-round session data set
CN120653950A
Question and answer model evaluation method and device
CN120706581A
Question-answering model evaluation method and device
CN120706581B