Model evaluation method, apparatus, and electronic device

By integrating model evaluation into game applications and using player roles to assess the performance of non-player roles, the inaccuracy of evaluation results caused by the limitations and subjectivity of users in user experience evaluation is solved, resulting in more objective and reliable evaluation results.

CN119782142BActive Publication Date: 2025-11-21NETEASE (HANGZHOU) NETWORK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411628585.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-13
Publication Date
2025-11-21
Estimated Expiration
2044-11-13

AI Technical Summary

Technical Problem

Existing user experience evaluation methods suffer from a lack of objectivity and credibility due to the limited number of users participating in the evaluation and the subjectivity of users.

Method used

将模型评测融入到游戏应用中,通过待评测模型控制非玩家角色执行游戏任务,并由玩家角色在无感知的情况下进行评估,收集评估数据以计算评测结果。

Benefits of technology

It improves the objectivity and credibility of the evaluation results, overcomes the limitation of the limited number of users participating in the evaluation, and reduces the influence of user subjectivity and personal bias.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119782142B_ABST
    Figure CN119782142B_ABST
Patent Text Reader

Abstract

The application discloses a model evaluation method, system and device, electronic equipment and computer readable storage medium. The method comprises the following steps: obtaining a game application preset for at least one evaluation dimension; controlling a non-player character to perform a first game task by a to-be-evaluated model; providing a second game task to each player character, wherein the second game task comprises evaluating the non-player character under each evaluation dimension according to the execution of the first game task by the non-player character; collecting evaluation data of the non-player character under each evaluation dimension by each player character, and calculating an evaluation result of the to-be-evaluated model under each evaluation dimension according to the evaluation data. The method solves the technical problem that the evaluation result lacks objectivity and credibility due to the limited number of users participating in the evaluation and the subjectivity of the users in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a model evaluation method, apparatus, electronic device, and computer-readable storage medium. Background Technology

[0002] With the rapid development of computer technology, especially the significant improvement in computing power and the arrival of the big data era, neural network models (hereinafter referred to as "models") have gradually become an indispensable technical tool in many fields, resulting in an increasing variety and number of models. Different models have different advantages, and model users usually evaluate multiple models based on their own business scenarios to select the model that meets their business needs.

[0003] Large language models are advanced artificial intelligence models. Their vast number of parameters allows them to support a wider range of application scenarios, making them a popular model type among application developers across various fields. Evaluation of large language models involves multiple aspects, with user experience evaluation primarily based on human assessments or user surveys. However, these evaluation methods often lack objectivity and credibility due to the limited number of participating users and user subjectivity.

[0004] Therefore, there is an urgent need for a model evaluation method that can overcome the influencing factors that affect the evaluation results in existing user experience evaluations, so as to make the model evaluation results objective and highly credible. Summary of the Invention

[0005] This application provides a model evaluation method, apparatus, electronic device, and computer-readable storage medium to solve the technical problem that existing user experience evaluation methods lack objectivity and credibility due to factors such as the limited number of users participating in the evaluation and the subjectivity of users.

[0006] In a first aspect, embodiments of this application provide a model evaluation method, the method comprising: acquiring a preset game application for at least one evaluation dimension, the evaluation dimension being used to characterize preset indicators for evaluating a model to be evaluated, the game application including at least one non-player character and multiple player characters; controlling the non-player character to execute a first game task in the application process of the game application through the model to be evaluated; in response to the completion of the first game task, providing a second game task to each player character, the second game task including evaluating the non-player character under each evaluation dimension based on the non-player character's execution of the first game task; collecting evaluation data of each player character on the non-player character under each evaluation dimension, and calculating the evaluation result of the model to be evaluated under each evaluation dimension based on the evaluation data.

[0007] Secondly, embodiments of this application provide a model evaluation system, the system comprising: a model supply module, a first model evaluation module, and a second model evaluation module; the model supply module is used to encapsulate multiple models to be evaluated into a unified interface and provide the interface to the first model evaluation module; the first model evaluation module is used to obtain a preset game application for at least one evaluation dimension, the evaluation dimension being used to characterize preset evaluation indicators for evaluating the model to be evaluated, the game application including at least one non-player character and multiple player characters; controlling the non-player character to execute a first game task in the application process of the game application through the model to be evaluated; in response to the completion of the first game task, providing a second game task to each player character, the second game task including, according to the non-player character's... The first game task execution status is evaluated on the non-player character under various evaluation dimensions; evaluation data of each player character on the non-player character under various evaluation dimensions is collected, and the evaluation results of the model to be evaluated under various evaluation dimensions are calculated based on the evaluation data; the second model evaluation module is used to connect the non-player character controlled by the model to be evaluated to a crowdsourcing platform, so that the crowdsourcing platform users evaluate the non-player character on various evaluation dimensions based on the execution status of the first game task; evaluation data of the crowdsourcing platform users on the non-player character under various evaluation dimensions is collected, and the crowdsourcing evaluation results of the model to be evaluated under various evaluation dimensions are calculated based on the evaluation data; the evaluation results and the crowdsourcing evaluation results are compared to provide credibility verification of the evaluation results.

[0008] Thirdly, embodiments of this application provide a model evaluation apparatus, comprising: a game application acquisition unit, a game task execution unit, a game task provisioning unit, and an evaluation result calculation unit; the game application acquisition unit is used to acquire a game application preset for at least one evaluation dimension, the evaluation dimension being used to characterize preset indicators for evaluating the model to be evaluated, the game application including at least one non-player character and multiple player characters; the game task execution unit is used to control the non-player character to execute a first game task in the application process of the game application through the model to be evaluated; the game task provisioning unit is used to provide a second game task to each player character in response to the completion of the first game task, the second game task including evaluating the non-player character under each evaluation dimension based on the non-player character's execution of the first game task; the evaluation result calculation unit is used to collect evaluation data of each player character on the non-player character under each evaluation dimension, and calculate the evaluation result of the model to be evaluated under each evaluation dimension based on the evaluation data.

[0009] Fourthly, embodiments of this application provide an electronic device, including: a memory and a processor; the memory is used to store one or more computer instructions; the processor is used to execute the one or more computer instructions to implement the above-described method.

[0010] Fifthly, embodiments of this application provide a computer-readable storage medium storing one or more computer instructions that, when executed by a processor, perform the above-described method.

[0011] Compared with existing technologies, the model evaluation method provided in this application includes: acquiring a pre-defined game application for at least one evaluation dimension, wherein the evaluation dimension is used to characterize the pre-defined indicators for evaluating the model to be evaluated, and the game application includes at least one non-player character and multiple player characters; controlling the non-player character to execute a first game task in the application process of the game application through the model to be evaluated; in response to the completion of the first game task, providing a second game task to each player character, wherein the second game task includes evaluating the non-player character under each evaluation dimension based on the non-player character's execution of the first game task; collecting the evaluation data of each player character on the non-player character under each evaluation dimension, and calculating the evaluation results of the model to be evaluated under each evaluation dimension based on the evaluation data. First, this method integrates model evaluation into the game application, issuing game tasks to both non-player and player characters. Player characters then evaluate the performance of non-player characters in executing game tasks (the first game task) while performing their own tasks (the second game task). This allows players to participate in model evaluation without being aware of it, improving the objectivity of the evaluation results. Second, by integrating model evaluation into the game application, and given the large number and diverse profiles of participating players, this method overcomes the limitation of a limited number of users involved in the evaluation. It also reduces the influence of user subjectivity, personal bias, self-selection bias, and personal comprehension bias on the evaluation results, comprehensively improving the credibility of the evaluation results. Therefore, the model evaluation method provided in this application is a method for unconsciously evaluating models by integrating model evaluation into game applications. It solves the technical problem of existing user experience evaluation methods lacking objectivity and credibility due to the limited number of participating users and user subjectivity. Attached Figure Description

[0012] Figure 1 This is an application system diagram of the model evaluation method provided in the embodiments of this application;

[0013] Figure 2 This is a flowchart of the model evaluation method provided in the first embodiment of this application;

[0014] Figure 3This is a schematic diagram of the model evaluation system provided in the second embodiment of this application;

[0015] Figure 4 This is a schematic diagram of the model evaluation device provided in the third embodiment of this application;

[0016] Figure 5 This is a schematic diagram of the structure of the electronic device provided in the fourth embodiment of this application. Detailed Implementation

[0017] Many specific details are set forth in the following description to provide a full understanding of this application. However, this application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this application; therefore, this application is not limited to the specific embodiments disclosed below.

[0018] With the rapid development of computer technology, especially the significant improvement in computing power and the arrival of the big data era, neural network models (hereinafter referred to as "models") have gradually become an indispensable technical tool in many fields. A model is a computer neural network that mimics the structure and function of the human brain. By learning complex patterns and relationships in data, it can perform tasks such as classification, regression, clustering, and prediction. As various business sectors demand models, the types and numbers of models are increasing. Different types of models have different advantages, giving model users more choices. Model users typically combine their own business scenarios, connect to multiple models, evaluate each model, and select one or more models that meet their business needs. This allows for model scheduling and service provision tailored to different user requirements during the business application phase.

[0019] Large Language Models (LLMs) are advanced artificial intelligence models built using deep learning techniques. They learn the grammatical, semantic, and stylistic characteristics of a language through training on massive amounts of text data, thereby achieving the understanding and generation of natural language. The core of LLMs lies in their sheer scale, possessing millions or even billions of model parameters. This massive scale allows LLMs to capture more complex language structures and patterns, thus supporting a wider range of applications, such as natural language understanding, natural language generation, dialogue, and translation. Therefore, LLMs have become a popular model type among application developers across various fields.

[0020] The evaluation of large language models involves multiple aspects, such as model performance, robustness, fairness, efficiency, and user experience. Among these, user experience evaluation is a comprehensive assessment of the large language model's linguistic logic and comprehension, response speed, and accuracy. It allows users to fully understand the model's performance in practical applications, thus holding a crucial position in model evaluation. Currently, user experience evaluation is primarily based on human assessments or user surveys. Human assessments involve subjective evaluations of the model's output by human reviewers, while user surveys collect user feedback on model performance and experience through questionnaires and other methods. While both methods provide users with user experience evaluation results to some extent, they each have certain limitations. For example, human evaluation methods have the following shortcomings: First, they are highly subjective, as human reviewers are inherently subjective and may differ significantly from one another, often leading to inconsistent evaluation results. Second, they are costly, requiring substantial human and time investment, especially for large-scale evaluations. Third, they have poor repeatability, as the subjectivity and human factors of human reviewers make it difficult to guarantee that evaluations can be repeated, and subsequent evaluations may not be able to fully replicate the conditions and results of the initial evaluation. Fourth, they carry the risk of bias, as the personal preferences and backgrounds of human reviewers may influence the evaluation results, increasing the risk of bias. Fifth, they are limited in scale, as human resources typically limit the number of models and scenarios that can be evaluated, making it difficult to cover all possible use cases. Sixth, they are time-consuming, as the entire evaluation process can be very time-consuming, limiting the frequency and timeliness of model evaluations. For example, user surveys have the following shortcomings: First, low response rate: users may be unwilling to spend time completing questionnaires, leading to a low response rate for model evaluation and affecting the representativeness and reliability of the evaluation results. Second, self-selection bias: only users with strong positive or negative experiences with the model are more likely to participate in the survey, which undoubtedly leads to self-selection bias in the evaluation results. Third, quality issues with feedback data: users may not fill out the questionnaire carefully or may give random answers, affecting the quality of the feedback data. Fourth, comprehension bias: different users may have different understandings of the questionnaire questions, which will lead to problems with the accuracy and consistency of the answers. Fifth, difficulty in quantitative analysis: subjective user feedback is difficult to analyze quantitatively, especially open-ended questions, making it difficult to systematically summarize and generalize. Sixth, feedback delay: user surveys are often conducted after users have experienced the model, and cannot reflect user experience in real time, resulting in timeliness issues. Seventh, limited depth and breadth: questionnaire surveys may not be able to explore all possible user experience issues in depth, making it difficult to comprehensively cover the diverse needs and feedback of users. In summary, existing evaluation methods suffer from a lack of representativeness and credibility in the evaluation results due to factors such as the limited number of participating users and user subjectivity.

[0021] In view of this, this application provides a model evaluation method, which is a method for unconsciously evaluating models by integrating model evaluation into game applications. First, this method integrates model evaluation into the game application, issuing game tasks to both non-player and player characters. Player characters then evaluate the performance of non-player characters in executing game tasks (the first game task) while performing the second game task. This allows players to participate in model evaluation unconsciously, improving the objectivity of the evaluation results. Second, by integrating model evaluation into the game application, and given the large number and diverse profiles of players participating in the game application, this method overcomes the limitation of a limited number of participating users and significantly reduces the influence of user subjectivity, personal bias, self-selection bias, and personal comprehension bias on the evaluation results, thus comprehensively improving the credibility of the evaluation results.

[0022] The model evaluation method, apparatus, electronic device, and computer-readable storage medium described in this application will be further described in detail below with reference to specific embodiments and accompanying drawings.

[0023] Figure 1 This is an application system diagram of the model evaluation method provided in the embodiments of this application. For example... Figure 1 As shown, the system includes a first user terminal 101, a second user terminal 102, a first server terminal 103, and a second server terminal 104. The first user terminal 101 and the second user terminal 102 can be any device such as a smartphone, tablet, laptop, desktop computer, or personal digital assistant (PDA). Specifically, the first user terminal 101 is the terminal device corresponding to the business application developer (i.e., the model caller), and the second user terminal 102 is the terminal device corresponding to the business application user (i.e., the user). The first server terminal 103 deploys multiple models, and the second server terminal 104 deploys business applications and model evaluation methods. During model evaluation, the second server 104 integrates various models deployed in the first server 103 into the business application and provides the business application to the second user terminal 102. The user generates usage data based on the business application on the second user terminal 102. The second user terminal 102 sends the generated usage data to the second server 104. Based on the received usage data, the second server 104 generates evaluation results for the models and sends the evaluation results to the first user terminal 101, so that business application developers can select models that meet the business application requirements based on the evaluation results.

[0024] The first embodiment of this application provides a model evaluation method, which is deployed on... Figure 1The second server 104 shown is used to provide model evaluation services to application developers. The method provided in this embodiment is particularly suitable for user experience evaluation of large language models.

[0025] Figure 2 This is a flowchart of the model evaluation method provided in this embodiment. The following is in conjunction with... Figure 2 The model evaluation method provided in this embodiment will be described in detail. The embodiments described below are used to explain the technical solutions of this application and are not intended to limit actual use.

[0026] like Figure 2 As shown, the model evaluation method provided in this embodiment includes the following steps S210 to S240:

[0027] Step S210: Obtain a preset game application for at least one evaluation dimension. The evaluation dimension is used to characterize the preset indicators for evaluating the evaluation model. The game application includes at least one non-player character and multiple player characters.

[0028] The evaluation dimensions can be understood as the metrics used to evaluate the performance of a model. A series of metrics are set from multiple angles or aspects, and the model is evaluated for each metric. This can comprehensively measure the performance of the model and help the model user to have a more comprehensive understanding of the model's advantages and disadvantages. User experience evaluation of large language models typically includes the following dimensions: First, model dialogue ability, assessing the coherence, fluency, consistency, and extensibility of the model's dialogue; Second, model persona consistency, assessing consistency in speech (i.e., language style and expression habits), behavior, knowledge, and knowledge illusion; Third, model scenario consistency, assessing the adaptability of the model persona to situations where there is a conflict between the model persona and the application scenario; Fourth, model worldview consistency, assessing the consistency between the current dialogue and the worldview background or the worldview background provided by the user; Fifth, the expressiveness of the model's output content, assessing the themes, decisions, plot diversity, persona expressiveness, and content expressiveness of the model's output content; Sixth, model subjective feeling, assessing the model's anthropomorphism, human-like performance, communication skills, subjective comfort, sense of companionship, and empathy; Seventh, the safety of the model's output content, assessing the absence of harmful guidance in the model's output content.

[0029] In one optional implementation, the evaluation dimensions described in this embodiment include at least one of the following dimensions: language interaction capability, language-role consistency, language-scene consistency, language content expressiveness, character anthropomorphism, and language content security.

[0030] The game application is pre-designed by the model caller (e.g., a game developer) for one or more evaluation dimensions. Specifically, the model caller designs a game application that can obtain evaluation data for the evaluation dimensions they want to understand. For example, if the model caller wants to understand the model's intelligence level, they would design a riddle-like game application to obtain evaluation data on the correctness of players' answers to the riddles output by the model, thereby understanding the intelligence level of the model being evaluated.

[0031] In this embodiment, the game application includes at least one preset non-player character and multiple player characters. The non-player character is used to evaluate the reception model, while the player characters are used to connect with real players participating in the game application. The purpose of acquiring the game application is to provide an evaluation scenario for the model evaluation.

[0032] Step S220: Control the non-player character to perform the first game task in the application process of the game application through the model to be evaluated.

[0033] In one optional implementation, the non-player character is controlled by the model to be evaluated to execute the first game task in the application process of the game application. Specifically, this may include: in response to the start command of the game application, connecting the model to be evaluated to the game application so that the model to be evaluated controls the non-player character to execute the first game task in the application process of the game application.

[0034] The opening instructions can be understood as players based on... Figure 1 The instruction information sent by the second user terminal 102 to the second server 104 to start a game match can be generated by the second user terminal 102 based on the player's initial actions. For example, the second user terminal 102 generates an initial action in response to the player opening the game application and clicking the initial action control, and sends this initial action to the second server 104. Upon receiving the initial action, the second server 104 runs the game application and, based on a model evaluation method, connects the model to be evaluated into the game application to control the non-player character within the game application.

[0035] In one alternative implementation, the model to be evaluated includes multiple models, and based on this, the game application also includes multiple non-player characters.

[0036] Controlling a non-player character to perform the first game task in the application process of a game application by the model to be evaluated can specifically include: connecting multiple models to be evaluated to the game application and establishing a corresponding relationship with multiple non-player characters, so that each model to be evaluated controls a non-player character to perform the first game task in the application process of the game application.

[0037] When there are multiple models to be evaluated, they can be integrated into the game application together, and each model can be mapped one-to-one with a non-player character in the game application, so that one model controls one non-player character to perform game tasks in the application process.

[0038] In one optional implementation, the mapping between multiple models to be evaluated and multiple non-player characters is established randomly. Specifically, multiple models to be evaluated are integrated into the game application, and each model is randomly assigned a non-player character, establishing a mapping between the model and the assigned non-player character. In another optional implementation, the mapping between multiple models to be evaluated and multiple non-player characters is established based on a preset mapping table. Specifically, a mapping table recording the binding relationships between multiple models to be evaluated and multiple non-player characters is pre-built. When multiple models to be evaluated are integrated into the game application, information is read from the mapping table, and each model to be evaluated is bound to a non-player character with a binding relationship, establishing a mapping between the model to be evaluated and the bound non-player character.

[0039] In one optional implementation, multiple models to be evaluated share a unified interface. The interface defines a set of rules and specifications to enable interaction and communication between different software components or systems. In this embodiment, the interface refers to the rules and specifications for integrating the models to be evaluated into the game application. When there are multiple models to be evaluated, they are pre-encapsulated into a unified interface to facilitate model integration. Optionally, multiple open-source and closed-source models are encapsulated into a unified interface, exposing model parameters such as persona descriptions, maximum response length, response diversity, memory rounds, and knowledge bases to maximize the customization needs of model callers. Optionally, the encapsulated unified model interface can be further encapsulated into a unified model interface installation package to enhance its reusability. Optionally, the encapsulated unified model interface installation package can be further converted into an AOP SDK (Software Development Kit) using AOP (Agent Oriented Programming) for model callers to use.

[0040] Based on this, multiple models to be evaluated are integrated into the game application, and corresponding relationships are established with multiple non-player characters. Specifically, this may include controlling the game application to call multiple models to be evaluated based on a unified interface, and binding each model to a non-player character. In this embodiment, each model to be evaluated can control the non-player character bound to it to execute the first game task in the application process of the game application.

[0041] The first game task can be understood as a game task provided to a non-player character. This task can be preset during the game application development phase or input by the player through the second user terminal 102. The model to be evaluated controls the non-player character to execute the first game task, interacting with the player character. This allows the player to perceive the performance level of the model to be evaluated from the execution of the first game task, thus providing evaluation data. In an optional implementation, the first game task can be a natural language interaction with the player character. For example, the player character poses a riddle, the model to be evaluated parses the riddle, and controls the non-player character to provide the answer. Another example is that the player character provides a hint, the model to be evaluated parses the hint based on the hint and the dialogue context, and controls the non-player character to provide supplementary content to the hint.

[0042] In step S230, in response to the completion of the first game task, a second game task is provided to each player character. The second game task includes evaluating the non-player characters in various evaluation dimensions based on their performance in the first game task.

[0043] The second game task can be understood as a game task provided to the player character, corresponding to the first game task. The player character provides evaluation data for the non-player character through the second game task. Specifically, the second game task evaluates the non-player character across various evaluation dimensions based on their performance in the first game task. This evaluation data serves as the basis for calculating the evaluation results of the model to be evaluated later. In an optional implementation, the first game task could be natural language interaction with the player character. In this case, the second game task would evaluate the non-player character across various evaluation dimensions based on their interaction with the player character. For example, if the first game task involves the non-player character guessing a riddle posed by the player character, the second game task could include evaluating the correctness of the player character's answer to the riddle. Or, if the first game task involves the non-player character supplementing a hint given by the player character, the second game task could include evaluating the anthropomorphism of the supplementary content provided by the player character.

[0044] Optionally, the performance of non-player characters in the first game task can include not only the result of their performance but also the process of their performance, depending on the preset evaluation dimensions. For example, if the evaluation dimension is the expressiveness of the model's output, then the basis for evaluating the non-player character is the result of their performance in the first game task, such as whether the answer they provide is correct. Alternatively, if the evaluation dimension is the model's reaction speed, then the basis for evaluating the non-player character is the process of their performance in the first game task, such as the time it takes for the non-player character to provide the answer after the player character provides the riddle.

[0045] Optionally, non-player characters can be evaluated by scoring them across various evaluation dimensions based on their performance in the first game task. For example, a non-player character could be scored 8 points for the expressiveness of the model's output and 5 points for the model's reaction speed. Alternatively, non-player characters can be evaluated by ranking them across various evaluation dimensions based on their performance in the first game task. For example, with 5 non-player characters corresponding to 5 models to be evaluated, the ranking of non-player characters based on the expressiveness of their output could be [Non-player character 2, Non-player character 5, Non-player character 1, Non-player character 3, Non-player character 4], and the ranking of non-player characters based on the model's reaction speed could be [Non-player character 2, Non-player character 1, Non-player character 4, Non-player character 3, Non-player character 5]. The specific implementation method is not limited here.

[0046] Step S240: Collect the evaluation data of each player character on the non-player character under each evaluation dimension, and calculate the evaluation results of the model to be evaluated under each evaluation dimension based on the evaluation data.

[0047] In one optional implementation, all operations and results performed by non-player characters and player characters during the game application's process are recorded in the game log. Based on this, the evaluation data of each player character for non-player characters across various evaluation dimensions can be collected from the game logs corresponding to the player accounts controlling each player character. This evaluation data can be a score for a non-player character in a certain evaluation dimension, or a ranking of multiple non-player characters in a certain evaluation dimension; there are no specific limitations.

[0048] Calculating the evaluation results of the model under evaluation in each evaluation dimension based on the evaluation data can involve summing or averaging multiple evaluation data points for each evaluation dimension. In this embodiment, the evaluation data obtained after processing multiple evaluation data points for a certain evaluation dimension is used as the evaluation result of the model under evaluation in that dimension. For example, in a game application, there are 10 player characters. These 10 player characters give non-player characters scores of 8, 7, 8, 10, 5, 7, 6, 8, 9, and 9 points respectively in the dimension of the expressiveness of the model's output content. Then, the evaluation result of the model under evaluation controlling the non-player character in the dimension of the expressiveness of the model's output content is 7.7 points (the average of the above ten scores).

[0049] In one optional implementation, after calculating the evaluation results of the model under evaluation across each evaluation dimension, the comprehensive evaluation result of the model can be calculated through methods such as summation, averaging, or weighted summation. For example, if the evaluation result of the model under evaluation is calculated as 8 points in the model's dialogue ability dimension, 5 points in the model's persona consistency dimension, 10 points in the expressiveness of the model's output content dimension, and 10 points in the security dimension of the model's output content, then the comprehensive evaluation result of the model under evaluation is 8.25 points (the average of the above four evaluation results).

[0050] In one optional implementation, when there are multiple evaluation models, the evaluation data of each player character on non-player characters under each evaluation dimension is collected, and the evaluation results of the model to be evaluated under each evaluation dimension are calculated based on the evaluation data. Specifically, this may include the following steps S241 to S243:

[0051] Step S241: For each model to be evaluated, collect the evaluation data of the non-player characters corresponding to the model under each evaluation dimension.

[0052] Step S242: Calculate the evaluation results of each model under each evaluation dimension based on the evaluation data corresponding to each model to be evaluated.

[0053] Step S243: Based on the evaluation results of each model to be evaluated under each evaluation dimension, sort the multiple models to be evaluated and generate the ranking results of the multiple models to be evaluated under each evaluation dimension.

[0054] When there are multiple models to be evaluated, the ranking results of multiple models under each evaluation dimension can also be calculated, that is, the ranking of multiple models under each evaluation dimension, so that the model caller can select and call the model that is more suitable for the application scenario based on the ranking.

[0055] For example, a game application designed for evaluating two model assessment dimensions—dialogue ability and output content expressiveness—is a supplementary content game. It features five player roles and four non-player roles: five player roles act as referees, and four non-player roles are participants. During the game, the five player roles take turns providing prompts, and the four non-player roles simultaneously supplement the prompts with relevant content. The five player roles then score the four non-player roles based on their supplementary information, evaluating them on both the dialogue ability and output content expressiveness dimensions. After the game ends, the game logs of the five player roles' accounts are retrieved. The scoring data for the four non-player roles across the two assessment dimensions is collected from these logs, and the evaluation results for the four models corresponding to the four non-player roles are calculated based on this data. Assuming that in terms of the calculated model dialogue ability dimension, the evaluation results for Model 1 are 5.5, Model 2 is 8.5, Model 3 is 3.7, and Model 4 is 9.0; and in terms of the model output content expressiveness dimension, the evaluation results for Model 1 are 8.5, Model 2 is 6.5, Model 3 is 5.7, and Model 4 is 8.0, then the ranking of the four models in the model dialogue ability dimension is [Model 4, Model 2, Model 1, Model 3], and the ranking of the models in the model output content expressiveness dimension is [Model 1, Model 4, Model 2, Model 3]. Based on the aforementioned ranking list, when the model caller wants to use a model with strong dialogue capabilities, they can call the model to be evaluated, model 4; when the model caller wants to use a model with strong output content expressiveness, they can call the model to be evaluated, model 1.

[0056] In an optional implementation, the method provided in this embodiment may further include the following steps S11 to S13:

[0057] Step S11 involves connecting the non-player character controlled by the model to be evaluated to the crowdsourcing platform, so that the crowdsourcing platform users can evaluate the non-player character based on the performance of the non-player character in the first game task under various evaluation dimensions.

[0058] Step S12: Collect evaluation data of non-player characters from crowdsourcing platform users under each evaluation dimension, and calculate the crowdsourcing evaluation results of the model to be evaluated under each evaluation dimension based on the evaluation data.

[0059] Step S13: Compare the crowdsourcing evaluation results with the evaluation results to provide verification of the credibility of the evaluation results.

[0060] Crowdsourcing platforms are an internet-based collaborative model that utilizes a large amount of human resources to complete specific tasks. These tasks can include data annotation, translation, design, programming, market research, and creative work. The core idea is to outsource tasks originally performed by company employees or specific teams to an unspecified number of members of the public. In this embodiment, the specific task is the evaluation of a model to be evaluated, and crowdsourcing platform users can be high-level players and / or professionals in model development and evaluation.

[0061] In one optional implementation, the execution status of the first game task and the non-player character's performance on the first game task can be sent to a crowdsourcing platform. This allows crowdsourcing platform users to evaluate the model corresponding to the non-player character across various evaluation dimensions based on the non-player character's performance on the first game task. The evaluation method is consistent with how the player character performs the second game task in the game application. For example, the first game task in the game application is for the non-player character to guess a riddle posed by the player character, and the second game task is for the player character to score the time taken to provide the answer and the correctness of the answer. After obtaining the player character's scoring data for the non-player character, and calculating the evaluation results of the model under evaluation in the dimensions of model reaction speed and model output content expressiveness based on the scoring data, the riddle posed by the player character, the answer provided by the non-player character, and the time taken to provide the answer are sent to the crowdsourcing platform. This allows crowdsourcing platform users to score the time taken to provide the answer and the correctness of the answer.

[0062] In another alternative implementation, a non-player character controlled by the model to be evaluated in the game application can be connected to the page of the crowdsourcing platform, so that the crowdsourcing platform users can perform the same operations as the player character in the game application from the perspective of the player character. For example, they can give riddles to the non-player character controlled by the model to be evaluated, and score the model to be evaluated under various evaluation dimensions based on the answer given by the non-player character and the process of giving the answer.

[0063] After the crowdsourcing platform users complete their evaluations, the platform will collect their evaluation data for the model under evaluation across various evaluation dimensions. Based on this data, the crowdsourcing evaluation results for the model under evaluation will be calculated across each evaluation dimension. The calculation method for the crowdsourcing evaluation results is the same as that for the player evaluation results (i.e., the evaluation results calculated based on the evaluation scores given by players), and will not be elaborated upon here.

[0064] In this embodiment, the purpose of calculating the crowdsourced evaluation results of the model to be evaluated is to serve as a comparison with the player evaluation results. On the one hand, it verifies the correctness of the player evaluation results, and on the other hand, it proves to the model caller that the player evaluation results are credible.

[0065] In an optional implementation, the model evaluation method provided in this embodiment may further include the following steps S21 to S22:

[0066] Step S21: Based on the evaluation data of each player character on the non-player character under each evaluation dimension, select training samples. The training samples include the first game task and the execution results of the first game task by non-player characters whose evaluation data is higher than the preset score threshold.

[0067] Step S22: Train the model to be trained based on the training samples.

[0068] The above steps are downstream processes of the model evaluation method provided in this embodiment. Their purpose is to select high-quality model output data as training data and optimize and train other models that are in the research and development stage or have insufficient performance.

[0069] In one optional implementation, the training data includes a first game task and high-quality execution results. Optionally, high quality is selected based on the player character's evaluation data. A preset score threshold can be used as the standard for measuring high quality. When the player character's evaluation data for the first game task execution result of a non-player character is higher than the preset score threshold, that execution result is considered a high-quality execution result and combined with the first game task as training data. For example, the player character scores 9 points for the correctness of the answer to non-player character 1's riddle and 3 points for the correctness of the answer to non-player character 2's riddle. The preset score threshold is 8 points. Therefore, the answer given by non-player character 1 is a high-quality execution result, and the riddle given by the player character and the answer given by non-player character 1 can be combined as training samples. Training the model with these training samples can improve the output quality of the model and optimize its performance.

[0070] Optionally, training the model to be trained based on the training samples may include the following steps: First, input the first game task from the training samples into the model to be trained, so that the model can parse the first game task based on the current model parameters and output the predicted execution result for the first game task; Second, compare the predicted execution result with the execution result in the training samples, and calculate the loss function between the two; Third, use gradient descent to update the current model parameters of the model to be trained based on minimizing the loss function; Fourth, iteratively train the model to be trained through the above steps until the loss function between the predicted execution result output by the model and the execution result in the training samples is reduced to the target value, thus completing the training of the model to be trained. Since the training samples are high-quality samples obtained through evaluation, training the model to be trained based on these training samples can result in higher output quality and stronger robustness of the model.

[0071] In an optional implementation, the game application provided in this embodiment is a sub-application embedded in a main game application, which includes a massively multiplayer online role-playing game (MMORPG). The method provided in this embodiment may further include: in response to the completion of the second game task, distributing rewards to the player accounts corresponding to each player character, exiting the game application, and entering the main game application, so that the player accounts can use the rewards in the main game application. This approach encourages players participating in the main game application to participate in the sub-application, further ensuring the number of users participating in the model evaluation.

[0072] In an optional implementation, the method provided in this embodiment may further include the following steps: broadcasting a first game task, the execution status of the first game task by non-player characters, and the evaluation data of player characters on non-player characters under each evaluation dimension. Optionally, the execution status of the game task by player characters and non-player characters is published on the game's social platform in the form of communication information. On the one hand, this can attract more players to participate in the game application, thereby further increasing the number of users participating in the model evaluation. On the other hand, it can provide reverse supervision of players' evaluation behavior, avoiding players from providing abnormal evaluation data and further enhancing the accuracy of the evaluation results.

[0073] The first embodiment described above provides an optional model evaluation method. First, this method integrates model evaluation into a game application, issuing game tasks to both non-player and player characters. Player characters then evaluate the performance of non-player characters in executing game tasks (the first game task) while performing their own tasks (the second game task). This allows players to participate in model evaluation without being aware of it, improving the objectivity of the evaluation results. Second, by integrating model evaluation into the game application, and given the large number and diverse profiles of participating players, this method overcomes the limitation of a limited number of participating users and significantly reduces the influence of user subjectivity, personal bias, self-selection bias, and personal comprehension bias on the evaluation results, thus comprehensively improving the credibility of the evaluation results. Third, this method provides high-quality training samples for the model to be trained, improving the model's ability to acquire knowledge. In summary, the model evaluation method provided in this embodiment is a method for unconsciously evaluating models by integrating model evaluation into game applications, ensuring the objectivity, accuracy, and credibility of the evaluation results.

[0074] It should be noted that the examples in the first embodiment are only for explaining the methods described in this application and are not intended to limit actual use. The model evaluation methods provided in this application include, but are not limited to, the methods described in the first embodiment.

[0075] The second embodiment of this application provides a model evaluation system. Figure 3This is a schematic diagram of the model evaluation system provided in this embodiment. The system includes: a model supply module 301, a first model evaluation module 302, and a second model evaluation module 303.

[0076] The model supply module 301 is used to encapsulate multiple models to be evaluated into a unified interface and provide the interface to the first model evaluation module.

[0077] Optionally, multiple open-source and closed-source models can be encapsulated into a unified interface, exposing model parameters such as persona descriptions, maximum response length, response diversity, memory iterations, and knowledge bases to maximize the customization needs of model callers. Optionally, the encapsulated unified model interface can be further encapsulated into a unified model interface installation package to enhance its reusability. Optionally, the encapsulated unified model interface installation package can be further converted into an AOP SDK (Software Development Kit) using AOP (Agent Oriented Programming).

[0078] The first model evaluation module 302 is used to acquire a pre-set game application for at least one evaluation dimension. The evaluation dimension is used to characterize the pre-set evaluation indicators for evaluating the model to be evaluated. The game application includes at least one non-player character and multiple player characters. The module controls the non-player character to execute a first game task in the application process of the game application through the model to be evaluated. In response to the completion of the first game task, a second game task is provided to each player character. The second game task includes evaluating the non-player character under each evaluation dimension based on the non-player character's execution of the first game task. The module collects the evaluation data of each player character for the non-player character under each evaluation dimension and calculates the evaluation results of the model to be evaluated under each evaluation dimension based on the evaluation data.

[0079] The first model evaluation module 302 is mainly used to evaluate the model to be evaluated based on game users (players) in game applications in order to obtain player evaluation results. This part has been described in detail in the first embodiment. For details, please refer to the first embodiment of this application. It will not be repeated here.

[0080] In one optional implementation, the first model evaluation module 302 interfaces with a game application system. The game application system deploys game applications for evaluating the model under evaluation, such as player-generated questions games and clue-filling games. The game application system also deploys corresponding support modules, such as a riddle question bank, a module for binding the model to non-player characters, a player character scoring module, and an evaluation score synchronization module, to support the operation of the game application and model evaluation.

[0081] The second model evaluation module 303 is used to connect non-player characters controlled by the model to be evaluated to the crowdsourcing platform, so that the crowdsourcing platform users can evaluate the non-player characters based on their performance in the first game task under various evaluation dimensions; collect the evaluation data of the crowdsourcing platform users on the non-player characters under various evaluation dimensions, and calculate the crowdsourcing evaluation results of the model to be evaluated under various evaluation dimensions based on the evaluation data; compare the evaluation results with the crowdsourcing evaluation results to provide credibility verification of the evaluation results.

[0082] The second model evaluation module 303 is mainly used to connect with the crowdsourcing platform to evaluate the model to be evaluated based on the crowdsourcing platform users (high-level players and / or professionals in model development and evaluation), and obtain the crowdsourcing evaluation results, which are used as the credibility verification of the player evaluation results.

[0083] The third embodiment of this application provides a model evaluation device. Figure 4 This is a schematic diagram of the model evaluation device provided in this embodiment.

[0084] like Figure 4 As shown, the model evaluation device provided in this embodiment includes: a game application acquisition unit 401, a game task execution unit 402, a game task provision unit 403, and an evaluation result calculation unit 404.

[0085] The game application acquisition unit 401 is used to acquire a game application preset for at least one evaluation dimension. The evaluation dimension is used to characterize the preset indicators for evaluating the model to be evaluated. The game application includes at least one non-player character and multiple player characters.

[0086] The game task execution unit 402 is used to control the non-player character to execute the first game task in the application process of the game application through the model to be evaluated.

[0087] Optionally, controlling the non-player character to perform the first game task in the application process of the game application through the model to be evaluated includes:

[0088] In response to the start command for the game application, the model to be evaluated is connected to the game application so that the model to be evaluated controls the non-player character to execute the first game task in the application process of the game application.

[0089] Optionally, the model to be evaluated includes multiple models, and the game application includes multiple non-player characters;

[0090] The step of controlling the non-player character to perform the first game task in the application process of the game application through the model to be evaluated includes:

[0091] Multiple models to be evaluated are connected to the game application and a corresponding relationship is established with multiple non-player characters, so that each model to be evaluated controls a non-player character to execute the first game task in the application process of the game application.

[0092] Optionally, multiple models to be evaluated may have a unified interface;

[0093] The step of connecting multiple models to be evaluated into the game application and establishing corresponding relationships with multiple non-player characters includes:

[0094] The game application is controlled to call multiple models to be evaluated based on the unified interface, and a non-player character is bound to each model to be evaluated.

[0095] The game task providing unit 403 is used to provide a second game task to each player character in response to the completion of the first game task. The second game task includes evaluating the non-player characters in various evaluation dimensions based on their performance on the first game task.

[0096] The evaluation result calculation unit 404 is used to collect evaluation data of each player character on the non-player character under each evaluation dimension, and calculate the evaluation results of the model to be evaluated under each evaluation dimension based on the evaluation data.

[0097] Optionally, the step of collecting evaluation data of each player character on the non-player character under each evaluation dimension, and calculating the evaluation results of the model to be evaluated under each evaluation dimension based on the evaluation data, includes:

[0098] For each of the models to be evaluated, collect evaluation data of each player character on the non-player character corresponding to the model under each evaluation dimension;

[0099] Based on the evaluation data corresponding to each of the models to be evaluated, calculate the evaluation results of each model under each evaluation dimension;

[0100] Based on the evaluation results of each model under each evaluation dimension, the multiple models under evaluation are sorted to generate a ranking result of the multiple models under each evaluation dimension.

[0101] Optionally, the device further includes a crowdsourcing evaluation unit; the crowdsourcing evaluation unit is used for:

[0102] The non-player character controlled by the model to be evaluated is connected to the crowdsourcing platform so that the crowdsourcing platform users can evaluate the non-player character in various evaluation dimensions based on the non-player character's performance in the first game task.

[0103] Collect evaluation data of non-player characters from crowdsourcing platform users under various evaluation dimensions, and calculate the crowdsourcing evaluation results of the model to be evaluated under various evaluation dimensions based on the evaluation data;

[0104] The crowdsourcing evaluation results are compared with the evaluation results to provide a verification of the credibility of the evaluation results.

[0105] Optionally, the apparatus further includes a model training unit; the model training unit is used for:

[0106] Based on the evaluation data of each player character on the non-player character under each evaluation dimension, training samples are selected. The training samples include the first game task and the execution results of the non-player character on the first game task whose evaluation data is higher than a preset score threshold.

[0107] The model to be trained is trained based on the training samples.

[0108] Optionally, the game application is a sub-application embedded in a main game application, the main game application including a massively multiplayer online role-playing game; the device further includes a reward distribution unit; the reward distribution unit is used for:

[0109] Upon completion of the second game task, rewards are distributed to the player accounts corresponding to each player character, and the game application is exited, while the main game application is entered, allowing the player accounts to use the rewards in the main game application.

[0110] Optionally, the first game task is to interact with the player character using natural language, and the second game task is to evaluate the non-player character based on the interaction between the non-player character and the player character using natural language, under various evaluation dimensions.

[0111] Optionally, the evaluation dimensions may include at least one of the following dimensions:

[0112] Language interaction ability, consistency between language and role, consistency between language and scene, expressiveness of language content, anthropomorphism of role, and security of language content.

[0113] Optionally, the apparatus further includes a data broadcasting unit; the data broadcasting unit is used for:

[0114] Broadcast the first game task, the non-player character's execution of the first game task, and the player character's evaluation data of the non-player character under each evaluation dimension.

[0115] The fourth embodiment of this application provides an electronic device. Figure 5 This is a schematic diagram of the structure of the electronic device provided in this embodiment.

[0116] like Figure 5 As shown, the electronic device provided in this embodiment includes: a memory 501 and a processor 502;

[0117] The memory 501 is used to store computer instructions for executing the model evaluation method;

[0118] The processor 502 is configured to execute computer instructions stored in the memory 501 to perform the following operations:

[0119] Obtain a game application with at least one preset evaluation dimension, wherein the evaluation dimension is used to characterize the preset indicators for evaluating the evaluation model, and the game application includes at least one non-player character and multiple player characters.

[0120] The non-player character is controlled to perform the first game task in the application process of the game application through the model to be evaluated.

[0121] In response to the completion of the first game task, a second game task is provided to each player character. The second game task includes evaluating the non-player characters in various evaluation dimensions based on their performance on the first game task.

[0122] Collect evaluation data of each player character on the non-player character under each evaluation dimension, and calculate the evaluation results of the model to be evaluated under each evaluation dimension based on the evaluation data.

[0123] Optionally, controlling the non-player character to perform the first game task in the application process of the game application through the model to be evaluated includes:

[0124] In response to the start command for the game application, the model to be evaluated is connected to the game application so that the model to be evaluated controls the non-player character to execute the first game task in the application process of the game application.

[0125] Optionally, the model to be evaluated includes multiple models, and the game application includes multiple non-player characters;

[0126] The step of controlling the non-player character to perform the first game task in the application process of the game application through the model to be evaluated includes:

[0127] Multiple models to be evaluated are connected to the game application and a corresponding relationship is established with multiple non-player characters, so that each model to be evaluated controls a non-player character to execute the first game task in the application process of the game application.

[0128] Optionally, the step of collecting evaluation data of each player character on the non-player character under each evaluation dimension, and calculating the evaluation results of the model to be evaluated under each evaluation dimension based on the evaluation data, includes:

[0129] For each of the models to be evaluated, collect evaluation data of each player character on the non-player character corresponding to the model under each evaluation dimension;

[0130] Based on the evaluation data corresponding to each of the models to be evaluated, calculate the evaluation results of each model under each evaluation dimension;

[0131] Based on the evaluation results of each model under each evaluation dimension, the multiple models under evaluation are sorted to generate a ranking result of the multiple models under each evaluation dimension.

[0132] Optional, also execute:

[0133] The non-player character controlled by the model to be evaluated is connected to the crowdsourcing platform so that the crowdsourcing platform users can evaluate the non-player character in various evaluation dimensions based on the non-player character's performance in the first game task.

[0134] Collect evaluation data of non-player characters from crowdsourcing platform users under various evaluation dimensions, and calculate the crowdsourcing evaluation results of the model to be evaluated under various evaluation dimensions based on the evaluation data;

[0135] The crowdsourcing evaluation results are compared with the evaluation results to provide a verification of the credibility of the evaluation results.

[0136] Optionally, multiple models to be evaluated may have a unified interface;

[0137] The step of connecting multiple models to be evaluated into the game application and establishing corresponding relationships with multiple non-player characters includes:

[0138] The game application is controlled to call multiple models to be evaluated based on the unified interface, and a non-player character is bound to each model to be evaluated.

[0139] Optional, also execute:

[0140] Based on the evaluation data of each player character on the non-player character under each evaluation dimension, training samples are selected. The training samples include the first game task and the execution results of the non-player character on the first game task whose evaluation data is higher than a preset score threshold.

[0141] The model to be trained is trained based on the training samples.

[0142] Optionally, the game application is a sub-application embedded in a main game application, the main game application including a massively multiplayer online role-playing game; it also executes:

[0143] Upon completion of the second game task, rewards are distributed to the player accounts corresponding to each player character, and the game application is exited, while the main game application is entered, allowing the player accounts to use the rewards in the main game application.

[0144] Optionally, the first game task is to interact with the player character using natural language, and the second game task is to evaluate the non-player character based on the interaction between the non-player character and the player character using natural language, under various evaluation dimensions.

[0145] Optionally, the evaluation dimensions may include at least one of the following dimensions:

[0146] Language interaction ability, consistency between language and role, consistency between language and scene, expressiveness of language content, anthropomorphism of role, and security of language content.

[0147] Optional, also execute:

[0148] Broadcast the first game task, the non-player character's execution of the first game task, and the player character's evaluation data of the non-player character under each evaluation dimension.

[0149] The fifth embodiment of this application provides a computer-readable storage medium, which includes computer instructions that, when executed by a processor, are used to implement the methods described in the embodiments of this application.

[0150] It should be noted that the relational terms such as "first" and "second" used in this document are only used to distinguish one entity or operation from another, and do not require or imply any actual relationship or order between these entities or operations. Furthermore, "including," "having," "containing," and other similar terms are synonymous, and the conclusion of any one or more items following any of the foregoing words is open-ended; none of the foregoing terms indicates that the one or more items have been exhaustively listed, or are limited to only one or more of the listed items.

[0151] When used herein, unless otherwise expressly stated, the term "or" includes all possible combinations except those that are impractical. For example, if expressed as a database may include A or B, then unless otherwise specified or impractical, it may include database A, or B, or A and B. As a second example, if expressed as a database may include A, B, or C, then unless otherwise specified or impractical, the database may include database A, or B, or C, or A and B, or A and C, or B and C, or A and B and C.

[0152] It is worth noting that the above embodiments can be implemented by hardware or software (program code), or a combination of hardware and software. If implemented by software, it can be stored in the above-described computer-readable medium. When executed by a processor, the software can perform the methods disclosed above. The computing units and other functional units described in this disclosure can be implemented by hardware or software, or a combination of hardware and software. Those skilled in the art will also understand that the above-described multiple modules / units can be combined into one module / unit, and each of the above-described modules / units can be further divided into multiple sub-modules / sub-units.

[0153] In the foregoing detailed description, embodiments have been described with reference to numerous specific details, which may vary depending on the implementation. Certain adaptations and modifications can be made to the embodiments. Other implementations will be readily apparent to those skilled in the art from the specific embodiments disclosed herein. This specification and examples are for illustrative purposes only, and the true scope and essence of this application are defined by the claims. The sequence of steps shown in the figures is also for illustrative purposes only and is not intended to limit to any particular step or order. Therefore, those skilled in the art will recognize that these steps can be performed in different orders when implementing the same method.

[0154] Exemplary embodiments are disclosed in the figures and detailed description of this application. However, many variations and modifications can be made to these embodiments. Accordingly, although specific terms are used, they are only general and descriptive and not for limiting purposes.

Claims

1. A model evaluation method, characterized in that, The method includes: Obtain a game application with at least one preset evaluation dimension, wherein the evaluation dimension is used to characterize the preset indicators for evaluating the evaluation model, and the game application includes at least one non-player character and multiple player characters. The non-player character is controlled by the model to be evaluated to perform the first game task in the application process of the game application; In response to the completion of the first game task, a second game task is provided to each player character. The second game task includes evaluating the non-player characters in various evaluation dimensions based on their performance on the first game task. Collect evaluation data of each player character on the non-player character under each evaluation dimension, and calculate the evaluation results of the model to be evaluated under each evaluation dimension based on the evaluation data.

2. The method according to claim 1, characterized in that, The models to be evaluated include multiple models, and the game application includes multiple non-player characters. The step of controlling the non-player character to perform the first game task in the application process of the game application through the model to be evaluated includes: Multiple models to be evaluated are connected to the game application and a corresponding relationship is established with multiple non-player characters, so that each model to be evaluated controls a non-player character to execute the first game task in the application process of the game application.

3. The method according to claim 2, characterized in that, The process of collecting evaluation data from each player character on the non-player character across various evaluation dimensions, and calculating the evaluation results of the model under evaluation across these dimensions based on the evaluation data, includes: For each of the models to be evaluated, collect evaluation data of each player character on the non-player character corresponding to the model under each evaluation dimension; Based on the evaluation data corresponding to each of the models to be evaluated, calculate the evaluation results of each model under each evaluation dimension; Based on the evaluation results of each model under each evaluation dimension, the multiple models under evaluation are sorted to generate a ranking result of the multiple models under each evaluation dimension.

4. The method according to claim 1, characterized in that, The method further includes: The non-player character controlled by the model to be evaluated is connected to the crowdsourcing platform so that the crowdsourcing platform users can evaluate the non-player character in various evaluation dimensions based on the non-player character's performance in the first game task. Collect evaluation data of non-player characters from crowdsourcing platform users under various evaluation dimensions, and calculate the crowdsourcing evaluation results of the model to be evaluated under various evaluation dimensions based on the evaluation data; The crowdsourcing evaluation results are compared with the evaluation results to provide a verification of the credibility of the evaluation results.

5. The method according to claim 2, characterized in that, All of the aforementioned models to be evaluated have a unified interface; The step of connecting multiple models to be evaluated into the game application and establishing corresponding relationships with multiple non-player characters includes: The game application is controlled to call multiple models to be evaluated based on the unified interface, and a non-player character is bound to each model to be evaluated.

6. The method according to claim 1, characterized in that, The method further includes: Based on the evaluation data of each player character on the non-player character under each evaluation dimension, training samples are selected. The training samples include the first game task and the execution results of the non-player character on the first game task whose evaluation data is higher than a preset score threshold. The model to be trained is trained based on the training samples.

7. The method according to claim 1, characterized in that, The game application is a sub-application embedded in a main game application, and the main game application includes a massively multiplayer online role-playing game; the method further includes: Upon completion of the second game task, rewards are distributed to the player accounts corresponding to each player character, and the game application is exited, while the main game application is entered, allowing the player accounts to use the rewards in the main game application.

8. The method according to claim 1, characterized in that, The first game task is to interact with the player character using natural language. The second game task is to evaluate the non-player character based on the interaction between the non-player character and the player character using natural language, under various evaluation dimensions.

9. The method according to claim 8, characterized in that, The evaluation dimensions shall include at least one of the following dimensions: Language interaction ability, consistency between language and role, consistency between language and scene, expressiveness of language content, anthropomorphism of role, and security of language content.

10. The method according to claim 1, characterized in that, The method further includes: Broadcast the first game task, the non-player character's execution of the first game task, and the player character's evaluation data of the non-player character under each evaluation dimension.

11. The method according to claim 1, characterized in that, The step of controlling the non-player character to perform the first game task in the application process of the game application through the model to be evaluated includes: In response to the start command for the game application, the model to be evaluated is connected to the game application so that the model to be evaluated controls the non-player character to execute the first game task in the application process of the game application.

12. A model evaluation system, characterized in that, The system includes: a model supply module, a first model evaluation module, and a second model evaluation module; The model supply module is used to encapsulate multiple models to be evaluated into a unified interface and provide the interface to the first model evaluation module. The first model evaluation module is used to acquire a preset game application for at least one evaluation dimension, wherein the evaluation dimension is used to characterize preset evaluation indicators for evaluating the model to be evaluated, and the game application includes at least one non-player character and multiple player characters; the module controls the non-player character to execute a first game task in the application process of the game application through the model to be evaluated; in response to the completion of the first game task, the module provides a second game task to each player character, wherein the second game task includes evaluating the non-player character under each evaluation dimension based on the non-player character's execution of the first game task; the module collects evaluation data of each player character for the non-player character under each evaluation dimension, and calculates the evaluation result of the model to be evaluated under each evaluation dimension based on the evaluation data; The second model evaluation module is used to connect the non-player character controlled by the model to be evaluated to a crowdsourcing platform, so that the crowdsourcing platform users can evaluate the non-player character in each evaluation dimension based on the non-player character's performance on the first game task; collect the evaluation data of the crowdsourcing platform users on the non-player character in each evaluation dimension, and calculate the crowdsourcing evaluation results of the model to be evaluated in each evaluation dimension based on the evaluation data; compare the evaluation results with the crowdsourcing evaluation results to provide credibility verification of the evaluation results.

13. A model evaluation device, characterized in that, The device includes: a game application acquisition unit, a game task execution unit, a game task provision unit, and an evaluation result calculation unit; The game application acquisition unit is used to acquire game applications preset for at least one evaluation dimension. The evaluation dimension is used to characterize preset indicators for evaluating the evaluation model. The game application includes at least one non-player character and multiple player characters. The game task execution unit is used to control the non-player character to execute the first game task in the application process of the game application through the model to be evaluated. The game task providing unit is configured to provide a second game task to each player character in response to the completion of the first game task. The second game task includes evaluating the non-player character in various evaluation dimensions based on the non-player character's performance on the first game task. The evaluation result calculation unit is used to collect evaluation data of each player character on the non-player character under each evaluation dimension, and calculate the evaluation results of the model to be evaluated under each evaluation dimension based on the evaluation data.

14. An electronic device, characterized in that, include: Memory, processor; The memory is used to store one or more computer instructions; The processor is configured to execute one or more computer instructions to implement the method as described in any one of claims 1-11.

15. A computer-readable storage medium storing one or more computer instructions thereon, characterized in that, When this instruction is executed by the processor, it performs the method as described in any one of claims 1-11.

Citation Information

Patent Citations

  • Game testing method and device, electronic equipment and storage medium

    CN112783781A

  • Game role generation method and device, equipment, storage medium and program product

    CN116966590A