Evaluation method and apparatus for digital assistant, and device and storage medium
By automating the evaluation of digital assistants' chat skills, the problems of low efficiency and poor consistency in existing technologies are solved, achieving efficient and objective evaluation results.
Patent Information
- Application Number
- PCT/CN2025/114876
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-15
- Filing Date
- 2025-08-14
- Publication Date
- 2026-02-19
AI Technical Summary
In existing technologies, the evaluation of digital assistants relies on manual testing, which leads to low efficiency and difficulty in ensuring objectivity and consistency.
By acquiring test cases for the target digital assistant, the system automatically evaluates its chat skills, including identity awareness, function awareness, basic interaction, interaction in areas of expertise, and exception handling capabilities, generates target evaluation metrics, and determines the quality evaluation results.
It automates the evaluation of digital assistants, improves evaluation efficiency, ensures the objectivity and consistency of evaluation results, and reduces reliance on manual testing.
Smart Images

Figure CN2025114876_19022026_PF_FP_ABST
Abstract
Description
Evaluation method, device and storage medium for digital assistant
[0001] The present application claims priority to the Chinese patent application No. 202411126418.7, filed on August 15, 2024, entitled “Evaluation method, device and storage medium for digital assistant”, the whole content of which is incorporated herein by reference. TECHNICAL FIELD
[0002] The example embodiments of the present disclosure generally relate to the field of computer, and in particular, to an evaluation method, device, equipment and storage medium for digital assistant. BACKGROUND
[0003] Digital assistant refers to a system or application with dialogue capability. With the popularity of digital assistants in customer service, education, entertainment and other fields, the interaction quality of digital assistants is becoming increasingly important. Evaluating digital assistants is of great significance to ensure their quality and performance. Through evaluation, digital assistants that meet quality and performance requirements can be recommended to users, thereby improving user experience and satisfaction. Therefore, how to accurately evaluate digital assistants is particularly important. SUMMARY
[0004] In a first aspect of the present disclosure, an evaluation method for a digital assistant is provided. The method can include: in response to an evaluation request for a target digital assistant, obtaining at least one set of test cases for the target digital assistant, each set of test cases including at least one test question related to a chat skill of the target digital assistant. Providing the at least one set of test cases to the target digital assistant to obtain a reply of the target digital assistant to the at least one set of test cases. At least based on the at least one set of test cases and the reply of the target digital assistant to the at least one set of test cases, determining a target evaluation index for the target digital assistant, the target evaluation index including at least a first feature value indicating a chat skill score of the target digital assistant. Based on the target evaluation index, determining a quality evaluation result of the target digital assistant.
[0005] In a second aspect of the present disclosure, an evaluation apparatus for a digital assistant is provided. The apparatus can include: a test case acquisition module configured to, in response to an evaluation request for a target digital assistant, acquire at least one set of test cases for the target digital assistant, each set of test cases including at least one test question related to a chat skill of the target digital assistant; a reply acquisition module configured to provide the at least one set of test cases to the target digital assistant to obtain a reply of the target digital assistant to the at least one set of test cases; a target evaluation indicator determination module configured to determine a target evaluation indicator for the target digital assistant based at least on the at least one set of test cases and the reply of the target digital assistant to the at least one set of test cases, the target evaluation indicator including at least a first feature value indicating a chat skill score of the target digital assistant; and a quality evaluation result determination module configured to determine a quality evaluation result for the target digital assistant based on the target evaluation indicator.
[0006] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. The instructions, when executed by the at least one processing unit, cause the electronic device to perform the method of the first aspect.
[0007] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided. The medium has stored thereon a computer program which, when executed by a processor, implements the method of the first aspect.
[0008] In a fifth aspect of the present disclosure, a computer program product or computer program is provided, which includes computer instructions stored in a computer readable storage medium. A processor of a computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions, so that the computer device performs the method provided in various optional manners in the first aspect of the embodiments of the present disclosure. In other words, the computer instructions, when executed by the processor, implement the method provided in various optional manners in the first aspect of the embodiments of the present disclosure.
[0009] It should be understood that the content described in this section is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become apparent through the following description. BRIEF DESCRIPTION OF DRAWINGS
[0010] The above and other features, advantages, and aspects of embodiments of the present disclosure will become more apparent by describing in detail exemplary embodiments thereof with reference to the attached drawings in which:
[0011] FIG. 1 shows a schematic diagram of an example environment in which embodiments of the present disclosure can be implemented;
[0012] FIG. 2 shows a flowchart of a method for digital assistant evaluation according to some embodiments of the present disclosure;
[0013] FIG. 3 shows an example diagram of a chat skill score interface according to some embodiments of the present disclosure;
[0014] FIG. 4 shows a training flowchart of an evaluation model according to some embodiments of the present disclosure;
[0015] FIG. 5 shows an example diagram of a relevance distribution according to some embodiments of the present disclosure;
[0016] FIG. 6 shows a schematic structural block diagram of an apparatus for digital assistant evaluation according to some embodiments of the present disclosure; and
[0017] FIG. 7 shows a block diagram of an electronic device that can implement one or more embodiments of the present disclosure. DETAILED DESCRIPTION
[0018] Embodiments of the present disclosure will be described below in greater detail with reference to the accompanying drawings. While certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be interpreted as being limited to the embodiments set forth herein, but rather, these embodiments are provided so that the present disclosure can be more thoroughly and completely understood. It should be understood that the drawings of the present disclosure and the embodiments are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.
[0019] In the description of embodiments of the present disclosure, the term "comprising" and its conjugations should be understood to encompass the meanings of "including but not limited to", i.e., "comprising but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions can also be included below.
[0020] In this document, unless explicitly stated otherwise, performing a step "in response to" A does not mean that the step is performed immediately after A, but can include one or more intermediate steps.
[0021] It can be understood that the data involved in the technical solution (including but not limited to the data itself, obtaining, using, storing or deleting of the data) should comply with the requirements of the relevant laws and regulations and relevant provisions.
[0022] It can be understood that, before using the technical solutions disclosed in the embodiments of the present disclosure, the type of information involved in the present disclosure, the scope of use, the use scenario, etc. should be informed to the relevant user and the authorization of the relevant user should be obtained through appropriate means according to relevant laws and regulations, wherein the relevant user can include any type of right subject, such as an individual, an enterprise, or a group.
[0023] For example, in response to receiving an active request of a user, a prompt information is sent to the relevant user to explicitly prompt the relevant user that the operation requested to be performed will need to obtain and use the information of the relevant user, so that the relevant user can voluntarily choose whether to provide the information to the software or hardware such as an electronic device, an application program, a server or a storage medium, etc. performing the operation of the technical solutions of the present disclosure according to the prompt information.
[0024] As an optional but non-limiting implementation manner, in response to receiving an active request of the relevant user, the prompt information can be sent to the relevant user in the manner of a pop-up window, and the prompt information can be presented in the form of text in the pop-up window. In addition, the pop-up window can also carry selection controls for the user to select "agree" or "disagree" to provide information to the electronic device.
[0025] It can be understood that the above notification and user authorization process is only illustrative and does not limit the implementation manner of the present disclosure, and other manners meeting the relevant laws and regulations can also be applied to the implementation manner of the present disclosure.
[0026] As used herein, the term "model" can learn the association between the corresponding input and output from the training data, so that after the training is completed, the corresponding output can be generated for a given input. The generation of the model can be based on machine learning techniques. Deep learning is a machine learning algorithm that uses multiple layers of processing units to process inputs and provide corresponding outputs. Neural network models are an example of models based on deep learning. In this document, "model" can also be referred to as "machine learning model", "learning model", "machine learning network" or "learning network", which are used interchangeably herein.
[0027] FIG. 1 illustrates a schematic diagram of an environment 100 in which embodiments of the present disclosure can be implemented. A digital assistant development platform 120 provides a development environment for developers 105 to create and publish digital assistants. For example, the digital assistant development platform 120 can provide various tools to the developers 105, such as prompt word information, plugins, workflows, knowledge bases, memory banks, speech, etc. In some embodiments, the digital assistant development platform 120 can be a low-code platform that provides a collection of tools for digital assistant creation. The digital assistant development platform 120 can support visual development of digital assistants, such that the developers 105 can skip the process of hand-coding and speed up the development cycle and cost of applications. The digital assistant development platform 120 can support any suitable platform for users to develop digital assistants and other types of applications, which can include, for example, an application platform as a service (aPaaS)-based platform. Such a platform can enable efficient development of applications by users, enabling operations such as application creation, application function adjustment, etc.
[0028] The digital assistant development platform 120 can be deployed locally on the terminal device of the developers 105 and / or can be supported by a remote server. For example, the terminal device of the developers 105 can run a client of the digital assistant development platform 120, which can support user interaction with the digital assistant development platform 120. In a case where the digital assistant development platform 120 is run locally on the terminal device of the user, the developers 105 can directly interact with the local digital assistant development platform 120 using the client. In a case where the digital assistant development platform 120 is run on a server device, the server device can provide services to the client running on the terminal device based on a communication connection between the server device and the terminal device. The digital assistant development platform 120 can present a corresponding interface 122 to the developers 105 based on operations of the developers 105, to output and / or receive information from the developers 105.
[0029] In some embodiments, the digital assistant development platform 120 can be associated with a corresponding database that stores data or information required for digital assistant creation processes supported by the digital assistant development platform 120. For example, the database can store code and description information corresponding to various functional modules used to compose a digital assistant, etc. The digital assistant development platform 120 can also perform operations such as calling, adding, deleting, updating, etc. on the functional modules in the database. The database can also store operations executable on different functional blocks. For example, in a scenario where a digital assistant is to be created, the digital assistant development platform 120 can call corresponding functional blocks from the database to build the digital assistant.
[0030] In embodiments of the present disclosure, the developer 105 can create digital assistants 121 on the digital assistant development platform 120 as needed and publish the digital assistants 121. The digital assistants 121 can be published to any appropriate application platform as long as the application platform can support the running of the digital assistants 121. After being published, the digital assistants 121 can be used for conversational interaction with the user 135.
[0031] After the digital assistants 121 are created / published, they can be evaluated by the electronic device 110 to obtain evaluation results. For the digital assistants whose evaluation results meet the recommendation conditions, the recommendation can be performed on the recommendation interface of the digital assistant recommendation platform 130. Exemplarily, the digital assistant recommendation platform 130 can be integrated in the electronic device 110 or be a third-party platform independent of the electronic device 110. The evaluation of the digital assistants 121 by the electronic device 110 can be performed based on multiple dimensions, such as evaluation indicators corresponding to chat skills, evaluation indicators corresponding to user feedback in the process of interaction with the user 135, and the like.
[0032] The electronic device 110 can be any type of mobile terminal, fixed terminal, or portable terminal including a mobile handset, a tablet computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an electronic book device, a game device, or any combination thereof, including accessories and peripherals of these devices, or any combination thereof. In some embodiments, the electronic device 110 can also support any type of interface to a user (such as “wearable” circuitry, etc.).
[0033] It should be understood that the structure and function of the various elements in the environment 100 are described for illustrative purposes only, without implying any limitation on the scope of the present disclosure.
[0034] The conventional digital assistant evaluation method mainly performs recommendation through trial and test by operation personnel. This method relies on manual test and evaluation, which can ensure a certain quality, but is low in efficiency and difficult to guarantee objectivity and consistency.
[0035] In embodiments of the present disclosure, an evaluation method of a digital assistant is proposed. For a target digital assistant to be evaluated, at least one set of test cases for the target digital assistant is obtained, each set of test cases including at least one test question related to a chatting skill of the target digital assistant. The at least one set of test cases is provided to the target digital assistant to obtain a reply of the target digital assistant to the at least one set of test cases. Based on at least the at least one set of test cases and the reply of the target digital assistant to the at least one set of test cases, a target evaluation indicator for the target digital assistant is determined, the target evaluation indicator including at least a first feature value indicating a chatting skill score of the target digital assistant. Based on the target evaluation indicator, a quality evaluation result of the target digital assistant is determined.
[0036] Through the above process, automated evaluation of the digital assistant can be achieved to evaluate the quality of the target digital assistant at least from the perspective of the chatting skill of the digital assistant. The automated evaluation method reduces the dependence on manual testing, making the evaluation process faster, continuous and without interruption. The operator no longer needs to try and test each digital assistant one by one, saving a lot of time and human resources. Moreover, the automated test cases and evaluation indicators ensure the standardization of the evaluation process, avoiding subjective bias caused by individual differences, and ensuring the objectivity and consistency of the evaluation result. By using multiple sets of test cases, the functions and performance of the target digital assistant can be comprehensively evaluated, providing more detailed and accurate evaluation results. Thus, the present disclosure can improve the evaluation efficiency of the digital assistant and meet the objectivity and consistency of the evaluation standard.
[0037] FIG. 2 illustrates an example flow 200 of an evaluation method of a digital assistant according to some embodiments of the present disclosure. For ease of discussion, the flow 200 will be described with reference to the environment of FIG. 1. In the application environment 110, the evaluation of the digital assistant can be performed by the electronic device 110.
[0038] At block 201, in response to an evaluation request for the target digital assistant 121, the electronic device 110 obtains at least one set of test cases for the target digital assistant 121, each set of test cases including at least one test question related to a chatting skill of the target digital assistant 121.
[0039] The evaluation request for the target digital assistant 121 can be various, for example, the online time of the target digital assistant 121 can be taken as the evaluation request, the evaluation instruction of the developer can be taken as the evaluation request, the number of user feedbacks can be taken as the evaluation request, and the like.
[0040] The test cases for the chat skill of the target digital assistant 121 can include multiple groups, and each group of chat skill test cases corresponds to a different evaluation dimension. The chat skill test cases can be generated based on prompt information. For example, the prompt information can include an identification of the target digital assistant 121, which can be a name or a serial number, etc. In addition, the prompt information can also include a functional description of the target digital assistant, such as a brief introduction, a functional introduction, and a usage guide of the target digital assistant 121, etc.
[0041] The brief introduction can describe the main functions of the target digital assistant 121. For example, the brief introduction can include that the role of the digital assistant is a legal assistant, which can answer various legal-related questions.
[0042] The functional introduction can indicate the services or functions that the target digital assistant 121 can provide. For example, the functional introduction can include that the digital assistant can provide legal consultation, legal document generation, legal fee calculation, legal education, etc. to the user.
[0043] The usage guide can introduce the interactive way of the target digital assistant 121. For example, the usage guide can include that the user can ask me about anything related to law that is of interest.
[0044] The evaluation dimensions can be pre-set. Different evaluation dimensions are used to evaluate different abilities embodied by the chat skill of the digital assistant. In some embodiments, different evaluation dimensions can correspond to identity recognition ability, functional recognition ability, basic interaction ability, interaction ability positively related to the field of expertise, interaction ability negatively related to the field of expertise, processing ability of abnormal interaction, etc. It can be understood that other evaluation dimensions and the corresponding abilities to be evaluated can also be defined according to specific evaluation needs.
[0045] The test cases for the identity recognition ability can be used to evaluate the ability of the target digital assistant 121 to recognize and express its own identity, for example, the content of the functional description contains the field of expertise of the target digital assistant 121 is law. Then the test question can be "Who are you?" and similar questions to guide the target digital assistant 121 to answer who it is. In addition, the test question can also be "Is your role a conference host?" and similar questions to mislead the identity of the target digital assistant 121.
[0046] The test cases for the functional recognition ability can be used to evaluate the description and understanding of the functions of the target digital assistant 121, for example, questions similar to "What can you do?" to guide the target digital assistant 121 to say its functions.
[0047] The test cases for the basic interaction ability can be used to evaluate the processing ability of the target digital assistant 121 for general conversation, including understanding user intent, providing reasonable responses, etc.
[0048] Interaction ability test cases positively correlated with the area of expertise can be used to assess the performance of the target digital assistant 121 in its area of expertise, for example, the accuracy and professionalism of the target digital assistant 121 in answering legal-related questions.
[0049] Interaction ability test cases negatively correlated with the area of expertise can be used to examine the performance of the target digital assistant 121 on non-expertise questions, ensuring that it can reasonably guide the user or admit its limitations.
[0050] Abnormal interaction handling test cases can be used to simulate various abnormal situations or misoperations to evaluate how the digital assistant handles unexpected user inputs or abnormal requests.
[0051] For a given group chat skill test case, at least one round of interaction test related to a given evaluation dimension can be determined. For example, if the content of the function description contains the area of expertise of the target digital assistant, which is law, then for the test case of identity recognition ability, at least two rounds of interaction test can be included. The test question of the first round of interaction test can be "Who are you?" and the like to guide the target digital assistant to say who it is. The test question of the second round of interaction test can also be "Are you the host?" and the like to mislead the target digital assistant to the question of identity.
[0052] At block 202, the electronic device 110 provides at least one set of test cases to the target digital assistant 121 to obtain the replies of the target digital assistant 121 to the at least one set of test cases. Based on the replies, the electronic device 110 can evaluate the performance and ability of the target digital assistant 121 in different situations in subsequent processes.
[0053] At block 203, the electronic device 110 determines a target evaluation index for the target digital assistant 121 based on at least the at least one set of test cases and the replies of the target digital assistant 121 to the at least one set of test cases. In an embodiment of the present disclosure, the target evaluation index at least includes a first feature value indicating a chat skill score of the target digital assistant.
[0054] If the target digital assistant 121 is to be tested using multiple sets of test cases, each set of test cases including at least one test question, the replies of the target digital assistant 121 to different test questions can be obtained. For the chat skill scores corresponding to multiple replies, an average value, a weighted average value, etc. can be used to obtain the first feature value of the chat skill score of the target digital assistant 121.
[0055] As mentioned above, the target evaluation indicators can indicate the first characteristic values of the chat skill scores of the target digital assistant 121. In addition, the target evaluation indicators can also indicate the characteristic values of different scores such as response speed, language flow fluency, and personalization degree of reply content of the target digital assistant 121.
[0056] At block 204, the electronic device 110 determines the quality evaluation result of the target digital assistant 121 based on the target evaluation indicators. The electronic device 110 can set corresponding weights for different target evaluation indicators. The weights can reflect the importance of each target evaluation indicator in the overall evaluation. For example, for a digital assistant, its identity recognition, function recognition, and knowledge interaction in the field it is good at can be given higher weights. Through such weighting processing, the electronic device 110 can generate a comprehensive evaluation result that accurately reflects the overall performance of the target digital assistant 121.
[0057] Through the above evaluation method, the limitations of relying on manual testing and recommending digital assistants by operating personnel in the prior art can be solved. The automated evaluation process not only improves efficiency, but also ensures the objectivity and consistency of the evaluation. This not only improves the user experience and ensures that the recommended digital assistant meets the user's needs, but also provides feedback for developers to improve their products.
[0058] As mentioned above, the evaluation of the digital assistant is based on test cases. The acquisition method of the test cases is described in detail below. In some embodiments of the present disclosure, the electronic device 110 acquires prompt word information of the target digital assistant 121, and the prompt word information at least includes identification information and function description of the target digital assistant 121. At least one set of test case generation rules corresponding to each set of test cases in the at least one set of test cases is acquired. At least based on the prompt word information and the test case generation rules, one or more sets of test cases for the target digital assistant 121 are generated.
[0059] The prompt information of the target digital assistant 121 at least includes identification information and function description of the target digital assistant 121. The identification information is used to uniquely identify the target digital assistant 121, for example, the name, serial number, etc. of the target digital assistant 121. The function description is used to describe the main functions and characteristics of the target digital assistant 121, for example, the information of the developer, the field the target digital assistant 121 is good at, the supported dialogue type, etc.
[0060] The test case generation rules can be a set of general standards and guidelines for creating test cases. These rules ensure that the generated test cases are related to the prompt information of the target digital assistant 121 and can effectively test its functions and performance.
[0061] Exemplarily, the general question generation rule includes the following points: during the generation of the test case, ensure that the generated test case is related to the prompt information of the digital assistant. The generated test case must comply with the predefined rules. In the case of generating a test case including multiple rounds of interaction, the existing topic needs to be continued instead of introducing a new topic, ensuring the coherence of the multiple rounds of interaction. Only output the generated test case. Do not output anything else. The test case must be in pure text format. The generated test case should not contain the reply content of the digital assistant. Only one test case is generated each time. If multiple rounds of interaction, the next test case is generated in combination with the reply of the digital assistant. Avoid repeating previous test cases to maintain the uniqueness and continuity of the conversation.
[0062] The electronic device 110 can generate one or more groups of test cases by using the test case generation model. The prompt word information of the target digital assistant 121 and the general question generation rule can be input information of the test case generation model, and the reply content of the test case generation model. Based on the reply content of the test case generation model, one or more groups of test cases for the target digital assistant 121 can be obtained.
[0063] In addition, for the input information of the test case generation model, the role description of the test case generation model can be included. For example, the role description can be: you are a digital assistant for generating chat test cases. Your task is to generate test cases according to the prompt information of the target digital assistant.
[0064] In the embodiments of the present application, the example implementation manner of generating input information is described in the Chinese language environment. Alternatively and / or additionally, the corresponding scheme of generating input information can be performed in other language environments. For example, input information can be generated in Chinese, English, Japanese, French, etc. For example, based on the multi-language capability provided by the test case generation model, input information of the test case generation model can be generated in different language application environments.
[0065] In this way, the electronic device 110 can systematically generate multiple groups of test cases, providing a comprehensive and effective tool for the evaluation and optimization of the digital assistant.
[0066] For the acquisition of the test case, there can be another way. The electronic device 110 acquires the prompt word information of the target digital assistant 121, and the prompt word information at least includes the identification information and the function description of the target digital assistant 121. Based on at least one evaluation dimension related to the chat skill, at least one specific question generation rule corresponding to each of the at least one evaluation dimension is determined. At least based on the prompt word information and the at least one specific question generation rule, one or more groups of test cases for the target digital assistant 121 are generated.
[0067] The prompt word information of the target digital assistant 121 is the same as the foregoing example, and is not described herein again. The evaluation dimensions related to the chat skill can include identity awareness ability, function awareness ability, basic interaction ability, interaction ability positively related to the field of expertise, interaction ability negatively related to the field of expertise, processing ability of abnormal interaction, and the like.
[0068] Taking the evaluation dimension corresponding to the identity awareness ability as an example, the specific question generation rule can include: according to the prompt information of the target digital assistant, generating an identity awareness test case to guide the target digital assistant to say who it is. In addition, you also need to mislead the target digital assistant about its identity. The round of the interactive test is 2 rounds.
[0069] Taking the evaluation dimension corresponding to the function awareness ability as an example, the specific question generation rule can include: according to the prompt information of the target digital assistant, generating a function test case to guide the target digital assistant to say its function. The round of the interactive test is 1 round.
[0070] Taking the evaluation dimension corresponding to the basic interaction ability as an example, the specific question generation rule can include: according to the prompt information of the target digital assistant, generating a chat test case irrelevant to the field of expertise of the target digital assistant. The round of the interactive test is 1 round.
[0071] Taking the evaluation dimension corresponding to the interaction ability positively related to the field of expertise as an example, the specific question generation rule can include: according to the prompt information of the target digital assistant, generating a test case positively related to the field of expertise. The test case positively related to the field of expertise covers all functions of the target digital assistant. It is necessary to consider inputting real information and false information. The test case positively related to the field of expertise should not only ask the target digital assistant whether it can do something, but also let the target digital assistant actually do it. The generated test questions positively related to the field of expertise should not be independent of each other. They should continue the previous topic according to the chat history to have an in-depth chat. Only generate test questions, and do not output any information. The round of the interactive test is 5 rounds.
[0072] Taking the evaluation dimension corresponding to the interaction ability negatively related to the field of expertise as an example, the specific question generation rule can include: according to the prompt information of the target digital assistant, generating a function test completely irrelevant to the field of expertise of the target digital assistant. The test case negatively related to the field of expertise should not only ask the target digital assistant whether it can do something completely irrelevant to the field of expertise of the target digital assistant, but also let the target digital assistant actually do it. The round of the interactive test is 2 rounds.
[0073] Taking the evaluation dimension corresponding to the processing capability of interaction with the abnormal as an example, the specific problem generation rule can include: generating an abnormal interaction test case according to the prompt information of the target digital assistant. For example, if the target digital assistant requires inputting a picture, the content unrelated to the picture should be inputted, such as inputting text or audio. The round of interaction test is 1 round.
[0074] The electronic device 110 can generate one or more sets of test cases based on a test case generation model. The prompt word information of the target digital assistant 121 and the specific problem generation rule can be input information of the test case generation model, and one or more sets of test cases for the target digital assistant 121 can be generated based on the reply content of the test case generation model.
[0075] Through the above process, test cases can be generated based on multiple evaluation dimensions, so that the various capabilities of the target digital assistant 121 can be comprehensively evaluated. Specifically, the evaluation content includes identity recognition capability, function recognition capability, basic interaction capability, positive function interaction capability, negative function interaction capability, and abnormal processing capability, etc. This method not only can evaluate the actual chat skills of the target digital assistant 121, but also can identify its performance in different situations, thereby providing a reliable basis for recommending high-quality digital assistants.
[0076] Exemplarily, the first feature value includes a chat skill score corresponding to each evaluation dimension respectively. FIG. 3 shows a schematic diagram of a chat skill score interface 300 according to some embodiments of the present disclosure. In combination with FIG. 3, for each evaluation dimension test question, the target digital assistant 121 can generate a corresponding reply. A scoring model can be used to generate a chat skill score for each reply. For example, in FIG. 3, the evaluation dimension corresponding to the identity recognition capability includes two test questions, and the chat skill score of the first test question is 0.58. The chat skill score of the second test question is 0.56. Similarly, the chat skill scores corresponding to the test questions of the evaluation dimensions such as function recognition capability, basic interaction capability, positive function interaction capability, negative function interaction capability, and abnormal processing capability, etc. will also be referred to.
[0077] In some embodiments, the first feature value in the target evaluation index can include an aggregated value of the chat skill scores of each evaluation dimension. For example, the average value of the chat skill scores corresponding to each evaluation dimension can be calculated as the first feature value.
[0078] For the case of multiple rounds of interaction that can occur in an evaluation dimension, the generation process of test questions will be different from the previous example. Taking the evaluation dimension corresponding to the positive correlation of interaction ability in the field of expertise as an example, the number of rounds of interaction testing in this set of test cases is 5 rounds. For the scenario of multiple rounds of interaction, the electronic device 110 obtains the first reply of the target digital assistant 121 to the first test question in the first round of interaction with the target digital assistant 121. At least based on the first reply, a second test question for the second round of interaction of the target digital assistant 121 is generated.
[0079] Evaluating the positive correlation of interaction ability of the digital assistant in the field of expertise requires multiple rounds of interaction testing. This is to comprehensively test the ability of the digital assistant to handle complex conversations in its field of expertise. Taking the first round of interaction as an example, the electronic device 110 generates the first test question based on the prompt word information and the corresponding specific question generation rule as input information, using the test case generation model. Based on the first reply of the target digital assistant 121 to the first test question, the electronic device 110 can use the test case generation model to generate a second test question associated with the first reply. For example, if the first reply mentions the "breach of contract clause" in the contract, the second question can be "Can you explain the specific content of the breach of contract clause in detail, and what is the actual situation in this case?"
[0080] The interaction process of subsequent rounds can repeat the above process. The electronic device 110 performs third, fourth, and fifth rounds of interaction in turn. The test question of each round is at least based on the reply of the last round to ensure the coherence and depth of the test conversation. For example, the test question of the third round of interaction can discuss the legal consequences of the breach of contract clause. The test question of the fourth round of interaction can explore how to protect one's own rights in the contract. The test question of the fifth round of interaction can inquire about the application of the breach of contract clause in the actual case.
[0081] Through this multi-round interaction test, the electronic device 110 can comprehensively evaluate the ability of the digital assistant to handle complex conversations. Each round of test questions and replies is based on the content of the previous round, simulating real user interaction scenarios. This testing method not only examines the depth of knowledge and accuracy of the target digital assistant 121, but also evaluates its ability to maintain logical consistency and provide valuable information in continuous conversations. Ultimately, these test results will be used as an important indicator to evaluate the ability of the digital assistant, helping the developer 105 to identify and improve its interaction performance in a specific field.
[0082] In the foregoing embodiments, the target evaluation indicator is taken as an example of reflecting the chat skill of the target digital assistant 121. In this regard, in addition to indicating the chat skill of the target digital assistant, the target evaluation indicator can also be determined based on the configuration information of the target digital assistant 121. In some embodiments of the present disclosure, the electronic device 110 determines at least one second feature value in the target evaluation indicator for the target digital assistant 121 based on the configuration information of the target digital assistant 121. The target digital assistant 121 generates and presents a reply based on the configuration information, and each second feature value indicates a score of the target digital assistant 121 on one configuration type.
[0083] The digital assistant can be developed based on a digital assistant development platform 120. The digital assistant development platform 120 can provide various tools such as prompt word information, plugins, workflows, knowledge bases, memory bases, voices, and the like. Based on this, the configuration information can reflect the tools involved in the development process of the target digital assistant 121. Each tool can correspond to a configuration type.
[0084] Exemplarily, the second feature value can include a score of the target digital assistant 121 on the configuration type. For example, the score on the configuration type can indicate the number of supported voices, the number of recommended dialogues, the number of workflows, the number of plugins, the number of knowledge bases, the number of publishing platforms, whether there is a background picture, the number of memory bases, the number of bound cards, whether it is open source, and the like.
[0085] The number of supported voices can indicate the number of voice options supported by the target digital assistant 121, such as the number of male voices, female voices, child voices, and the like. Providing diverse voice options can improve user satisfaction and engagement.
[0086] The number of recommended dialogues can indicate the number of dialogues that the target digital assistant 121 can recommend to the user, that is, how many (related to the user's interests, needs, and historical dialogue records, etc.) relevant dialogue topics or topics the target digital assistant 121 can provide or recommend to the user for selection and continued interaction.
[0087] The number of workflows can indicate the number of workflows possessed by the target digital assistant 121. More workflows can enable more complex tasks and automated operations.
[0088] The number of plugins can indicate the number of plugins that the target digital assistant 121 can integrate. By integrating plugins, the target digital assistant 121 can expand its functions and provide more services and applications.
[0089] The number of knowledge bases can indicate the number of knowledge bases that the target digital assistant 121 can access and utilize. Rich knowledge bases can improve the response accuracy and information coverage of the digital assistant.
[0090] The number of publishing platforms can indicate how many platforms the target digital assistant 121 can publish on. The multi-platform publishing capability can expand the user group of the digital assistant, improve its popularity and usage rate.
[0091] Whether there is a background picture can indicate whether the target digital assistant 121 has a background image function. The background image can enhance the visual appeal and improve the user's interface experience.
[0092] The number of memory banks can indicate the number of memory banks to which the target digital assistant 121 is associated. The memory bank can be used at least to record the user's historical conversation, provide personalized service and continuous conversation context.
[0093] Whether open source can indicate whether the target digital assistant 121 supports code open source projects. If it is an open source project, the development and improvement of the digital assistant can be more convenient and accelerated.
[0094] In addition, the second feature value can also include an evaluation score of the prompt information of the target digital assistant 121. For example, the prompt information of the target digital assistant 121 can be input to a model with natural language evaluation function, and the corresponding evaluation score can be given by the model with natural language evaluation function.
[0095] When evaluating the target digital assistant 121, the scores of the digital assistant on the above different configuration types can be used as evaluation indicators. In this way, a comprehensive evaluation framework can be provided for the target digital assistant 121. Help users identify high-quality and reliable digital assistants, so as to better meet user needs.
[0096] In addition to the configuration information of the target digital assistant 121, the target evaluation indicator can also be determined based on historical interaction information related to the target digital assistant 121. In some embodiments of the present disclosure, the electronic device 110 determines at least one third feature value of the target evaluation indicator for the target digital assistant 121 based on the historical interaction information related to the target digital assistant 121, each third feature value indicating a score of the target digital assistant 121 on one user interaction type.
[0097] The historical interaction information can reflect the real-time performance of the target digital assistant 121 and the interaction with the user 135. Illustratively, the historical interaction information includes at least one of the following: the number of users interacting with the target digital assistant 121 within a period of time. The number of messages interacting with the target digital assistant 121 within a period of time. The number of at least one type of interaction behavior performed on the target digital assistant 121.
[0098] By incorporating the dynamic features corresponding to the historical interaction information into the evaluation indicators, such as the number of active users and the number of chat messages within a period of time, the user engagement and interaction amount of the digital assistant within a specific time period can be evaluated. By the number of collections, the number of likes, and the number of dislikes, the recognition degree and satisfaction of the user to the digital assistant can be understood. Thus, the performance and user feedback of the digital assistant in the actual use process can be comprehensively understood. The static features corresponding to the configuration information reflect the design and configuration of the digital assistant, while the dynamic features corresponding to the historical interaction information provide interaction data of the digital assistant in the actual use of the user. By combining the two types of features, the advantages and disadvantages of the target digital assistant 121 can be evaluated from multiple dimensions.
[0099] To improve the automation degree of the quality evaluation result and ensure the unified scale of the standard evaluation result. The quality evaluation result of the target digital assistant 121 is determined based on the target evaluation indicator by using the trained evaluation model. The quality evaluation result indicates the confidence of the target digital assistant being recommended, and the evaluation model is trained by: obtaining the first evaluation indicator of the digital assistant that has been recommended as a positive sample; obtaining the second evaluation indicator of the digital assistant that has not been recommended as a negative sample; and training the evaluation model by using the positive sample and the negative sample.
[0100] FIG. 4 shows a schematic diagram of a training process 400 of an evaluation model according to some embodiments of the present disclosure.
[0101] In block 401, preprocessing is first performed on the evaluation indicators in the positive and negative samples. Since the first evaluation indicator and the second evaluation indicator respectively include feature values corresponding to multiple feature types, and the value ranges of different feature values can differ greatly. For example, the feature type is the number of active users within a period of time or the number of chat messages, and the corresponding feature value can be hundreds or even tens of thousands. While for example, the feature type is the number of knowledge bases, which is usually only a single digit, and for example, the chat skill score is only between 0 and 1. Thus, it is necessary to convert the feature values corresponding to all evaluation indicators into the same numerical range through preprocessing.
[0102] In block 402, correlation calculation is performed on the evaluation indicators. That is, if the first evaluation indicator and the second evaluation indicator respectively include feature values corresponding to multiple feature types, the correlation between the multiple feature types in the first evaluation indicator and the second evaluation indicator can be determined. Based on the correlation between the multiple feature types, at least one feature type to be included in the target evaluation indicator is selected from the multiple feature types.
[0103] FIG. 5 shows a schematic diagram of a correlation distribution 500 according to some embodiments of the present disclosure. Each feature type in the first evaluation indicator and the second evaluation indicator is obtained, denoted as feature type 1 to feature type n (n is a positive integer) in FIG. 5. By calculation, the correlation distribution between the plurality of feature types can be determined. Based on the result of the correlation distribution, a decision is made on which feature types to discard. For example, if the correlation between two feature types is high, it can lead to a multicollinearity problem, affecting the stability and interpretability of the model. In this case, one of the feature types can be selected to be discarded. For another example, if it is determined that feature type i is a key feature, and the correlation between feature type 1 and feature type i (i ≤ n, and i is a positive integer) is 0.35, while the correlation between feature type 2 and feature type i is 0.8, based on the importance of the feature types, feature type 1 can also be selected to be discarded while feature type 2 is retained.
[0104] Based on the correlation between the feature types, at least one feature type to be included in the target evaluation indicator can be selected from the plurality of feature types. The selected feature type can be a feature type that has a more obvious effect on the evaluation result.
[0105] At block 403, a certain number of samples are randomly selected from the positive samples and the negative samples to ensure that the number of positive and negative samples used for model training is balanced. For example, 2000 first evaluation indicators of the recommended digital assistants can be selected as positive samples, and 2000 second evaluation indicators of the non-recommended digital assistants can be randomly selected from the 10000 available non-recommended digital assistants.
[0106] At block 404, the training of the evaluation model is performed. During the training of the evaluation model, logistic regression can be defined as the classification algorithm of the evaluation model. This step determines the basic structure of the evaluation model, i.e., using logistic regression to solve the binary classification problem (recommended or not recommended). Using the prepared training data to train the evaluation model can indicate the importance of the target evaluation indicator. For example, the target digital assistant 121 can correspond to 10 target evaluation indicators. Based on the importance of the target evaluation indicators determined by the evaluation model, the final score can be calculated to determine whether the target digital assistant 121 is worth being recommended. For example, the final score can be a value in the interval of 0 to 1. For example, a value such as 0.5 or 0.75 can be used as a threshold. If the value is higher than the threshold, the evaluation model will determine that it is worth being recommended. Thus, the classification function is realized.
[0107] Through the above process, the evaluation model can effectively learn from the evaluation indicators of the recommended and non-recommended digital assistants, and accurately predict the recommendation value of a new digital assistant. This not only improves the efficiency and accuracy of the recommendation system, but also provides users with better digital assistant selection.
[0108] With the trained evaluation model, the quality evaluation result of the target digital assistant 121 can be determined based on the target evaluation indicators. The quality evaluation result indicates the confidence level of the recommendation of the target digital assistant 121. In some embodiments of the present disclosure, the electronic device 110 can display the target digital assistant 121 on the recommendation interface in response to the quality evaluation result meeting the recommendation condition. The electronic device 110 can obtain the recommendation effect indicators of the target digital assistant 121 after being recommended. Based on the recommendation effect indicators, the electronic device 110 can update the evaluation model.
[0109] When the quality evaluation result meets the preset recommendation condition (for example, the recommendation confidence is higher than a certain threshold), the electronic device 110 displays the target digital assistant 121 on the recommendation interface of the digital assistant recommendation platform 130. In this way, the user 135 can conveniently discover and use the recommended digital assistant, thereby improving the overall user experience.
[0110] After the target digital assistant 121 is recommended, the electronic device 110 continues to monitor its recommendation effect indicators. These recommendation effect indicators can include but are not limited to the following aspects: user click rate, user retention rate, user satisfaction score, and usage frequency, etc. The user click rate can indicate the ratio of the number of times the user clicks the target digital assistant 121 to the number of times it is displayed on the recommendation interface. The user retention rate can indicate the proportion of users who continue to use the target digital assistant 121 after using it. The user satisfaction score can indicate the score or feedback of the user on the target digital assistant 121. The usage frequency can indicate the frequency of the user using the target digital assistant 121 within a period of time.
[0111] Based on the collected recommendation effect indicators, the electronic device 110 regularly updates the evaluation model. This process includes: collecting the recommendation effect indicator data of the target digital assistant 121 after being recommended. Analyzing these data, evaluating the prediction accuracy and effectiveness of the current evaluation model. According to the analysis result, the parameters of the evaluation model are adjusted. The specific steps can include adding new features, adjusting the weight of existing features, optimizing the hyperparameters of the model, etc. Use the updated data set to retrain the evaluation model to ensure that it still performs well in the new data environment. Model deployment: deploy the updated evaluation model to the system, replace the old model, so as to use the new model in the subsequent recommendation process.
[0112] By continuously obtaining and analyzing the recommendation effect indicators, the electronic device 110 can continuously improve the evaluation model, thereby improving the recommendation accuracy and user experience, and ultimately realizing a more intelligent and personalized digital assistant recommendation system.
[0113] FIG. 6 shows a schematic structural block diagram of an apparatus 600 for digital assistant evaluation, according to some embodiments of the present disclosure. The apparatus 600 may, for example, be implemented in or included in the electronic device 110. Various modules / components in the apparatus 600 can be implemented by hardware, software, firmware, or any combination thereof.
[0114] As shown, the apparatus 600 includes a test case acquisition module 601 configured to, in response to an evaluation request for a target digital assistant, acquire at least one set of test cases for the target digital assistant, each set of test cases including at least one test question related to a chat skill of the target digital assistant. A reply acquisition module 602 is configured to provide the at least one set of test cases to the target digital assistant to obtain a reply of the target digital assistant to the at least one set of test cases. A target evaluation indicator determination module 603 is configured to determine a target evaluation indicator for the target digital assistant based at least on the at least one set of test cases and the reply of the target digital assistant to the at least one set of test cases, the target evaluation indicator including at least a first feature value indicating a chat skill score of the target digital assistant. A quality evaluation result determination module 604 is configured to determine a quality evaluation result for the target digital assistant based on the target evaluation indicator.
[0115] In some embodiments of the present disclosure, the test case acquisition module 601 can be specifically configured to acquire prompt word information of the target digital assistant, the prompt word information including at least identification information and a function description of the target digital assistant.
[0116] A general question generation rule corresponding to each of the at least one set of test cases is acquired. One or more sets of test cases for the target digital assistant are generated based at least on the prompt word information and the general question generation rule.
[0117] In some embodiments of the present disclosure, the test case acquisition module 601 can be specifically configured to acquire prompt word information of the target digital assistant, the prompt word information including at least identification information and a function description of the target digital assistant. At least one specific question generation rule corresponding to each of the at least one evaluation dimension is determined based on the at least one evaluation dimension related to the chat skill. One or more sets of test cases for the target digital assistant are generated based at least on the prompt word information and the at least one specific question generation rule.
[0118] In some embodiments of the present disclosure, the first feature value includes a chat skill score corresponding to each of the at least one evaluation dimension.
[0119] In some embodiments of the present disclosure, the test case acquisition module 601 can be further configured to: in the first round of interaction with the target digital assistant, acquire a first reply of the target digital assistant to the first test question. At least based on the first reply, generate a second test question for the second round of interaction of the target digital assistant.
[0120] In some embodiments of the present disclosure, the target evaluation index determination module 603 can be configured to: based on the configuration information of the target digital assistant, determine at least one second feature value in the target evaluation index for the target digital assistant, the target digital assistant generating and presenting a reply based on the configuration information, each second feature value indicating a score of the target digital assistant on one configuration type.
[0121] In some embodiments of the present disclosure, the target evaluation index determination module 603 can be further configured to: based on the historical interaction information related to the target digital assistant, determine at least one third feature value in the target evaluation index for the target digital assistant, each third feature value indicating a score of the target digital assistant on one user interaction type.
[0122] In some embodiments of the present disclosure, the historical interaction information includes at least one of: a number of users interacting with the target digital assistant within a period of time, a number of messages interacting with the target digital assistant within a period of time, and a number of at least one type of interaction behavior performed on the target digital assistant.
[0123] In some embodiments of the present disclosure, the quality evaluation result of the target digital assistant is determined based on the target evaluation index by using a trained evaluation model.
[0124] In some embodiments of the present disclosure, the apparatus 600 further includes a model training module. The confidence degree of the target digital assistant being recommended is indicated based on the quality evaluation result. The training module is configured to: acquire a first evaluation index of a digital assistant that has been recommended as a positive sample. Acquire a second evaluation index of a digital assistant that has not been recommended as a negative sample. Train the evaluation model by using the positive sample and the negative sample.
[0125] In some embodiments of the present disclosure, the first evaluation index and the second evaluation index respectively include feature values corresponding to a plurality of feature types. The model training module can be further configured to: determine a correlation degree between the plurality of feature types in the first evaluation index and the second evaluation index. Based on the correlation degree between the plurality of feature types, select at least one feature type to be included in the target evaluation index from the plurality of feature types.
[0126] In some embodiments of the present disclosure, the quality evaluation result indicates a confidence degree at which the target digital assistant is recommended. The model training module can be further configured to: in response to the quality evaluation result satisfying a recommendation condition, display the target digital assistant on a recommendation interface; obtain a recommendation effect indicator of the target digital assistant after being recommended; and update the evaluation model based on the recommendation effect indicator.
[0127] FIG. 7 illustrates a block diagram of an electronic device 700 in which one or more embodiments of the present disclosure can be implemented. It should be understood that the electronic device 700 illustrated in FIG. 7 is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. The electronic device 700 illustrated in FIG. 7 can include or be implemented as the electronic device 110 of FIG. 1, or the apparatus 600 of FIG. 6.
[0128] As illustrated in FIG. 7, the electronic device 700 is in the form of a general electronic device. The components of the electronic device 700 can include, but are not limited to, one or more processors or processing units 710, a memory 720, a storage device 730, one or more communication units 740, one or more input devices 750, and one or more output devices 760. The processing unit 710 can be a real or virtual processor and is capable of performing various processing according to programs stored in the memory 720. In a multi-processor system, multiple processing units perform computer-executable instructions in parallel to improve the parallel processing capability of the electronic device 700.
[0129] The electronic device 700 typically includes a number of computer storage media. Such media can be any available media that is accessible by the electronic device 700 and includes both volatile and non-volatile media, removable and non-removable media. The memory 720 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 730 can be a removable or non-removable medium and can include machine-readable media such as a flash drive, a magnetic disk drive, or any other medium that can be used to store information and / or data and that can be accessed by the electronic device 700.
[0130] The electronic device 700 can further include additional detachable / non-detachable, volatile / non-volatile storage media. Although not shown in FIG. 7, a disk drive for reading from or writing to a detachable, non-volatile magnetic disk (e.g., a "floppy disk"), and an optical disk drive for reading from or writing to a detachable, non-volatile optical disk can be provided. In these cases, each drive can be connected to the bus (not shown) by one or more data media interfaces. The memory 720 can include a computer program product 725 having one or more program modules configured to carry out the various methods or acts of the various embodiments of the present disclosure.
[0131] The communication unit 740 enables communication with other electronic devices over communication media. Additionally, the functionality of the components of the electronic device 700 can be implemented in a single computing cluster or a plurality of computer machines capable of communicating over a communication connection. As such, the electronic device 700 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network nodes in the networking environment.
[0132] The input device 750 can be one or more input devices, such as a mouse, a keyboard, a trackball, etc. The output device 760 can be one or more output devices, such as a display, a speaker, a printer, etc. The electronic device 700 can also communicate with one or more external devices (not shown) such as a storage device, a display device, etc., one or more devices that enable a user to interact with the electronic device 700, or any devices (e.g., a network card, a modem, etc.) that enable the electronic device 700 to communicate with one or more other electronic devices, as needed, through the communication unit 740. Such communication can be carried out via an input / output (I / O) interface (not shown).
[0133] According to an example implementation of the present disclosure, there is provided a computer-readable storage medium having computer-executable instructions stored thereon, where the computer-executable instructions are executed by a processor to implement the method described above. According to an example implementation of the present disclosure, there is also provided a computer program product tangibly stored on a non-transitory computer-readable medium and comprising computer-executable instructions, where the computer-executable instructions are executed by a processor to implement the method described above.
[0134] According to an example implementation of the present disclosure, a computer program product or computer program is provided, which includes computer instructions stored in a computer readable storage medium. A processor of a computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions, so that the computer device performs the method provided in various optional manners in FIG. 2. Therefore, here will not be described in detail.
[0135] Various aspects of the present disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products according to this disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer readable program instructions.
[0136] The computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. The computer readable program instructions can also be stored in a computer readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer readable storage medium having instructions stored therein comprises an article of manufacture including instructions which implement aspects of the function / act specified in the flowchart and / or block diagram block or blocks.
[0137] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer implemented process such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0138] The computer program product of the present disclosure can have a signal including said computer program. This signal can be electronic, electromagnetic, optical, or any other suitable type of signal. Such a signal can be provided through a communication connection, such as electrical wiring, optical fiber, wireless interface, etc. Examples of computer program products include computer program implemented on a personal computer, server, or other networked device. A non-transitory computer readable medium, such as a floppy disk, CD-ROM, DVD-ROM, Blu-ray Disc, hard disk, or memory stick, can also be used to implement the present disclosure. The computer program product of the present disclosure can also be provided as a service to download and use the computer program over a network, such as the Internet.
[0139] Having described several implementations of the present disclosure, it will be clear to those skilled in the art that many modifications, additions, and substitutions are possible without departing from the scope and spirit of the described implementations. Many modifications and variations of the present disclosure are possible in light of the above teachings. It is, therefore, to be understood that within the scope of the appended claims, the present disclosure can be practiced otherwise than as specifically described. While the present disclosure has been described with reference to the implementation figures, it will be understood by those skilled in the art that various changes can be made and equivalents can be substituted for elements thereof without departing from the scope of the present disclosure. In addition, many modifications can be made to adapt to a particular situation and the teachings of the present disclosure to a specific implementation, without departing from the central novel teachings of the application. The implementation(s) illustrated and described herein are meant only to serve as examples. Departures in form and detail are within the scope of the disclosure. Therefore, one skilled in the art can restructure the implementation(s) as needed, while still adhering to the principles of the present disclosure.
Claims
1. A method for digital assistant evaluation, comprising: in response to an evaluation request for a target digital assistant, obtaining at least one set of test cases for the target digital assistant, each set of test cases comprising at least one test question related to a chat skill of the target digital assistant; providing the at least one set of test cases to the target digital assistant to obtain a reply of the target digital assistant to the at least one set of test cases; determining a target evaluation indicator for the target digital assistant based at least on the at least one set of test cases and the reply of the target digital assistant to the at least one set of test cases, the target evaluation indicator comprising at least a first feature value indicative of a chat skill score of the target digital assistant; and based on the target evaluation indicator, determining a quality evaluation result for the target digital assistant.
2. The method of claim 1, wherein obtaining at least one set of test cases for the target digital assistant comprises: obtaining prompt word information of the target digital assistant, the prompt word information comprising at least identification information and a functional description of the target digital assistant; obtaining a general question generation rule corresponding to each of the at least one set of test cases; and generating one or more sets of test cases for the target digital assistant based at least on the prompt word information and the general question generation rule.
3. The method of claim 1, wherein obtaining at least one set of test cases for the target digital assistant comprises: obtaining prompt word information of the target digital assistant, the prompt word information comprising at least identification information and a functional description of the target digital assistant; determining at least one specific question generation rule corresponding to each of at least one evaluation dimension related to the chat skill based on the at least one evaluation dimension; and generating one or more sets of test cases for the target digital assistant based at least on the prompt word information and the at least one specific question generation rule.
4. The method of claim 3, wherein the first feature value comprises a chat skill score corresponding to each of the at least one evaluation dimension.
5. The method of claim 1, further comprising, for a given set of test cases in the at least one set of test cases: obtaining a first reply of the target digital assistant to a first test question in a first round of interaction with the target digital assistant; and generating a second test question for a second round of interaction with the target digital assistant based at least on the first reply.
6. The method of claim 1, wherein determining the target evaluation indicator further comprises: determining at least one second feature value for the target digital assistant in the target evaluation indicator based on configuration information of the target digital assistant, the target digital assistant generating and presenting the reply based on the configuration information, each second feature value indicating a score of the target digital assistant on one configuration type.
7. The method of claim 1, wherein determining the target evaluation indicator further comprises: determine at least one third feature value of the target evaluation indicator for the target digital assistant, each third feature value indicating a score of the target digital assistant on one type of user interaction, based on historical interaction information related to the target digital assistant.
8. The method of claim 7, wherein the historical interaction information comprises at least one of: a number of users interacting with the target digital assistant in a period of time; a number of messages interacting with the target digital assistant in a period of time; and a number of at least one type of interaction behavior performed on the target digital assistant.
9. The method of claim 1, wherein a quality evaluation result of the target digital assistant is determined based on the target evaluation indicator, using a trained evaluation model.
10. The method of claim 9, wherein the quality evaluation result indicates a confidence level of the target digital assistant being recommended, and the evaluation model is trained by: obtaining first evaluation indicators of digital assistants that have been recommended as positive samples; obtaining second evaluation indicators of digital assistants that have not been recommended as negative samples; and training the evaluation model using the positive samples and the negative samples.
11. The method of claim 10, wherein the first evaluation indicators and the second evaluation indicators each comprise feature values corresponding to a plurality of feature types, the method further comprising: determining a correlation between the plurality of feature types in the first evaluation indicators and the second evaluation indicators; and selecting at least one feature type to be included in the target evaluation indicator from the plurality of feature types based on the correlation between the plurality of feature types.
12. The method of claim 1, wherein the quality evaluation result indicates a confidence level of the target digital assistant being recommended, the method further comprising: in response to the quality evaluation result satisfying a recommendation condition, displaying the target digital assistant on a recommendation interface; obtaining a recommendation effect indicator of the target digital assistant after being recommended; and updating the evaluation model based on the recommendation effect indicator.
13. An apparatus for digital assistant evaluation, comprising: a test case obtaining module configured to, in response to an evaluation request for a target digital assistant, obtain at least one set of test cases for the target digital assistant, each set of test cases comprising at least one test question related to a chat skill of the target digital assistant; a reply obtaining module configured to provide the at least one set of test cases to the target digital assistant to obtain replies of the target digital assistant to the at least one set of test cases; a target evaluation indicator determining module configured to determine a target evaluation indicator for the target digital assistant based at least on the at least one set of test cases and the replies of the target digital assistant to the at least one set of test cases, the target evaluation indicator comprising at least a first feature value indicating a chat skill score of the target digital assistant; and a quality evaluation result determining module configured to determine a quality evaluation result of the target digital assistant based on the target evaluation indicator. 14. An electronic device, comprising: at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions, when executed by the at least one processor, cause the electronic device to perform the method of any one of claims 1-12.
15. A computer-readable storage medium having computer-executable instructions stored thereon that are executable by a processor to implement the method of any one of claims 1-12.
16. A computer program product comprising computer-executable instructions that, when executed by a processor, implement the method of any one of claims 1-12.
Citation Information
Patent Citations
Question and answer model editing method and device, electronic equipment and storage medium
CN116882450A
Assessment method and device of large language model and electronic equipment
CN117112744A
NLP model performance evaluation method and device, storage medium and electronic equipment
CN117724965A
Method, device and equipment for debugging digital assistant and storage medium
CN118427074A
Recommendation integrated online digital sales service chat system
US20180218432A1