Method, device, equipment and medium for evaluating model
By receiving and comparing the responses of different machine learning models and providing evaluation controls, the problem of sparse feedback on machine learning model evaluation is solved, and more accurate model optimization is achieved.
Patent Information
- Application Number
- CN202510408298.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, the evaluation feedback of machine learning models is sparse, making it difficult to effectively collect the user's satisfaction with different models, resulting in difficulty in model optimization.
By receiving and presenting responses from different machine learning models and providing evaluation controls, users can compare and evaluate individual responses to determine the evaluation of the model.
Improve the quantity and quality of user reviews, helping to better optimize machine learning models to match user needs.
Smart Images

Figure CN120336471A_ABST
Abstract
Description
Technical Field
[0001] Implementations of the present disclosure generally relate to the field of machine learning, and particularly to methods, apparatuses, devices, and computer-readable storage media for evaluating machine learning models. Background Art
[0002] In the development process of machine learning models, it is necessary to evaluate the performance of different models. For example, multiple different models can be developed, and it is desired to understand the satisfaction levels of different users with the outputs of each model. Currently, technical solutions for evaluating models have been proposed. For example, different models can be provided to different users, and feedback from different users on the responses from the models (e.g., whether they are satisfied with the responses) can be received. However, in actual use, a large number of users may only read the responses from the models without submitting feedback, which results in only a very small number of user feedback being collected. The feedback is sparse and average, making it difficult to be used for later optimization of the models. Summary of the Invention
[0003] In a first aspect of the present disclosure, a method for evaluating a model is provided. In this method, in response to receiving a user request, a first response and a second response to the user request are received. The first response is determined using a first model, and the second response is determined using a second model. The first response and the second response, as well as evaluation controls for evaluating the first response and the second response, are presented. Based on the evaluation actions associated with the evaluation controls, evaluations of the first model and the second model are determined.
[0004] In a second aspect of the present disclosure, an apparatus for evaluating a model is provided. The apparatus includes: a receiving module configured to, in response to receiving a user request, receive a first response and a second response to the user request, where the first response is determined using a first model and the second response is determined using a second model; a presenting module configured to present the first response and the second response, as well as evaluation controls for evaluating the first response and the second response; and a determining module configured to determine evaluations of the first model and the second model based on the evaluation actions associated with the evaluation controls.
[0005] In a third aspect of the present disclosure, an electronic device is provided. The electronic device includes: at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor. The instructions, when executed by the at least one processor, cause the electronic device to execute the method according to the first aspect of the present disclosure.
[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. The computer program, when executed by a processor, causes the processor to implement the method according to the first aspect of the present disclosure.
[0007] In a fifth aspect of the present disclosure, there is provided a computer program product comprising a computer program, wherein the computer program, when executed by a processor, implements the method according to the first aspect of the present disclosure.
[0008] It should be understood that the content described in this section is not intended to limit the key features or important features of the implementation manners of the present disclosure, nor to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In the following, with reference to the accompanying drawings and the following detailed description, the above and other features, advantages and aspects of the implementation manners of the present disclosure will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:
[0010] Figure 1 A block diagram of an application environment according to an implementation manner of the present disclosure is shown;
[0011] Figure 2 A block diagram of a model for evaluation according to some implementation manners of the present disclosure is shown;
[0012] Figure 3 A block diagram of determining a response according to some implementation manners of the present disclosure is shown;
[0013] Figure 4 A block diagram of an interface for switching responses according to some implementation manners of the present disclosure is shown;
[0014] Figure 5 A block diagram of an interface for switching responses according to some implementation manners of the present disclosure is shown;
[0015] Figure 6 A block diagram of an interface including a switched interface according to some implementation manners of the present disclosure is shown;
[0016] Figure 7 A block diagram of an interface for selecting an evaluation reason according to some implementation manners of the present disclosure is shown;
[0017] Figure 8 A block diagram of an interface for submitting various types of user requests according to some implementation manners of the present disclosure is shown;
[0018] Figure 9 A flowchart of a method for evaluating a model according to some implementation manners of the present disclosure is shown;
[0019] Figure 10 A block diagram of an apparatus for evaluating a model according to some implementation manners of the present disclosure is shown; and
[0020] Figure 11 A block diagram of a device capable of implementing multiple implementations of the present disclosure is shown. Detailed implementation manners
[0021] The implementations of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain implementations of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the implementations set forth herein. On the contrary, these implementations are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and implementations of the present disclosure are only for exemplary purposes and are not intended to limit the protection scope of the present disclosure.
[0022] In the description of the implementations of the present disclosure, the term "including" and its like should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "an implementation" or "the implementation" should be understood as "at least one implementation". The term "some implementations" should be understood as "at least some implementations". There may also be other explicit and implicit definitions hereinafter. As used herein, the term "model" may represent the association relationship between various data. For example, the above-mentioned association relationship can be obtained based on various technical solutions known currently and / or to be developed in the future.
[0023] It can be understood that the data involved in the technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the corresponding laws, regulations and related provisions.
[0024] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner according to the relevant laws and regulations.
[0025] For example, when responding to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested by it will require the acquisition and use of the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server or a storage medium that executes the operation of the technical solution of the present disclosure according to the prompt message.
[0026] As an optional but non-limiting implementation manner, in response to receiving an active request from a user, the manner of sending a prompt message to the user may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry selection controls for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0027] It can be understood that the above notification and user authorization acquisition process is only illustrative and does not limit the implementation manner of the present disclosure. Other manners that comply with relevant laws and regulations can also be applied to the implementation manner of the present disclosure.
[0028] The term "in response to" used herein represents a state in which a corresponding event occurs or a condition is satisfied. It will be understood that the execution timing of subsequent actions executed in response to the event or condition and the time when the event occurs or the condition is established are not necessarily strongly correlated. For example, in some cases, the subsequent action can be executed immediately when the event occurs or the condition is established; in other cases, the subsequent action can be executed after a period of time after the event occurs or the condition is established.
[0029] Example environment
[0030] Multiple different models can be developed, and it is desired to understand the satisfaction levels of different users with the outputs of each model. Currently, technical solutions for evaluating models have been proposed. For example, different models can be provided to different users, and feedback from different users on the responses from the models (such as whether they are satisfied with the responses) can be received. Figure 1 FIG. 100 is a block diagram showing an application environment according to an implementation manner of the present disclosure. As Figure 1 shown, an application 130 can be provided, and a user can submit a user request 110 in the application 130. For example, "introduce the tourism overview of XX island". The application 130 can call a machine learning model 140 to answer questions from the user. Specifically, the application 130 can forward the user request 110 to the machine learning model 140 and present a response 120 from the machine learning model 140 to the user request 110.
[0031] To better meet user needs, different machine learning models 140 can be developed and the evaluations of each user on the machine learning models can be determined, so as to update the machine learning models in a direction that makes the responses more matching the user needs. For example, model a and model b can be developed, and provided as Figure 1The page shown is used to determine the feedback of different users on the outputs of Model A and Model B. Specifically, different models can be provided to different users. For example, Model A can be provided to User A and Model B can be provided to User B. At this time, all questions from User A are replied by Model A, and all questions from User B are replied by Model B.
[0032] Users can use control 130 to submit "satisfied" and control 132 to submit "dissatisfied". However, in actual use, a large number of users may only read the responses from the model without submitting feedback, which results in only a very small number of user feedback being collected. The feedback is sparse and average, making it difficult to be used for subsequent model optimization. At this time, it is expected to obtain the evaluation of the machine learning model by a more effective method.
[0033] Summary of model evaluation
[0034] To at least partially address the deficiencies in the prior art, according to an implementation of the present disclosure, a method for evaluating a model is proposed. Refer to Figure 2 Describe the overview according to an implementation of the present disclosure, which Figure 2 shows a block diagram 200 for evaluating a model according to some implementations of the present disclosure. As Figure 2 shown, a user can input a user request 210 in an application 230. The application 230 can call multiple machine learning models, for example, machine learning models 240-1, …, 240-2 (collectively referred to as machine learning models 240). In response to receiving the user request 210, multiple responses to the user request 210 can be received, for example, a first response and a second response. Here, the first response (for example, response 220) is determined using a first model (for example, machine learning model 240-1), and the second response is determined using a second model (for example, machine learning model 240-2).
[0035] According to some implementations of the present disclosure, the first response, the second response, and evaluation controls (for example, at least any one of controls 230-1, 230-2, 230-3, 230-4) for evaluating the first response and the second response can be presented. The above controls can be collectively referred to as control 230. The user can input their evaluation of each response via control 230. Subsequently, based on the evaluation actions associated with the evaluation controls (for example, the user clicks on the control), the evaluation of the first model and the second model can be determined.
[0036] According to some implementations of the present disclosure, in the initial stage, since the user has not viewed the two answers yet, the control 230 can be in an invalid state at this time (for example, displayed in gray scale and cannot be clicked). In this way, it is possible to prevent the user from randomly submitting an evaluation without fully reading the two answers. After determining that the user has viewed the two answers, the control 230 can be in a valid state (for example, normally displayed and can be clicked), and at this time, the user can use the control 230 to select the answer that better meets their own needs.
[0037] By using the implementations of the present disclosure, responses from different machine learning models can be presented to the user. At this time, the user can view and compare different responses, and then evaluate each response. Compared with the prior art solutions that only present responses from a single machine learning model to the user, when faced with multiple responses, the user is more inclined to select the response that better matches their own expectations from the multiple responses. In this way, evaluations from a large number of users can be received, increasing the number of obtained evaluations and further optimizing the machine learning model.
[0038] Detailed process of model evaluation
[0039] The overview of some implementations according to the present disclosure has been described. In the following, more details about obtaining evaluations are provided. According to some implementations of the present disclosure, an application can call multiple machine learning models to perform different types of tasks. Specifically, the application can include different call branches for performing corresponding tasks. At this time, it is possible to first determine the call branch corresponding to the user request, and then determine the first model and the second model for providing responses.
[0040] Specifically, during the process of receiving the first response and the second response, the first model and the second model can be determined based on the intent associated with the user request; and the first response from the first model and the second response from the second model are received. By using some implementations of the present disclosure, multiple machine learning models can be organized in a more refined manner, so that each model can provide a response that better matches the user request. Refer to Figure 3 for more details. The Figure 3 shows a block diagram 300 for determining responses according to some implementations of the present disclosure.
[0041] As Figure 3As shown, it can receive a user request 210 and determine a user intention 310 based on the user request 210. The user intention 310 can be determined in various ways. Here, the user intention can include a query intention 312, a generation intention 314, …, and a translation intention 316, etc. For example, a pre-trained intention recognition model can be used to determine the user intention. Here, the intention recognition model can be trained based on training samples (including the user request and the true value sample of the intention). Alternatively and / or additionally, the user intention can be determined based on text processing. For example, if keywords such as “search” or “query” are detected in the user request 210, the user intention can be determined as the query intention 312; if keywords such as “generate” or “create” are detected in the user request 210, the user intention can be determined as the generation intention 314; if keywords such as “Chinese to English translation” or “translation” are detected in the user request 210, the user intention can be determined as the translation intention 316, etc.
[0042] As shown at block 320, subsequent steps can be determined based on the user intention. For example, at model invocation 322, the first model and the second model to be invoked can be determined from the invocation branches that match the user intention. The user request 210 can be input to the first model and the second model respectively, and then responses 328 (e.g., including a first response and a second response) from each model can be received. Alternatively and / or additionally, in addition to the user request 210, context data associated with the user request 210 can be input to each model respectively. For example, one or more user requests before the user request 210 and the corresponding responses can be used as context data. At this time, the model will process the current user request in the context language environment. In this way, the machine learning model can understand the user request in a more accurate manner, so as to output a more accurate response. Further, multiple responses can be provided to the user so that the user can select a response that better matches their own needs. Using some implementation manners of the present disclosure, a more accurate user evaluation can be obtained, which is convenient for further optimizing the machine learning model in the later stage.
[0043] According to some implementation manners of the present disclosure, some user requests may involve other tools. For example, the user request “introduce the tourism overview of XX island” is a query intention, and this user request involves multiple steps: 1) call a search tool to search for relevant articles on the tourism of XX island; 2) generate a summary based on the found relevant articles. At this time, it can be further determined that the user intention involves a tool intention 324, and during the tool invocation 326, a search tool can be called to search for relevant articles on the tourism of XX island. The search results can be returned to the model invocation 322, and a machine learning model can be used to summarize the content of the relevant articles, and then a response 328 can be provided.
[0044] In accordance with some implementations of the present disclosure, the first response and the second response can be presented in multiple layouts. Specifically, based on the display parameters of the terminal device on which the application is running, the layout for presenting the first response and the second response can be determined; and the first response and the second response can be presented based on the layout. Here, the display parameters can include, for example: the display mode of the terminal device and / or the size of the display area. For example, the first response and the second response can be presented based on the display mode of the terminal device on which the application is running and / or the size of the display area.
[0045] In accordance with some implementations of the present disclosure, if it is determined that the display mode is portrait mode and / or the display area is small, the response can be presented at the same position in the display area (e.g., the position where response 220 is located) in the manner shown in Figure 2 . The user can click on controls 212 and 214 to switch between the first response and the second response. In this way, multiple responses can be presented in a limited screen display space, facilitating the user to view and compare multiple responses. Alternatively and / or additionally, if it is determined that the display mode is landscape mode and / or the display area is large, the first response and the second response can be presented at different positions on the terminal device in a side-by-side manner. In this way, it is convenient for the user to view and compare multiple responses. Alternatively and / or additionally, the display layout can be determined based on the data volume of the response. If the response includes a large number of words, the response can be displayed in the manner shown in Figure 2 . If the response includes a small number of words, the first response and the second response can be presented at different positions on the terminal device in a side-by-side manner.
[0046] In accordance with some implementations of the present disclosure, during the process of presenting the first response and the second response, the target response among the first response and the second response and the switching control can be presented in the initial stage; and in response to receiving a switching action associated with the switching control, the other response different from the target response among the first response and the second response can be presented. Using some implementations of the present disclosure, different responses can be presented in a limited display area, thereby facilitating the user to read and compare each response, and thus providing an evaluation of each response.
[0047] See Figure 2 and Figure 4 for more details. In the initial stage, an interface as shown in Figure 2 can be presented. At this time, by default, the first response (i.e., response 220) is presented first. The switching control (e.g., control 214) can be presented, and in response to receiving a switching action for control 214, a page as shown in Figure 4 can be presented, which Figure 4 shows block diagram 400 of an interface for switching responses according to some implementations of the present disclosure. As shown in Figure 4As shown, the control 214 is pressed, and the response 410 can be presented at a position corresponding to the response 220 in Figure 2 . The response 410 can be generated by the second model.
[0048] According to some implementations of the present disclosure, the first response and the second response are presented at the same position, and the switching control includes a first switching control for presenting the first response and a second switching control for presenting the second response. Specifically, as Figure 4 shown, the control 212 can correspond to the first response, and the control 214 can correspond to the second response. The user can click the control 212 to present the response 220 as Figure 2 shown, and the user can click the control 214 to present the response 410 as Figure 4 shown. At this time, each response has a corresponding switching control, so that the user can present the desired response according to their own needs. With some implementations of the present disclosure, more responses can be displayed in a limited display space.
[0049] It should be understood that although Figure 4 shows an example of setting a switching control for each response, alternatively and / or additionally, a single switching control can be presented. Refer to Figure 5 for more details. The Figure 5 shows a block diagram 500 of an interface for switching responses according to some implementations of the present disclosure. As Figure 5 shown, in response to receiving the user request 210, the interface as Figure 5 shown can be presented. At the initial stage, the response 220 from the first model can be presented, the control 510 for switching answers can be presented, and the control 520 for evaluating the response 220 can be presented. The user can click the control 510 to present another response, and the user can click the control 520 to give an evaluation of the response 220.
[0050] It should be understood that before the user clicks and views all of the multiple responses, the control 520 may be in an inactive state. For example, it may be displayed in gray and cannot be clicked. Alternatively and / or additionally, after the user clicks and views the multiple responses, the control 520 may be in an active state. For example, it may be normally displayed and can be clicked. It should be understood that the user may click the evaluation control due to accidental operation without viewing each response, or the user may randomly click the evaluation control. With some implementations of the present disclosure, the above situations can be effectively avoided, thereby improving the accuracy of user evaluation. According to some implementations of the present disclosure, the predicted time length for the user to read each response can be estimated based on the data volume of the response, and when the user's reading time exceeds the predicted time length, the control 520 is converted to an active state. In this way, the risk of the user submitting an inaccurate evaluation can be reduced.
[0051] According to some implementations of the present disclosure, in response to receiving an interaction action for the control 510, a page as Figure 6 shown may be presented, and the Figure 6 figure shows a block diagram 600 including the switched interface according to some implementations of the present disclosure. As Figure 6 shown, the response 410 from the second model may be presented, the control 510 for switching the answer may be presented, and the control 610 for evaluating the response 410 may be presented. The user can click the control 510 to present another answer, and the user can click the control 610 to give an evaluation of the response 410.
[0052] According to some implementations of the present disclosure, during the process of presenting the first response and the second response, the presentation order of the first response and the second response may be determined based on a predetermined condition; and the first response and the second response are presented in the presentation order. Specifically, the response to be presented first may be randomly determined. At this time, the response presented first in the application is randomly determined. Alternatively and / or additionally, the response to be presented first may be determined in an alternating manner. Assuming that the first response from the first model is presented first in the previous round, the second response from the second model may be presented first in the current round, and the first response from the first model may be presented first in the subsequent rounds, and so on.
[0053] Alternatively and / or additionally, the presentation order can be determined in a polling manner. For example, the presentation order can be determined in a polling manner at the server in the chronological order of receiving user requests. Here, the user requests can be requests from different users. For the first user request from the first user, the first response from the first model can be presented first; for the second user request from the second user, the second response from the second model can be presented first; for the third user request from the third user, the first response from the first model can be presented first, and so on. With some implementations of the present disclosure, the first response and the second response can be presented in any order. In this way, it can prompt the user to browse and compare each response, and avoid the problem that the user always selects the first presented response, thereby improving the accuracy of the evaluation.
[0054] According to some implementations of the present disclosure, the time lengths for the user to read each response can be determined, and then whether the evaluation submitted by the user is valid can be determined. Specifically, in response to determining that the first time length for presenting the first response and the second time length for presenting the second response satisfy a predetermined time length, the evaluation can be marked as a valid evaluation. Alternatively and / or additionally, in response to determining that the first time length for presenting the first response and the second time length for presenting the second response do not satisfy the predetermined time length, the evaluation can be marked as an invalid evaluation. Here, the predetermined time length can be determined based on the data volume of the response. For example, if the response involves a large number of words, the predetermined time length can be set to a higher value; if the response involves a small number of words, the predetermined time length can be set to a lower value. With some implementations of the present disclosure, only user evaluations with a reasonable reading speed within a reasonable time range can be considered, thereby excluding incorrect evaluations caused by too fast reading speed.
[0055] According to some implementations of the present disclosure, the machine learning model can consider the context of the user request. Alternatively and / or additionally, the user can create a new session. At this time, in response to receiving a creation request for creating a new session, the context associated with the user request can be cleared. With some implementations of the present disclosure, the user is allowed to exclude historical context data during the conversation process, so as to more accurately test the answering ability of the machine learning model without context interference.
[0056] According to some implementations of the present disclosure, the evaluation can be provided in a more refined manner. Specifically, selection controls associated with the evaluation can be presented; and based on the selection action associated with the selection control, the evaluation reason associated with the evaluation can be determined. See Figure 7 Describe more details related to providing a refined evaluation. The Figure 7 shows a block diagram 700 of an interface for selecting an evaluation reason according to some implementations of the present disclosure. As Figure 7As shown, controls for submitting a refined evaluation can be presented. Assume that the user has submitted an evaluation: Answer 1 is better than Answer 2. Controls 710-1, 710-2, 710-3, and 710-4 (collectively referred to as control 710) for submitting a refined evaluation can be presented. It should be understood that Figure 7 only exemplary controls are schematically shown. Alternatively and / or additionally, the controls can include other content. For example, the user can be prompted to submit the reasons for selecting Answer 1, including but not limited to: accurate information, meeting requirements, etc. Using some implementations of the present disclosure, a refined evaluation can be determined, thereby obtaining richer evaluation data for further optimizing the machine learning model.
[0057] According to some implementations of the present disclosure, the first model and the second model can be multimodal models. Refer to Figure 8 for more details about the inputs and outputs of the machine learning model, the Figure 8 block diagram 800 of an interface for submitting various types of user requests according to some implementations of the present disclosure is shown. As Figure 8 shown, the application 230 can provide a welcome message 810, and the user can input data of different modalities in the input box 820. Further, the user can click on controls 811 to 818 to perform corresponding types of tasks.
[0058] Specifically, control 811 can specify a writing task. The user can input the topic of the article, the writing language, the word limit, etc. in the input box 820. Alternatively and / or additionally, an image, audio, or video, etc. can be input, and the machine learning model is required to write an article based on the input content. Control 812 can specify an image generation task from text. The user can input a description related to the image in the input box 820 to instruct the machine learning model to generate an image. Control 813 can specify a search task. The user can input keywords, etc. in the input box 820 to search for corresponding results. Control 814 can specify a reading summary task. The user can input the content of the article, a file, or the access address of a web article, etc. in the input box 820 to obtain the summary content of the article. Control 815 can specify an audio generation task. The user can input requirements for the audio or an audio file for reference during the generation process, etc. in the input box 820 to instruct the machine learning model to generate audio. Control 816 can specify a video generation task. The user can input a description related to the video in the input box 820 to instruct the machine learning model to generate a video. Control 817 can specify a translation task. The user can input the text or file to be translated and the target language in the input box 820 to specify that the machine learning model generates a translation result in the target language. Alternatively and / or additionally, control 818 can specify other types of tasks.
[0059] According to some implementations of the present disclosure, to avoid interference with the normal use of the machine learning model during the evaluation process, the methods described above may be executed at a predetermined frequency. For example, it may be specified that the evaluation process is executed only once a day (or a smaller number of times, such as three times, etc.). According to some implementations of the present disclosure, regardless of whether the user submits an evaluation or whether the evaluation submitted by the user is valid, the evaluation interface may be provided to the user only once.
[0060] According to some implementations of the present disclosure, if the user just starts using the application and inputs the first user request, to reduce interference with the user, it may be prohibited to call the evaluation method described above during the processing of the first user request. In this way, the problem of disturbing the normal use of the user at the initial stage of using the application and causing the user to reduce the interest in using it can be alleviated.
[0061] According to some implementations of the present disclosure, the evaluation method described above may be called during the processing of subsequent user requests. For example, it may be estimated based on user feedback whether the user is satisfied with the response of the model, and the evaluation method described above is executed when it is determined that the degree of satisfaction is low. Specifically, it is assumed that it is found that the user has input multiple similar user requests for a certain topic, for example, "XX island tourism", "XX island strategy", "accommodation and food on XX island", etc., which indicates that the user may not be very satisfied with the response provided by the model. At this time, the evaluation method described above may be provided so that the user can compare the responses from different models and thus select a response that better matches their own needs.
[0062] According to some implementations of the present disclosure, after receiving a valid evaluation from the user, the model selected by the user in the evaluation may be used to serve the subsequent requests of the user. In this way, the user can be supported to actively select a machine learning model that better matches their own needs. Assume that the application has been calling the first model to answer questions, and during the evaluation process, the user believes that the second response of the second model is better. At this time, the second response from the second model may be used to update the context data of the session. In this way, a response that better meets the user's own needs can be provided to the user without disturbing the user. Alternatively and / or additionally, during the subsequent use of the user, the second model may be called to process subsequent user requests.
[0063] According to some implementations of the present disclosure, the content presented in the application can be determined based on the user's actions. For example, assume that the user clicks the switch control and submits an evaluation after browsing two responses. At this time, it can be considered that the evaluation is valid and the response selected by the user is kept presented on the page. Another example, assume that the user clicks the switch control and does not submit an evaluation after browsing two responses. At this time, the current content can continue to be presented on the screen, and voting cannot be collected at this time. Another example, assume that the user does not click the switch control and continues to operate (such as submitting a subsequent user request, clearing the chat history, not typing for a long time, etc.). At this time, the current content can continue to be presented on the screen. Another example, assume that the user starts a new session, then the context is cleared. It should be understood that although a valid evaluation is not obtained in some of the above examples, it is still considered that the evaluation method has been executed, and the number of executable times within a predetermined time period can be reduced by one. Assume that the evaluation method is only allowed to be executed once a day, then the evaluation interface is not provided to the user on that day. In this way, evaluation data can be collected with minimal interruption to the user.
[0064] According to some implementations of the present disclosure, the received evaluations can be used to determine various evaluation parameters, so as to support further optimization of the machine learning model. For example, evaluation parameters as shown in Table 1 below can be generated based on the user evaluations.
[0065] Table 1 Examples of Evaluation Parameters
[0066] Date Model Rank Score Lower bound Upper bound Win Draw Neither good Vote Length 0101 M1 5 X1 X2 X3 X4 X5 X6 X7 X8 0101 M2 5 Y1 Y2 Y3 Y4 Y5 Y6 Y7 Y8
[0067] As shown in Table 1, the date represents the date when the evaluation is collected, the model represents the name of the model, the ranking represents the ranking according to the confidence interval of the model, the score represents the score of the model determined based on the user evaluation, the lower bound represents the lower bound of the confidence interval of the evaluation, the upper bound represents the upper bound of the confidence interval of the evaluation, the wins represent the number of times the model wins, the ties represent the number of times the model ties (i.e., the user selects "both are good"), the "both are bad" represents the number of times the user selects "both are bad", the votes represent the total number of votes participating in the evaluation, the length represents the number of words in the response output by the model, and so on. Further, data annotation can be guided based on the above evaluation parameters, and the labeled data can be used to fine-tune the machine learning model. In this way, the fine-tuned machine learning model can be made to better match the true needs of the user.
[0068] It should be understood that although only a situation with two machine learning models is schematically shown above, alternatively and / or additionally, there may be three or more models. Responses from three or more models can be presented to the user so that the user can evaluate each model. Alternatively and / or additionally, in order to avoid too many responses interfering with the normal usage process of the user, two models can be selected from multiple models, and only two responses from the two models are presented each time. For different user requests, different responses from different models can be presented. For example, for user request 1, responses from model 1 and model 2 can be presented; for user request 2, responses from model 3 and model 4 can be presented; for user request 3, responses from model 1 and model 3 can be presented, and so on.
[0069] According to some implementations of the present disclosure, an evaluation scheme based on crowdsourcing testing is provided. Compared with the prior art solutions that provide different models to a specific user group, the technical solution of the present disclosure can widely disperse the evaluation tasks among a large number of users, reduce the interference to individual users, and can collect real user needs in a more effective way. Further, compared with the prior art solutions that perform manual evaluation for predefined user requests, the technical solution of the present disclosure can evaluate the accuracy of the responses of the model to uncertain user requests from a large number of users. In this way, the coverage of the testing process can be greatly improved. Further, compared with the automatic evaluation that can only provide objective dimension judgments, the crowdsourcing testing evaluation can use the feedback from a large number of users to evaluate the advantages and disadvantages of the model, thereby providing data support for further optimizing the machine learning model in the future.
[0070] Example process
[0071] Figure 9 A flowchart of a method 900 for evaluating a model according to some implementations of the present disclosure is shown. At block 910, in response to receiving a user request, a first response and a second response to the user request are received, where the first response is determined using a first model and the second response is determined using a second model. At block 920, the first response and the second response, and an evaluation control for evaluating the first response and the second response are presented. At block 930, an evaluation of the first model and the second model is determined based on an evaluation action associated with the evaluation control.
[0072] According to some implementations of the present disclosure, presenting the first response and the second response includes: presenting a target response among the first response and the second response and a switching control; and in response to receiving a switching action associated with the switching control, presenting another response different from the target response among the first response and the second response.
[0073] According to some implementations of the present disclosure, the first response and the second response are presented at the same position, and the switching control includes a first switching control for presenting the first response and a first switching control for presenting the first response.
[0074] According to some implementations of the present disclosure, presenting the first response and the second response includes: determining a presentation order of the first response and the second response based on a predetermined condition; and presenting the first response and the second response in the presentation order.
[0075] According to some implementations of the present disclosure, receiving the first response and the second response includes: determining a first model and a second model based on an intention associated with a user request; and receiving the first response from the first model and the second response from the second model.
[0076] According to some implementations of the present disclosure, the method further includes at least any one of the following: in response to determining that a first time length for presenting the first response and a second time length for presenting the second response satisfy a predetermined time length, marking the evaluation as a valid evaluation; or in response to determining that the first time length for presenting the first response and the second time length for presenting the second response do not satisfy the predetermined time length, marking the evaluation as an invalid evaluation.
[0077] According to some implementations of the present disclosure, the method further includes: in response to receiving a creation request for creating a new session, clearing the context associated with the user request.
[0078] According to some implementations of the present disclosure, presenting the first response and the second response by the method further includes: determining a layout for presenting the first response and the second response based on display parameters of the terminal device; and presenting the layout of the first response and the second response based on the layout.
[0079] According to some implementations of the present disclosure, the method further includes: presenting a selection control associated with the evaluation; and determining an evaluation reason associated with the evaluation based on a selection action associated with the selection control.
[0080] According to some implementations of the present disclosure, the first model and the second model are multimodal models, and the method further includes: executing the method at a predetermined frequency.
[0081] Example device and equipment
[0082] Figure 10A block diagram of an apparatus 1000 for evaluating a model according to some implementations of the present disclosure is shown. The apparatus 1000 includes: a receiving module 1010 configured to receive a first response and a second response to a user request in response to receiving the user request, where the first response is determined using a first model and the second response is determined using a second model; a presenting module 1020 configured to present the first response and the second response, and an evaluation control for evaluating the first response and the second response; and a determining module 1030 configured to determine an evaluation of the first model and the second model based on an evaluation action associated with the evaluation control.
[0083] According to some implementations of the present disclosure, the presenting module 1020 is further configured to: present a target response among the first response and the second response and a switching control; and in response to receiving a switching action associated with the switching control, present another response different from the target response among the first response and the second response.
[0084] According to some implementations of the present disclosure, the first response and the second response are presented at the same position, and the switching control includes a first switching control for presenting the first response and a first switching control for presenting the first response.
[0085] According to some implementations of the present disclosure, the presenting module 1020 is further configured to: determine a presentation order of the first response and the second response based on a predetermined condition; and present the first response and the second response in the presentation order.
[0086] According to some implementations of the present disclosure, the receiving module 1010 is further configured to: determine the first model and the second model based on an intention associated with the user request; and receive the first response from the first model and the second response from the second model.
[0087] According to some implementations of the present disclosure, the apparatus 1000 further includes a processing module configured to: in response to determining that a first time length for presenting the first response and a second time length for presenting the second response satisfy a predetermined time length, mark the evaluation as a valid evaluation; or in response to determining that the first time length for presenting the first response and the second time length for presenting the second response do not satisfy the predetermined time length, mark the evaluation as an invalid evaluation.
[0088] According to some implementations of the present disclosure, the apparatus 1000 further includes a processing module configured to: in response to receiving a creation request for creating a new session, clear the context associated with the user request.
[0089] According to some implementations of the present disclosure, the presentation module 1020 is further configured to: determine a layout for presenting the first response and the second response based on the display parameters of the terminal device; and present the layout of the first response and the second response based on the layout.
[0090] According to some implementations of the present disclosure, the apparatus 1000 further includes a processing module configured to: present a selection control associated with an evaluation; and determine an evaluation reason associated with the evaluation based on a selection action associated with the selection control.
[0091] According to some implementations of the present disclosure, the first model and the second model are multimodal models, and the apparatus further includes a processing module configured to: execute the method at a predetermined frequency.
[0092] Figure 11 A block diagram of a device 1100 capable of implementing multiple implementations of the present disclosure is shown. It should be understood that Figure 11 The illustrated computing device 1100 is merely exemplary and should not impose any limitation on the functionality and scope of the implementations described herein. Figure 11 The illustrated computing device 1100 can be used to implement the methods described above.
[0093] As Figure 11 shown, the computing device 1100 is in the form of a general-purpose computing device. The components of the computing device 1100 may include, but are not limited to, one or more processors 1110, a memory 1120, a storage device 1130, one or more communication units 1140, one or more input devices 1150, and one or more output devices 1160. The processor 1110 can be an actual or virtual processor and can execute various processes according to the programs stored in the memory 1120. In a multi-processor system, multiple processors execute computer-executable instructions in parallel to improve the parallel processing ability of the computing device 1100.
[0094] The computing device 1100 generally includes multiple computer storage media. Such media can be any accessible media available to the computing device 1100, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 1120 can be a volatile memory (e.g., registers, caches, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 1130 can be a removable or non-removable medium and can include a machine-readable medium, such as a flash drive, a magnetic disk, or any other medium that can be used to store information and / or data (e.g., training data for training) and can be accessed within the computing device 1100.
[0095] The computing device 1100 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in Figure 11 , a disk drive for reading from or writing to a removable, non-volatile disk (such as a "floppy disk") and an optical disk drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to the bus (not shown) by one or more data media interfaces. The memory 1120 may include a computer program product 1125 having one or more program modules configured to execute the various methods or actions of the various implementations of the present disclosure.
[0096] The communication unit 1140 enables communication with other computing devices via a communication medium. Additionally, the functionality of the components of the computing device 1100 may be implemented in a single computing cluster or multiple computer machines capable of communicating via a communication link. Thus, the computing device 1100 may operate in a networked environment using a logical connection to one or more other servers, network personal computers (PCs), or another network node.
[0097] The input device 1150 may be one or more input devices such as a mouse, keyboard, trackball, etc. The output device 1160 may be one or more output devices such as a display, speaker, printer, etc. The computing device 1100 may also communicate with one or more external devices (not shown) as needed via the communication unit 1140, such as storage devices, display devices, etc., communicate with one or more devices that enable a user to interact with the computing device 1100, or communicate with any device that enables the computing device 1100 to communicate with one or more other computing devices (e.g., network card, modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).
[0098] According to an implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, where the computer-executable instructions are executed by a processor to implement the method described above. According to an implementation of the present disclosure, a computer program product is also provided, the computer program product being tangibly stored on a non-transitory computer-readable medium and including computer-executable instructions, and the computer-executable instructions being executed by a processor to implement the method described above. According to an implementation of the present disclosure, a computer program product is provided, on which a computer program is stored, and when the program is executed by a processor, the method described above is implemented.
[0099] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0100] These computer-readable program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, when executed by the processor of the computer or other programmable data processing apparatus, create a means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable medium storing the instructions comprises a manufacture including instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0101] The computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, such that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0102] The flowcharts and block diagrams in the figures illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various implementations of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions.
[0103] The various implementations of the present disclosure have been described above. The above description is exemplary and not exhaustive, and is also not limited to the disclosed implementations. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described implementations. The choice of terminology used herein is intended to best explain the principles of the implementations, the practical application, or the improvement of the technology in the market, or to enable other ordinary skilled persons in the art to understand the various implementation manners disclosed herein.
Claims
1. A method for evaluating a model, comprising: In response to receiving a user request, receiving a first response and a second response to the user request, where the first response is determined using a first model and the second response is determined using a second model; Presenting the first response and the second response, and an evaluation control for evaluating the first response and the second response; And Based on an evaluation action associated with the evaluation control, determining an evaluation for the first model and the second model.
2. The method according to claim 1, wherein presenting the first response and the second response comprises: Presenting a target response among the first response and the second response and a switching control; And In response to receiving a switching action associated with the switching control, presenting another response different from the target response among the first response and the second response.
3. The method according to claim 2, wherein the first response and the second response are presented at the same position, and the switching control includes a first switching control for presenting the first response and a first switching control for presenting the first response.
4. The method according to claim 1, wherein presenting the first response and the second response comprises: Determining a presentation order of the first response and the second response based on a predetermined condition; And Presenting the first response and the second response in accordance with the presentation order.
5. The method according to claim 1, wherein receiving the first response and the second response comprises: Determining the first model and the second model based on an intention associated with the user request; And Receiving a first response from the first model and a second response from the second model.
6. The method according to claim 1, further comprising at least any one of the following: In response to determining that a first time length for presenting the first response and a second time length for presenting the second response satisfy a predetermined time length, marking the evaluation as a valid evaluation; or In response to determining that the first time length for presenting the first response and the second time length for presenting the second response do not satisfy the predetermined time length, marking the evaluation as an invalid evaluation.
7. The method according to claim 1, further comprising: In response to receiving a creation request for creating a new session, clearing the context associated with the user request.
8. The method according to claim 1, wherein presenting the first response and the second response further comprises: Determining a layout for presenting the first response and the second response based on display parameters of a terminal device; And Presenting the first response and the second response based on the layout.
9. The method according to claim 1, further comprising: Presenting a selection control associated with the evaluation; And Based on a selection action associated with the selection control, determining an evaluation reason associated with the evaluation.
10. The method according to claim 1, wherein the first model and the second model are multimodal models, and the method further comprises: Executing the method at a predetermined frequency.
11. An apparatus for evaluating a model, comprising: A receiving module, configured to receive a first response and a second response to the user request in response to receiving the user request, where the first response is determined using a first model and the second response is determined using a second model; A presenting module, configured to present the first response and the second response, and an evaluation control for evaluating the first response and the second response; And A determining module, configured to determine an evaluation of the first model and the second model based on an evaluation action associated with the evaluation control.
12. An electronic device, comprising: At least one processor; And At least one memory, the at least one memory being coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions, when executed by the at least one processor, causing the electronic device to perform the method according to any one of claims 1 to 10.
13. A computer-readable storage medium, having stored thereon computer instructions, which, when executed by a processor, cause the processor to implement the method according to any one of claims 1 to 10.
14. A computer instruction product, comprising computer instructions, wherein the computer instructions, when executed by a processor, implement the method according to any one of claims 1 to 10.