Methods, apparatus, equipment and storage media for displaying evaluation results of prompt word templates
By intuitively displaying the local and global usage effects of the prompt word template on the evaluation result display interface, the problems of low efficiency and poor accuracy in the display of evaluation results in the prior art are solved, and efficient evaluation and optimization of prompt word templates are achieved.
Patent Information
- Application Number
- CN202310833408.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-07
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2043-07-07
AI Technical Summary
In existing technologies, the evaluation results of prompt word templates are displayed inefficiently and with poor accuracy, causing developers to spend a lot of time reviewing the text content.
By displaying the evaluation results of the prompt word template in different areas according to different evaluation dimensions on the evaluation results display interface, the local and global usage effects are intuitively displayed using evaluation icons, and dynamic area adjustment and manual intervention functions are provided to support intuitive comparison and optimization of evaluation results.
It improves the evaluation efficiency and accuracy of prompt word templates, simplifies the operation process for developers, and enhances the efficiency of human-computer interaction.
Smart Images

Figure CN119336420B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence (AI) technology, and in particular to a method, apparatus, device and storage medium for displaying evaluation results of prompt word templates. Background Technology
[0002] A prompt is input provided to an AI model (such as a large AI model used for natural language processing) to guide and stimulate the model to generate corresponding responses or content. A prompt template is a template text that includes variables and fixed text content. By inputting specific values for the variables, prompts can be dynamically generated. It should be understood that because prompts are crucial for obtaining accurate and high-quality AI model outputs, developers need to carefully design optimal prompt templates to guide AI models to better complete tasks.
[0003] In related technologies, AI models are typically used to evaluate the prompt word templates designed by developers, and the evaluation results are displayed on the terminal interface in the form of a text list. Developers can view the evaluation results of each prompt word template one by one to determine the impact of different prompt word templates on the output of the AI model, and then design better prompt word templates.
[0004] However, the above method of displaying evaluation results using a text list is ineffective. Developers need to spend a lot of time reviewing the text content, resulting in low evaluation efficiency and poor accuracy of the prompt word template. Summary of the Invention
[0005] This application provides a method, apparatus, device, and storage medium for displaying evaluation results of prompt word templates, which can effectively improve the display effect of evaluation results and enhance the evaluation efficiency and accuracy of prompt word templates. The technical solution is as follows:
[0006] Firstly, a method for displaying the evaluation results of a prompt word template is provided, the method comprising:
[0007] Obtain a set of use cases and at least one prompt word template. The set of use cases includes multiple use cases. The prompt word template is used to generate prompt words based on the use cases. The prompt words are used as input to an artificial intelligence (AI) model to obtain the output result corresponding to the prompt words.
[0008] Based on the first use case in the set of use cases and at least one of the prompt word templates, a display area corresponding to each prompt word template is displayed in the first area of the evaluation result display interface. In each of the display areas, a first evaluation identifier for the corresponding prompt word template for the first use case is displayed. The first evaluation identifier indicates the usage effect of the prompt word template obtained based on the first use case.
[0009] The second area of the evaluation result display interface displays a second evaluation identifier for each of the prompt word templates for the set of use cases. Each second evaluation identifier indicates the usage effect of the prompt word template obtained based on the corresponding use case in the set of use cases. The first area is adjacent to the second area.
[0010] Using the above method, when evaluating the prompt word templates of the AI model, the evaluation results of the prompt word templates are displayed in different areas of the evaluation result display interface in the form of evaluation labels according to different evaluation dimensions. This allows relevant personnel to intuitively understand the local usage effect of using different prompt word templates on the same use case and the global usage effect of using different prompt word templates on the entire set of use cases. This approach effectively improves the display effect of evaluation results, enabling relevant personnel to quickly and intuitively understand the evaluation results, improving human-computer interaction efficiency, and thus enhancing the evaluation efficiency and accuracy of prompt word templates.
[0011] In some embodiments, the method further includes:
[0012] The first use case is filled into at least one of the prompt word templates to obtain at least one first prompt word;
[0013] Input at least one of the first prompt words into the AI model to obtain output results corresponding to at least one of the first prompt words;
[0014] Based on the difference between the output result and the reference result corresponding to at least one of the first prompt words, an evaluation result of at least one of the prompt word templates is determined, and the evaluation result is used to determine the first evaluation identifier.
[0015] In some embodiments, determining the evaluation result of at least one of the prompt word templates based on the difference between the output result and the reference result corresponding to at least one of the first prompt words includes at least one of the following:
[0016] Based on the matching degree between the output result and the reference result corresponding to at least one of the first prompt words, the evaluation result of at least one prompt word template is determined;
[0017] The evaluation result of at least one of the prompt word templates is determined based on the similarity between the output result and the reference result corresponding to at least one of the first prompt words.
[0018] The above methods provide multiple ways to determine the evaluation results of prompt word templates, making it easier for developers to set them according to actual needs.
[0019] In some embodiments, the first evaluation identifier includes at least one of a text identifier and a graphic identifier; the second evaluation identifier includes at least one of a text identifier and a graphic identifier.
[0020] In some embodiments, the first evaluation identifier includes at least one of the following: a symbol for representing the effect of use, text for representing the effect of use, and a numerical value for representing the effect of use; the second evaluation identifier includes at least one of the following: a symbol for representing the effect of use, text for representing the effect of use, and a numerical value for representing the effect of use.
[0021] In some embodiments, different first evaluation identifiers indicate different usage effects of the prompt word template obtained based on the first use case; different second evaluation identifiers indicate different usage effects of the prompt word template obtained based on the corresponding use case in the use case set.
[0022] The above methods provide a variety of evaluation icon styles, which, when displayed on the evaluation results display interface, allow relevant personnel to intuitively understand the effectiveness of the prompt word templates and improve human-computer efficiency.
[0023] In some embodiments, the method further includes:
[0024] In response to a drag operation performed on the target control, the evaluation result display interface shows that the areas of the first region and the second region change as the drag operation occurs.
[0025] In some embodiments, the display of changes in the areas of the first and second regions in response to a drag operation on the target control includes any of the following:
[0026] In response to a first drag operation performed on the target control, the area of the first region is reduced and the area of the second region is increased, wherein the first drag operation refers to a drag operation in which the drag direction is directed toward the first region;
[0027] In response to a second drag operation performed on the target control, the area of the first region is increased and the area of the second region is decreased. The second drag operation refers to a drag operation in which the drag direction is directed towards the second region.
[0028] The above method provides a function for dynamically adjusting the area, which can meet the personalized needs of relevant personnel in viewing assessment results. Moreover, in some scenarios, the displayed content can be dynamically updated according to changes in the area, thereby making full use of interface space, simplifying the operation of viewing assessment results, and improving assessment efficiency and user experience.
[0029] In some embodiments, displaying the corresponding prompt word template for the first evaluation identifier of the first use case in each of the display areas includes:
[0030] In each of the display areas, the corresponding prompt word template is displayed for the first evaluation identifier of the first use case and the output result corresponding to at least one first prompt word, which is generated based on the first use case and the prompt word template.
[0031] In some embodiments, the method further includes:
[0032] The second area displays at least one of the prompt word templates for the set of use cases, with each third evaluation identifier indicating the overall usage effect of at least one of the prompt word templates obtained based on the corresponding use case in the set of use cases.
[0033] In some embodiments, the method further includes:
[0034] In response to a trigger operation performed on the display area corresponding to the target prompt word template in the first region, the output result and reference result corresponding to the target prompt word are displayed, wherein the target prompt word is generated based on the first use case and the target prompt word template;
[0035] In response to the evaluation identifier update operation performed on the target prompt word template, the first evaluation identifier of the target prompt word template for the first use case is updated.
[0036] In some embodiments, the evaluation identifier update operation is any of the following:
[0037] The operation performed on the evaluation identifier adjustment track, which indicates multiple different evaluation identifiers;
[0038] The operation performed on the evaluation identifier adjustment option, which indicates multiple different evaluation identifiers.
[0039] The above method provides a function for manual intervention of evaluation results, enabling relevant personnel to update the corresponding evaluation labels based on the evaluation results displayed on the interface.
[0040] In some embodiments, the method further includes:
[0041] In response to the prompt word template comparison operation performed in the first area, the evaluation results of the first prompt word template and the evaluation results of the second prompt word template are displayed;
[0042] In response to a trigger operation performed on the display area corresponding to the first prompt word template, or in response to a trigger operation performed on the display area corresponding to the second prompt word template, update the first evaluation identifier of the first prompt word template and the second prompt word template for the first use case.
[0043] The above method provides a comparison operation for the evaluation results of prompt word templates. That is, relevant personnel can make a horizontal comparison between any two prompt word templates to select the prompt word template that meets the requirements, and provide a reference for further optimization of prompt word templates.
[0044] Secondly, embodiments of this application provide an evaluation result display device for prompt word templates. The device includes at least one functional module for executing the evaluation result display method for prompt word templates as provided in the first aspect or any possible implementation of the first aspect.
[0045] Thirdly, embodiments of this application provide a computing device, which includes a processor and a memory; the memory is used to store at least one piece of program code, which is loaded by the processor and executed as described in the first aspect or any possible implementation of the first aspect, to display the evaluation result of the prompt word template.
[0046] Fourthly, embodiments of this application provide a computer-readable storage medium for storing at least one piece of program code, which implements the method for displaying the evaluation results of the prompt word template provided in the first aspect or any possible implementation thereof. The storage medium includes, but is not limited to, volatile memory, such as random access memory, and non-volatile memory, such as flash memory, hard disk drive (HDD), and solid state drive (SSD).
[0047] Fifthly, embodiments of this application provide a computer program product that, when run on a computing device, enables the computing device to implement the method for displaying evaluation results of the prompt word template provided in the first aspect or any possible implementation of the first aspect. The computer program product can be a software installation package; when the aforementioned method for displaying evaluation results of the prompt word template needs to be implemented, the computer program product can be downloaded and executed on the computing device. Attached Figure Description
[0048] Figure 1 This is a schematic diagram of an implementation environment provided in an embodiment of this application;
[0049] Figure 2 This is a schematic diagram of the hardware structure of a computing device provided in an embodiment of this application;
[0050] Figure 3 This is a schematic diagram of the structure of a computing device cluster provided in an embodiment of this application;
[0051] Figure 4 This is a schematic diagram illustrating a connection method for a computing device cluster provided in an embodiment of this application;
[0052] Figure 5 This is a flowchart illustrating a method for displaying the evaluation results of a prompt word template, as provided in an embodiment of this application.
[0053] Figure 6 This is a schematic diagram of an evaluation result display interface provided in an embodiment of this application;
[0054] Figure 7 This is a schematic diagram of another evaluation result display interface provided in an embodiment of this application;
[0055] Figure 8 This is a schematic diagram of another evaluation result display interface provided in the embodiments of this application;
[0056] Figure 9 This is a schematic diagram of another evaluation result display interface provided in an embodiment of this application;
[0057] Figure 10 This is a schematic diagram of another evaluation result display interface provided in an embodiment of this application;
[0058] Figure 11 This is a schematic diagram of an evaluation identifier update operation provided in an embodiment of this application;
[0059] Figure 12 This is a schematic diagram of a prompt word template comparison operation provided in an embodiment of this application;
[0060] Figure 13 This is a schematic diagram of the structure of a prompt word template evaluation result display device provided in an embodiment of this application. Detailed Implementation
[0061] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be further described in detail below with reference to the accompanying drawings. It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the prompt word templates involved in this application were obtained under fully authorized conditions.
[0062] To facilitate understanding, the relevant terms and concepts involved in this application will be introduced below.
[0063] Artificial Intelligence (AI): The basic principle of AI is to combine massive amounts of data with powerful computing capabilities and intelligent algorithms to build an AI model that solves specific problems. This AI model can automatically summarize and learn potential patterns or features from the data, thereby achieving a way of thinking that is close to that of humans. AI models, also known as AI algorithms (or AI operators), are a collective term for mathematical algorithms built based on the principles of AI. They are also the foundation for using AI to solve specific problems, such as deep learning (DL) models.
[0064] AI Model Training and Inference: Any AI model needs to be trained before it can solve a specific technical problem. AI model training involves using a specified initial model to calculate training data, and then adjusting the parameters of the initial model based on the calculation results. This process allows the model to gradually learn certain patterns and acquire specific functions. Once trained and possessing stable functionality, the AI model can be used for inference. AI model training is the process of using the trained AI model to calculate input data and obtain predictive inference results.
[0065] The AI Infrastructure Development Platform is a one-stop AI development platform for users. It provides capabilities across the entire AI development process, including data preprocessing, model building and training, model management, model deployment, data optimization, and model optimization updates. Users can develop AI models and deploy and manage AI applications based on this platform. The various capabilities within the platform can be integrated for comprehensive AI workflow use, or they can be provided as independent functions.
[0066] AI Services in the Cloud Sector: There are two main types of AI services in the cloud sector: Platform-as-a-Service (PaaS) for AI infrastructure development platforms and Software-as-a-Service (SaaS) for AI application cloud services. For the first type, public cloud service providers leverage their ample underlying resources and upper-layer AI algorithm capabilities to offer customers an AI infrastructure development platform. This platform includes built-in AI development frameworks and various AI algorithms, allowing customers to quickly build and develop AI models or applications tailored to their specific needs. For the second type, AI application cloud services, public cloud service providers offer ready-to-use, universal AI application cloud services through their cloud platforms, enabling customers to use AI capabilities in various application scenarios with zero barriers to entry.
[0067] AI large models are deep neural networks composed of a large number of layers and parameters. Due to their powerful predictive capabilities, they are widely used in various fields, such as natural language processing, computer vision, speech synthesis and speech recognition, autonomous driving, etc.
[0068] A prompt is an input provided to an AI model (such as a large AI model used for natural language processing) to guide and stimulate the model to generate a corresponding response or content. For example, a prompt can be a complete text (or sentence) that can be used as input to a large language model (LLM), enabling the model to output a corresponding inference result based on the prompt.
[0069] A cue word template is a template text that includes variables and fixed text content. By inputting specific values for the variables, cue words can be dynamically generated. For example, in practical applications, by adjusting the specific values of the variables in the cue word template, multiple different cue words can be dynamically generated. Illustratively, taking an AI model for sentiment analysis as an example, the cue word template is:
Determine which of the three tags "positive," "neutral," or "negative" this comment {variable} belongs to?
Determine which of the three tags "positive," "neutral," or "negative" this comment {high cost-performance ratio, can run games smoothly} belongs to?
[0070] Use case: Also known as a usage scenario, it is a description of how a system responds to external requests in software engineering or systems engineering. In the embodiments of this application, the use case indicates the intention to ask a question to the AI model. It can serve as the variable value of the prompt word template to dynamically generate different prompt words. In other words, combining different use cases with the same prompt word template can generate multiple different prompt words. For example, taking the AI model for sentiment analysis as an example, the use case is: "1. High cost-effectiveness, can run games smoothly. 2. High order processing efficiency. 3. Fast logistics." If the prompt word template is: "[Determine which of the three labels 'positive,' 'neutral,' and 'negative' this comment {1. High cost-effectiveness, can run games smoothly. 2. High order processing efficiency. 3. Fast logistics} belongs to which of the three labels 'positive,' 'neutral,' and 'negative'?", then combining the use case with the prompt word template, or filling the prompt word template with the use case, yields the prompt word: "[Determine which of the three labels 'positive,' 'neutral,' and 'negative' this comment {1. High cost-effectiveness, can run games smoothly. 2. High order processing efficiency. 3. Fast logistics} belongs to?".
[0071] The application scenarios and implementation environment involved in this application are described below.
[0072] The technical solution provided in this application can be applied to scenarios where prompt word templates for AI models are evaluated. Evaluating the prompt word templates for AI models involves combining use cases of the AI model with a pre-designed prompt word template to generate prompt words. These prompt words are then input into the AI model to obtain corresponding output results. Next, by comparing the output results with reference results, the evaluation result of the prompt word template is obtained. It should be understood that the smaller the difference between the output results and reference results indicated by the evaluation result, the more accurate the AI model's output results are when using the prompt word template to generate prompt words; that is, the better the effectiveness of the prompt word template. Therefore, the evaluation result of the prompt word template can indicate its effectiveness, thereby providing guidance for relevant personnel to further optimize the prompt word templates.
[0073] refer to Figure 1 , Figure 1 This is a schematic diagram of an implementation environment provided in an embodiment of this application. For example... Figure 1 As shown, the implementation environment includes a terminal 101 and a server 102. The terminal 101 is directly or indirectly connected to the server 102 via a wireless network or a wired network.
[0074] Terminal 101 has a display function, capable of showing the evaluation results of the prompt word template. Terminal 101 is at least one of a smartphone, game console, desktop computer, augmented reality terminal, tablet computer, e-book reader, and laptop computer. Terminal 101 has an application installed and running that supports displaying evaluation results. This application can be a client, a browser, etc., and this application is not limited to these. Taking the application as a client as an example, a user (such as a developer) can input a set of use cases and a designed prompt word template through the client, triggering Terminal 101 to evaluate the prompt word template and display the corresponding evaluation results on the client interface. The user can then further optimize the prompt word template based on the evaluation results displayed on the client interface. Illustratively, Terminal 101 is a terminal used by a user, and the application running on Terminal 101 contains a user account logged in, also referred to as an object.
[0075] Server 102 provides background services for applications running on terminal 101. Server 102 can be a standalone physical server, a server cluster consisting of multiple physical servers, a distributed file system, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. Taking server 102 as a cloud server as an example, a cloud server refers to a service based on hardware and software resources, providing computing, networking, and storage capabilities. Through the network "cloud," massive amounts of data are processed and analyzed remotely before being returned to the user, featuring characteristics such as large scale, distributed nature, virtualization, high availability, scalability, on-demand service, and security. In this embodiment, the services provided by the cloud server are also referred to as AI services in the cloud domain (see the foregoing content for details, which will not be repeated here).
[0076] In this embodiment, terminal 101 can independently execute the method for displaying the evaluation results of the prompt word template involved in the following method embodiments, or it can work in conjunction with server 102. For example, terminal 101 sends the obtained use cases and prompt word templates to server 102, server 102 evaluates the prompt word templates and feeds back the evaluation results to terminal 101, and terminal 101 displays the evaluation results. This application does not limit this.
[0077] The aforementioned terminal 101 can refer to one of a plurality of terminals, or a collection of a plurality of terminals; server 102 can refer to one of a plurality of servers, or a collection of a plurality of servers. This application embodiment does not limit the number or type of each type of device in the implementation environment. Furthermore, the aforementioned implementation environment can also be understood as an AI basic development platform, upon which users can complete the development of AI models (including the evaluation of prompt word templates) and the deployment and management of AI applications.
[0078] In some embodiments, the aforementioned wireless or wired networks use standard communication technologies and / or protocols. Networks include, but are not limited to, data center networks, storage area networks (SANs), local area networks (LANs), metropolitan area networks (MANs), wide area networks (WANs), mobile, wired or wireless networks, private networks, or any combination of virtual private networks. In some implementations, technologies and / or formats, including Hypertext Markup Language (HTML), Extensible Markup Language (XML), etc., are used to represent data exchanged over the network. Furthermore, conventional encryption technologies such as Secure Sockets Layer (SSL), Transport Layer Security (TLS), Virtual Private Networks (VPNs), and Internet Protocol Security (IPsec) can be used to encrypt all or part of the link. In other embodiments, custom and / or dedicated data communication technologies can be used to replace or supplement the aforementioned data communication technologies.
[0079] The hardware structure of the terminal and server in the above implementation environment is described below.
[0080] This application provides a computing device that can be configured as the aforementioned terminal or server. (Illustratively, refer to...) Figure 2 , Figure 2 This is a schematic diagram of the hardware structure of a computing device provided in an embodiment of this application. Figure 2As shown, the computing device 200 includes a memory 201, a processor 202, a communication interface 203, and a bus 204. The memory 201, processor 202, and communication interface 203 are interconnected via the bus 204.
[0081] The memory 201 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto. Illustratively, taking the computing device configured as the aforementioned terminal 101 as an example, the memory 201 is used to store at least a piece of program code. When the program code stored in the memory 201 is executed by the processor 202, the processor 202 is used to perform the steps performed by the terminal in the following method embodiments.
[0082] Processor 202 can be a network processor (NP), a central processing unit (CPU), an application-specific integrated circuit (ASIC), or an integrated circuit used to control the execution of the program in this application. Processor 202 can be a single-core processor or a multi-core processor. There can be one or more processors 202. Communication interface 203 uses a transceiver module, such as a transceiver, to enable communication between computing device 200 and other devices or communication networks. For example, data can be acquired through communication interface 203.
[0083] The memory 201 and the processor 202 can be set separately or integrated together.
[0084] Bus 204 may include a pathway for transmitting information between various components of computing device 200 (e.g., memory 201, processor 202, communication interface 203).
[0085] This application also provides a computing device cluster, which includes multiple computing devices. The multiple computing devices include servers and terminals. The servers can be central servers, edge servers, or local servers in a local data center. The terminals can be desktop computers, laptops, or smartphones, etc. (See reference...) Figure 3 , Figure 3 This is a schematic diagram of the structure of a computing device cluster provided in an embodiment of this application. For example... Figure 3 As shown, the computing device cluster includes multiple computing devices 200. The memories 201 of the multiple computing devices 200 in the computing device cluster may store the same instructions for executing the method for displaying the evaluation results of the prompt word template. In some embodiments, the memories 201 of the multiple computing devices 200 in the computing device cluster may also each store a portion of the instructions for executing the method for displaying the evaluation results of the prompt word template. In other words, the combination of multiple computing devices 200 is used to jointly execute the instructions for displaying the evaluation results of the prompt word template.
[0086] In some embodiments, multiple computing devices 200 in a computing device cluster can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc. Figure 4 This is a schematic diagram illustrating a connection method for a computing device cluster provided in an embodiment of this application. For example... Figure 4 As shown, the two computing devices 200 are connected via a network. Specifically, they are connected to the network through the communication interface 203 in each computing device 200. It should be understood that... Figure 4 The functions of the computing device 200 shown can also be performed by multiple computing devices 200.
[0087] The following describes the method for displaying the evaluation results of the prompt word template provided in this application.
[0088] Figure 5 This is a flowchart illustrating a method for displaying the evaluation results of a prompt word template provided in an embodiment of this application. For example... Figure 5 As shown, the method is described using a terminal in the above implementation environment as an example. The method includes the following steps 501 to 507.
[0089] 501. The terminal obtains a set of use cases and at least one prompt word template. The set of use cases includes multiple use cases. The prompt word template is used to generate prompt words based on the use cases. The prompt words are used as input to the AI model to obtain the output results corresponding to the prompt words.
[0090] In this embodiment, the AI model refers to an AI model based on natural language processing, and this application does not limit the functions provided by the AI model. For example, the AI model may be used to provide functions such as sentiment analysis, machine translation, question answering, language generation, article creation, or text classification. Illustratively, the terminal has a client application with display capabilities installed and running. The terminal can respond to various operations performed on the client interface and display corresponding content on the client interface. For example, developers can perform various operations on the client interface to trigger the terminal to evaluate the prompt word templates of the AI model.
[0091] Schematic illustration: A client running on the terminal provides a prompt word template evaluation function for an AI model. For example, the terminal displays a prompt word template evaluation interface for the AI model, and in response to a set of use cases and at least one prompt word template input on the prompt word template evaluation interface, it obtains the corresponding set of use cases and at least one prompt word template. The prompt word template evaluation interface can provide multiple input methods such as keyboard input and voice input, which are not limited in this application.
[0092] Furthermore, in this embodiment, the prompt word template obtained by the terminal can be an initial prompt word template designed by the developer, or a new prompt word template obtained by arbitrarily combining multiple initial prompt word templates (the combination process can be done manually by the developer or automatically by the terminal, without limitation). It should be understood that the purpose of evaluating the prompt word template in this application is to select a prompt word template that better meets the requirements (i.e., guides the AI model to output results that better meet expectations), or to further optimize the designed prompt word template. Therefore, at least one prompt word template obtained by the terminal can be regarded as at least one candidate set. For any candidate set, the candidate set can be an initial prompt word template designed by the developer, or a new prompt word template obtained by combining multiple initial prompt word templates. Thus, after the terminal obtains the evaluation results of at least one prompt word template (i.e., at least one candidate set), it can display the evaluation results, making it easier for the developer to select the prompt word template with the best performance, or to further optimize the designed prompt word template, thereby improving the evaluation efficiency.
[0093] The following example, using an AI model to provide sentiment analysis functionality, illustrates the set of use cases and prompt word templates involved in this application.
[0094] The use case set includes multiple use cases (see the terminology section above for a detailed explanation of use cases, which will not be repeated here). It should be understood that for the same use case, combining it with different prompt word templates will generate different prompt words, and thus the final output of the AI model will often differ. For example, the use case set obtained by the terminal includes use case 1 and use case 2:
[0095] Use Case 1: "1. High cost-performance ratio, can run games smoothly. 2. High order processing efficiency. 3. Fast logistics";
[0096] Use Case 2: "1. Screen lag. 2. Low order processing efficiency. 3. Poor after-sales service attitude."
[0097] The prompt word template is used to generate prompt words based on the use case (for a detailed explanation of prompt words and prompt word templates, please refer to the aforementioned terminology introduction, which will not be repeated here). It should be understood that prompt word templates are usually designed by developers, and this application does not limit the number of prompt word templates. For example, the terminal obtains three prompt word templates for the AI model:
[0098] Prompt template a: [Determine which of the three tags this comment {variable} belongs to: "positive", "neutral", or "negative"?]
[0099] Prompt Template B: [I will provide you with a comment. Please determine the corresponding sentiment tag based on the comment content. Possible sentiment tags are positive, neutral, and negative. Do not make up other tags. You can refer to the following examples: "Neighborhood, beautiful, big keyboard" is positive; "How to buy this book?" is neutral; "The room has a strong smell, poor service, bad attitude" is negative. The comment is: {variable}. What is the sentiment tag of this comment?]
[0100] Prompt template c: [Please determine the sentiment tag of this comment based on the keywords in {variable}. Sentiment tags include positive, neutral, and negative. Do not make up other tags.]
[0101] Furthermore, combining use cases with prompt word templates can generate corresponding prompt words, as shown below:
[0102] Based on use case 1 and prompt word template ac, the following prompt words 1-a to 1-c can be generated:
[0103] Hint 1-a: [Determine which of the three tags—"positive," "neutral," or "negative"—this review {1. High cost-performance ratio, runs games smoothly. 2. High order processing efficiency. 3. Fast logistics} belongs to?]
[0104] Prompt 1-b: [I will provide you with a comment. Please determine the corresponding sentiment tag based on the comment content. Possible sentiment tags are positive, neutral, and negative. Do not make up other tags. You can refer to the following examples: "Neighborhood, beautiful, big keyboard" is positive; "How to buy this book?" is neutral; "The room has a strong smell, poor service, and bad attitude" is negative. The comment is: {1. High cost performance, can run games smoothly. 2. High order processing efficiency. 3. Fast logistics}. What is the sentiment tag of this comment?]
[0105] Hint 1-c: [Based on the keywords in this comment {1. High cost-performance ratio, runs games smoothly. 2. High order processing efficiency. 3. Fast logistics}, determine the sentiment tag of this comment. Sentiment tags include positive, neutral, and negative. Do not make up other tags.]
[0106] Based on use case 2 and prompt word template ac, the following prompt words 2-a to 2-c can be generated:
[0107] Hint 2-a: [Determine which of the three tags—"positive," "neutral," or "negative"—this comment {1. Lagging screen. 2. Low order processing efficiency. 3. Average after-sales service attitude} belongs to?]
[0108] Prompt 2-b: [I will provide you with a comment. Please determine the corresponding sentiment tag based on the comment content. Possible sentiment tags are positive, neutral, and negative. Do not make up other tags. You can refer to the following examples: "Neighborhood, beautiful, big keyboard" is positive; "How to buy this book?" is neutral; "The room has a strong smell, poor service, poor attitude" is negative. The comment is: {1. Screen lag. 2. Low order processing efficiency. 3. Average after-sales service attitude}. What is the sentiment tag of this comment?]
[0109] Hint 2-c: [Based on the keywords in this comment {1. Lagging screen. 2. Low order processing efficiency. 3. Average after-sales service}, determine the sentiment tag of this comment. Sentiment tags include positive, neutral, and negative. Do not make up other tags.]
[0110] It should be understood that the above description of the use case set and prompt word template is for illustrative purposes only and does not constitute a limitation of this application.
[0111] 502. The terminal obtains the evaluation result of at least one prompt word template based on the use case set and at least one prompt word template.
[0112] In this embodiment of the application, for any one use case in the use case set, the terminal combines the use case with each prompt word template to generate a corresponding prompt word, and inputs the generated prompt word into the AI model to obtain the corresponding output result. Then, by comparing the difference between the output result corresponding to the prompt word and the reference result, the evaluation result of each prompt word template is obtained.
[0113] The reference result corresponding to the prompt word can be understood as the standard result or expected result corresponding to the prompt word. For example, taking the example "prompt word 1-a" in step 501 above, the reference result corresponding to this prompt word is "positive". If "prompt word 1-a" is input into the AI model, the output result obtained is also "positive". Correspondingly, the evaluation result of prompt word template a is "match", indicating that using prompt word template a to generate prompt words can guide the AI model to output the expected result, that is, it indicates that the use of prompt word template a meets the requirements. It should be noted that the reference result corresponding to the prompt word can be obtained by the terminal in response to the input operation implemented on the prompt word template evaluation interface during the execution of step 501, or by the terminal in response to the input operation implemented on the prompt word template evaluation interface during the execution of this step 502, or it can be obtained from the server, etc. This application does not limit the method and timing of the terminal obtaining the reference result.
[0114] The following describes the process of obtaining the evaluation result of at least one prompt word template for the terminal, using any one of the use cases in the use case set (hereinafter referred to as the first use case) as an example, including the following steps 1 to 3:
[0115] Step 1: Fill the first use case into at least one prompt word template to obtain at least one first prompt word.
[0116] In this process, the terminal fills each prompt word template with the first use case, thus obtaining the first prompt word for each prompt word template when used in the first use case. Furthermore, this process can be referred to the example of prompt words in step 501 above, and will not be repeated here.
[0117] Step 2: Input at least one first prompt word into the AI model to obtain the output result corresponding to at least one first prompt word.
[0118] The terminal inputs each first prompt word into the AI model, and the AI model performs calculations through its built-in logic to obtain the output result corresponding to each first prompt word.
[0119] Step 3: Based on the difference between the output result and the reference result corresponding to at least one first prompt word, determine the evaluation result of at least one prompt word template.
[0120] The evaluation result of the prompt word template is used to determine the first evaluation identifier of the prompt word template for the first use case. The first evaluation identifier indicates the effectiveness of the prompt word template based on the first use case (the relevant content about the first evaluation identifier will be introduced in detail in subsequent step 503, and will not be repeated here). For any first prompt word, the terminal determines the evaluation result of the prompt word template corresponding to the first prompt word based on the difference between the output result corresponding to the first prompt word and the reference result. Illustratively, the difference between the output result corresponding to the first prompt word and the reference result can be reflected by information such as the matching degree and similarity between the two. Several implementation methods of step 3 are introduced below, including at least one of the following:
[0121] The first method involves determining the evaluation result of at least one prompt word template based on the matching degree between the output result corresponding to at least one first prompt word and the reference result.
[0122] In this process, for any given first prompt word, the terminal extracts the keywords from the output result corresponding to the first prompt word, matches the keywords with the reference result corresponding to the first prompt word, and determines the evaluation result of the prompt word template corresponding to the first prompt word based on the matching result. For example, if the matching result indicates a successful match, the evaluation result of the prompt word template corresponding to the first prompt word is determined to be "matched"; if the matching result indicates a failed match, the evaluation result of the prompt word template corresponding to the first prompt word is determined to be "not matched". For instance, continuing with the example "prompt word 1-a" in step 501, the reference result corresponding to this prompt word is "positive". If "prompt word 1-a" is input into the AI model, and the output result is also "positive", then the evaluation result of prompt word template a is "matched"; if the output result is "neutral", then the evaluation result of the prompt word template is "not matched". It should be noted that the above method of determining whether the output result matches the reference result through keywords is only an example, and this application is not limited to this. For example, the terminal directly compares the output result with the reference result; if they are the same, the evaluation result is "matched", and so on.
[0123] The second method involves determining the evaluation result of at least one prompt word template based on the similarity between the output result corresponding to at least one first prompt word and the reference result.
[0124] In this process, for any given first prompt word, the terminal determines the evaluation result of the prompt word template corresponding to the first prompt word based on the similarity between the output result corresponding to the first prompt word and the reference result in terms of sentence structure vectors. Illustratively, the evaluation result can be expressed as a score or percentage, and this application is not limited to this. For example, continuing with the example "prompt word 1-a" in step 501, the reference result corresponding to this prompt word is "positive." If "prompt word 1-a" is input into the AI model, and the output result is also "positive," then the evaluation result of prompt word template a is "similarity score 100 points" or "similarity 100%"; if the output result is "neutral," then the evaluation result of the prompt word template is "similarity score 0 points" or "similarity 0%." It should be noted that the above method of determining the similarity between the output result and the reference result through sentence structure vectors is only an illustrative example, and this application is not limited to this.
[0125] Furthermore, the above step 502 is described using the terminal as an example. In some embodiments, the process of determining the evaluation result in step 502 is performed by the server. That is, the terminal sends an evaluation request to the server based on the use case set and at least one prompt word template. The server determines the evaluation result of at least one prompt word template based on the evaluation request and sends the evaluation result to the terminal. In this way, the terminal's computing resources can be saved.
[0126] After steps 501 and 502 above, for at least one prompt word template of the AI model, the terminal obtains the evaluation result of each prompt word template, which provides technical support for the subsequent display of the corresponding evaluation results. The process of how the terminal displays the evaluation results is described below through steps 503 to 507.
[0127] 503. Based on the evaluation results of the first use case in the use case set and at least one prompt word template, the terminal displays the display area corresponding to each prompt word template in the first area of the evaluation result display interface. In each display area, the first evaluation identifier of the corresponding prompt word template for the first use case is displayed. The first evaluation identifier indicates the usage effect of the prompt word template obtained based on the first use case.
[0128] In this embodiment, the first use case refers to any one of the use cases in the use case set. For example, the first use case is the use case ranked at a specified position in the use case set, such as the first or last one, etc., but this application is not limited to this. The evaluation result display interface includes a first area, on which a display area corresponding to each prompt word template is displayed. Based on the evaluation results of the first use case and each prompt word template, the terminal displays the first evaluation identifier of the corresponding prompt word template for the first use case in the display area corresponding to each prompt word template. The evaluation result display interface can be automatically triggered by the terminal, for example, after the terminal executes steps 501 and 502, it automatically jumps from the prompt word template evaluation interface to the evaluation result display interface. The evaluation result display interface can also be triggered by the terminal in response to a triggering operation implemented on the client interface, for example, the terminal displays an "Evaluation Result View" control on the prompt word template evaluation interface, and in response to a triggering operation implemented on the control, it jumps to the evaluation result display interface. This application does not limit the triggering method of the evaluation result display interface. Furthermore, the terminal can jump to the evaluation result display interface after obtaining the evaluation results of all prompt word templates, or it can jump to the evaluation result display interface after obtaining the evaluation results of some prompt word templates, and continuously update the evaluation result display interface based on the evaluation results obtained subsequently. This application does not limit this.
[0129] In some embodiments, the first evaluation identifier includes at least one of a text identifier and a pattern identifier. For example, taking the evaluation result of a prompt word template determined by matching degree, if the evaluation result is "match," the first evaluation identifier includes the text identifier "match" and the pattern identifier "√" to indicate "match"; if the evaluation result is "not match," the first evaluation identifier includes the text identifier "not match" and the pattern identifier "×" to indicate "not match." For example, taking the evaluation result of a prompt word template determined by similarity, if the evaluation result is "similarity score 88," the first evaluation identifier includes the text identifier "score 88" and the pattern identifier "88." Of course, different colored pattern identifiers can be used to distinguish different evaluation results, and this application is not limited to this. In other words, the first evaluation identifier can also be understood to include at least one of the following: a symbol for indicating the effect of use, text for indicating the effect of use, and a numerical value for indicating the effect of use, etc. It should be understood that any identifier capable of distinguishing different effects of use can be applied to the embodiments of this application, and this application does not limit this.
[0130] Furthermore, different first evaluation identifiers indicate different usage effects of the prompt word template obtained from the first use case. Illustratively, taking the evaluation result of the prompt word template determined by matching degree as an example, if the evaluation result is "match," the first evaluation identifier includes the text identifier "match" and a pictorial identifier "√" to represent "match," with the pictorial identifier being a black circle; if the evaluation result is "not match," the first evaluation identifier includes the text identifier "not match" and a pictorial identifier "×" to represent "not match," with the pictorial identifier being a gray circle. For another example, taking the evaluation result of the prompt word template determined by similarity as an example, different colors can be used to distinguish different first evaluation identifiers; that is, a red circle is used for scores higher than 80, a green circle is used for scores between 60 and 80, and a gray circle is used for scores lower than 60. Within different score ranges, the intensity of the color can also reflect the score level, but this application is not limited to this.
[0131] It should be understood that the various implementations of the first evaluation identifier can be combined arbitrarily. In practical applications, the style of the first evaluation identifier can be configured according to requirements, and this application does not limit this.
[0132] In some embodiments, the display area corresponding to each prompt word template in the first region can also display the output result of each prompt word. That is, the terminal displays the first evaluation identifier of the corresponding prompt word template for the first use case and the output result corresponding to at least one first prompt word in each display area. In this way, it can assist relevant personnel in evaluating the prompt word templates and improve evaluation efficiency.
[0133] The following is for reference. Figures 6 to 8 The first area of the evaluation results display interface will be introduced.
[0134] Figure 6 This is a schematic diagram of an evaluation result display interface provided in an embodiment of this application. For example... Figure 6As shown, the evaluation result display interface includes a first area 601, and the use case set includes 20 use cases. Taking the 19th use case as the first use case as an example, the terminal displays multiple prompt word templates 1 to 3 in the first area 601, and displays the first evaluation identifier of the corresponding prompt word template for the first use case in each display area. It should be understood that since the multiple prompt word templates 1 to 3 are also prompt word templates designed by the developers, the developers often select prompt word templates that meet the requirements or make further optimizations. Therefore, these prompt word templates can also be understood as multiple candidate sets (see the aforementioned step 501 for details, which will not be repeated here). In addition, the terminal can display the use case content of the first use case, such as "1. High cost performance, can run the game smoothly. 2. High order processing efficiency. 3. Fast logistics", and can also display the template content of the prompt word template and the generated prompt word content, etc. This application does not limit this (subsequent). Figures 7 to 8 Similarly, (I won't elaborate further). Illustratively, Figure 6 In the above, for prompt word template 1, the first evaluation indicator is displayed: the text indicator "Match" and the graphic indicator "√", with the graphic indicator filled in black; for prompt word template 2, the first evaluation indicator is displayed: the text indicator "Mismatch" and the graphic indicator "×", with the graphic indicator filled with lines; for prompt word template 3, since its evaluation result is being calculated, the text indicator "Evaluating" and a dynamic graphic indicator indicating that the evaluation is in progress are displayed. It should be understood that the above graphic indicators can also be filled in other ways, and any filling method that can distinguish different usage effects can be applied to this application.
[0135] Figure 7 This is a schematic diagram of another evaluation result display interface provided in an embodiment of this application. For example... Figure 7 As shown, the evaluation result display interface includes a first area 701, and the use case set includes 20 use cases. Taking the 19th use case as the first use case as an example, the terminal displays multiple display areas corresponding to prompt word templates 1 to 3 in the first area 701, and displays the first evaluation identifier for the corresponding prompt word template for the first use case in each display area. For prompt word template 1, the first evaluation identifier is displayed: the text identifier "Score 88" and the image identifier "88", with the image identifier filled in light red; for prompt word template 2, the first evaluation identifier is displayed: the text identifier "Score 97" and the image identifier "97", with the image identifier filled in dark red; for prompt word template 3, since its evaluation result is being calculated, the text identifier "Evaluating" and the dynamic image identifier indicating that it is being evaluated are displayed. It should be understood that the above image identifiers can also be filled in other ways, and any filling method that can distinguish different usage effects can be applied to this application.
[0136] Figure 8 This is a schematic diagram of another evaluation result display interface provided in an embodiment of this application. For example... Figure 8 As shown, the evaluation result display interface includes a first area 801, and the use case set includes 20 use cases. Taking the 19th use case as the first use case as an example, the terminal displays multiple display areas corresponding to prompt word templates 1 to 3 in the first area 801, and displays the first evaluation identifier for the corresponding prompt word template for the first use case in each display area. For prompt word template 1, the first evaluation identifier is displayed: the text identifier "Similarity 80%" and the pattern identifier "80%", with the pattern identifier filled in green; for prompt word template 2, the first evaluation identifier is displayed: the text identifier "Similarity 100%" and the pattern identifier "100%", with the pattern identifier filled in red; for prompt word template 3, since its evaluation result is being calculated, the text identifier "Evaluating" and the dynamic pattern identifier indicating that it is being evaluated are displayed. It should be understood that the above pattern identifiers can also be filled in other ways, and any filling method that can distinguish different usage effects can be applied to this application.
[0137] It should be noted that the above Figures 6 to 8 The layout of the evaluation result display interface and the style of the first evaluation mark shown are merely illustrative examples and do not constitute a limitation on this application. For example, multiple prompt word templates can be arranged horizontally or vertically. For example, the graphic mark in the first evaluation mark can be a circle, a rectangle, a rhombus, etc.
[0138] After step 503 above, in the first area of the evaluation result display interface, the evaluation result of each prompt word template for the first use case is displayed in the form of an evaluation icon. This makes it easy for relevant personnel to intuitively understand the local usage effect of using different prompt word templates on the same use case, thereby improving the efficiency of human-computer interaction and thus improving the evaluation efficiency and accuracy of prompt word templates.
[0139] 504. The terminal displays the second evaluation identifier for each prompt word template for the use case set in the second area of the evaluation result display interface. Each second evaluation identifier indicates the usage effect of the prompt word template obtained based on the corresponding use case in the use case set. The first area and the second area are adjacent.
[0140] In this embodiment, the evaluation result display interface further includes a second area. For example, the first area and the second area are adjacent vertically or horizontally; this application does not limit this. It should be understood that step 504 and step 503 described above can be executed simultaneously. Furthermore, the second evaluation identifier is similar to the aforementioned first evaluation identifier, and therefore will not be repeated here. That is, the second evaluation identifier includes at least one of text identifiers and graphic identifiers; in other words, the second evaluation identifier includes at least one of the following: a symbol for indicating the effect of use, text for indicating the effect of use, and a numerical value for indicating the effect of use. Moreover, different second evaluation identifiers indicate different effects of the prompt word templates obtained based on the corresponding use cases in the use case set.
[0141] In some embodiments, the terminal displays the second evaluation identifier of each prompt word template for multiple use cases in the use case set in the second area in the form of a two-dimensional table. In this way, the global effect of using different prompt word templates on the entire use case set can be clearly displayed, which facilitates quick comparison by relevant personnel and improves evaluation efficiency. In other embodiments, the terminal displays the comprehensive evaluation information of each prompt word template for multiple use cases in the use case set in the second area. For example, if the use case set includes 19 use cases, after applying prompt word template 1 to these 19 use cases, 15 of the evaluation results are "matched" and 4 are "not matched". This information is displayed as comprehensive evaluation information in the second area.
[0142] In some embodiments, the second area can also display at least one third evaluation identifier for at least one prompt word template for multiple use cases in the use case set. Each third evaluation identifier indicates the overall usage effect of at least one prompt word template obtained based on the corresponding use case in the use case set. This approach assists relevant personnel in conducting an overall evaluation of the prompt word templates, improving evaluation efficiency. For example, for a certain use case, if the evaluation result for prompt word template 1 is "similarity score 80", the evaluation result for prompt word template 2 is "similarity score 90", and the evaluation result for prompt word template 3 is "similarity score 70", then these three prompt word templates are merged into a total prompt word template, resulting in an overall evaluation result of "average similarity score 80". Correspondingly, the third evaluation identifier is the graphic identifier "80". It should be understood that this is merely an example and does not constitute a limitation of this application; the third evaluation identifier can also be displayed in other ways. Furthermore, the second area can simultaneously display the second evaluation identifier and the third evaluation identifier, or only one of them; this application does not limit this.
[0143] In some embodiments, in response to a triggering operation performed on a second use case in the second area, the terminal displays at least one prompt word template and a fourth evaluation identifier for each prompt word template for the second use case in the first area. The fourth evaluation identifier indicates the effectiveness of the prompt word template based on the second use case, where the second use case refers to other use cases in the use case set besides the first use case. That is, relevant personnel can trigger operations on use cases in the second area to switch the use cases displayed in the first area; in other words, the second area provides navigation functionality so that relevant personnel can view the corresponding evaluation results by selecting different use cases.
[0144] The following is a reference. Figures 6 to 8 The second area of the evaluation results display interface will be introduced.
[0145] like Figure 6 As shown, the evaluation result display interface also includes a second area 602. The terminal displays the second evaluation identifiers of the 19 processed use cases for prompt word templates 1 to 3 in the second area 602 in the form of a two-dimensional table, and displays the corresponding comprehensive evaluation information. That is, for prompt word template 1, the number of use cases with an evaluation result of "match" is 15, and the number of use cases with an evaluation result of "not match" is 4; for prompt word template 2, the number of use cases with an evaluation result of "match" is 17, and the number of use cases with an evaluation result of "not match" is 2; for prompt word template 3, the number of use cases with an evaluation result of "match" is 17, and the number of use cases with an evaluation result of "not match" is 1.
[0146] like Figure 7 As shown, the evaluation result display interface also includes a second area 702. The terminal displays the second evaluation identifiers for the 19 processed use cases using prompt word templates 1 to 3 in the second area 702 in the form of a two-dimensional table, and displays the corresponding comprehensive evaluation information. Specifically, for prompt word template 1, the average similarity score for the 19 use cases is 81 (it should be understood that the average value is only an example; tolerances, etc., can also be used, and this application is not limited to this); for prompt word template 2, the average similarity score for the 19 use cases is 97; and for prompt word template 3, the average similarity score for the 18 use cases (the evaluation result for the 19th use case is being processed) is 67. In addition, prompt information is displayed in the second area 702 to remind relevant personnel of the meaning of the second evaluation identifiers. For example, multiple pattern identifiers of 0, 20, 40, 60, 80, and 100 are used as prompt information to indicate different color types and different shades of color corresponding to different scores. It should be understood that the various second evaluation identifiers in the figure are distinguished by different colors; the specific colors are not drawn in the figure but explained through bottom annotations.
[0147] like Figure 8As shown, the evaluation result display interface also includes a second area 802. The terminal displays the second evaluation identifiers for the 19 processed use cases using prompt word templates 1 to 3 in a two-dimensional table format within this area. It also displays the corresponding comprehensive evaluation information: for prompt word template 1, the average similarity percentage for the 19 use cases is 81%; for prompt word template 2, the average similarity percentage is 97%; and for prompt word template 3, the average similarity score for the 18 use cases (the evaluation result for the 19th use case is currently being processed) is 67%. Additionally, the second area 802 displays prompt information to remind relevant personnel of the meaning of the second evaluation identifiers. For example, multiple graphic identifiers of 0, 20, 40, 60, 80, and 100 are used as prompts to indicate different scores corresponding to different color types and shades. It should be understood that the various second evaluation identifiers in the diagram are distinguished by different colors; however, the specific colors are not drawn in the diagram and are explained through bottom annotations.
[0148] Following step 504 above, in the second area of the evaluation result display interface, the evaluation results of each prompt word template for each use case in the use case set are displayed in the form of evaluation icons. This allows relevant personnel to intuitively understand the global effect of using different prompt word templates across the entire use case set, improving human-computer interaction efficiency and thus enhancing the evaluation efficiency and accuracy of the prompt word templates. Furthermore, as can be seen from step 503 above, this application divides the evaluation result display interface into two adjacent areas, displaying the evaluation results of the prompt word templates according to different evaluation dimensions in different areas of the evaluation result display interface in the form of evaluation icons. This allows relevant personnel to quickly and intuitively understand the entire evaluation result.
[0149] Based on steps 503 and 504 above, this application provides various display types for evaluation identifiers, such as evaluation identifiers based on "match" or "not match," or evaluation identifiers based on similarity scores or percentages. In some embodiments, the terminal determines the evaluation identifier type that matches the evaluation result of the prompt word template. That is, the terminal can automatically select the corresponding evaluation identifier type based on the evaluation result of the prompt word template, thereby improving display efficiency. For example, if the evaluation result is determined by matching degree, the evaluation identifier uses an evaluation identifier based on "match" or "not match"; if the evaluation result is determined by similarity, the evaluation identifier uses an evaluation identifier based on similarity scores or percentages. This application is not limited to this. Of course, the type of evaluation identifier can also be manually set by relevant personnel.
[0150] In some embodiments, the evaluation result display interface also provides display area adjustment functions for the first and second regions and manual intervention functions for the evaluation results, so as to further improve human-computer interaction efficiency and evaluation efficiency. These functions will be introduced below through steps 505 to 507.
[0151] 505. In response to a drag operation performed on the target control, the terminal displays the changes in the area of the first and second regions on the evaluation result display interface as the drag operation occurs.
[0152] In this embodiment, the target control can be displayed as a dividing line between the first and second areas, or as a bar at any position on the evaluation result display interface; this application is not limited to these forms. Displaying the change in area between the first and second areas on the evaluation result display interface as the drag operation occurs means dynamically displaying the area change process of the first and second areas as the drag operation is performed. Illustratively, in response to a first drag operation on the target control, the terminal shrinks the area of the first area and increases the area of the second area; the first drag operation refers to a drag operation with the drag direction pointing towards the first area. In response to a second drag operation on the target control, the terminal increases the area of the first area and shrinks the area of the second area; the second drag operation refers to a drag operation with the drag direction pointing towards the second area.
[0153] The following is for reference. Figure 9 Here is an example to illustrate step 505.
[0154] Figure 9 This is a schematic diagram of another evaluation result display interface provided in an embodiment of this application. For example... Figure 9 As shown in Figure (a), the target control is displayed between the first and second areas using a dividing line. In response to a first drag operation on the target control (an upward drag), the terminal shrinks the area of the first area and increases the area of the second area; in response to a second drag operation on the target control (a downward drag), the terminal increases the area of the first area and shrinks the area of the second area. Figure 9As shown in Figure (b), the target control is displayed in the form of a bar and is positioned at any location on the evaluation result display interface (the position of the target control can be adjusted as needed). This target control can be understood as an adjustment track. In response to a first drag operation on the target control, i.e., dragging upwards, the terminal reduces the area of the first region and increases the area of the second region; in response to a second drag operation on the target control, i.e., dragging downwards, the terminal increases the area of the first region and reduces the area of the second region. It should be understood that the form of the target control shown in the figure is merely illustrative and does not constitute a limitation of this application. For example, the target control could also be an edit box for adjusting the area of a region, allowing the terminal to dynamically adjust the areas of the first and second regions in response to the area entered in the target control, and so on.
[0155] In some embodiments, the terminal dynamically updates the displayed content of the first and second regions when the areas of the first and second regions change with dragging operations. Specifically, if the area of the first region increases, additional information is displayed in the first region. For example, if the area of the first region increases to a preset threshold or the ratio of the areas of the first and second regions is greater than a preset threshold, the output result corresponding to each first prompt word is displayed. If the area of the second region increases, additional information is displayed in the second region. For example, if the area of the second region increases to a preset threshold or the ratio of the areas of the second and first regions is greater than a preset threshold, the use case content for each use case is displayed. Conversely, if the area of the first region decreases, the information already displayed in the first region is reduced. For example, if the area of the first region decreases to a preset threshold or the ratio of the areas of the first and second regions is less than a preset threshold, the text identifier in the first evaluation identifier is hidden, while the graphic identifier is retained. If the area of the second region shrinks, the information displayed in the second region is reduced. For example, if the area of the second region shrinks to a preset threshold or the ratio of the area of the second region to the area of the first region is less than a preset threshold, all prompt word templates are treated as a single, unified prompt word template. At least one prompt word template is displayed in the second region as a third evaluation identifier for multiple use cases in the use case set. Each third evaluation identifier indicates the overall effectiveness of at least one prompt word template obtained based on the corresponding use case in the use case set. This method of dynamically updating the displayed content according to changes in the area fully utilizes the interface space, simplifies the operation of viewing evaluation results, and meets the personalized needs for viewing evaluation results. This not only improves evaluation efficiency but also enhances the user experience.
[0156] The following is for reference. Figure 10 Here is an example illustrating another implementation of step 505.
[0157] Figure 10This is a schematic diagram of another evaluation result display interface provided in an embodiment of this application. For example... Figure 10 As shown in Figure (a), the evaluation result display interface includes a first area 1001 and a second area 1002. The first area 1001 includes a display area for each prompt word template, displaying the first evaluation identifier for the corresponding prompt word template for the first use case in each display area. The second area 1002 displays the second evaluation identifier for each prompt word template for multiple use cases in the use case set. A target control 1003 "boundary line" is displayed between the first area 1001 and the second area 1002. In response to a downward drag operation on the target control 1003, the terminal increases the area of the first area and decreases the area of the second area. If the area of the first area increases to a preset threshold or the ratio of the area of the first area to the area of the second area is greater than a preset threshold, the evaluation result display interface is as follows. Figure 10 As shown in Figure (b), in the first area 1001, the display area for each prompt word template displays the first evaluation identifier of the corresponding prompt word template for the first use case and the output result corresponding to at least one first prompt word. For example, taking the sentiment analysis function shown in step 501 above, "positive" corresponds to output value 1, "neutral" corresponds to output value 2, and "negative" corresponds to output value 3. Then, the output value of prompt word template 1 is 1, the output value of prompt word template 2 is 3, and prompt word template 3 is being calculated and is being output. In the second area 1002, the third evaluation identifier of each prompt word template for multiple use cases in the use case set is displayed. Each third evaluation identifier indicates the comprehensive usage effect of all prompt word templates obtained based on the corresponding use cases in the use case set. Taking the evaluation result determined by matching degree as an example, the same third evaluation identifier can be displayed for each use case to provide a navigation function. That is, relevant personnel can view the corresponding evaluation results by selecting different use cases. This process refers to step 504 above and will not be repeated. Of course, in some embodiments, taking the evaluation result determined by similarity as an example, the corresponding third evaluation identifier can be displayed for each use case based on the average similarity score of all prompt word templates for a certain use case. This application is not limited to this. It should be understood that the above Figure 10 The evaluation results display interface shown in Figure (b) can also be understood as a current use case viewing mode, that is, relevant personnel can view the horizontal comparison of a certain use case under different prompt word templates in detail, so as to compare the usage effects of different prompt word templates.
[0158] 506. The terminal responds to the trigger operation performed on the display area corresponding to the target prompt word template in the first area, and displays the output result and reference result corresponding to the target prompt word. The target prompt word is generated based on the first use case and the target prompt word template.
[0159] In this embodiment, the target prompt word template refers to any prompt word template. The triggering operation performed on the display area corresponding to the target prompt word template can be a click operation or a click operation on the prompt word template viewing control in the first area; this application is not limited to these. In response to the triggering operation performed on the display area corresponding to the target prompt word template, the terminal displays the output result and reference result corresponding to the target prompt word in the first area of the evaluation result display interface for relevant personnel to view.
[0160] 507. In response to the evaluation identifier update operation performed on the target prompt word template, the terminal updates the first evaluation identifier of the target prompt word template for the first use case.
[0161] In this embodiment, the terminal provides a manual intervention function for the evaluation results. That is, relevant personnel can update the first evaluation identifier of the target prompt word template based on the output and reference results of the target prompt word displayed by the terminal, thereby updating the first evaluation identifier of the target prompt word template for the first use case. This is achieved through an evaluation identifier update operation performed on the target prompt word template. Indicatively, the evaluation identifier update operation is any of the following:
[0162] The first method involves adjusting the evaluation indicator track, which indicates multiple different evaluation indicators.
[0163] The first area displays an evaluation label adjustment track. The two ends of the track indicate the upper and lower limits for adjusting the evaluation labels. Indicatively, the evaluation label update operation includes at least one of the following: triggering an operation at any position on the evaluation label adjustment track; or sliding along the track to reach any position. For example, the two ends of the track display prompts indicating "compliant" for the upper limit and "non-compliant" for the lower limit. This means the evaluation labels indicated by the track are not fixed, and their form can be customized according to actual needs (e.g., different labels can be distinguished by color depth). Users can select any evaluation label by triggering an operation at any position or by sliding along the track. The terminal then updates according to the corresponding evaluation label. This method satisfies the personalized needs of users for updating evaluation labels, thereby improving the user experience.
[0164] The second type involves operations performed on the evaluation label adjustment option, which indicates multiple different evaluation labels.
[0165] The first area displays multiple assessment indicator adjustment options, each corresponding to a different assessment indicator. For example, the adjustment options are 100, 80, 60, 40, 20, and 0. These options indicate multiple fixed assessment indicators, and users can trigger an action on a specific option, causing the terminal to update according to the corresponding assessment indicator. This approach allows users to intuitively understand the available assessment indicators, thereby improving human-computer interaction efficiency.
[0166] Indicatively, for reference Figure 11 , Figure 11 This is a schematic diagram illustrating an evaluation identifier update operation provided in an embodiment of this application. For example... Figure 11 As shown in Figure (a), the first area 1101 displays an assessment indicator adjustment track. The track has upper limit "compliant" and lower limit "non-compliant" prompts at both ends. Personnel can select any assessment indicator by triggering an operation at any position or by sliding an operation on the assessment indicator adjustment track. Figure 11 As shown in Figure (b), the first area 1101 displays multiple evaluation indicator adjustment options, namely 100, 80, 60, 40, 20, and 0. Relevant personnel can trigger an operation on a specific adjustment option to select the corresponding evaluation indicator. Additionally, the first area 1101 also displays a prompt word template switching control. The terminal can respond to triggering an operation on this prompt word template switching control to switch the currently displayed prompt word template in the first area 1101. It should be understood that the evaluation indicator update operation shown in the figure is merely illustrative and does not constitute a limitation of this application. For example, the evaluation result display interface may also display an edit box for adjusting the evaluation indicator, and the terminal can respond to the evaluation indicator entered in the edit box to update the corresponding evaluation indicator, and so on.
[0167] In some embodiments, the evaluation result display interface also provides an evaluation result comparison function for prompt word templates. That is, relevant personnel can select the prompt word template that meets the requirements more based on the evaluation results of multiple prompt word templates displayed on the terminal, and then update the evaluation identifier of the corresponding prompt word template. Illustratively, this process includes: in response to a prompt word template comparison operation performed in the first area, displaying the evaluation results of the first prompt word template and the second prompt word template; in response to a trigger operation performed on the display area corresponding to the first prompt word template, or in response to a trigger operation performed on the display area corresponding to the second prompt word template, updating the first evaluation identifier of the first and second prompt word templates for the first use case. The prompt word template comparison operation can be a trigger operation on the display area corresponding to a certain prompt word template (e.g., double-click or long-press), or it can be a trigger operation on the prompt word template comparison control in the first area; this application is not limited to these.
[0168] Indicatively, for reference Figure 12 , Figure 12 This is a schematic diagram illustrating a prompt word template comparison operation provided in an embodiment of this application. For example... Figure 12 As shown, the first area 1201 displays the evaluation results of the first and second prompt word templates. In response to a trigger operation on the display area corresponding to either prompt word template, the first evaluation identifiers of both prompt word templates are updated. For example, if the first evaluation identifier of the first prompt word template was originally the text identifier "match" and the first evaluation identifier of the second prompt word template was originally "match", if a trigger operation is performed on the display area corresponding to the first prompt word template, the first evaluation identifier of the first prompt word template remains unchanged, and the first evaluation identifier of the second prompt word template is updated to "not match". That is, relevant personnel can select one of the two prompt word templates to indicate that the selected prompt word template is more suitable than the unselected prompt word template. In some embodiments, a first control "Almost" is also displayed on the first area 1201. In response to a trigger operation on this first control, the terminal updates the first evaluation identifiers of the two prompt word templates. For example, the first evaluation identifier of the first prompt word template was originally the text identifier "Match," and the first evaluation identifier of the second prompt word template was originally "Mismatch." If a trigger operation is performed on the first control, the first evaluation identifier of the first prompt word template remains unchanged, and the first evaluation identifier of the second prompt word template is updated to "Match," or the first evaluation identifier of the second prompt word template remains unchanged, and the first evaluation identifier of the first prompt word template is updated to "Mismatch." This can be set according to actual needs, and this application does not limit it. It should be understood that the above... Figure 12The evaluation results display interface shown can also be understood as a prompt word template comparison mode. That is, relevant personnel can compare any two prompt word templates horizontally to select the prompt word template that meets the requirements, providing a reference for further optimization of the prompt word template.
[0169] It should be noted that steps 505 to 507 are optional steps. That is, the terminal may execute at least one of steps 505 to 507, or may not execute any of them. In addition, this application does not limit the execution order of steps 505 to 507.
[0170] In summary, the method for displaying the evaluation results of the prompt word templates provided in this application, when evaluating the prompt word templates of the AI model, displays the evaluation results of the prompt word templates in different areas of the evaluation result display interface in the form of evaluation identifiers according to different evaluation dimensions. This allows relevant personnel to intuitively understand the local usage effect of using different prompt word templates on the same use case and the global usage effect of using different prompt word templates on the entire set of use cases. This approach effectively improves the display effect of the evaluation results, enabling relevant personnel to quickly and intuitively understand the evaluation results, improving human-computer interaction efficiency, and thus enhancing the evaluation efficiency and accuracy of the prompt word templates.
[0171] Figure 13 This is a schematic diagram of a device for displaying the evaluation results of a prompt word template, as provided in an embodiment of this application. This device can achieve some or all of the functions of the aforementioned terminal through software, hardware, or a combination of both. For example... Figure 13 As shown, the device includes an acquisition module 1301 and a display module 1302.
[0172] The acquisition module 1301 is used to acquire a set of use cases and at least one prompt word template. The set of use cases includes multiple use cases. The prompt word template is used to generate prompt words based on the use cases. The prompt words are used as input to the AI model to obtain the output results corresponding to the prompt words.
[0173] The display module 1302 is used to display the display area corresponding to each prompt word template in the first area of the evaluation result display interface based on the first use case in the use case set and at least one prompt word template. In each display area, the first evaluation identifier of the corresponding prompt word template for the first use case is displayed. The first evaluation identifier indicates the usage effect of the prompt word template obtained based on the first use case.
[0174] The display module 1302 is also used to display the second evaluation identifier of each prompt word template for the use case set in the second area of the evaluation result display interface. Each second evaluation identifier indicates the usage effect of the prompt word template obtained based on the corresponding use case in the use case set. The first area and the second area are adjacent.
[0175] In some embodiments, the apparatus further includes an evaluation result determination module, configured to:
[0176] Fill the first use case into at least one prompt word template to obtain at least one first prompt word;
[0177] Input at least one first prompt word into the AI model and obtain the output result corresponding to at least one first prompt word;
[0178] Based on the difference between the output result and the reference result corresponding to at least one first prompt word, the evaluation result of at least one prompt word template is determined, and the evaluation result is used to determine the first evaluation identifier.
[0179] In some embodiments, the evaluation result determination module is used for at least one of the following:
[0180] Based on the matching degree between the output result and the reference result corresponding to at least one first prompt word, determine the evaluation result of at least one prompt word template;
[0181] The evaluation result of at least one prompt word template is determined based on the similarity between the output result and the reference result corresponding to at least one first prompt word.
[0182] In some embodiments, the first evaluation identifier includes at least one of a text identifier and a graphic identifier; the second evaluation identifier includes at least one of a text identifier and a graphic identifier.
[0183] In some embodiments, the first evaluation identifier includes at least one of the following: a symbol for indicating the effect of use, text for indicating the effect of use, and a numerical value for indicating the effect of use; the second evaluation identifier includes at least one of the following: a symbol for indicating the effect of use, text for indicating the effect of use, and a numerical value for indicating the effect of use.
[0184] In some embodiments, different first evaluation identifiers indicate different usage effects of prompt word templates obtained based on first use cases; different second evaluation identifiers indicate different usage effects of prompt word templates obtained based on corresponding use cases in the use case set.
[0185] In some embodiments, the display module 1302 is further configured to:
[0186] In response to a drag operation performed on the target control, the area of the first and second regions in the evaluation results display interface changes as the drag operation occurs.
[0187] In some embodiments, the display module 1302 is used for any of the following:
[0188] In response to the first drag operation performed on the target control, the area of the first region is reduced and the area of the second region is increased. The first drag operation refers to a drag operation in which the drag direction is directed toward the first region.
[0189] In response to a second drag operation performed on the target control, the area of the first region is increased and the area of the second region is decreased. The second drag operation refers to a drag operation in which the drag direction is directed towards the second region.
[0190] In some embodiments, the display module 1302 is used for:
[0191] Each display area shows the first evaluation identifier for the first use case and the output result corresponding to at least one first prompt word for each prompt word template. The first prompt word is generated based on the first use case and the prompt word template.
[0192] In some embodiments, the display module 1302 is further configured to: display at least one prompt word template for a set of use cases in a second area as a third evaluation identifier, each third evaluation identifier indicating the overall effect of using at least one prompt word template based on the corresponding use case in the set of use cases.
[0193] In some embodiments, the display module 1302 is further configured to:
[0194] In response to a trigger operation performed on the display area corresponding to the target prompt word template in the first region, the output result and reference result corresponding to the target prompt word are displayed. The target prompt word is generated based on the first use case and the target prompt word template.
[0195] In response to the evaluation identifier update operation performed on the target prompt word template, update the first evaluation identifier of the target prompt word template for the first use case.
[0196] In some embodiments, the evaluation identifier update operation is any of the following: an operation performed on an evaluation identifier adjustment track, which indicates multiple different evaluation identifiers; or an operation performed on an evaluation identifier adjustment option, which indicates multiple different evaluation identifiers.
[0197] In some embodiments, the display module 1302 is further configured to:
[0198] In response to the prompt template comparison operation performed in the first area, the evaluation results of the first prompt template and the evaluation results of the second prompt template are displayed;
[0199] In response to a trigger operation performed on the display area corresponding to the first prompt word template, or in response to a trigger operation performed on the display area corresponding to the second prompt word template, update the first evaluation identifier of the first prompt word template and the second prompt word template for the first use case.
[0200] When evaluating the prompt word templates of the AI model using the aforementioned device, the evaluation results are displayed in different areas of the evaluation result display interface according to different evaluation dimensions, in the form of evaluation labels. This allows relevant personnel to intuitively understand the local usage effect of using different prompt word templates on the same use case and the global usage effect of using different prompt word templates on the entire set of use cases. This method effectively improves the display effect of the evaluation results, enabling relevant personnel to quickly and intuitively understand the evaluation results, improving human-computer interaction efficiency, and thus enhancing the evaluation efficiency and accuracy of prompt word templates.
[0201] Furthermore, in the above-described apparatus, both the acquisition module 1301 and the display module 1302 can be implemented in software or in hardware. For example, the implementation of the acquisition module 1301 will be described below. Similarly, the implementation of other modules can refer to the implementation of the acquisition module 1301.
[0202] As an example of a software functional unit, module 1301 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, or a container. Further, the aforementioned computing instance may be one or more. For example, module 1301 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Further, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one or more geographically proximate data centers. Typically, a region may include multiple AZs.
[0203] Similarly, multiple hosts / virtual machines / containers used to run this code can be distributed within the same Virtual Private Cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Communication between two VPCs within the same region, as well as between VPCs in different regions, requires a communication gateway to be set up within each VPC to enable interconnection between VPCs.
[0204] As an example of a hardware functional unit, the acquisition module 1301 may include at least one computing device. Alternatively, the acquisition module 1301 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be implemented using a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), generic array logic (GAL), or any combination thereof.
[0205] The multiple computing devices included in the acquisition module 1301 can be distributed in the same region or in different regions. Similarly, the multiple computing devices included in the acquisition module 1301 can be distributed in the same Availability Zone (AZ) or in different AZs. Likewise, the multiple computing devices included in the acquisition module 1301 can be distributed in the same Virtual Private Cloud (VPC) or in multiple VPCs. These multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.
[0206] It should be noted that the above-described embodiment of the prompt word template evaluation result display device is only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. Furthermore, the prompt word template evaluation result display device and the prompt word template evaluation result display method embodiment belong to the same concept, and their specific implementation process is detailed in the method embodiment, which will not be repeated here.
[0207] In this application, the terms "first," "second," etc., are used to distinguish identical or similar items with substantially the same function and purpose. It should be understood that there is no logical or temporal dependency between "first," "second," and "nth," nor does it limit the quantity or execution order. It should also be understood that although the following description uses the terms "first," "second," etc., to describe various elements, these elements should not be limited by the terms. These terms are merely used to distinguish one element from another. For example, without departing from the various examples described, a first use case can be referred to as a second use case, and similarly, a second use case can be referred to as a first use case. Both the first and second use cases can be use cases, and in some cases, they can be separate and distinct use cases.
[0208] In this application, the term "at least one" means one or more, and the term "multiple" means two or more. For example, multiple use cases refer to two or more use cases.
[0209] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0210] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, in the form of program structure information. This program structure information includes one or more program instructions. When these program instructions are loaded and executed on a computing device, the processes or functions according to the embodiments of this application are generated, in whole or in part.
[0211] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0212] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
Claims
1. A method for displaying the evaluation results of a prompt word template, characterized in that, The method includes: Obtain a set of use cases and at least one prompt word template. The set of use cases includes multiple use cases. The prompt word template is used to generate prompt words based on the use cases. The prompt words are used as input to an artificial intelligence (AI) model to obtain the output result corresponding to the prompt words. Based on the first use case in the set of use cases and at least one of the prompt word templates, a display area corresponding to each prompt word template is displayed in the first area of the evaluation result display interface. In each of the display areas, a first evaluation identifier for the corresponding prompt word template for the first use case is displayed. The first evaluation identifier indicates the usage effect of the prompt word template obtained based on the first use case. The second area of the evaluation result display interface displays a second evaluation identifier for each of the prompt word templates for the set of use cases. Each second evaluation identifier indicates the usage effect of the prompt word template obtained based on the corresponding use case in the set of use cases. The first area is adjacent to the second area.
2. The method according to claim 1, characterized in that, The method further includes: The first use case is filled into at least one of the prompt word templates to obtain at least one first prompt word; Input at least one of the first prompt words into the AI model to obtain output results corresponding to at least one of the first prompt words; Based on the difference between the output result and the reference result corresponding to at least one of the first prompt words, an evaluation result of at least one of the prompt word templates is determined, and the evaluation result is used to determine the first evaluation identifier.
3. The method according to claim 2, characterized in that, The determination of the evaluation result of at least one of the prompt word templates based on the difference between the output result and the reference result corresponding to at least one of the first prompt words includes at least one of the following: Based on the matching degree between the output result and the reference result corresponding to at least one of the first prompt words, the evaluation result of at least one prompt word template is determined; The evaluation result of at least one of the prompt word templates is determined based on the similarity between the output result and the reference result corresponding to at least one of the first prompt words.
4. The method according to any one of claims 1 to 3, characterized in that, The first evaluation identifier includes at least one of text identifiers and graphic identifiers; The second evaluation identifier includes at least one of text identifiers and graphic identifiers.
5. The method according to any one of claims 1 to 3, characterized in that, The first evaluation identifier includes at least one of the following: a symbol for representing the effect of use, text for representing the effect of use, and a numerical value for representing the effect of use; The second evaluation identifier includes at least one of the following: a symbol for indicating the effect of use, text for indicating the effect of use, and a numerical value for indicating the effect of use.
6. The method according to any one of claims 1 to 3, characterized in that, Different first evaluation identifiers indicate different usage effects of the prompt word templates obtained based on the first use case. Different second evaluation identifiers indicate different usage effects of the prompt word templates obtained based on the corresponding use cases in the use case set.
7. The method according to any one of claims 1 to 3, characterized in that, The method further includes: In response to a drag operation performed on the target control, the evaluation result display interface shows that the areas of the first region and the second region change as the drag operation occurs.
8. The method according to claim 7, characterized in that, The response to a drag operation performed on the target control, displaying changes in the areas of the first and second regions on the evaluation result display interface as a result of the drag operation, includes any of the following: In response to a first drag operation performed on the target control, the area of the first region is reduced and the area of the second region is increased, wherein the first drag operation refers to a drag operation in which the drag direction is directed toward the first region; In response to a second drag operation performed on the target control, the area of the first region is increased and the area of the second region is decreased. The second drag operation refers to a drag operation in which the drag direction is directed towards the second region.
9. The method according to any one of claims 1, 2, 3 or 8, characterized in that, The step of displaying the corresponding prompt word template for the first evaluation identifier of the first use case in each of the display areas includes: In each of the display areas, the first evaluation identifier of the corresponding prompt word template for the first use case and the output result corresponding to at least one first prompt word are displayed. The first prompt word is generated based on the first use case and the prompt word template.
10. The method according to any one of claims 1, 2, 3 or 8, characterized in that, The method further includes: The second area displays at least one of the prompt word templates for the set of use cases, with each third evaluation identifier indicating the overall usage effect of at least one of the prompt word templates obtained based on the corresponding use case in the set of use cases.
11. The method according to any one of claims 1, 2, 3 or 8, characterized in that, The method further includes: In response to a trigger operation performed on the display area corresponding to the target prompt word template in the first region, the output result and reference result corresponding to the target prompt word are displayed, wherein the target prompt word is generated based on the first use case and the target prompt word template; In response to the evaluation identifier update operation performed on the target prompt word template, the first evaluation identifier of the target prompt word template for the first use case is updated.
12. The method according to claim 11, characterized in that, The evaluation identifier update operation is any one of the following: The operation performed on the evaluation identifier adjustment track, which indicates multiple different evaluation identifiers; The operation performed on the evaluation identifier adjustment option, which indicates multiple different evaluation identifiers.
13. The method according to any one of claims 1, 2, 3, 8 or 12, characterized in that, The method further includes: In response to the prompt word template comparison operation performed in the first area, the evaluation results of the first prompt word template and the evaluation results of the second prompt word template are displayed; In response to a trigger operation performed on the display area corresponding to the first prompt word template, or in response to a trigger operation performed on the display area corresponding to the second prompt word template, update the first evaluation identifier of the first prompt word template and the second prompt word template for the first use case.
14. A device for displaying the evaluation results of a prompt word template, characterized in that, The device includes: An acquisition module is used to acquire a set of use cases and at least one prompt word template. The set of use cases includes multiple use cases, and the prompt word template is used to generate prompt words based on the use cases. The prompt words are used as input to an AI model to obtain the output result corresponding to the prompt words. The display module is used to display a display area corresponding to each prompt word template in a first area of the evaluation result display interface based on a first use case in the use case set and at least one prompt word template. In each display area, a first evaluation identifier of the corresponding prompt word template for the first use case is displayed. The first evaluation identifier indicates the usage effect of the prompt word template obtained based on the first use case. The display module is further configured to display a second evaluation identifier for each of the prompt word templates for the set of use cases in the second area of the evaluation result display interface. Each second evaluation identifier indicates the usage effect of the prompt word template obtained based on the corresponding use case in the set of use cases. The first area and the second area are adjacent.
15. The apparatus according to claim 14, characterized in that, The device further includes an evaluation result determination module, used for: The first use case is filled into at least one of the prompt word templates to obtain at least one first prompt word; Input at least one of the first prompt words into the AI model to obtain output results corresponding to at least one of the first prompt words; Based on the difference between the output result and the reference result corresponding to at least one of the first prompt words, an evaluation result of at least one of the prompt word templates is determined, and the evaluation result is used to determine the first evaluation identifier.
16. The apparatus according to claim 15, characterized in that, The evaluation result determination module is used for at least one of the following: Based on the matching degree between the output result and the reference result corresponding to at least one of the first prompt words, the evaluation result of at least one prompt word template is determined; The evaluation result of at least one of the prompt word templates is determined based on the similarity between the output result and the reference result corresponding to at least one of the first prompt words.
17. The apparatus according to any one of claims 14 to 16, characterized in that, The first evaluation identifier includes at least one of text identifiers and graphic identifiers; The second evaluation identifier includes at least one of text identifiers and graphic identifiers.
18. The apparatus according to any one of claims 14 to 16, characterized in that, The first evaluation identifier includes at least one of the following: a symbol for representing the effect of use, text for representing the effect of use, and a numerical value for representing the effect of use; The second evaluation identifier includes at least one of the following: a symbol for indicating the effect of use, text for indicating the effect of use, and a numerical value for indicating the effect of use.
19. The apparatus according to any one of claims 14 to 16, characterized in that, Different first evaluation identifiers indicate different usage effects of the prompt word templates obtained based on the first use case. Different second evaluation identifiers indicate different usage effects of the prompt word templates obtained based on the corresponding use cases in the use case set.
20. The apparatus according to any one of claims 14 to 16, characterized in that, The display module is also used for: In response to a drag operation performed on the target control, the evaluation result display interface shows that the areas of the first region and the second region change as the drag operation occurs.
21. The apparatus according to claim 20, characterized in that, The display module is used for any of the following: In response to a first drag operation performed on the target control, the area of the first region is reduced and the area of the second region is increased, wherein the first drag operation refers to a drag operation in which the drag direction is directed toward the first region; In response to a second drag operation performed on the target control, the area of the first region is increased and the area of the second region is decreased. The second drag operation refers to a drag operation in which the drag direction is directed towards the second region.
22. The apparatus according to any one of claims 14, 15, 16 or 21, characterized in that, The display module is used for: In each of the display areas, the first evaluation identifier of the corresponding prompt word template for the first use case and the output result corresponding to at least one first prompt word are displayed. The first prompt word is generated based on the first use case and the prompt word template.
23. The apparatus according to any one of claims 14, 15, 16 or 21, characterized in that, The display module is also used for: The second area displays at least one of the prompt word templates for the set of use cases, with each third evaluation identifier indicating the overall usage effect of at least one of the prompt word templates obtained based on the corresponding use case in the set of use cases.
24. The apparatus according to any one of claims 14, 15, 16 or 21, characterized in that, The display module is also used for: In response to a trigger operation performed on the display area corresponding to the target prompt word template in the first region, the output result and reference result corresponding to the target prompt word are displayed, wherein the target prompt word is generated based on the first use case and the target prompt word template; In response to the evaluation identifier update operation performed on the target prompt word template, the first evaluation identifier of the target prompt word template for the first use case is updated.
25. The apparatus according to claim 24, characterized in that, The evaluation identifier update operation is any one of the following: The operation performed on the evaluation identifier adjustment track, which indicates multiple different evaluation identifiers; The operation performed on the evaluation identifier adjustment option, which indicates multiple different evaluation identifiers.
26. The apparatus according to any one of claims 14, 15, 16, 21 or 25, characterized in that, The display module is also used for: In response to the prompt word template comparison operation performed in the first area, the evaluation results of the first prompt word template and the evaluation results of the second prompt word template are displayed; In response to a trigger operation performed on the display area corresponding to the first prompt word template, or in response to a trigger operation performed on the display area corresponding to the second prompt word template, update the first evaluation identifier of the first prompt word template and the second prompt word template for the first use case.
27. A computing device, characterized in that, The computing device includes a processor and a memory, the memory being used to store at least one piece of program code, the at least one piece of program code being loaded by the processor and executed as the method for displaying the evaluation results of the prompt word template as described in any one of claims 1 to 13.
28. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store at least one piece of program code, which is used to execute the method for displaying the evaluation results of the prompt word template as described in any one of claims 1 to 13.
29. A computer program product, characterized in that, When the computer program product is run on a computing device, the computing device performs the evaluation result display method of the prompt word template as described in any one of claims 1 to 13.
Citation Information
Patent Citations
Artificial intelligence model test method and device, electronic equipment and storage medium
CN114492764A
Predictive model scoring to optimize test case order in real time
US9495642B1