User experience evaluation method and system

By analyzing emotional data during user interactions with AI products, this approach addresses the issues of low accuracy and reliability caused by manual evaluation in existing technologies, achieving more accurate user experience assessments applicable to various AI product application scenarios.

CN120848726APending Publication Date: 2025-10-28ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510947441.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-09
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing user experience evaluation methods rely on manual methods, which are easily affected by subjective human factors, resulting in low accuracy and reliability, low user participation, and a disconnect from actual user experience.

Method used

By obtaining interaction data from the process of users interacting with AI products, analyzing users' positive and negative feedback emotions, and determining the evaluation results of user experience based on interaction emotions.

Benefits of technology

It improves the effectiveness and reliability of user experience evaluation, aligns with actual user application scenarios and feelings, covers personalized user characteristics, and achieves effective and reliable user experience evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120848726A_ABST
    Figure CN120848726A_ABST
Patent Text Reader

Abstract

The invention provides a user experience evaluation method and system, and the method comprises the steps: obtaining the interaction data of an interaction process between a user and an AI product, analyzing the interaction emotion of the user in the interaction process based on the interaction data, and determining an evaluation result of the experience of the user on the AI product according to the interaction emotion. The method can effectively fit the actual user application scene and the actual user feeling, and relatively effectively covers the complex personalized characteristics of the user. And the effectiveness and the reliability of evaluation can also be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of artificial intelligence (AI) technology, and in particular to a method and system for evaluating user experience. Background Technology

[0002] With the development of AI technology, AI products are being applied more and more widely in various fields, such as healthcare, financial services, education, and entertainment. Therefore, how to evaluate user experience with AI products in order to continuously improve user satisfaction has become a pressing issue.

[0003] In related technologies, user experience evaluation is primarily achieved through human intervention. For example, training evaluation teams to understand business needs enables them to annotate the mind maps of AI product models, mainly for rapid evaluation of service model iterations. Therefore, the evaluation content focuses more on the model's strengths, such as its professional accuracy.

[0004] However, the aforementioned evaluation methods require manual implementation, making them susceptible to subjective human factors and resulting in lower accuracy and reliability. Furthermore, because the evaluation focuses on the model itself, it also suffers from low user participation and a disconnect from actual user experience.

[0005] It should be noted that the above-mentioned related technologies are only information known to the inventor personally, and do not mean that the above information had entered the public domain before the application date of this specification, nor do they mean that it can be considered prior art in this specification. Summary of the Invention

[0006] This specification provides a method and system for evaluating user experience to avoid at least one of the aforementioned technical problems.

[0007] Firstly, this specification provides a user experience evaluation method, which is applied to the evaluation of users' experience with AI products, including:

[0008] Obtain interaction data of the user's interaction process with the AI ​​product;

[0009] The interaction data is used to analyze the user's emotional responses during the interaction process, wherein the emotional responses include positive feedback and / or negative feedback; and

[0010] The evaluation result of the user's experience with the AI ​​product is determined based on the interactive emotions.

[0011] Secondly, this specification provides a user experience evaluation system, which is applied to evaluating users' experience with AI products, including:

[0012] At least one storage medium storing at least one instruction set for evaluating the user experience of using AI products;

[0013] At least one processor is communicatively connected to the at least one storage medium, wherein when the at least one processor is running, it reads the at least one instruction set and executes the evaluation method as described in the first aspect according to the instructions of the at least one instruction set.

[0014] Thirdly, this specification provides a computer-readable non-transitory storage medium, wherein the computer-readable non-transitory storage medium stores at least one instruction set, which is executed by at least one processor to implement the method as described in the first aspect.

[0015] As can be seen from the above technical solutions, the user experience evaluation method and system provided in this specification, by taking a user-centric approach and determining the user's emotional engagement during the interaction with the AI ​​product based on the interaction data generated during the interaction, thereby determining the user's evaluation of the AI ​​product experience. This effectively aligns with actual user application scenarios and user experiences, and more effectively covers the complex and personalized characteristics of users. It avoids the technical problems mentioned in the related technologies above. It can improve the effectiveness and reliability of the evaluation, such as enabling an effective and reliable assessment of whether the user experience is "good" or "bad."

[0016] The user experience evaluation methods and other system functionalities provided in this specification are partially listed in the following description. The ingenious aspects of the user experience evaluation methods and systems provided in this specification can be fully explained through practice or by using the methods, devices, and combinations described in the detailed examples below. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of this specification, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 A schematic diagram illustrating an application scenario of the user experience evaluation method provided in the embodiments of this specification;

[0019] Figure 2 A schematic diagram of the structure of a user experience evaluation system provided in the embodiments of this specification;

[0020] Figure 3 A flowchart illustrating a user experience evaluation method provided in one embodiment of this specification;

[0021] Figure 4 A flowchart illustrating a user experience evaluation method provided in another embodiment of this specification;

[0022] Figure 5 This is a schematic diagram illustrating the principle of the user experience evaluation method provided in the embodiments of this specification. Detailed Implementation

[0023] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this specification. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this specification as detailed in the appended claims.

[0024] It should be understood that the terms “comprising” and “having”, and any variations thereof, in the embodiments of this specification are intended to cover but not exclude inclusion. For example, a product or device that includes a series of components is not necessarily limited to those components that are explicitly listed, but may include other components that are not explicitly listed or that are inherent to such product or device.

[0025] The term "and / or" in the embodiments of this specification describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.

[0026] In the embodiments of this specification, the term "multiple" refers to two or more, and other quantifiers are similar.

[0027] The terms “first,” “second,” “third,” “initial,” “target,” etc., used in this specification are used to distinguish similar or related objects or entities and do not necessarily imply a specific order or sequence, unless otherwise indicated. It should be understood that such terms can be used interchangeably where appropriate, for example, in situations where implementation can proceed in a sequence other than those given in the embodiments illustrated or described in this specification.

[0028] As used in this specification, the term "unit / module" means any known or subsequently developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code capable of performing the functions associated with that element.

[0029] To avoid at least one of the technical problems mentioned in the background section above, this specification proposes a technical concept developed through inventive effort: starting from the user's feelings when using AI products, analyzing the user's emotional data during the process of using AI products, and evaluating the user's user experience of using AI products based on the user's emotional data.

[0030] For example, the process of a user using an AI product can be considered as the interaction process between the user and the AI ​​product. First, interaction data during the interaction process can be obtained.

[0031] Then, the user's emotional state during the interaction can be analyzed based on this interaction data. For example, it is possible to analyze the positive and / or negative emotional states of the user when using the AI ​​product.

[0032] Positive feedback emotions can be understood as the relatively "good" (or positive) emotions experienced by users during the use of AI products, such as active feedback, smooth cooperation, and in-depth content engagement. Negative feedback emotions can be understood as the relatively "bad" (or negative) emotions experienced by users during the use of AI products, such as dissatisfaction, confusion, needing further guidance, or even abandoning the product altogether.

[0033] Ultimately, user experience can be evaluated based on positive and / or negative feedback emotions to obtain an assessment of the user's experience using AI products from the user's perspective (such as the user's own feelings).

[0034] The technical solution provided in this specification is based on the aforementioned technical concept. As described above, the technical solution provided in this specification, from the user's perspective, determines the user's emotional engagement during the AI ​​interaction process based on the interaction data generated between the user and the AI ​​product. This emotional engagement is then used to determine the user's evaluation of the AI ​​product's experience. This effectively aligns with actual user application scenarios and user experiences, and more effectively covers the complex and personalized characteristics of users. It avoids the technical problems mentioned in the related technologies above. It improves the effectiveness and reliability of the evaluation, enabling an effective and reliable assessment of whether the user experience is "good" or "bad."

[0035] To facilitate readers' understanding of this manual, the application scenarios of this manual are introduced below.

[0036] The technical solutions provided in this specification are applicable to scenarios requiring user experience evaluation. For example, they can be applied to scenarios where user experience with AI products is evaluated. These AI products include: AI dialogue products, image and video analysis AI products, speech recognition and synthesis AI products, natural language processing AI products, recommendation system AI products, predictive analytics AI products, automated robot AI products, and intelligent decision support system AI products, etc.

[0037] Taking AI products as examples, such as chatbots, virtual assistants, and conversational AI platforms:

[0038] The interaction process between a user and an AI-powered conversational product can be considered a dialogue. The interaction data can be the dialogue data between the user and the AI-powered conversational product.

[0039] Correspondingly, the evaluation system can obtain dialogue data and analyze the interactive emotions (such as positive feedback emotions and / or negative feedback emotions) exhibited by users during the dialogue based on the dialogue data, so as to ultimately determine the evaluation result of the user's experience with the AI ​​dialogue product during this dialogue based on the interactive emotions.

[0040] Furthermore, the focus of analyzing interactive sentiment can vary depending on the scenario. For example, for security-related AI conversational products, greater emphasis can be placed on analyzing negative feedback sentiment. Conversely, for sales-related AI conversational products, the analysis of both positive and negative feedback sentiment can be given equal importance.

[0041] It should be noted that the above examples are only used to illustrate the application scenarios to which the technical solutions in this specification can be applied, and should not be construed as limiting the application scenarios.

[0042] Figure 1 This diagram illustrates an application scenario of the user experience evaluation method (hereinafter referred to as the evaluation method) according to embodiments of this specification. The evaluation method of this specification can be applied to, for example... Figure 1 Scenario 100 is shown. (e.g.) Figure 1 As shown, scenario 100 may include target user 101, client 102, server 103, and network 104.

[0043] The target user 101 can be the user who triggers the evaluation of the user experience. For example, the target user 101 can perform a targeted action on the client 102 to trigger the evaluation of the user experience.

[0044] Client 102 may be an electronic device that provides interactive functionality to target user 101. For example, client 102 may provide an interactive interface to target user 101, where target user 101 can perform target operations. In some embodiments, client 102 executes the evaluation method described herein in response to detecting an operation triggered by target user 101 to evaluate user experience. In this case, client 102 may store data or instructions for executing the evaluation method described herein, and may execute or be used to execute the data or instructions. In some embodiments, client 102 may include a hardware device with data processing capabilities and the necessary programs required to drive the hardware device to execute the evaluation method described herein.

[0045] In some embodiments, client 102 may include a mobile device, tablet, laptop, built-in device in a motor vehicle, or similar content, or any combination thereof. In some embodiments, the mobile device may include a smart home device, a smart mobile device, a virtual reality device, an augmented reality device, or similar device, or any combination thereof. In some embodiments, smart home devices may include a smart TV, a desktop computer, etc., or any combination thereof. In some embodiments, smart mobile devices may include a smartphone, a personal digital assistant, a gaming device, a navigation device, etc., or any combination thereof. In some embodiments, built-in devices in a motor vehicle may include an in-vehicle computer, an in-vehicle television, etc.

[0046] In some embodiments, client 102 may have one or more applications (APPs) installed. APPs provide target user 101 with the ability and interface to interact with the outside world via network 104. APPs include, but are not limited to: web browser APPs, search APPs, chat APPs, shopping APPs, video APPs, financial management APPs, instant messaging tools, email clients, social media platform software, etc.

[0047] like Figure 1 As shown, client 102 can establish a communication connection with server 103. Server 103 can communicate with one client 102 or multiple clients 102. In some embodiments, client 102 can interact with server 103 via network 104 to receive or send messages, etc.

[0048] Server 103 can be a server that provides various services. For example, server 103 can be a cloud server or a local server. Server 103 can communicate with one client 102 and receive data sent by that client 102, or it can communicate with multiple clients 102 and receive data sent by each client 102.

[0049] In some embodiments, the evaluation methods described herein can be executed on server 103. In this case, server 103 may store data or instructions for executing the evaluation methods described herein, and may execute or be used to execute the data or instructions. Server 103 may include hardware devices with data processing capabilities and the necessary programs required to drive the hardware devices.

[0050] In addition, the evaluation method described in this specification can be triggered based on the target operation of the target user 101, or it can be triggered based on the interaction between the target user 101 and the AI ​​product by the client 102 or the server 103.

[0051] Network 104 is a medium used to provide a communication connection between client 102 and server 103. Network 104 can facilitate the exchange of information or data. Figure 1 As shown, client 102 and server 103 can connect to network 104 respectively and transmit information or data to each other through network 104.

[0052] In some embodiments, network 104 can be any type of wired or wireless network, or a combination thereof. For example, network 104 may include a cable network, a wired network, a fiber optic network, a telecommunications network, an intranet, the Internet, a local area network (LAN), a wide area network (WAN), a wireless local area network (WLAN), a metropolitan area network (MAN), a public switched telephone network (PSTN), a Bluetooth network™, a ZigBee™ short-range wireless network, a near field communication (NFC) network, or a similar network.

[0053] In some embodiments, network 104 may include one or more network access points. For example, network 104 may include wired or wireless network access points, such as base stations or internet switching points, through which one or more components of client 102 and server 103 can connect to network 104 to exchange data or information.

[0054] It is worth noting that, Figure 1The number of clients 102, servers 103, and networks 104 shown is merely illustrative. Depending on implementation needs, there can be any number of clients 102, servers 103, and networks 104. Furthermore, the evaluation method provided in this specification can be executed entirely on client 102, entirely on server 103, or partially on client 102 and partially on server 103.

[0055] That is to say, Figure 1 and targeting Figure 1 The above description is only used to illustrate the possible application scenarios for which the evaluation methods in this specification may be applicable, and should not be construed as limiting the application scenarios.

[0056] Figure 2 A hardware structure diagram of an evaluation system 200 provided according to an embodiment of this specification is shown. The evaluation system 200 can perform the evaluation methods described in this specification. The evaluation methods are described in other parts of this specification. When the evaluation method is executed on client 102, the evaluation system 200 can be client 102. When the evaluation method is executed on server 103, the evaluation system 200 can be server 103. When the evaluation method is executed partly on client 102 and partly on server 103, the evaluation system 200 can be a system including client 102 and server 103.

[0057] like Figure 2 As shown, the evaluation system 200 may include at least one storage medium 203 and at least one processor 202. In some embodiments, the evaluation system 200 may also include a communication port 204 and an internal communication bus 201. The evaluation system 200 may also include I / O components 205.

[0058] The internal communication bus 201 can connect to different system components. For example, the internal communication bus 201 can connect to storage medium 203, processor 202, communication port 204, and I / O component 205.

[0059] I / O component 205 supports input / output between evaluation system 200 and other components.

[0060] Communication port 204 is used to evaluate data communication between system 200 and the outside world. For example, communication port 204 can be used to evaluate data communication between system 200 and network 104. Communication port 204 can be a wired communication port or a wireless communication port.

[0061] Storage medium 203 may include a data storage device. The data storage device may be a non-transitory storage medium or a temporary storage medium. For example, the data storage device may include one or more of a disk 2031, a read-only storage medium (ROM) 2032, or a random access storage medium (RAM) 2033. Storage medium 203 also includes at least one instruction set stored in the data storage device. The instruction set includes computer program code, which may include programs, routines, objects, components, data structures, procedures, modules, etc., that execute the evaluation methods provided in this specification.

[0062] At least one processor 202 may be communicatively connected to at least one storage medium 203. At least one processor 202 is used to execute at least one instruction set described above. When the evaluation system 200 is running, at least one processor 202 reads the at least one instruction set and, according to the instructions of the at least one instruction set, executes the evaluation method provided in this specification. Processor 202 may execute all steps included in the evaluation method. Processor 202 may be in the form of one or more processors. In some embodiments, processor 202 may include one or more hardware processors, such as microcontrollers, microprocessors, reduced instruction set computers (RISC), application-specific integrated circuits (ASICs), application-specific instruction set processors (ASIPs), central processing units (CPUs), graphics processing units (GPUs), physical processing units (PPUs), microcontroller units, digital signal processors (DSPs), field-programmable gate arrays (FPGAs), advanced RISC machines (ARMs), programmable logic devices (PLDs), any circuit or processor capable of performing one or more functions, or any combination thereof.

[0063] For illustrative purposes only, only one processor 202 is shown in the accompanying drawings of the evaluation system 200. However, it should be noted that the evaluation system 200 in this specification may also include multiple processors. Therefore, the operation and / or method steps disclosed in this specification may be executed by one processor or by multiple processors in combination. For example, if the processor 202 of the evaluation system 200 described in this specification executes steps A and B, it should be understood that steps A and B may also be executed jointly or separately by two different processors 202 (e.g., the first processor executes step A, the second processor executes step B, or the first and second processors jointly execute steps A and B).

[0064] Please see Figure 3 , Figure 3 This is a flowchart illustrating a user experience evaluation method provided in one embodiment of this specification. The evaluation method can be applied to evaluating user experience with AI products (see the examples above for details, which will not be repeated here). Furthermore, Figure 3 The entity implementing the evaluation method shown can be an evaluation system. For a description of the evaluation system, please refer to the example above; it will not be repeated here.

[0065] like Figure 3 As shown, the method includes the following steps S301 to S303:

[0066] S301: Obtain interaction data of the user's interaction process with the AI ​​product.

[0067] Interaction data can be understood as the content generated during the interaction between a user and an AI product. From the user's perspective, interaction data includes, but is not limited to, user input (such as text, voice, etc.), click behavior, dwell time, and browsing path. From the AI ​​product's perspective, interaction data includes, but is not limited to, the AI ​​product's response content.

[0068] Based on the above examples, let's take an AI product as the AI ​​dialogue device, specifically a financial management robot applied in the financial services field:

[0069] Users can ask a financial management robot to recommend financial products. The robot can inquire about the user's needs and preferences and recommend products based on the user's answers. Afterwards, the user may ask the robot more detailed and in-depth questions about the recommendations, which the robot can answer one by one.

[0070] The question-and-answer process between a user and a financial management chatbot can be considered an interactive process. The dialogue data generated during this interaction, such as the user's questions and the chatbot's answers, can be called interactive data.

[0071] In addition, interaction data can also include relatively indirect user data, such as the time users spend on the chat page with the financial management robot; or, for example, user click behavior and browsing behavior (such as duration) of links and images recommended by the financial management robot.

[0072] Understandably, to avoid tedious explanations, this manual will not repeat identical or similar content. For example, this manual primarily uses AI products, specifically financial management robots, as an illustrative example. For other forms of AI products and their applications in other fields, please refer to the examples corresponding to financial management robots; these will not be listed here.

[0073] In some embodiments, interaction data may also be referred to as user input samples. After obtaining user input samples, the evaluation system can first clean the user input samples to ensure that the cleaned data has higher accuracy and reliability.

[0074] This embodiment does not limit the cleaning method, which can be determined by the evaluation system based on requirements, historical records, experiments, etc. Examples include denoising, standardization, and screening. Denoising can be understood as removing irrelevant or interfering information. Standardization can be understood as unifying text format, such as case conversion and punctuation processing. Screening can be understood as filtering out samples that do not meet the requirements according to certain rules or conditions.

[0075] S302: Analyze user interaction emotions during the interaction process based on interaction data, wherein interaction emotions include positive feedback emotions and / or negative feedback emotions.

[0076] Interactive sentiment can be understood as the emotional information of users during the interaction process, as reflected in interaction data. Interactive sentiment reflects the user's experience and feelings during the interaction. It can be positive feedback (such as satisfaction and pleasure) or negative feedback (such as dissatisfaction and confusion).

[0077] Continuing with the example above, regarding the feedback from the financial management robot, if the user replies "Oh, I see, I understand!", then the evaluation system can relatively identify this as positive feedback from the user.

[0078] Conversely, if a user replies with "Really? That doesn't seem right. I still don't understand," then the evaluation system can determine the user's negative feedback sentiment.

[0079] S303: Determine the user's evaluation results of the AI ​​product experience based on interactive emotions.

[0080] The evaluation results can be understood as an assessment of the overall user experience with AI products based on the analyzed emotional responses to those interactions. These results can be used to guide improvements to AI products or optimization of services.

[0081] Continuing with the above example, the evaluation system can assess a user's overall experience of using a financial robot to recommend financial products based on the user's positive and / or negative feedback emotions.

[0082] Furthermore, the evaluation system can also improve or optimize the financial management robot based on the evaluation results. For example, for negative feedback emotions in the interaction, the evaluation system can make targeted improvements to the financial management robot. Conversely, for positive feedback emotions in the interaction, the evaluation system can strengthen the financial management robot.

[0083] Based on the above analysis of S301 to S303, it can be seen that in this embodiment, by taking a user-centric approach and determining the user's emotional engagement during the interaction with the AI ​​product based on the interaction data generated during the interaction, the user's evaluation of the AI ​​product's experience is determined. This effectively aligns with actual user application scenarios and user experiences, and more effectively covers the complex and personalized characteristics of users. It avoids the technical problems mentioned in the related technologies. It can improve the effectiveness and reliability of the evaluation, such as enabling an effective and reliable assessment of whether the user experience is "good" or "bad".

[0084] In some embodiments, S302 may include the following steps 11 and 12:

[0085] Step 11: Determine the user's interaction feedback signals during the interaction process based on the interaction data, wherein the interaction feedback signals include emotional feedback signals and / or behavioral feedback signals.

[0086] For example, the evaluation system extracts key information from the interaction data, identifying emotional and behavioral feedback signals that reflect the user's emotional state. For instance, the evaluation system can utilize natural language processing techniques to parse the content of the interaction data and can combine this with the user's behavioral logs for comprehensive analysis.

[0087] The evaluation system can combine emotional and behavioral feedback signals to form a comprehensive assessment of the user's emotional state. For example, if a user uses positive language in a conversation but simultaneously exhibits a quick withdrawal behavior, the evaluation system may further analyze this contradiction to determine the true emotional state.

[0088] Interactive feedback signals can be understood as signals extracted from interactive data that reflect the user's feelings about the interactive process. These can include emotional feedback signals and / or behavioral feedback signals. This specification primarily uses emotional and behavioral feedback signals as examples to illustrate interactive feedback signals.

[0089] Emotional feedback signals can be understood as information that directly or indirectly reflects a user's emotional state. This includes information about a user's emotional state based on emotional vocabulary ("very satisfied", "unsatisfied") and tone intensity in interaction data.

[0090] For example, the emotional tone conveyed by users during interaction can reflect their emotional attitude and is a relatively intuitive indicator of their response and satisfaction with AI products. The emotional feedback signal can be determined based on factors such as "whether it occurs" and "the probability of its occurrence," but this embodiment does not limit the intensity of the emotion.

[0091] The emotional feedback signal can include positive emotional feedback signals and / or negative emotional feedback signals.

[0092] For example, regarding positive emotional feedback signals (primarily reflecting the user's acceptance sentiment):

[0093] When AI products respond well to user expectations, users typically exhibit positive emotions such as satisfaction, trust, pleasure, and gratitude. These positive emotions often appear when the task is clearly understood and progresses smoothly, directly reflecting whether the AI ​​product's response has been "accepted." Especially in multi-turn conversations, the consistent expression of positive emotions can serve as a key signal of a stable and improving user experience and a willingness to collaborate.

[0094] For negative emotional feedback signals (mainly reflecting user resistance):

[0095] When a user's intent cannot be clearly interpreted by an AI product, and the user enters a fatigue cycle of "clarification-misunderstanding-rewriting," negative emotions such as confusion, anger, disappointment, denial, dissatisfaction, and repeated clarifications may occur. Emotional tone is a direct signal to measure whether the user has a poor experience during the interaction process, especially in multi-turn dialogues.

[0096] In some embodiments, the evaluation system may utilize sentiment analysis / emotion recognition models to analyze the emotions conveyed by the user's text and audio / video (speech and / or video). The polarity and intensity of the output emotions are combined to determine positive and negative emotional feedback signals.

[0097] For example, emotional feedback signals may include signals fused from text-based and audio / video emotional feedback signals using a multimodal approach. This allows for the determination of a user's emotional state during interaction with an AI product from multiple dimensions, including text, audio, and video. This improves the accuracy and reliability of evaluation results based on emotional feedback signals.

[0098] At the text level, the evaluation system can use emotion recognition models such as BERT and RoBERTa to classify the text content into emotions, thereby identifying positive and negative emotional feedback signals. At the audio and video level, the evaluation system can use audio and video emotion models to perform multimodal fusion by analyzing tone of voice, speech rate, etc., or by using computer vision expression analysis models, to identify positive and negative emotional feedback signals.

[0099] Furthermore, based on the above analysis, both positive and negative emotions can encompass different emotion types. For example, positive emotions include satisfaction, trust, joy, and gratitude. Negative emotions include denial, confusion, impatience, and fatigue.

[0100] Correspondingly, for positive emotions, the assessment system can determine the positive emotion feedback signal based on the probability / confidence value of each type of positive emotion appearing in the interaction data (specifically, it can be represented by a scoring method). For negative emotions, the assessment system can determine the negative emotion feedback signal based on the probability / confidence value of each type of negative emotion appearing in the interaction data (specifically, it can be represented by a scoring method).

[0101] Behavioral feedback signals can be understood as emotional tendencies inferred from user behavior patterns. For example, frequent page switching may indicate confusion or dissatisfaction; a user staying on a particular page for a long time may indicate strong interest, and so on.

[0102] In the interaction between users and AI products, behavioral feedback signals can reflect whether users are willing to collaborate with the AI ​​product to advance tasks, and whether they demonstrate compliance with Standard Operating Procedures (SOPs) and a results-oriented approach during the interaction. Users can convey their level of satisfaction through specific operational behaviors. Therefore, behavioral feedback signals can be analyzed from the perspectives of willingness to collaborate, SOP implementation, and action results.

[0103] Therefore, in some embodiments, the behavioral feedback signal may include at least one of: collaboration willingness feedback signal, process progress (i.e., SOP progress) feedback signal, and behavioral result feedback signal. This allows for the determination of behavioral feedback signals from multiple dimensions, enabling them to accurately, effectively, and reliably represent the user's experience. This, in turn, improves the accuracy and reliability of the evaluation.

[0104] The collaboration willingness feedback signal can be understood as whether the user, during the interaction with the AI ​​product, demonstrates a willingness and attitude to jointly advance the task. This includes proactively engaging in deeper interaction, asking follow-up questions, continuously inquiring, expanding on the current topic, and being willing to cooperate in clarifying or proactively supplementing more information. In other words, the collaboration willingness feedback signal reflects the degree of cooperation between the user and the AI ​​product, indicating whether they are in a collaborative state.

[0105] Collaboration willingness feedback signals can include positive and negative signals. Positive collaboration willingness feedback signals indicate that the user is willing to continue collaborating. Negative collaboration willingness feedback signals indicate that the user ignores or refuses to collaborate.

[0106] Positive feedback signals of willingness to collaborate also have corresponding collaboration types, such as proactively asking follow-up questions or explicitly expressing interest. Negative feedback signals of willingness to collaborate also have corresponding collaboration types, such as no response or cold treatment, impatience or denial, or explicitly terminating the interaction.

[0107] Process progression feedback signals can be understood as whether, during the interaction with an AI product, the user continues to advance the content depth, enrich the information dimensions, or introduce new stages based on the current dialogue structure. It is a crucial indicator of whether the current interaction process is evolving in a clearer, more focused, and more goal-oriented direction. Process progression feedback signals reflect the user's acceptance and depth of understanding of the AI ​​product's process guidance, emphasizing structural continuity and interactive progression.

[0108] Process progress feedback signals can include positive and negative signals. Positive feedback signals indicate that the user is cooperating with the process. Negative feedback signals indicate that the user has questions about the process, causing the Standard Operating Procedure (SOP) to stall or be interrupted.

[0109] Positive process progress feedback signals also have corresponding progress types, such as information increments, requests for expanded content, demands for in-depth details, and the introduction of the next stage's intent. Negative process progress feedback signals also have corresponding progress types, such as clarifying intent, controlling quality, correcting the AI ​​product's understanding, regeneration and stalling, and stagnation.

[0110] Behavioral outcome feedback signals can be understood as whether, during the interaction between a user and an AI product, the user demonstrates a clear intention to perform an action, expresses a willingness to take a certain action, or has already completed an action. Emphasizing outcome-oriented behavior is the final judgment of behavioral signals and a crucial indicator of whether the interaction generates tangible value. Behavioral outcome feedback signals include the user's verbal expression of intent and actual actions such as clicking / jumping / copying / favoriting / placing in a purchase / adding to a favorites list, etc.

[0111] Behavioral outcome feedback signals can include positive and negative signals. Positive signals indicate that the user has clearly demonstrated an intention to act or has actually taken action. Negative signals indicate that the user has not taken any action or has explicitly denied the possibility of taking any action.

[0112] Positive behavioral feedback signals also have outcome types, such as execution intention, copying and saving, sharing and forwarding, verification attempts, and adding to favorites / selections. Negative behavioral feedback signals also have outcome types, such as explicitly denying behavioral intention, behavioral hesitation, or procrastination.

[0113] In some embodiments, the evaluation system can analyze interaction data through dialogue sequence modeling to analyze user interaction behavior. For example, the evaluation system can use recurrent neural networks (LSTM / GRU) or Transformer dialogue models to learn user speaking and response patterns; and calculate information such as user response speed and question ending features to assess engagement. Alternatively, user dialogue sequences can be encoded as vectors, and algorithms such as cosine similarity or dynamic time warping (DTW) can be used to compare the current session with typical progression / stagnation scenarios to identify positive and negative behavioral feedback signals.

[0114] In addition, in some embodiments, the evaluation system can also determine the behavioral feedback signal (including positive behavioral feedback signal and negative behavioral feedback signal, and can be specifically represented by a scoring method) from three dimensions: collaboration willingness feedback signal, process progress feedback signal, and behavioral result feedback signal.

[0115] In some embodiments, step 11 above may include the following sub-step 111 and sub-step 112:

[0116] Sub-step 111: Identify the interaction data to obtain emotional data representing user emotions and behavioral data representing user behavior.

[0117] Emotional data can be understood as data identified from interaction data that reflects a user's emotional state. For example, emotional data can come from a user's language expressions (such as the vocabulary used and tone of voice) or voice characteristics (such as changes in intonation).

[0118] Behavioral data can be understood as data identified from interaction data that reflects user behavior patterns. Behavioral data can include data on how users interact with AI products (such as clicking, scrolling, and dwell time), browsing paths, and usage frequency.

[0119] Sub-step 112: Determine emotional feedback signals based on emotional data, and determine behavioral feedback signals based on behavioral data.

[0120] For example, based on the above analysis, the evaluation system can determine the type of emotional data (such as satisfaction in positive emotions, negation in negative emotions, etc.) and determine the emotional feedback signal (such as positive emotional feedback signal, negative emotional feedback signal) based on the type of emotional data.

[0121] The assessment system can first determine the categories of behavioral data, such as collaboration willingness, process advancement, and behavioral outcome; then determine the types of behavioral data under each category (such as proactive questioning or follow-up questioning under the collaboration willingness category, information increment under the process advancement category, and execution intention under the behavioral outcome category); finally, based on the types of behavioral data under each category, determine the emotional feedback signals (positive collaboration willingness feedback signal, negative collaboration willingness feedback signal, positive process advancement feedback signal, negative process advancement feedback signal, positive behavioral outcome feedback signal, and negative behavioral outcome feedback signal).

[0122] Based on the above analysis of sub-steps 111 and 112, it can be seen that in this embodiment, the evaluation system can more comprehensively and accurately capture the user's actual emotional state by identifying emotional data and behavioral data from the interaction data, that is, by performing dual analysis of emotional data and behavioral data, avoiding the bias caused by single-dimensional analysis, and improving the accuracy and reliability of the determined emotional feedback signals and behavioral feedback signals.

[0123] Step 12: Determine the interactive emotion based on the interactive feedback signals.

[0124] For example, the assessment system can determine the emotional state of an interaction based on emotional feedback signals and behavioral feedback signals.

[0125] Based on the above analysis of steps 11 and 12, it can be seen that in this embodiment, the evaluation system analyzes the emotional and behavioral feedback signals reflected by the user based on interaction data to determine the user's interactive emotions from both emotional and behavioral dimensions. This can improve the accuracy and reliability of the determined interactive emotions.

[0126] To facilitate the reader's understanding of the technical solutions provided in this specification, the following is combined with... Figure 4 The evaluation methods provided in this manual are described in more detail. Figure 4 This is a flowchart illustrating a user experience evaluation method provided in another embodiment of this specification. Figure 4 As shown, the method includes the following steps S401 to S407:

[0127] S401: Obtain interaction data of the user's interaction process with the AI ​​product.

[0128] Similarly, this embodiment will not repeat the same or similar technical features as those in the examples above. For example, regarding the implementation principle of S401, please refer to the description of S301 in the examples above.

[0129] S402: Map the interactive data to the initial prompt.

[0130] The initial prompt can be understood as a preliminary prompt form transformed from interactive data that the AI ​​product can comprehend.

[0131] For example, combining Figure 5 It can be seen that after obtaining user input samples, the user input samples can be cleaned and the cleaned user input samples can be mapped to the initial prompt.

[0132] S403: Filter the initial prompt based on preset quality assessment dimensions to obtain a target prompt that meets preset quality requirements.

[0133] Preset quality assessment dimensions can be understood as standards or criteria used to evaluate the quality of the initial prompt, which may include one or more aspects such as clarity, structure, and adaptability.

[0134] Clarity can be understood as determining the clarity of the initial prompt by examining whether there are ambiguities such as unclear references, ambiguous meanings, or blurred boundaries. It is primarily used to determine whether the initial prompt clearly conveys the intent.

[0135] Correspondingly, the evaluation system filters the initial prompts based on clarity, eliminating those that do not clearly convey the intent (such as those that do not meet the clarity quality requirements), and identifying the prompts that clearly convey the intent as the target prompts.

[0136] Structure can be understood as determining the structure of the initial prompt based on factors such as whether it is complete and whether it has a logical hierarchy. It is primarily used to determine the logic and organization of the initial prompt.

[0137] Correspondingly, the evaluation system filters the initial prompts based on their structure, eliminating those that lack logic and organization (e.g., prompts that do not meet the quality requirements for logic and organization), and identifying those that do have logic and organization as the target prompts.

[0138] Adaptability can be understood as determining the clarity of the initial prompt based on whether it aligns with task requirements and possesses executability (such as instruction closure and specific objectives). For example, it is primarily used to determine whether the initial prompt is suitable for the current scenario or the corresponding user.

[0139] Correspondingly, the evaluation system filters the initial prompt based on its suitability, which can filter out prompts that are not suitable for the current scenario or the corresponding user (such as prompts that do not meet the quality requirements of the current scenario or the corresponding user), so as to determine the prompts that are suitable for the current scenario or the corresponding user as the target prompts.

[0140] A target prompt can be understood as the final version of the prompt after filtering and optimization. For example, it's a prompt that meets certain quality requirements after filtering the initial prompt based on preset quality assessment dimensions. Relatively speaking, a target prompt satisfies the preset quality requirements and is more suitable for subsequent sentiment analysis or other processing.

[0141] Similarly, the preset quality requirements can be determined by the evaluation system based on requirements, historical records, experiments, etc., and this embodiment does not impose any limitations.

[0142] Based on the above analysis, it can be seen that the evaluation system can filter the initial prompt from one dimension of clarity, structure, and adaptability, or it can filter the initial prompt from multiple dimensions.

[0143] If the initial prompt is filtered from multiple dimensions, the evaluation system can score the initial prompt from multiple dimensions based on a pre-built prompt parsing model to determine whether to retain the initial prompt or filter it out.

[0144] For example, each dimension has its own corresponding weight (similarly, this embodiment does not limit the specific weight size). The prompt parsing model can score the initial prompt from each dimension based on each weight. Relatively speaking, if the score is high (such as reaching the preset score), it indicates that the quality of the initial prompt is high, and the evaluation system determines the initial prompt as the target prompt; conversely, if the score is low (such as below the preset score), it indicates that the quality of the initial prompt is low, and the evaluation system filters out the initial prompt.

[0145] For example, the evaluation system can score the initial prompt based on Equation 1, Equation 1:

[0146] Prompt Score =(α×Clarity+β×Structure+γ×FitForTask)×Adjust(t)

[0147] Wherein, Prompt_Score is the initial prompt score, α, β, and γ are the corresponding weights, Clarity is the clarity, Structure is the structure, FitForTask is the fit, and Clarity, Structure, and FitForTask ∈ [0, 1]. Adjust(t) is a dynamic time factor that reflects whether subsequent rounds compensate for the initial ambiguity (the initialization process is used to appropriately reduce the weight of the evaluation bear in the optimization learning).

[0148] It should be understood that Equation 1 only describes the three aspects of clarity, structure, and adaptability. It is a possible way for the evaluation system to score the initial prompt, but it should not be interpreted as a limitation on the way the evaluation system scores the initial prompt.

[0149] Furthermore, this embodiment does not limit the structure or training method of the prompt parsing model, as long as the quality of the prompt can be determined from at least one of clarity, structure, and adaptability.

[0150] Continuing with the examples above and Figure 5 S403 can be understood as a process of filtering initial prompts from a quality perspective. Initial prompts that pass the filter (e.g., meet preset quality requirements) are identified as target prompts, and subsequent operations such as S404 are performed. Initial prompts that fail the filter (e.g., do not meet preset quality requirements) are filtered out to prevent them from entering subsequent operations such as S404.

[0151] S404: Analyze the user's interactive emotions during the interaction process based on the target prompt, where interactive emotions include positive feedback emotions and / or negative feedback emotions.

[0152] Correspondingly, after obtaining a relatively high-quality target prompt, the evaluation system can perform interactive sentiment analysis based on the target prompt.

[0153] Continuing with the examples above and Figure 5 The evaluation system analyzes the emotional impact of interactions based on the selected target prompts. This analysis can include both positive and negative feedback emotional impact.

[0154] In addition, regarding the implementation principle of the evaluation system's analysis of user interaction emotions based on the target prompt during the interaction process, please refer to the description of the analysis of user interaction emotions based on interaction data in the example above, which will not be repeated here.

[0155] Based on the above analysis of S402 to S404, it can be seen that in this embodiment, the evaluation system filters the initial prompt from at least one of preset quality evaluation dimensions (such as clarity, structure, and adaptability) to obtain a target prompt that meets preset quality requirements. Interactive sentiment analysis is then performed based on this relatively high-quality target prompt. This improves the accuracy and reliability of interactive sentiment analysis; it also avoids interaction failures due to prompt issues (such as avoiding misjudging the AI ​​product's response as an interaction failure due to unclear user intent or distorted expression); it provides clues to the causes of anomalies in the evaluation results, serving as a source variable for experience analysis; and it allows for "controllable early warning" intervention, improving the AI ​​product's fault tolerance strategy and reducing the occurrence of negative user experiences.

[0156] S405: Determine the user's interaction experience level based on the interaction sentiment, wherein the interaction experience level includes the probability of positive interaction experience and / or the probability of negative interaction experience.

[0157] Interaction experience level can be understood as a measure of the probability of a user's overall experience during interaction with an AI product. For example, interaction experience level can reflect the probability of a user's overall satisfaction or dissatisfaction with the interaction process.

[0158] Positive Interaction Probability (PIP) can be understood as the probability that a user will exhibit positive emotions such as positivity and satisfaction during an interaction, indicating a good user experience.

[0159] Negative Interaction Probability (NIP) can be understood as the probability that a user will exhibit negative emotions such as dissatisfaction or confusion during an interaction, indicating that there are problems or obstacles in the user experience.

[0160] Continuing with the examples above and Figure 5 The analysis of positive feedback sentiment can include positive feedback analysis of the probability of positive interactive experiences (such as...). Figure 5 The PIP positive feedback analysis shown here. Analysis of negative feedback sentiment can include negative feedback analysis of the probability of a negative interaction experience (e.g., ...). Figure 5 (NIP negative feedback analysis shown in the figure).

[0161] In other words, in this embodiment, the evaluation system focuses more on the probabilistic assessment of the AI ​​product's impact on user experience, and there is no absolute correlation between the probabilities of positive and negative interaction experiences. For example, a user may have positive feedback on some aspects of the AI ​​product and negative feedback on others, resulting in a high probability of both positive and negative interaction experiences.

[0162] In other embodiments, the probability of a negative interaction experience can also be determined based on negative feedback emotion, combined with factors such as the nature of the target prompt, the frequency of user guidance during the interaction, and the frequency of corrective actions.

[0163] The nature of the target prompt can be understood as some key characteristics or attributes exhibited by the target prompt during the interaction between the user and the AI ​​product. These characteristics determine whether the target prompt can effectively guide the AI ​​product to generate high-quality, expected output.

[0164] In other words, the properties of the target prompt describe what characteristics the target prompt should possess in order for AI products to better understand and execute the corresponding tasks.

[0165] S406: Determine the evaluation results based on the interactive experience level.

[0166] For example, the evaluation system can combine the probability of positive interaction experience and the probability of negative interaction experience to determine the evaluation result.

[0167] In some embodiments, the evaluation system can determine the interaction experience level (probability of positive interaction experience, probability of negative interaction experience) as the evaluation result.

[0168] In other embodiments, based on the above analysis, the probability of a positive interactive experience can be determined based on positive emotional feedback signals and positive behavioral feedback signals. Furthermore, positive emotional feedback signals and positive behavioral feedback signals can be represented by scores.

[0169] For example, the evaluation system can score positive emotion feedback signals based on Equation 2 to obtain the corresponding score Emotion_PIP_Score, Equation 2:

[0170] Emotion_PIP_Score=∑(P i ×w i ), i∈{Satisfaction, Trust, Pleasure, Gratitude}

[0171] Among them, P i Let w represent the probability of emotion i, where Satisfaction represents satisfaction, Satisfaction represents trust, Pleasure represents joy, and Gratitude represents gratitude. i Let i be the weight of emotion i.

[0172] Similarly, Equation 2 only describes the emotional types of satisfaction, trust, pleasure, and gratitude. It may be a way for the evaluation system to score positive emotional feedback signals, but it should not be interpreted as a limitation on how the evaluation system scores positive emotional feedback signals.

[0173] The evaluation system can score the positive behavior feedback signal based on Equation 3 to obtain the corresponding score Behavior_PIP_Score, Equation 3:

[0174] Behavior_PIP_Score=(w1×Coop_PIP+w2×SOP_PIP+w3×SOP_PIP)

[0175] Any of Coop_PIP, SOP_PIP, and Action_PIP can be represented by ∑(fi×p). i ) indicates, such as Coop PIP =∑(fi×p i fi represents whether the corresponding behavior occurs (can be 0 or 1, 1 if it occurs), p i For the corresponding behavioral confidence model, Coop_PIP is at least one of the following: positive collaboration intention behavior, SOP_PIP is positive process advancement behavior, and SOP_PIP is positive behavior outcome behavior, and w1, w2, and w3 are the corresponding weights.

[0176] Similarly, Equation 3 only describes the types of behaviors, such as positive collaborative intention behavior, positive process advancement behavior, and positive behavior result behavior, and it is a possible way for the evaluation system to score positive behavior feedback signals. It should not be interpreted as a limitation on the way the evaluation system scores positive behavior feedback signals.

[0177] In some embodiments, the evaluation system can determine the probability of a positive interaction experience based on positive emotional feedback signals and / or positive behavioral feedback signals.

[0178] For example, the probability of a positive interactive experience can also be represented by a score. Positive emotional feedback signals and positive behavioral feedback signals each have their own corresponding weights. By weighting the weights and scores of the positive emotional and behavioral feedback signals, the probability of a positive interactive experience can be obtained as a score.

[0179] Similarly, the probability of a positive interactive experience can be determined based on positive emotional feedback signals and positive behavioral feedback signals. Furthermore, positive emotional feedback signals and positive behavioral feedback signals can be represented by scores.

[0180] For example, the evaluation system can determine the score PIP_score corresponding to the probability of a positive interactive experience based on Equation 4, Equation 4:

[0181] PIP_score=W1×Emotion_PIP_Score+W2×Behavior_PIP_Score

[0182] Where W1 and W2 are the corresponding weights, Emotion_PIP_Score is the score corresponding to the positive emotion feedback signal, and Behavior_PIP_Score is the score corresponding to the positive behavior feedback signal.

[0183] Similarly, the evaluation system can determine the probability of a negative interactive experience based on negative emotional feedback signals and / or negative behavioral feedback signals.

[0184] For example, the probability of a negative interactive experience can also be represented by a score. Negative emotional feedback signals and negative behavioral feedback signals each have their own corresponding weights. By weighting the weights and scores of the negative emotional and behavioral feedback signals, the probability of a negative interactive experience can be obtained in the form of a score.

[0185] For example, the evaluation system can score negative emotional feedback signals based on Equation 5 to obtain the corresponding score Emotion_NIP_Score, Equation 5:

[0186] Emotion_NIP_Score=∑(P×w), j∈{Disapproval, Confusion, Impatience, Fatigue, Disappointment, Anger, Sarcastic}

[0187] Where P is the probability of emotion j, w is the weight of emotion j, Disapproval is negation, Disapproval is confusion, Impatience is impatience, Fatigue is fatigue, Disappointment is disappointment / speechlessness, Anger is anger, and Sarcastic is sarcasm.

[0188] Similarly, Equation 5 only describes the emotional types of negativity, confusion, impatience, fatigue, disappointment / speechlessness, anger, and sarcasm. It is a way for the evaluation system to score negative emotional feedback signals, but it should not be interpreted as a limitation on how the evaluation system scores negative emotional feedback signals.

[0189] In some embodiments, when the score corresponding to the negative emotion feedback signal is low (e.g., below a certain preset threshold), that is, when the user's negative emotions are too high, the evaluation system or other systems (such as AI products) can insert "clarification guidance" or "proactive prompts" to avoid the user experience "collapsing".

[0190] The evaluation system can score the negative behavior feedback signal based on Equation 6 to obtain the Behavior_NIP_Score corresponding to the negative behavior feedback signal, Equation 6:

[0191] Behavior_PIP_Score=(w1×Coop_NIP+w2×SOP_NIP+w3×SOP_NIP)

[0192] Any one of Coop_NIP, SOP_NIP, and Action_NIP can be represented by ∑(fi×p). i ) indicates, such as Coop NIP =∑(f i ×p i ), f i To indicate whether the corresponding behavior has occurred (can be 0 or 1, 1 for occurrence), p i For the corresponding behavioral confidence model, Coop_NIP represents at least one of the following: negative collaboration intention behavior, SOP_PIP represents negative process advancement behavior, and SOP_PIP represents negative behavioral outcome behavior. w1, w2, and w3 are the corresponding weights.

[0193] Similarly, Equation 5 only describes the types of behaviors such as negative collaborative intention behavior, negative process advancement behavior, and negative behavior outcome behavior. It is a possible way for the evaluation system to score negative behavior feedback signals, but it should not be interpreted as a limitation on the way the evaluation system scores negative behavior feedback signals.

[0194] The evaluation system can determine the score NIP_score corresponding to the probability of a positive interactive experience based on Equation 7, Equation 7:

[0195] NIP_score=W1′×Emotion_NIP_Score+W2′×Behavior_NIP_Score

[0196] Where W1′ and W2′ are the corresponding weights, Emotion_NIP_Score is the score corresponding to the negative emotion feedback signal, and Behavior_NIP_Score is the score corresponding to the negative behavior feedback signal.

[0197] In some other embodiments, the evaluation system may combine the probability of positive interaction experience and the probability of negative interaction experience to determine the evaluation result.

[0198] For example, the evaluation system will fit and sum the probabilities of positive and negative interactive experiences, which are represented by scores, to obtain the evaluation result.

[0199] In some other embodiments, the evaluation system can also combine the target prompt score to determine the experience score corresponding to the evaluation result.

[0200] For example, the Interaction Experience Score corresponding to the evaluation result can be represented by Equation 8, Equation 8:

[0201] (Interaction Experience Score)=W1×PromptScore+W2×Emotion_PIP_Score+W3×Behavior_PIP_Score-(W4×Emotion_NIP_Score+W5×Behavior_NIP_Score)

[0202] Wherein, W1 to W5 are the corresponding dynamic weighting factors, PromptScore is the score corresponding to the initial prompt, Emotion_PIP_Score is the score corresponding to the positive emotion feedback signal, Behavior_PIP_Score is the score corresponding to the positive behavior feedback signal, Emotion_NIP_Score is the score corresponding to the negative emotion feedback signal, and Behavior_NIP_Score is the score corresponding to the negative behavior feedback signal.

[0203] Based on the above analysis of S405 and S406, it can be seen that in this embodiment, the evaluation system can more accurately evaluate the user's interactive experience through detailed sentiment analysis and experience level calculation, thereby improving the credibility of the evaluation results.

[0204] S407: Optimize AI products based on the evaluation results. Optimization includes improving the product experience (if necessary) and optimizing the corresponding models for the AI ​​products.

[0205] Continuing with the examples above and Figure 5 After obtaining the evaluation results, the evaluation system (or other systems or staff, etc.) can use the evaluation results to make feedback and optimizations.

[0206] For example, the model corresponding to the AI ​​product can be optimized (such as...). Figure 5 The model optimization shown can also be used to optimize the user experience of AI products (such as...). Figure 5 (See the product experience optimization shown).

[0207] Model optimization can include model correction and model enhancement. For example, based on the above analysis, the evaluation system can correct the parameters of the model corresponding to the AI ​​product based on users' negative interaction feelings, such as negative feedback sentiment and the probability of negative interaction experience, in order to minimize users' negative interaction feelings with the AI ​​product.

[0208] The evaluation system can enhance the parameters of the model corresponding to the AI ​​product based on the user's positive interaction experience, such as positive feedback emotion and positive interaction experience probability, for example, through the activation function (or reward function), so as to maximize the user's positive interaction experience with the AI ​​product.

[0209] Product experience optimization can include optimizing communication language. For example, by optimizing the product experience, AI products can use more concise, polite (or playful, etc.) language to interact with users.

[0210] In some embodiments, the evaluation system can first trace the reasons for optimizing the AI ​​product based on the evaluation results, and then optimize based on those reasons. This ensures that the evaluation results are not only comprehensive and accurate, but also interpretable.

[0211] For example, based on the evaluation results, the evaluation system (or staff, etc.) can track the reasons why an AI product needs optimization. The evaluation system can output and display these reasons. Accordingly, staff can selectively optimize the AI ​​product based on the displayed reasons. Alternatively, staff can also provide feedback on the evaluation results (such as analysis of the reasons, other possible causes, etc.), and the evaluation system can combine the evaluation results with the staff's feedback to optimize the AI ​​product.

[0212] Based on the above analysis, it can be seen that in this embodiment, the evaluation system optimizes the AI ​​product by combining the evaluation results, which can continuously improve the AI ​​product and form a virtuous cycle.

[0213] In addition, combining the above analysis and Figure 5 As can be seen, the analysis of interactive sentiment (including NIP negative feedback analysis and PIP positive feedback analysis) is the core content of the technical solution provided in this specification.

[0214] In some embodiments, the analysis of interactive sentiment can be achieved through a network model.

[0215] For example, a sentiment analysis network model (or interactive experience evaluation model) can be pre-trained by an evaluation system or other system (such as a training system), and then the evaluation system performs interactive sentiment analysis based on the sentiment analysis network model, specifically including NIP negative feedback analysis and PIP positive feedback analysis.

[0216] For example, interactive sentiment is obtained by analyzing interactive data using a sentiment analysis network model; the sentiment analysis network model is obtained by training a basic network model based on a collected training dataset, which includes interactive data samples between sample users and AI products; the interactive data samples are used to enable the basic network model to learn the ability to analyze positive and / or negative feedback sentiment samples corresponding to the interactive data samples.

[0217] Similarly, this embodiment will not repeat the same or similar technical features as those in the examples above.

[0218] For example, regarding the understanding of interaction data samples, please refer to the description of interaction data in the above examples; regarding the understanding of positive feedback sentiment samples, please refer to the understanding of positive feedback sentiment in the above examples; regarding the understanding of positive-negative feedback sentiment samples, please refer to the understanding of negative feedback sentiment in the above examples, and so on.

[0219] Furthermore, this embodiment does not limit the structure of the basic network model. For the training principles of the basic network model, please refer to the analysis of the interaction data in the above example.

[0220] However, during the training process, taking the training system as an example, after the training system obtains the predicted values ​​of positive and negative interaction experience probabilities based on the basic network model, it can construct a loss function by combining the ground truth values ​​of positive and negative interaction experience probabilities. The parameters of the basic network model are then adjusted with the goal of minimizing the loss function, until the sentiment analysis network model is obtained.

[0221] It is worth noting that this embodiment does not limit the structure of the sentiment analysis network model, and the sentiment analysis network model is not limited to a single network model. The specific model can be determined based on historical records, (scenario) requirements, experiments, etc.

[0222] For example, a sentiment analysis network model can be a model that has both PIP (Positive Influence Perception) and NIP (Negative Influence Perception) capabilities. A sentiment analysis network model can also include: a network model for determining the probability of positive interactive experiences and a network model for determining the probability of negative interactive experiences.

[0223] In other embodiments, in addition to evaluating user experience from the perspective of user perception as described above, the evaluation system can also combine the perspective of AI products to jointly evaluate user experience.

[0224] For example, in some embodiments, the evaluation system can test the interactive performance of AI products based on preset indicators to obtain performance test results. These preset indicators include, but are not limited to, accuracy, progress rate, and efficiency. Accuracy may include professional accuracy in intent recognition, accuracy in language expression, etc. Progress rate may include multi-turn dialogue progress rate. Efficiency can be determined from dimensions such as time consumption.

[0225] Based on this, the evaluation system can combine performance test results with evaluation results determined based on interactive sentiment to determine the final evaluation result.

[0226] For example, for different scenarios, the evaluation system can assign corresponding weights to the performance detection results and evaluation results, and determine the final evaluation result through weighted summation.

[0227] It is worth noting that the models (such as the basic network model) mentioned in this manual can be Large Language Models (LLMs). In this manual, Large Language Models can also be simply referred to as Large Models. A Large Language Model is a natural language processing model based on deep learning techniques, typically with billions to hundreds of billions or even more parameters, possessing powerful language understanding and generation capabilities. Large Language Models can employ the Transformer architecture or its variants (such as GPT, BERT, etc.), which utilizes an attention mechanism to achieve global modeling of sequential data, efficiently handling long-distance dependencies and thus performing excellently in natural language tasks. Large Language Models learn the statistical features and semantic relationships of language through pre-training on large-scale corpora, giving them outstanding generalization capabilities. The core capabilities of Large Language Models include, but are not limited to: understanding contextual semantics, generating coherent and grammatically correct text, performing logical reasoning, and handling multi-task scenarios. Their usage typically includes two modes: direct inference and fine-tuning. In direct reasoning mode, users guide the large language model to generate specific outputs by designing prompts. Prompts can be task descriptions or instructions in text form, used to stimulate the large language model's semantic understanding and generation capabilities. In fine-tuning mode, the large language model is further trained on small-scale datasets within a specific domain to optimize its performance on specific tasks. The powerful generalization ability and flexibility of large language models make them an important tool in the field of artificial intelligence, providing efficient and accurate solutions for automated text generation and understanding.

[0228] In some embodiments, large language models can also understand and generate data from other modalities (such as visual and audio data). In this case, large language models can also be called multimodal large language models (MLLMs). MLLMs provide a richer and more natural interactive experience by integrating multiple types of input and output, such as text, images, and sound. The core advantage of MLLMs lies in their ability to process and understand information from different modalities and fuse this information to complete complex tasks. For example, MLLMs can analyze an image and generate descriptive text, or generate a corresponding image based on a text description. This cross-modal understanding and generation capability makes MLLMs widely applicable across multiple fields.

[0229] It should be noted that the key technologies of large language models can be found in the detailed description in the paper "A Survey of Large Language Models" (paper number: arXiv:2303.18223v16, published on March 11, 2025, public link: https: / / doi.org / 10.48550 / arXiv.2303.18223), and will not be repeated here.

[0230] It is worth noting that the above examples are merely illustrative of possible implementations of the evaluation method described in this specification, and should not be construed as limiting the implementation of the evaluation method described in this specification. For example, based on the above technical concept, some of the technical features described above can be combined to obtain new embodiments; new technical features can be added to the above examples to obtain new embodiments; some technical features can be removed from the above examples to obtain new embodiments; some technical features in the above examples can be replaced with other technical features; some technical features and their order in the above examples can be adjusted to obtain new embodiments, and so on, which will not be listed here.

[0231] Based on the above-described technical concept, this specification also provides a computer-readable non-transitory storage medium storing at least one instruction set, wherein when the at least one instruction set is executed by a processor, the steps of the evaluation method described in this specification are implemented.

[0232] In some possible implementations, various aspects of this specification can also be implemented as a program product comprising program code. When the program product is run on the evaluation system 200, the program code causes the evaluation system 200 to perform the steps of the evaluation method described in this specification. The program product for implementing the above method may employ a portable compact disc read-only memory (CD-ROM) containing program code and may run on the evaluation system 200. However, the program product of this specification is not limited thereto. In this specification, a readable storage medium may be any tangible medium containing or storing a program that may be used by or in conjunction with an instruction execution system. The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. The computer-readable storage medium may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable storage medium may also be any readable medium other than a readable storage medium that can send, propagate, or transmit programs for use by or in connection with an instruction execution system, apparatus, or device. Program code contained on a readable storage medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof. Program code for performing the operations described herein can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java and C++, and conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on evaluation system 200, partially on evaluation system 200, as a standalone software package, partially on evaluation system 200 and partially on a remote evaluation system, or entirely on remote evaluation system 200.

[0233] It should be noted that the collection, storage, use, processing, transmission, provision, and disclosure of user-related information (such as interactive data) involved in the technical solutions of this specification all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0234] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0235] In summary, after reading this detailed disclosure, those skilled in the art will understand that the foregoing detailed disclosure is presented by way of example only and is not restrictive. Although not explicitly stated herein, those skilled in the art will understand that this specification requires various reasonable changes, improvements, and modifications to the embodiments. These changes, improvements, and modifications are intended to be made by this specification and are within the spirit and scope of the exemplary embodiments described herein.

[0236] Furthermore, certain terms in this specification have been used to describe embodiments of this specification. For example, "an embodiment," "an embodiment," and / or "some embodiments" mean that a particular feature, structure, or characteristic described in connection with that embodiment may be included in at least one embodiment of this specification. Therefore, it is to be emphasized and understood that two or more references to "an embodiment" or "an embodiment" or "alternative embodiment" in various parts of this specification do not necessarily refer to the same embodiment. Moreover, specific features, structures, or characteristics may be suitably combined in one or more embodiments of this specification.

[0237] It should be understood that in the foregoing description of the embodiments in this specification, various features are combined in a single embodiment, drawing, or description for the purpose of simplifying the description and to aid in understanding a feature. However, this does not mean that the combination of these features is necessary, and those skilled in the art, upon reading this specification, may readily identify some of the devices as separate embodiments. That is, the embodiments in this specification can also be understood as an integration of multiple secondary embodiments. And the content of each secondary embodiment is valid even if it contains fewer than all the features of a single foregoing disclosed embodiment.

[0238] Every patent, patent application, publication of a patent application, and other material cited herein, such as articles, books, specifications, publications, documents, and literature (excluding any related historical examination documents), is referenced for all purposes relevant to this document, including in the specification and claims herein. However, in the event of any inconsistency or conflict between the descriptions, definitions, and / or terms used in the foregoing and those used herein, the descriptions, definitions, and / or terms used herein shall prevail.

[0239] Finally, it should be understood that the embodiments disclosed herein are illustrative of the principles of the embodiments described in this specification. Other modified embodiments are also within the scope of this specification. Therefore, the embodiments disclosed in this specification are merely examples and not limitations. Those skilled in the art can implement the applications described in this specification using alternative configurations based on the embodiments in this specification. Therefore, the embodiments in this specification are not limited to the embodiments precisely described in the applications.

Claims

1. A user experience evaluation method, said evaluation method being applied to the evaluation of users' experience with AI products, comprising: Obtain interaction data of the user's interaction process with the AI ​​product; The interaction data is used to analyze the user's emotional responses during the interaction process, wherein the emotional responses include positive feedback and / or negative feedback; and The evaluation result of the user's experience with the AI ​​product is determined based on the interactive emotions.

2. The method according to claim 1, wherein, The analysis of the user's emotional interaction during the interaction process based on the interaction data includes: Based on the interaction data, the user's interaction feedback signals during the interaction process are determined, wherein the interaction feedback signals include emotional feedback signals and / or behavioral feedback signals; and The interactive emotion is determined based on the interactive feedback signal.

3. The method according to claim 2, wherein, The behavioral feedback signals include at least one of the following: collaboration willingness feedback signals, process progress feedback signals, and behavioral result feedback signals.

4. The method according to claim 2, wherein, The emotion feedback signal includes a multimodal fusion signal based on text emotion feedback signal and audio / video emotion feedback signal.

5. The method according to claim 2, wherein, Determining the user's interaction feedback signals during the interaction process based on the interaction data includes: The interaction data is identified to obtain emotional data representing the user's emotions and behavioral data representing the user's behavior; and The emotional feedback signal is determined based on the emotional data, and the behavioral feedback signal is determined based on the behavioral data.

6. The method according to any one of claims 1 to 4, wherein, The step of determining the user's evaluation result of the AI ​​product based on the interactive emotion includes: The user's interaction experience level is determined based on the interactive sentiment, wherein the interaction experience level includes the probability of a positive interaction experience and / or the probability of a negative interaction experience; and The evaluation result is determined based on the level of interactive experience.

7. The method according to claim 6, wherein, The positive feedback emotion includes positive emotional feedback signals and / or positive behavioral feedback signals; the negative feedback emotion includes negative emotional feedback signals and / or negative behavioral feedback signals; determining the user's interaction experience level based on the interaction emotion includes: The probability of the positive interactive experience is determined based on the positive emotional feedback signal and / or the positive behavioral feedback signal; and The probability of the negative interactive experience is determined based on the negative emotional feedback signal and / or the negative behavioral feedback signal.

8. The method according to any one of claims 1 to 4, wherein, The method further includes: The AI ​​product is optimized based on the evaluation results, wherein the optimization includes optimizing the model corresponding to the AI ​​product.

9. The method according to any one of claims 1 to 4, wherein, The analysis of the user's emotional interaction during the interaction process based on the interaction data includes: Map the interactive data to an initial prompt; The initial prompt is filtered based on preset quality assessment dimensions to obtain a target prompt that meets preset quality requirements; and The interactive sentiment is analyzed based on the target prompt.

10. The method according to any one of claims 1 to 4, wherein, The interactive sentiment is obtained by analyzing the interactive data using a sentiment analysis network model; The sentiment analysis network model is obtained by training a basic network model based on a collected training dataset, which includes interaction data samples between sample users and the AI ​​product. The interaction data samples are used to enable the basic network model to learn the ability to analyze positive feedback sentiment samples and / or negative feedback sentiment samples corresponding to the interaction data samples.

11. The method according to any one of claims 1 to 4, wherein, The AI ​​product includes an AI dialogue product; the interaction data includes dialogue data between the user and the AI ​​dialogue product.

12. A user experience evaluation system, said evaluation system being applied to evaluating users' experience with AI products, comprising: At least one storage medium storing at least one instruction set for evaluating the user experience of using AI products; At least one processor is communicatively connected to the at least one storage medium, wherein when the at least one processor is running, it reads the at least one instruction set and executes the evaluation method as described in any one of claims 1 to 11 according to the instructions of the at least one instruction set.