Voice interaction device experience evaluation method and device, equipment and storage medium

By evaluating the interaction process of voice interaction devices from multiple dimensions, including consistency of voice content, consistency of form, and consistency of guidance, the problem of low accuracy of experience evaluation in existing technologies has been solved, and a more comprehensive improvement in user experience has been achieved.

CN116403568BActive Publication Date: 2026-01-09HEFEI IFLYTEK TOYCLOUD TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310421198.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-14
Publication Date
2026-01-09
Estimated Expiration
2043-04-14

AI Technical Summary

Technical Problem

In existing technologies, the accuracy of experience evaluation for voice interaction products is low, and a comprehensive evaluation cannot be conducted systematically, resulting in poor user experience.

Method used

By identifying the interaction data of the target voice interaction device, including voice data, appearance data, interaction guidance data, and user behavior data, the evaluation results of each node in the interaction process and the overall evaluation results are assessed. A multi-dimensional evaluation method is adopted, including the evaluation of voice content consistency, voice form consistency, voice guidance consistency, user response nodes, device response nodes, and device guidance nodes.

Benefits of technology

It enabled more systematic and accurate experience evaluation, improved the user experience of voice interaction devices, and identified and improved weak links through evaluation of the entire process and each node, thereby improving product performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116403568B_ABST
    Figure CN116403568B_ABST
Patent Text Reader

Abstract

The application provides a voice interaction device experience evaluation method, device, equipment and storage medium, and relates to the technical field of experience evaluation. The method comprises the following steps: determining a target voice interaction device to be evaluated and interaction data corresponding to the target voice interaction device; determining a plurality of experience evaluation results of the target voice interaction device in an interaction process based on the interaction data, wherein the plurality of experience evaluation results comprise node evaluation results of each interaction node in the interaction process and / or an overall evaluation result of the interaction process, and the interaction process comprises an interaction initiation node, a user response node, a device response node and a device guidance node; and determining an experience evaluation result of the target voice interaction device based on the plurality of experience evaluation results. The application realizes more comprehensive and systematic experience evaluation by performing experience evaluation on the whole interaction process and each interaction node, thereby improving the experience evaluation accuracy of the voice interaction device.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of experience evaluation, in particular to an experience evaluation method, device and equipment of a voice interaction device and a storage medium. BACKGROUND

[0002] With the rapid development of technology, there are more and more voice interaction products. At present, the voice interaction products are mostly multi-round interaction, which can interact with the user for multiple times after being woken up once without the need of being woken up again. In order to improve the product performance and enhance the product experience, the voice interaction products need to be evaluated.

[0003] At present, the user experience effect of the voice interaction product is usually evaluated by some single-point indicators, such as evaluating the voice recognition accuracy of the voice interaction product or evaluating the answer accuracy of the voice interaction product. However, the voice interaction product cannot be evaluated systematically by the single-point indicators, which reduces the accuracy of the experience evaluation of the voice interaction product. SUMMARY

[0004] The present application provides an experience evaluation method, device and equipment of a voice interaction device and a storage medium, which solves the defect of low accuracy of the experience evaluation of the voice interaction product in the prior art and realizes a systematic and accurate experience evaluation method.

[0005] The present application provides an experience evaluation method of a voice interaction device, which comprises:

[0006] determining a target voice interaction device to be evaluated and interaction data corresponding to the target voice interaction device;

[0007] determining a plurality of experience evaluation results of the target voice interaction device in an interaction process based on the interaction data, wherein the plurality of experience evaluation results comprise node evaluation results of each interaction node in the interaction process and / or an overall evaluation result of the interaction process, and the interaction process comprises an interaction initiation node, a user response node, a device response node and a device guidance node;

[0008] determining an experience evaluation result of the target voice interaction device based on the plurality of experience evaluation results.

[0009] According to the experience evaluation method of the voice interaction device provided by the present application, the interaction data comprises voice data of the target voice interaction device, and the overall evaluation result comprises a voice content consistency evaluation result.

[0010] The voice content consistency evaluation result is determined based on the following steps:

[0011] determine a first voice content of first voice data in the voice data and a second voice content of second voice data in the voice data, the first voice data and the second voice data being data corresponding to a same interaction process;

[0012] determine the voice content consistency evaluation result based on a content consistency evaluation result of the first voice content and the second voice content.

[0013] According to the voice interaction device experience evaluation method provided by the application, the interaction data includes voice data of the target voice interaction device, and the overall evaluation result includes a voice form consistency evaluation result.

[0014] The voice form consistency evaluation result is determined based on the following steps:

[0015] determine a third voice content and a tone feature of third voice data in the voice data;

[0016] determine the voice form consistency evaluation result based on a form consistency evaluation result of the third voice content and the tone feature;

[0017] Alternatively, the interaction data further includes appearance form data of the target voice interaction device, and the voice form consistency evaluation result is determined based on the following steps:

[0018] determine a fourth voice content of fourth voice data in the voice data and a first emotional feature of first appearance form data in the appearance form data, the fourth voice data and the first appearance form data being data corresponding to a same interaction node;

[0019] determine the voice form consistency evaluation result based on a form consistency evaluation result of the fourth voice content and the first emotional feature.

[0020] According to the voice interaction device experience evaluation method provided by the application, the interaction data includes voice data of the target voice interaction device and interaction guide data of the target voice interaction device, and the overall evaluation result includes a voice guide consistency evaluation result.

[0021] The voice guide consistency evaluation result is determined based on the following steps:

[0022] determine a fifth voice content of fifth voice data in the voice data and a first guide content of first interaction guide data in the interaction guide data, the interaction node corresponding to the first interaction guide data being a next interaction node corresponding to the fifth voice data;

[0023] The speech guidance consistency evaluation result is determined based on consistency evaluation results of the fifth speech content and the first guidance content.

[0024] According to the speech interaction device experience evaluation method provided by the application, the interaction data includes user behavior data of interaction with the target speech interaction device;

[0025] The node evaluation result of the user response node is determined based on the following steps:

[0026] The understanding degree evaluation result of the user response node is determined based on first response content of first user behavior data in the user behavior data, the first user behavior data being data corresponding to the user response node;

[0027] The node evaluation result of the user response node is determined based on the understanding degree evaluation result.

[0028] According to the speech interaction device experience evaluation method provided by the application, the first user behavior data includes first user speech data, and the understanding degree evaluation result of the user response node is determined based on first response content of first user behavior data in the user behavior data, including:

[0029] The understanding degree evaluation result of the user response node is determined based on speech content of the first user speech data.

[0030] According to the speech interaction device experience evaluation method provided by the application, the interaction data includes user behavior data of interaction with the target speech interaction device and speech data of the target speech interaction device;

[0031] The node evaluation result of the device response node is determined based on the following steps:

[0032] Second response content of second user behavior data in the user behavior data and sixth speech content of sixth speech data in the speech data are determined, the interaction node corresponding to the sixth speech data being the next interaction node corresponding to the second user behavior data;

[0033] The response correlation degree evaluation result of the device response node is determined based on correlation degrees of the second response content and the sixth speech content;

[0034] The node evaluation result of the device response node is determined based on the response correlation degree evaluation result.

[0035] According to the speech interaction device experience evaluation method provided by the application, the interaction data includes user behavior data of interaction with the target speech interaction device;

[0036] The node evaluation result of the device guide node is determined based on the following steps:

[0037] A guide rationality evaluation result of the device guide node is determined based on an interest degree feature of third user behavior data in the user behavior data, the third user behavior data being data corresponding to the device guide node, and the third user behavior data including second user voice data and / or user video data.

[0038] The node evaluation result of the device guide node is determined based on the guide rationality evaluation result.

[0039] According to the experience evaluation method of the voice interaction device provided by the application, the interaction initiation node includes a device initiation node, and the node evaluation result of the device initiation node includes an initiation content evaluation result.

[0040] The interaction data includes user behavior data of a user interacting with the target voice interaction device, and the initiation content evaluation result is determined based on the following steps:

[0041] A second emotion feature of fourth user behavior data in the user behavior data is used to determine the initiation content evaluation result, the fourth user behavior data being data corresponding to the device initiation node, and the fourth user behavior data including third user voice data and / or user video data.

[0042] Alternatively, the interaction data includes interaction guide data of the target voice interaction device, and the initiation content evaluation result is determined based on the following steps:

[0043] Based on environmental information in which the target voice interaction device is located and a first initiation content of second interaction guide data in the interaction guide data, a first matching degree evaluation result of the environmental information and the first initiation content is determined, the second interaction guide data being data corresponding to the device initiation node.

[0044] The initiation content evaluation result is determined based on the first matching degree evaluation result.

[0045] Alternatively, the interaction data includes interaction guide data of the target voice interaction device, and the initiation content evaluation result is determined based on the following steps:

[0046] Based on user personal information of a user interacting with the target voice interaction device and a second initiation content of third interaction guide data in the interaction guide data, a second matching degree evaluation result of the user personal information and the second initiation content is determined, the third interaction guide data being data corresponding to the device initiation node.

[0047] Based on the second matching degree evaluation result, the provoking content evaluation result is determined.

[0048] According to the voice interaction device experience evaluation method provided by the application, the interaction data includes user behavior data of interaction with the target voice interaction device, the interaction provoking node includes a device provoking node, and the node evaluation result of the device provoking node includes a provoking timing evaluation result.

[0049] The provoking timing evaluation result is determined based on the following steps:

[0050] Based on fifth user behavior data in the user behavior data, the provoking timing evaluation result is determined.

[0051] The fifth user behavior data is data corresponding to the device provoking node, and the fifth user behavior data includes fourth user voice data and / or user video data.

[0052] The application further provides a voice interaction device experience evaluation device, comprising:

[0053] A determination module is configured to determine a target voice interaction device to be evaluated and interaction data corresponding to the target voice interaction device.

[0054] An evaluation module is configured to determine, based on the interaction data, a plurality of experience evaluation results of the target voice interaction device in an interaction process, the plurality of experience evaluation results including node evaluation results of each interaction node in the interaction process and / or an overall evaluation result of the interaction process, and the interaction process including an interaction provoking node, a user response node, a device response node and a device guiding node.

[0055] An evaluation module is configured to determine, based on the plurality of experience evaluation results, an experience evaluation result of the target voice interaction device.

[0056] The application further provides an electronic device including a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor implements the voice interaction device experience evaluation method according to any of the above when executing the program.

[0057] The application further provides a non-transitory computer readable storage medium having a computer program stored thereon, and the computer program is executable on a processor to implement the voice interaction device experience evaluation method according to any of the above.

[0058] The present invention provides a method, apparatus, device, and storage medium for evaluating the experience of a voice interaction device. Based on the interaction data corresponding to the target voice interaction device, it determines several experience evaluation results of the target voice interaction device in the interaction process. These several experience evaluation results include node evaluation results of each interaction node in the interaction process and / or the overall evaluation result of the interaction process. The interaction process includes interaction initiation nodes, user response nodes, device response nodes, and device guidance nodes. By evaluating the experience of the entire interaction process and each interaction node, a more comprehensive and systematic experience evaluation is achieved. Based on these several experience evaluation results, the experience evaluation results of the voice interaction device are determined more systematically, thereby improving the accuracy of the experience evaluation of the voice interaction device and enhancing the user experience of the voice interaction device. Attached Figure Description

[0059] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0060] Figure 1 A flowchart illustrating the experience evaluation method for the voice interaction device provided by the present invention;

[0061] Figure 2 A schematic diagram of the interaction process in which the device provided by the present invention actively initiates an interactive form;

[0062] Figure 3 This is a schematic diagram of the interaction process for a user-initiated interaction format provided by the present invention;

[0063] Figure 4 A schematic diagram of the structure of the experience evaluation device for the voice interaction device provided by the present invention;

[0064] Figure 5 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0065] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0066] With the rapid development of science and technology, there are more and more voice interaction products, such as Tmall Genie, Xiao Ai, and the like. Among them, Tmall Genie and Xiao Ai are passive voice interaction products, which are initiated by the user and responded by the product. In addition, there are also some active voice interaction products that can actively initiate some interactions with the user, and of course, they also support the user to actively initiate interactions and the product to respond. At present, most voice interaction products are multi-round interactions, especially for active voice interaction products, which can interact with the user multiple times after being woken up once without the need to be woken up again. In order to improve the performance of the product and enhance the user experience, it is necessary to evaluate the experience of the voice interaction product.

[0067] At present, the user experience effect of the voice interaction product, especially the user experience effect of the active voice interaction product, is usually evaluated by some single-point indicators, such as evaluating the voice recognition accuracy of the voice interaction product, or evaluating the answer accuracy of the voice interaction product, or evaluating the practicality of the voice interaction product, or evaluating the naturalness of the voice interaction product, or evaluating the fluency of the voice interaction product. However, evaluating the voice interaction product by single-point indicators cannot systematically evaluate the voice interaction product, resulting in a decrease in the accuracy of the experience evaluation of the voice interaction product.

[0068] To solve the above problems, the present application provides the following embodiments. Figure 1 The flowchart of the experience evaluation method of the voice interaction device provided by the present application is shown in Figure 1 The experience evaluation method of the voice interaction device includes the following steps:

[0069] Step 110, determining the target voice interaction device to be evaluated and the interaction data corresponding to the target voice interaction device.

[0070] In the embodiments of the present application, the execution subject of the experience evaluation method provided by the present application can be other electronic devices or the target voice interaction device.

[0071] Here, the target voice interaction device is a device to be evaluated for user experience effect. The target voice interaction device can be a voice interaction robot, an intelligent child companion robot, or other devices with voice interaction function, which is not limited in the embodiments of the present application.

[0072] In an embodiment, the target voice interaction device is an active voice interaction device, which can actively initiate some interactions with the user and also supports the user to actively initiate interactions and the device to respond.

[0073] In another embodiment, the target voice interactive device is a passive voice interactive device, i.e., it can be actively interacted by the user, and the device responds to the interaction.

[0074] Here, the interaction data is related data generated by the user interacting with the target voice interactive device. The interaction data can be collected or directly obtained by the execution subject of the embodiment of the present application, or obtained by collecting data of other devices. The interaction data can include but is not limited to at least one of the following: voice data of the target voice interactive device, appearance form data of the target voice interactive device, interaction guide data of the target voice interactive device, user behavior data interacting with the target voice interactive device, and the like.

[0075] Among them, the voice data is the data corresponding to the voice played by the target voice interactive device. The voice data can include but is not limited to: audio data, text data, and the like. The text data includes the text content corresponding to the audio data.

[0076] Among them, the appearance form data is the data corresponding to the appearance form displayed by the target voice interactive device, such as the data corresponding to the appearance form displayed by the voice interactive robot. In other words, the appearance form data is the data corresponding to the lively performance of the target voice interactive device. The appearance form data can be used to represent the mood of the target voice interactive device.

[0077] In an embodiment, the appearance form data can be video data, i.e., the video data is obtained by photographing the target voice interactive device.

[0078] In another embodiment, the appearance form data can be data of the target voice interactive device itself, i.e., the target voice interactive device displays the appearance form based on the appearance form data.

[0079] Among them, the interaction guide data is the data corresponding to the guide content guided by the target voice interactive device. The interaction guide data is used to represent the guide type and the guide content corresponding to the guide type. The guide type can include but is not limited to at least one of the following: chatting, playing games, resource playing, querying, and the like. For example, the guide content corresponding to chatting is chatting content, the guide content corresponding to playing games is the guided game content, the guide content corresponding to resource playing is the guided resource playing content, and the guide content corresponding to querying is the guided query method and query content.

[0080] The user behavior data is behavior data corresponding to a user interacting with the target voice interaction device. The user behavior data can include, but is not limited to, at least one of the following: user voice data, user video data, and the like. The user voice data can include, but is not limited to, audio data, text data, and the like; the text data includes text content corresponding to the audio data. The user voice data is used to represent the speaking content of the user, the speaking tone of the user, the speaking intonation of the user, and the like. The user video data is used to represent the emotion of the user, the interest level of the user, the positivity of the user, and the like.

[0081] In step 120, based on the interaction data, a plurality of experience evaluation results of the target voice interaction device in an interaction process are determined, the plurality of experience evaluation results including node evaluation results of each interaction node in the interaction process and / or an overall evaluation result of the interaction process, and the interaction process including an interaction initiation node, a user response node, a device response node, and a device guidance node.

[0082] Here, the interaction process is an interaction process of the target voice interaction device, and each interaction node in the interaction process includes an interaction initiation node, a user response node, a device response node, and a device guidance node. The interaction initiation node includes a device initiation node or a user initiation node. The device initiation node is a node in which the target voice interaction device initiatively initiates interaction; the user initiation node is a node in which the user initiatively initiates interaction, i.e., the user initiatively initiates interaction, and the target voice interaction device responds. The device initiation node can initiate interaction types such as chatting, playing games, playing resources, and querying. The user response node is a response node when the user interacts with the target voice interaction device. The device response node is a response node when the target voice interaction device interacts with the user. The device guidance node is a guidance node in which the target voice interaction device guides the user to the next step.

[0083] It should be noted that the active voice interaction device includes two interaction forms (a device initiatively initiates interaction form and a user initiatively initiates interaction form), i.e., the interaction initiation node thereof includes a device initiation node and a user initiation node. Therefore, the experience evaluation can be separately performed for the two interaction forms, so that the experience evaluation method of the embodiment of the present application is combined with product characteristics. The experience evaluation method can be selected according to the product characteristics of the target voice interaction device, such as the experience evaluation process corresponding to no device guidance node for a single-round interaction target voice interaction device. In addition, the difference between the device initiatively initiates interaction form and the user initiatively initiates interaction form lies in the initiation object of the new interaction.

[0084] For ease of understanding, reference is made to Figure 2The interaction process corresponding to the interaction form initiated by the device includes a device initiation node, a user response node, a device response node, and a device guidance node. For example, after the device initiation node, the user response node is waited for, after the user response node, the device response node is performed, after the device response node, the device guidance node is performed, after the device guidance node, the user response node is waited for, and the user response node is performed. Based on the above interaction process, the overall evaluation result of the interaction process, the node evaluation result of the device initiation node, the node evaluation result of the user response node, the node evaluation result of the device response node, and the node evaluation result of the device guidance node can be obtained.

[0085] For ease of understanding, reference is made to Figure 3 The interaction process corresponding to the interaction form initiated by the user includes a user initiation node, a device response node, a device guidance node, and a user response node. For example, after the user initiation node, the device response node is waited for, after the device response node, the device guidance node is performed, after the device guidance node, the user response node is performed, after the user response node, the device response node is performed, and the device response node is performed. Based on the above interaction process, the overall evaluation result of the interaction process, the node evaluation result of the user initiation node, the node evaluation result of the device response node, the node evaluation result of the device guidance node, and the node evaluation result of the user response node can be obtained.

[0086] Here, the number of the experience evaluation results can be one or more, and the experience evaluation results can only include the node evaluation result of the interaction node or only include the overall evaluation result of the interaction process. The node evaluation result of the interaction node can include, but is not limited to, at least one of the following: the node evaluation result of the interaction initiation node, the node evaluation result of the user response node, the node evaluation result of the device response node, and the node evaluation result of the device guidance node. The experience evaluation result can be represented by a score, and can also be represented by other means. The node evaluation result is the experience evaluation result of the corresponding interaction node, and the overall evaluation result is the experience evaluation result of the entire interaction process.

[0087] Specifically, different experience evaluation results require different interaction data, and then based on the first interaction data in the interaction data, the target experience evaluation result is determined, and the first interaction data is the data corresponding to the target experience evaluation result.

[0088] In some specific embodiments, the first interaction data is evaluated to obtain a target experience evaluation result based on a mapping relationship between the sample interaction data in the labeled data and the labeled sample experience evaluation result thereof. The mapping relationship can specifically be an experience evaluation model obtained through model training, or an experience evaluation rule obtained through correlation mining of the sample interaction data and the labeled sample experience evaluation result thereof, and the present embodiment does not make a specific limitation thereto.

[0089] The mapping relationship is obtained by adjusting the model parameters of the experience evaluation model or the representation manner of the experience evaluation rule, so that the experience evaluation result of the sample interaction data is as similar as possible to the sample experience evaluation result. The model parameters of the experience evaluation model or the representation manner of the experience evaluation rule are continuously adjusted through the plurality of labeled data, so that the mapping relationship obtained thereby is more accurate, and the experience evaluation result obtained through experience evaluation is more reliable when applied to experience evaluation.

[0090] For example, the experience evaluation model is obtained through training based on the following steps: based on the labeled data, the experience evaluation model is iteratively optimized to obtain the experience evaluation model. Based on this, the first interaction data is input into the experience evaluation model to obtain the target experience evaluation result output by the experience evaluation model.

[0091] Step 130, determining the experience evaluation result of the target voice interaction device based on the plurality of experience evaluation results.

[0092] Specifically, the plurality of experience evaluation results are aggregated to obtain the experience evaluation result of the target voice interaction device. The experience evaluation result can be represented by a score, i.e., a user experience effect score, and can also be represented by other means. The aggregation method can include but is not limited to weighted aggregation, addition, etc.

[0093] In an embodiment, the plurality of experience evaluation results are weighted aggregated to obtain the experience evaluation result based on the weights corresponding to the experience evaluation results. Different interaction nodes correspond to different weights, i.e., the weights corresponding to the interaction initiation node, the user response node, the device response node and the device guidance node are different, and the weights corresponding to the node evaluation result and the overall evaluation result are also different.

[0094] It should be noted that if a certain interaction node or link score is relatively low, the reason for the low score can be analyzed in depth, and targeted optimization can be performed to improve the product experience effect. For example, the challenge content in the device challenge node is not reasonable, such as the topic of hibernation in spring, and the challenge time range of the hibernation topic can be set to solve it. It can be understood that the weak points of the product are obtained by scoring the interaction process of the target voice interaction device, and the reasons can be analyzed in depth and improved to improve the product experience.

[0095] The experience evaluation method of the voice interaction device provided by the embodiment of the application determines a plurality of experience evaluation results of the target voice interaction device in the interaction process based on the interaction data corresponding to the target voice interaction device, and the plurality of experience evaluation results include node evaluation results of each interaction node in the interaction process and / or overall evaluation results of the interaction process, and the interaction process includes an interaction challenge node, a user response node, a device response node, and a device guide node. Then, by evaluating the experience of the whole process and each interaction node of the interaction process, a more comprehensive and systematic experience evaluation is realized, so that the experience evaluation result of the voice interaction device is more systematically determined based on the plurality of experience evaluation results, thereby improving the experience evaluation accuracy of the voice interaction device, and further improving the user experience of the voice interaction device.

[0096] Based on the above embodiment, in the method, the interaction data includes voice data of the target voice interaction device, and the overall evaluation result includes a voice content consistency evaluation result.

[0097] The voice content consistency evaluation result is determined based on the following steps:

[0098] Determine a first voice content of first voice data in the voice data and a second voice content of second voice data in the voice data, the first voice data and the second voice data being data corresponding to the same interaction process;

[0099] Based on the content consistency evaluation result of the first voice content and the second voice content, the voice content consistency evaluation result is determined.

[0100] Here, the voice content consistency evaluation result is used to represent the consistency degree of the speaking content of the target voice interaction device before and after the interaction process with a user. The voice content consistency evaluation result can be represented by a score, and of course it can also be represented by other means.

[0101] It should be noted that the first voice data and the second voice data are data corresponding to the same interaction process, that is, the first voice data and the second voice data should be two data under the same interaction process.

[0102] Here, the first voice content is the text content corresponding to the first voice data (i.e., the speaking content of the target voice interactive device). If the first voice data is audio data, the audio data is converted into text content; if the first voice data is text data, the text data is directly taken as the text content.

[0103] Here, the second voice content is the text content corresponding to the second voice data (i.e., the speaking content of the target voice interactive device). If the second voice data is audio data, the audio data is converted into text content; if the second voice data is text data, the text data is directly taken as the text content.

[0104] Here, the content consistency evaluation result is used to represent the consistency degree of the two voice contents. The content consistency evaluation result can be represented by a score, and of course can also be represented by other means.

[0105] In some specific embodiments, based on the mapping relationship between the sample first voice content, the sample second voice content, and the annotated sample content consistency evaluation result in the annotated data, the content consistency evaluation of the first voice content and the second voice content is performed to obtain the content consistency evaluation result. The mapping relationship can be embodied as a content consistency evaluation model obtained through model training, or can be embodied as a content consistency evaluation rule obtained through association mining of annotated data, and the embodiments of the present application do not make specific limitations thereto.

[0106] The mapping relationship is obtained by adjusting the model parameters of the content consistency evaluation model or the representation method of the content consistency evaluation rule, so that the content consistency evaluation result of the content consistency evaluation of the sample first voice content and the sample second voice content is as similar as possible to the sample content consistency evaluation result. Here, the model parameters of the content consistency evaluation model or the representation method of the content consistency evaluation rule are constantly adjusted through multiple annotated data, which can make the mapping relationship obtained thereby more accurate, and further make the content consistency evaluation result obtained by the content consistency evaluation more reliable when applied to the content consistency evaluation.

[0107] For example, the content consistency evaluation model is obtained based on the following steps: based on the annotated data, the content consistency evaluation model is iteratively optimized to obtain the content consistency evaluation model. Based on this, the first voice content and the second voice content are input into the content consistency evaluation model to obtain the content consistency evaluation result output by the content consistency evaluation model. Further, the content consistency evaluation model includes a semantic understanding layer to perform content consistency evaluation on the semantics of the two voice contents.

[0108] Specifically, the content consistency evaluation result can be directly taken as the speech content consistency evaluation result, or the speech content consistency evaluation result can be obtained by further processing the content consistency evaluation result.

[0109] In an embodiment, in one interactive process, multiple content consistency evaluation results can be obtained, and a speech content consistency evaluation result is determined based on an aggregation result of the multiple content consistency evaluation results. The aggregation result can be obtained by weighted aggregation of the multiple content consistency evaluation results, or can be obtained by addition processing of the multiple content consistency evaluation results.

[0110] The experience evaluation method of the speech interactive device provided in the embodiments of the present application determines a first speech content of first speech data in speech data of a target speech interactive device, and a second speech content of second speech data in the speech data, and determines a speech content consistency evaluation result based on a content consistency evaluation result of the first speech content and the second speech content, and the first speech data and the second speech data are data corresponding to one interactive process, so as to evaluate the consistency degree of the speech content of the target speech interactive device before and after the interactive process with one user, and further evaluate the consistency of the whole interactive process, thereby improving the experience evaluation accuracy of the speech interactive device, and further improving the speech interactive device based on the speech content consistency evaluation result, avoiding the conflict of the speech content before and after one interactive process, and avoiding the problem of contradiction before and after one interactive process, thereby improving the user experience of the speech interactive device.

[0111] Based on any of the above embodiments, the interactive data includes speech data of the target speech interactive device, and the overall evaluation result includes a speech form consistency evaluation result; the speech form consistency evaluation result is determined based on the following steps:

[0112] determining a third speech content and a tone feature of third speech data in the speech data;

[0113] determining the speech form consistency evaluation result based on a form consistency evaluation result of the third speech content and the tone feature.

[0114] Here, the speech form consistency evaluation result is used to represent the consistency degree of the speech content and the form performance of the target speech interactive device in the interactive process with one user. The speech form consistency evaluation result can be represented by a score, and of course can also be represented by other ways. In the embodiments of the present application, the speech form consistency evaluation result is used to represent the consistency degree of the speech content and the corresponding tone of the target speech interactive device.

[0115] Here, the third voice content is the text content corresponding to the third voice data (i.e., the speaking content of the target voice interactive device). If the third voice data is audio data, the audio data is converted into text content; if the third voice data is text data, the text data is directly taken as the text content.

[0116] Here, the tone feature is used to represent the tone of the target voice interactive device when speaking, such as an angry tone, a sad tone, and the like. The tone feature is a vivid representation of the target voice interactive device. The tone feature is determined through the audio data corresponding to the third voice data.

[0117] In some specific embodiments, the tone feature is extracted from the third voice data based on a mapping relationship between the sample audio data in the labeled data and the sample tone feature labeled for the sample audio data. The mapping relationship can specifically be a tone feature extraction model obtained through model training, or a tone feature extraction rule obtained by correlating and mining the labeled data, and the embodiments of the present application do not make specific limitations thereon. The specific acquisition process of the mapping relationship can refer to the acquisition process of other mapping relationships, which will not be repeated here.

[0118] Here, the form consistency evaluation result is used to represent the consistency degree of the voice content and the tone feature of the same voice data. The form consistency evaluation result can be represented by a score, and of course can also be represented by other means.

[0119] In some specific embodiments, the form consistency evaluation result is obtained by performing form consistency evaluation on the third voice content and the tone feature based on a mapping relationship between the sample third voice content, the sample tone feature, and the sample form consistency evaluation result labeled for the sample third voice content and the sample tone feature. The mapping relationship can specifically be a form consistency evaluation model obtained through model training, or a form consistency evaluation rule obtained by correlating and mining the labeled data, and the embodiments of the present application do not make specific limitations thereon.

[0120] The acquisition process of the mapping relationship includes: adjusting the model parameters of the form consistency evaluation model or the representation method of the form consistency evaluation rule, so that the form consistency evaluation result of the form consistency evaluation on the sample third voice content and the sample tone feature is as similar as possible to the sample form consistency evaluation result. Here, the model parameters of the form consistency evaluation model or the representation method of the form consistency evaluation rule are constantly adjusted through multiple labeled data, which can make the mapping relationship obtained thereby more accurate, and thus the form consistency evaluation result obtained by the form consistency evaluation is more reliable when applied to the form consistency evaluation.

[0121] For example, the morphology consistency evaluation model is trained based on the following steps: based on the labeled data, the morphology consistency evaluation model is iteratively optimized to obtain the morphology consistency evaluation model. Based on this, the third voice content and the tone feature are input into the morphology consistency evaluation model to obtain the morphology consistency evaluation result output by the morphology consistency evaluation model.

[0122] Specifically, the morphology consistency evaluation result can be directly used as the voice morphology consistency evaluation result, or the voice morphology consistency evaluation result can be obtained by further processing the morphology consistency evaluation result.

[0123] In an embodiment, in an interaction process, multiple morphology consistency evaluation results can be obtained, and the voice morphology consistency evaluation result is determined based on the aggregation result of the multiple morphology consistency evaluation results. The aggregation result can be obtained by weighted aggregation of the multiple morphology consistency evaluation results, or the multiple morphology consistency evaluation results can be added to obtain the aggregation result.

[0124] The experience evaluation method of the voice interaction device provided by the embodiment of the application determines the third voice content of the third voice data in the voice data of the target voice interaction device and the tone feature of the third voice data, and determines the voice morphology consistency evaluation result based on the morphology consistency evaluation result of the third voice content and the tone feature, thereby evaluating the consistency degree of the speaking content and the speaking tone of the target voice interaction device in the interaction process with a user, and further evaluating the consistency of the overall interaction process, thereby improving the experience evaluation accuracy of the voice interaction device, and further improving the voice interaction device based on the voice morphology consistency evaluation result, avoiding the conflict between the speaking content and the lively performance, such as the target voice interaction device saying that it is very happy, but its tone is sad, avoiding the problem of contradiction in an interaction process, and thereby improving the user experience of the voice interaction device.

[0125] Based on any of the above embodiments, in the method, the interaction data includes voice data of the target voice interaction device, and the overall evaluation result includes a voice morphology consistency evaluation result; the interaction data further includes appearance morphology data of the target voice interaction device, and the voice morphology consistency evaluation result is determined based on the following steps:

[0126] Determine the fourth voice content of the fourth voice data in the voice data and the first emotional feature of the first appearance morphology data in the appearance morphology data, the fourth voice data and the first appearance morphology data being data corresponding to the same interaction node;

[0127] Determine the voice morphology consistency evaluation result based on the morphology consistency evaluation result of the fourth voice content and the first emotional feature.

[0128] Here, the speech form consistency evaluation result is used to represent the consistency degree of the speech content and the form performance of the target voice interactive device in the interaction process with a user. The speech form consistency evaluation result can be represented by a score, and of course can also be represented by other means. In the embodiment of the present application, the speech form consistency evaluation result is used to represent the consistency degree of the speech content and the corresponding emotion of the target voice interactive device.

[0129] It should be noted that the fourth voice data and the first appearance form data are data corresponding to the same interaction node, that is, the fourth voice data and the first appearance form data should be two data under the same interaction node, so that the form consistency degree of the two can be determined.

[0130] Here, the fourth voice content is the text content corresponding to the fourth voice data (i.e. the speech content of the target voice interactive device). If the fourth voice data is audio data, the audio data is converted into text content; if the fourth voice data is text data, the text data is directly taken as text content.

[0131] Here, the first emotion feature is used to represent the emotion displayed by the appearance when the target voice interactive device speaks, such as anger, sadness, and the like. The first emotion feature is the lifelike performance of the target voice interactive device. The first emotion feature is determined by the video data corresponding to the first appearance form data, or by the data of the target voice interactive device itself corresponding to the first appearance form data.

[0132] In some specific embodiments, the first emotion feature is extracted from the first appearance form data based on the mapping relationship between the sample video data in the labeled data and the labeled sample first emotion feature. The mapping relationship can be embodied as an emotion feature extraction model obtained through model training, or can be embodied as an emotion feature extraction rule obtained by association mining of the labeled data, and the present embodiment does not make specific limitation thereto. The specific acquisition process of the mapping relationship can refer to the acquisition process of other mapping relationships, which will not be repeated here.

[0133] Here, the form consistency evaluation result is used to represent the consistency degree of the speech content and the emotion feature at the same time. The form consistency evaluation result can be represented by a score, and of course can also be represented by other means.

[0134] In some specific embodiments, based on a mapping relationship between the sample fourth voice content, the sample first emotion feature, and the labeled sample morphological consistency evaluation result, the morphological consistency evaluation of the fourth voice content and the emotion feature is performed to obtain a morphological consistency evaluation result. The mapping relationship can be embodied as a morphological consistency evaluation model obtained through model training, or can be embodied as a morphological consistency evaluation rule obtained through association mining of the labeled data, and the embodiments of the present application do not make specific limitations thereto.

[0135] The mapping relationship is obtained by adjusting the model parameters of the morphological consistency evaluation model or the representation method of the morphological consistency evaluation rule, so that the morphological consistency evaluation result of the sample fourth voice content and the sample first emotion feature is as similar as possible to the sample morphological consistency evaluation result. Here, the model parameters of the morphological consistency evaluation model or the representation method of the morphological consistency evaluation rule are constantly adjusted through multiple labeled data, so that the mapping relationship obtained thereby is more accurate, and thus the morphological consistency evaluation result obtained by the morphological consistency evaluation is more reliable when applied to the morphological consistency evaluation.

[0136] For example, the morphological consistency evaluation model is obtained through training based on the following steps: based on the labeled data, the morphological consistency evaluation model is iteratively optimized to obtain the morphological consistency evaluation model. Based on this, the fourth voice content and the emotion feature are input into the morphological consistency evaluation model to obtain the morphological consistency evaluation result output by the morphological consistency evaluation model.

[0137] Specifically, the morphological consistency evaluation result can be directly used as the voice morphological consistency evaluation result, or the voice morphological consistency evaluation result can be obtained by further processing the morphological consistency evaluation result.

[0138] In an embodiment, in an interaction process, multiple morphological consistency evaluation results can be obtained, and thus the voice morphological consistency evaluation result is determined based on the aggregation result of the multiple morphological consistency evaluation results. The aggregation result can be obtained by weighted aggregation of the multiple morphological consistency evaluation results, or can be obtained by addition processing of the multiple morphological consistency evaluation results. The multiple morphological consistency evaluation results can include the morphological consistency evaluation result of the third voice content and the tone feature.

[0139] The experience evaluation method of the voice interaction device provided by the embodiment of the present application determines the fourth voice content of the fourth voice data in the voice data of the target voice interaction device and the first emotional feature of the first appearance and morphology data in the appearance and morphology data, so as to determine the voice and morphology consistency evaluation result based on the morphology consistency evaluation result of the fourth voice content and the first emotional feature, determine the consistency degree of the speaking content and emotion of the target voice interaction device in the interaction process of the target voice interaction device with a user, and then evaluate the consistency of the overall interaction process, thereby improving the experience evaluation accuracy of the voice interaction device, and then improving the voice interaction device based on the voice and morphology consistency evaluation result, avoiding the conflict between the speaking content and the lively performance, such as the target voice interaction device saying that it is very happy, but its expression is sad, avoiding the contradictory problem in an interaction process, and thereby improving the user experience of the voice interaction device.

[0140] Based on any of the above embodiments, in the method, the interaction data includes voice data of the target voice interaction device and interaction guide data of the target voice interaction device, and the overall evaluation result includes a voice guide consistency evaluation result; the voice guide consistency evaluation result is determined based on the following steps:

[0141] determining the fifth voice content of the fifth voice data in the voice data and the first guide content of the first interaction guide data in the interaction guide data, the interaction node corresponding to the first interaction guide data being the next interaction node corresponding to the fifth voice data;

[0142] determining the voice guide consistency evaluation result based on the guide consistency evaluation result of the fifth voice content and the first guide content.

[0143] Here, the voice guide consistency evaluation result is used to represent the consistency degree of the speaking content of the target voice interaction device and the guide content of the next step in the interaction process of the target voice interaction device with a user. The voice guide consistency evaluation result can be represented by a score, and of course can also be represented by other means.

[0144] It should be noted that the interaction node corresponding to the first interaction guide data is the next interaction node corresponding to the fifth voice data, that is, the guide content of the next step corresponding to the fifth voice content is the first guide content.

[0145] Here, the fifth voice content is the text content corresponding to the fifth voice data (i.e., the speaking content of the target voice interaction device). If the fifth voice data is audio data, the audio data is converted into text content; if the fifth voice data is text data, the text data is directly taken as text content.

[0146] Here, the first guidance content is used to represent the guidance type, and the guidance content corresponding to the guidance type. The guidance type can include but is not limited to at least one of the following: chatting, playing games, resource playing, querying, etc. For example, the first guidance content is the guidance content corresponding to chatting, i.e. chatting content, or the first guidance content is the guidance content corresponding to playing games, i.e. guided game content, or the first guidance content is the guidance content corresponding to resource playing, i.e. guided resource playing content, or the first guidance content is the guidance content corresponding to querying, i.e. guided querying mode and querying content, etc.

[0147] Here, the guidance consistency evaluation result is used to represent the consistency degree of the voice content and the guidance content. The guidance consistency evaluation result can be represented by a score, and of course can also be represented by other means.

[0148] In some specific embodiments, based on the mapping relationship between the sample fifth voice content, the sample first guidance content and the sample guidance consistency evaluation result in the labeled data, the guidance consistency evaluation of the fifth voice content and the first guidance content is performed to obtain the guidance consistency evaluation result. The mapping relationship can be embodied as a guidance consistency evaluation model obtained by model training, or can be embodied as a guidance consistency evaluation rule obtained by association mining of the labeled data, and the embodiments of the present application do not make specific limitations.

[0149] The mapping relationship is obtained by adjusting the model parameters of the guidance consistency evaluation model or the representation method of the guidance consistency evaluation rule, so that the guidance consistency evaluation result of the sample fifth voice content and the sample first guidance content is as similar as possible to the sample guidance consistency evaluation result. Here, by constantly adjusting the model parameters of the guidance consistency evaluation model or the representation method of the guidance consistency evaluation rule through multiple labeled data, the mapping relationship obtained thereby can be more accurate, and thus the guidance consistency evaluation result obtained by the guidance consistency evaluation is more reliable when applied to the guidance consistency evaluation.

[0150] For example, the guidance consistency evaluation model is obtained by training based on the following steps: based on the labeled data, the guidance consistency evaluation model is iteratively optimized to obtain the guidance consistency evaluation model. Based on this, the fifth voice content and the first guidance content are input into the guidance consistency evaluation model to obtain the guidance consistency evaluation result output by the guidance consistency evaluation model.

[0151] Specifically, the guidance consistency evaluation result can be directly used as the voice guidance consistency evaluation result, or the guidance consistency evaluation result can be further processed to obtain the voice guidance consistency evaluation result.

[0152] In an embodiment, in an interaction process, multiple guidance consistency evaluation results can be obtained, and a voice guidance consistency evaluation result is determined based on an aggregation result of the multiple guidance consistency evaluation results. The aggregation result can be obtained by weighted aggregation of the multiple guidance consistency evaluation results, or can be obtained by addition processing of the multiple guidance consistency evaluation results.

[0153] It should be noted that the overall evaluation result can be obtained by aggregating the voice content consistency evaluation result and / or the voice form consistency evaluation result and / or the voice guidance consistency evaluation result. The aggregation processing mode can include but is not limited to weighted aggregation, addition, etc.

[0154] The experience evaluation method of the voice interaction device provided by the embodiment of the application determines the fifth voice content of the fifth voice data in the voice data of the target voice interaction device and the first guide content of the first interaction guide data in the interaction guide data, determines a voice guidance consistency evaluation result based on the guidance consistency evaluation result of the fifth voice content and the first guide content, and the interaction node corresponding to the first interaction guide data is the next interaction node corresponding to the fifth voice data, thereby evaluating the consistency degree of the speaking content of the target voice interaction device and the next guide content in the interaction process of the target voice interaction device with a user, and further evaluating the consistency of the overall interaction process, thereby improving the experience evaluation accuracy of the voice interaction device, and further improving the voice interaction device based on the voice guidance consistency evaluation result, avoiding the conflict between the speaking content and the next guide in an interaction process, such as the target voice interaction device saying that it wants to listen to the user, and then not giving the user an opportunity to speak and continuing to speak, avoiding the occurrence of contradictory problems in an interaction process, and thereby improving the user experience of the voice interaction device.

[0155] Based on any of the above embodiments, in the method, the interaction data includes user behavior data of interaction with the target voice interaction device; and the node evaluation result of the user response node is determined based on the following steps:

[0156] determining an understanding degree evaluation result of the user response node based on first response content of first user behavior data in the user behavior data, the first user behavior data being data corresponding to the user response node;

[0157] determining the node evaluation result of the user response node based on the understanding degree evaluation result.

[0158] Here, the first response content is the response content of the user when interacting with the target voice interaction device. The first response content can include but is not limited to voice content, video content, etc.

[0159] Here, the understanding degree evaluation result is used to represent the understanding degree of the user when interacting with the target voice interaction device, i.e., the understanding degree of the user is determined through the first response content of the user, and then the easy-to-understand degree of the target voice interaction device is determined based on the understanding degree of the user. The understanding degree evaluation result can be represented by a score, and of course can also be represented by other means.

[0160] It should be noted that the words spoken by the target voice interaction device and the guided content guided in the entire interaction process should be easy to understand and consistent with the cognitive level of the user at the age stage, and the user should understand how to use the product, such as how to wake up, the opening opportunity, etc. Based on this, it can be observed whether the user will not respond, answer irrelevant questions, or explicitly indicate that he does not understand due to not understanding by observing the user's response to the content of the target voice interaction device when using.

[0161] In an embodiment, the first user behavior data includes first user voice data, and the understanding degree evaluation result of the user response node is determined based on the voice content of the first user voice data. The specific execution process of this embodiment is described below, which will not be repeated here.

[0162] In another embodiment, the first user behavior data includes user video data, and the understanding degree evaluation result of the user response node is determined based on the video content of the user video data. That is, the understanding degree evaluation result of the user can be determined through the user's gestures, expressions, body movements, etc., such as the user's direct intention to show that he does not understand or does not understand during the interaction process.

[0163] In some specific embodiments, the understanding degree evaluation result is obtained by evaluating the understanding degree of the first response content based on the mapping relationship between the sample first response content and the sample understanding degree evaluation result in the labeled data. The mapping relationship can be embodied as an understanding degree evaluation model obtained by model training, or can be embodied as an understanding degree evaluation rule obtained by association mining of the labeled data, and the embodiments of the present application do not make specific limitations.

[0164] The mapping relationship is obtained by adjusting the model parameters of the understanding degree evaluation model or the representation method of the understanding degree evaluation rule, so that the understanding degree evaluation result of the understanding degree evaluation of the first response content is as similar as possible to the sample understanding degree evaluation result. Here, the model parameters of the understanding degree evaluation model or the representation method of the understanding degree evaluation rule are constantly adjusted through multiple labeled data, which can make the mapping relationship obtained thereby more accurate, and then the understanding degree evaluation result obtained by the understanding degree evaluation is more reliable when applied to the understanding degree evaluation.

[0165] For example, the understanding degree evaluation model is trained based on the following steps: based on the labeled data, the understanding degree evaluation model is iteratively optimized to obtain the understanding degree evaluation model. Based on this, the first response content is input into the understanding degree evaluation model to obtain the understanding degree evaluation result output by the understanding degree evaluation model.

[0166] Specifically, the understanding degree evaluation result can be directly used as the node evaluation result of the user response node, or the understanding degree evaluation result can be further processed to obtain the node evaluation result of the user response node.

[0167] In an embodiment, in an interaction process, multiple understanding degree evaluation results can be obtained, and based on the aggregation result of the multiple understanding degree evaluation results, the node evaluation result of the user response node is determined. The aggregation result can be obtained by weighted aggregation of the multiple understanding degree evaluation results, or the multiple understanding degree evaluation results can be added.

[0168] It should be noted that the node evaluation result of the user response node can also be determined based on other types of evaluation, which will not be described here.

[0169] The experience evaluation method of the voice interaction device provided by the embodiment of the application determines the understanding degree evaluation result of the user response node based on the first response content of the first user behavior data in the user behavior data, thereby determining the node evaluation result of the user response node based on the understanding degree evaluation result, and the first user behavior data is the data corresponding to the user response node, thereby evaluating the understanding degree of the user in the user response node in the interaction process, and further evaluating the node of the user response node, thereby improving the experience evaluation accuracy of the voice interaction device, and further improving the voice interaction device based on the understanding degree evaluation result, avoiding the user's difficulty in understanding, such as the user not responding after the voice interaction device finishes speaking, avoiding the problem of the user not understanding in an interaction process, and thereby improving the user experience of the voice interaction device.

[0170] Based on any of the above embodiments, in the method, the first user behavior data includes first user voice data, and the determination of the understanding degree evaluation result of the user response node based on the first response content of the first user behavior data in the user behavior data includes:

[0171] Determining the understanding degree evaluation result of the user response node based on the voice content of the first user voice data.

[0172] Here, the voice content is the response content of the user when interacting with the target voice interaction device.

[0173] In some specific embodiments, the understanding degree evaluation result of the voice content is obtained based on a mapping relationship between sample voice content in the labeled data and its labeled sample understanding degree evaluation result. The mapping relationship can be embodied as an understanding degree evaluation model obtained through model training, or an understanding degree evaluation rule obtained through correlation mining of the labeled data, and the embodiments of the application do not make specific limitations thereon. The specific obtaining process of the mapping relationship can refer to the obtaining process of other mapping relationships, which will not be repeated here.

[0174] For example, the understanding degree evaluation result is divided into three categories, each of which corresponds to different scores. The first category is no response, that is, the user does not respond after the target voice interactive device finishes speaking due to not understanding, and does not make a sound; the second category is answering a different question, that is, the user makes an irrelevant answer to the content asked by the target voice interactive device, such as the target voice interactive device asking the user what he ate in the morning, and the user answering that he went to school, which is a typical answering a different question; and the third category is explicitly expressing that he cannot understand, that is, the user directly expresses his intention of not understanding or not comprehending in the interaction process.

[0175] The experience evaluation method of the voice interactive device provided by the embodiments of the application determines the understanding degree evaluation result of the user response node through the voice content of the first user voice data, thereby evaluating the understanding degree evaluation result of the user in the user response node in the interaction process, and further evaluating the user response node, thereby improving the experience evaluation accuracy of the voice interactive device, and further improving the voice interactive device based on the understanding degree evaluation result, avoiding user difficulty in understanding, and avoiding the problem of user not understanding in an interaction process, thereby improving the user experience of the voice interactive device.

[0176] Based on any of the above embodiments, in the method, the interaction data includes user behavior data interacting with the target voice interactive device and voice data of the target voice interactive device; and the node evaluation result of the device response node is determined based on the following steps:

[0177] The second response content of the second user behavior data in the user behavior data and the sixth voice content of the sixth voice data in the voice data are determined, and the interaction node corresponding to the sixth voice data is the next interaction node corresponding to the second user behavior data;

[0178] Based on the correlation degree between the second response content and the sixth voice content, a response correlation degree evaluation result of the device response node is determined;

[0179] Based on the response correlation degree evaluation result, a node evaluation result of the device response node is determined.

[0180] It should be noted that the interaction node corresponding to the sixth voice data is the next interaction node corresponding to the second user behavior data, that is, the sixth voice data is response data for the second user behavior data.

[0181] Here, the second response content is the response content of the user when interacting with the target voice interaction device. The second response content can include but is not limited to voice content, video content, and the like.

[0182] Here, the sixth voice content is the text content corresponding to the sixth voice data (i.e., the speaking content of the target voice interaction device). If the sixth voice data is audio data, the audio data is converted into text content; if the sixth voice data is text data, the text data is directly taken as text content.

[0183] Here, the relevance degree is used to represent the relevance degree of the voice content of the device and the response content of the user.

[0184] Here, the response relevance evaluation result is used to represent the relevance degree of the voice content and the response content of the user. The response relevance evaluation result can be represented by a score, and of course can also be represented by other means.

[0185] It should be noted that the reply of the target voice interaction device to the response content of the user should be relevant. Based on this, the reply of the target voice interaction device can be evaluated, for example, the reply of the target voice interaction device and the response content of the user can be divided into several dimensions such as strong correlation, correlation, not relevant but reasonable, and completely irrelevant for evaluation.

[0186] In an embodiment, the second user behavior data includes user voice data, and the response relevance evaluation result is determined based on the relevance degree of the voice content of the user voice data and the sixth voice content.

[0187] In some specific embodiments, based on the mapping relationship between the sample voice content, the sample sixth voice content, and the labeled sample response relevance evaluation result in the labeled data, the voice content of the user voice data and the sixth voice content are evaluated for the response relevance evaluation result. The mapping relationship can be embodied as a response relevance evaluation model obtained by model training, or can be embodied as a response relevance evaluation rule obtained by association mining of the labeled data, and the embodiments of the present application do not make specific limitations. The specific acquisition process of the mapping relationship can refer to the acquisition process of other mapping relationships, which will not be repeated here.

[0188] Exemplarily, the response relevance evaluation result is divided into four categories, each of which corresponds to a different score. For example, the target voice interactive device asks the user "Have you eaten?", the user answers "Not yet"; if the target voice interactive device replies "You won't be hungry if you eat", it is the first category, irrelevant; if the target voice interactive device replies "You won't be hungry until you are full", it is the second category, irrelevant but reasonable; if the target voice interactive device replies "I haven't eaten either", it is the third category, relevant; if the target voice interactive device replies "Why haven't you eaten?", it is the fourth category, strongly relevant.

[0189] In another embodiment, the second user behavior data includes user video data, and the response relevance evaluation result is determined based on the relevance between the video content of the user video data and the sixth voice content.

[0190] In some specific embodiments, the video content of the user video data and the sixth voice content are evaluated for the response relevance based on a mapping relationship between the sample video content, the sample sixth voice content and the annotated sample response relevance evaluation result in the annotated data. The mapping relationship can specifically be a response relevance evaluation model obtained through model training, or a response relevance evaluation rule obtained through association mining of the annotated data, and the present embodiment does not make a specific limitation thereon. The specific obtaining process of the mapping relationship can refer to the obtaining process of other mapping relationships, which will not be repeated here.

[0191] Specifically, the response relevance evaluation result can be directly used as the node evaluation result of the device response node, or the response relevance evaluation result can be further processed to obtain the node evaluation result of the device response node.

[0192] In an embodiment, in an interaction process, multiple response relevance evaluation results can be obtained, and thus the node evaluation result of the device response node is determined based on the aggregation result of the multiple response relevance evaluation results. The aggregation result can be obtained by weighted aggregation of the multiple response relevance evaluation results, or by addition processing of the multiple response relevance evaluation results.

[0193] It should be noted that the node evaluation result of the device response node can also be determined based on other types of evaluation, which will not be repeated here.

[0194] The experience evaluation method of the voice interaction device provided by the embodiment of the present application is based on the second response content of the second user behavior data in the user behavior data and the sixth voice content of the sixth voice data in the voice data to determine the correlation degree of the second response content and the sixth voice content, thereby determining the response correlation degree evaluation result of the device response node based on the correlation degree, and the interaction node corresponding to the sixth voice data is the next interaction node corresponding to the second user behavior data, thereby evaluating the response correlation degree evaluation result of the device in the device response node in the interaction flow, and then the node evaluation of the device response node can be performed, thereby improving the experience evaluation accuracy of the voice interaction device, and then the voice interaction device is improved based on the response correlation degree evaluation result, the reply correlation of the voice interaction device is improved, and the user experience of the voice interaction device is improved.

[0195] Based on any of the above embodiments, in the method, the interaction data includes user behavior data interacting with the target voice interaction device; and the node evaluation result of the device guide node is determined based on the following steps:

[0196] Based on the interest degree feature of the third user behavior data in the user behavior data, a guide rationality degree evaluation result of the device guide node is determined, the third user behavior data is data corresponding to the device guide node, and the third user behavior data includes second user voice data and / or user video data;

[0197] Based on the guide rationality degree evaluation result, a node evaluation result of the device guide node is determined.

[0198] Here, the interest degree feature is used to represent the interest degree of the user when the target voice interaction device guides. The interest degree feature is determined by the user video data corresponding to the third user behavior data, or by the second user voice data corresponding to the third user behavior data. For example, whether the user is interested can be evaluated by the user's language expression, gesture, expression, body movement, etc. If the user does not show disinterest, it means that the user is interested. For example, when the user says that he does not want to chat, the target voice interaction device should guide the user to play games, listen to resources, or ask the user what he wants to do after responding that he does not want to chat.

[0199] In some embodiments, the interest degree feature is extracted from the third user behavior data based on a mapping relationship between the third user behavior data and the annotated interest degree feature of the sample in the annotated data. The mapping relationship can be embodied as an interest degree feature extraction model obtained through model training, or embodied as an interest degree feature extraction rule obtained through association mining of the annotated data, and the present embodiment does not make a specific limitation thereon. The specific obtaining process of the mapping relationship can refer to the obtaining process of other mapping relationships, which will not be repeated here.

[0200] Here, the guidance rationality evaluation result is used to represent the rationality of the next step guidance of the target voice interactive device. The guidance rationality evaluation result can be represented by a score, and of course can also be represented by other ways.

[0201] It should be noted that, after the target voice interactive device makes a targeted empathic reply to the response content of the user, in order to avoid a cold scene and meet the product characteristics, the next step should be guided, such as continuing to chat on a certain topic, guiding to play a game, listening to resources, etc. The rationality of the next step guidance can be judged by whether it meets the way of the human-like interaction. If the user is interested in the current content, the next step can guide the user to continue to stay in the current module and further interact, if the user is not interested, the next topic or other guidance types (such as playing a game, playing resources, etc.) or asking the user's intention are timely transferred.

[0202] In an embodiment, the third user behavior data includes second user voice data, and the guidance rationality evaluation result of the device guidance node is determined based on the interest degree feature of the second user voice data. That is, the language expression of the user is used to determine the guidance rationality evaluation result of the device guidance node.

[0203] In another embodiment, the third user behavior data includes user video data, and the guidance rationality evaluation result of the device guidance node is determined based on the interest degree feature of the user video data. That is, the gesture (such as the hand gesture), expression (such as the disgusted expression, the bored expression), body movement, etc. of the user are used to determine the guidance rationality evaluation result of the device guidance node.

[0204] In some embodiments, the mapping relationship between the sample interest degree feature in the labeled data and the labeled sample guide reasonable degree evaluation result is used to guide the reasonable degree evaluation of the sample interest degree feature to obtain the guide reasonable degree evaluation result. The mapping relationship can be embodied by a guide reasonable degree evaluation model obtained through model training, or embodied by a guide reasonable degree evaluation rule obtained through correlation mining of the labeled data, and the embodiments of the present application do not make specific limitations thereon. The specific obtaining process of the mapping relationship can refer to the obtaining process of other mapping relationships, which will not be repeated here.

[0205] Specifically, the guide reasonable degree evaluation result can be directly used as the node evaluation result of the device guide node, or the guide reasonable degree evaluation result can be further processed to obtain the node evaluation result of the device guide node.

[0206] In an embodiment, in an interaction process, multiple guide reasonable degree evaluation results can be obtained, and the node evaluation result of the device guide node is determined based on the aggregation result of the multiple guide reasonable degree evaluation results. The aggregation result can be obtained by weighted aggregation of the multiple guide reasonable degree evaluation results, or obtained by addition processing of the multiple guide reasonable degree evaluation results.

[0207] It should be noted that the node evaluation result of the device guide node can also be determined based on other types of evaluation, which will not be repeated here.

[0208] The experience evaluation method of the voice interaction device provided by the embodiments of the present application determines the guide reasonable degree evaluation result of the device guide node based on the interest degree feature of the third user behavior data in the user behavior data, thereby determining the node evaluation result of the device guide node based on the guide reasonable degree evaluation result, and the third user behavior data is the data corresponding to the device guide node, thereby evaluating the guide reasonable degree of the device in the device guide node in the interaction process, and further evaluating the node of the device guide node, thereby improving the experience evaluation accuracy of the voice interaction device, and further improving the voice interaction device based on the guide reasonable degree evaluation result, improving the next guide rationality of the voice interaction device, and improving the user experience of the voice interaction device.

[0209] Based on any of the above embodiments, in the method, the interaction trigger node includes a device trigger node, the node evaluation result of the device trigger node includes a trigger content evaluation result; the interaction data includes user behavior data interacting with the target voice interaction device, and the trigger content evaluation result is determined based on the following steps:

[0210] The second emotional feature in the fourth user behavior data corresponding to the device provoking node is used to determine the provoking content evaluation result.

[0211] Here, the provoking content evaluation result is used to represent the appropriateness of the provoking content in the active provoking interaction of the target voice interactive device. The provoking content evaluation result can be represented by a score, and of course can also be represented by other ways.

[0212] It should be noted that when the target voice interactive device actively provokes interaction, the provoking content should be in harmony with the user's emotion, and the most suitable provoking content should conform to the user's emotion at the time, such as comforting when the user is sad. It is inappropriate for the user to chat and interact happily when the user is sad.

[0213] It should be noted that the fourth user behavior data is the data corresponding to the device provoking node, so that the appropriateness of the provoking content of the device provoking node can be determined through the fourth user behavior data.

[0214] Here, the second emotional feature is used to represent the emotion exhibited by the user when the target voice interactive device provokes interaction, such as anger, disgust, happiness, and the like. The second emotional feature is determined through the user video data corresponding to the fourth user behavior data, or through the third user voice data corresponding to the fourth user behavior data.

[0215] In some specific embodiments, based on the mapping relationship between the sample fourth user behavior data in the labeled data and the labeled sample second emotional feature, the second emotional feature is extracted from the fourth user behavior data to obtain the second emotional feature. The mapping relationship can be embodied as a second emotional feature extraction model obtained through model training, or can be embodied as a second emotional feature extraction rule obtained through association mining of labeled data, and the embodiments of the present application do not make specific limitations. The specific acquisition process of the mapping relationship can refer to the acquisition process of other mapping relationships, which will not be repeated here.

[0216] In an embodiment, the fourth user behavior data includes third user voice data, and the second emotional feature based on the third user voice data is used to determine the provoking content evaluation result. That is, the language expression of the user is used to determine the provoking content evaluation result, such as determining the emotional feature based on the voice content, or determining the emotional feature based on the tone feature corresponding to the audio data.

[0217] In another embodiment, the fourth user behavior data comprises user video data, and the second emotional feature of the user video data is used to determine the provoking content evaluation result. That is, the emotional feature is determined by gestures, expressions (such as sad expressions, happy expressions), body movements, and the like of the user.

[0218] In some specific embodiments, the second emotional feature is used to determine the provoking content evaluation result based on a mapping relationship between the second emotional feature and a sample provoking content evaluation result of a sample in the labeled data. The mapping relationship can be embodied by a provoking content evaluation model obtained through model training, or can be embodied by a provoking content evaluation rule obtained by association mining of the labeled data, and the embodiments of the present application do not make specific limitations thereon. The specific acquisition process of the mapping relationship can refer to the acquisition process of other mapping relationships, which will not be repeated here.

[0219] In some embodiments, the interaction data further comprises interaction guidance data of the target voice interaction device, the second emotional feature of the fourth user behavior data in the user behavior data, and the third provoking content of the fourth interaction guidance data in the interaction guidance data are determined, and the provoking content evaluation result is determined based on the matching degree of the second emotional feature and the third provoking content. The collection time of the fourth user behavior data and the fourth interaction guidance data is the same, or the collection time of the fourth user behavior data is before the fourth interaction guidance data. The matching degree is used to represent the matching degree of the third provoking content and the current emotion of the user, for example, if the user is sad, the matching degree is higher if the user is sad, and the matching degree is lower if the user is happy.

[0220] In some specific embodiments, the matching degree is predicted based on a mapping relationship between the second emotional feature, the third provoking content, and a sample matching degree of a sample in the labeled data. The specific acquisition process of the mapping relationship can refer to the acquisition process of other mapping relationships, which will not be repeated here.

[0221] In some specific embodiments, the matching degree is used to determine the provoking content evaluation result based on a mapping relationship between the matching degree and a sample provoking content evaluation result of a sample in the labeled data. The specific acquisition process of the mapping relationship can refer to the acquisition process of other mapping relationships, which will not be repeated here.

[0222] It should be noted that the node evaluation result of the device provoking node can also be determined based on other types of evaluation, which will not be repeated here.

[0223] The experience evaluation method of the voice interaction device provided by the embodiment of the present application is based on the second emotional feature of the fourth user behavior data in the user behavior data, determines the provoking content evaluation result, and determines the node evaluation result of the device provoking node based on the provoking content evaluation result. The fourth user behavior data is the data corresponding to the device provoking node, so that the provoking content evaluation result of the device in the device provoking node in the interaction process is evaluated, and the node evaluation of the device provoking node is further performed, thereby improving the experience evaluation accuracy of the voice interaction device, and the voice interaction device is improved based on the provoking content evaluation result, improving the suitability of the provoking content of the voice interaction device, thereby improving the user experience of the voice interaction device.

[0224] Based on any of the above embodiments, in the method, the interaction provoking node includes a device provoking node, and the node evaluation result of the device provoking node includes a provoking content evaluation result; the interaction data includes interaction guide data of the target voice interaction device, and the provoking content evaluation result is determined based on the following steps:

[0225] Based on the environmental information in which the target voice interaction device is located and the first provoking content of the second interaction guide data in the interaction guide data, a first matching degree evaluation result of the environmental information and the first provoking content is determined, and the second interaction guide data is data corresponding to the device provoking node;

[0226] Based on the first matching degree evaluation result, the provoking content evaluation result is determined.

[0227] It should be noted that when the target voice interaction device actively provokes interaction, the provoking content should be consistent with the environmental information in which the user is currently located, and the most suitable provoking content should be consistent with the environment (such as the scene) at that time. For example, during the Spring Festival, the most inappropriate provoking content is to talk about hibernation in spring.

[0228] Here, the environmental information is used to represent the scene in which the target voice interaction device is located, and the environmental information can include but is not limited to at least one of the following: a holiday, a weekday or a weekend, the weather, the time (such as morning, noon, afternoon, evening, etc.), the geographical location, etc.

[0229] Here, the first matching degree evaluation result is used to represent the matching degree of the first provoking content and the environment in which the user is currently located. For example, during the Spring Festival, the matching degree is higher when talking about the Spring Festival, and the matching degree is lower when talking about hibernation in spring.

[0230] In some embodiments, the mapping relationship between the sample environment information, the sample first provoking content, and the labeled sample first matching degree evaluation result in the labeled data is used to evaluate the matching degree of the environment information and the first provoking content to obtain the first matching degree evaluation result. The specific process of obtaining the mapping relationship can refer to the process of obtaining other mapping relationships, which will not be repeated here.

[0231] Specifically, the first matching degree evaluation result can be directly used as the provoking content evaluation result, or the first matching degree evaluation result can be further processed to obtain the provoking content evaluation result.

[0232] It should be noted that the node evaluation result of the device provoking node can also be determined based on other types of evaluation, which will not be repeated here.

[0233] The experience evaluation method of the voice interaction device provided by the embodiments of the present application determines the provoking content evaluation result based on the environment information in which the target voice interaction device is located and the first provoking content of the second interaction guide data in the interaction guide data, thereby determining the node evaluation result of the device provoking node based on the provoking content evaluation result, so that the node evaluation of the device provoking node can be performed, thereby improving the experience evaluation accuracy of the voice interaction device, and further improving the provoking content suitability of the voice interaction device based on the provoking content evaluation result, thereby improving the user experience of the voice interaction device.

[0234] Based on any of the above embodiments, in the method, the interaction provoking node includes a device provoking node, the node evaluation result of the device provoking node includes a provoking content evaluation result, the interaction data includes interaction guide data of the target voice interaction device, and the provoking content evaluation result is determined based on the following steps:

[0235] Based on the user personal information of the user interacting with the target voice interaction device and the second provoking content of the third interaction guide data in the interaction guide data, a second matching degree evaluation result of the user personal information and the second provoking content is determined, and the third interaction guide data is data corresponding to the device provoking node.

[0236] Based on the second matching degree evaluation result, the provoking content evaluation result is determined.

[0237] It should be noted that when the target voice interaction device actively provokes interaction, the provoking content should be consistent with the current personal information of the user, and the best provoking content should be consistent with the user's current user situation, such as being consistent with the user's preferences, and the worst provoking content is completely inconsistent with the user's current user situation, such as being inappropriate for deviating from the user's preferences.

[0238] Here, the user personal information is used to represent the user situation of the user, which can include but is not limited to at least one of the following: preferences, age, historical emotions, and the like.

[0239] Here, the second matching degree evaluation result is used to represent the matching degree of the second provoking content and the current user personal information of the user.

[0240] In some embodiments, based on the mapping relationship between the sample user personal information, the sample second provoking content and the labeled sample second matching degree evaluation result in the labeled data, the matching degree evaluation of the user personal information and the second provoking content is performed to obtain the second matching degree evaluation result. The specific acquisition process of the mapping relationship can refer to the acquisition process of other mapping relationships, which will not be repeated here.

[0241] Specifically, the second matching degree evaluation result can be directly used as the provoking content evaluation result, or the second matching degree evaluation result can be further processed to obtain the provoking content evaluation result.

[0242] It should be noted that the node evaluation result of the device provoking node can also be determined based on other types of evaluation, which will not be repeated here.

[0243] The experience evaluation method of the voice interaction device provided by the embodiment of the application determines the provoking content evaluation result based on the user personal information of the user interacting with the target voice interaction device and the second provoking content of the third interaction guide data in the interaction guide data, so as to determine the node evaluation result of the device provoking node based on the provoking content evaluation result, so that the node evaluation of the device provoking node can be performed, the experience evaluation accuracy of the voice interaction device is improved, and the voice interaction device is improved based on the provoking content evaluation result, the suitability of the provoking content of the voice interaction device is improved, and the user experience of the voice interaction device is improved.

[0244] Based on any of the above embodiments, in the method, the interaction data includes user behavior data of a user interacting with the target voice interaction device, the interaction provoking node includes a device provoking node, and the node evaluation result of the device provoking node includes a provoking timing evaluation result; the provoking timing evaluation result is determined based on the following steps:

[0245] Determine the provoking timing evaluation result based on fifth user behavior data in the user behavior data;

[0246] The fifth user behavior data is data corresponding to the device provoking node, and the fifth user behavior data includes fourth user voice data and / or user video data.

[0247] Here, the evaluation result of the provoking time is used to represent the appropriateness of the provoking time in the active provoking interaction of the target voice interaction device. The evaluation result of the provoking time can be represented by a score, and of course can also be represented by other ways.

[0248] It should be noted that the provoking time can be set according to the characteristics of the voice interaction device, such as setting the time when the device is turned on, woken up, disturbed by a loud sound, and the like, at which the user is currently interested or is likely to interact with the device. The behavior of the user after the provoking can be used as a judgment standard for the appropriateness of the provoking time, such as whether the user responds and the response attitude (positive response or negative attitude such as disgust, complaint, etc.) to investigate. If the user responds positively, such as positive evaluation or interacts with the device for a certain length of time, it is determined that the provoking time is appropriate, and if the user has a negative attitude such as disgust, complaint, or does not respond, the provoking time is inappropriate.

[0249] Specifically, based on the response attitude feature of the fifth user behavior data in the user behavior data, the evaluation result of the provoking time is determined.

[0250] Here, the response attitude feature is used to represent the response attitude exhibited by the user when the target voice interaction device provokes the interaction, such as a positive attitude, a negative attitude (disgust, complaint, etc.). The response attitude feature is determined by the user video data corresponding to the fifth user behavior data, or by the fourth user voice data corresponding to the fifth user behavior data.

[0251] In some embodiments, based on the mapping relationship between the sample fifth user behavior data in the labeled data and the sample response attitude feature labeled by the mapping relationship, the response attitude feature of the fifth user behavior data is extracted to obtain the response attitude feature. The specific acquisition process of the mapping relationship can refer to the acquisition process of other mapping relationships, which will not be described here.

[0252] In an embodiment, the fifth user behavior data includes fourth user voice data, and the evaluation result of the provoking time is determined based on the response attitude feature of the fourth user voice data. That is, the evaluation result of the provoking time is determined by the language expression of the user, such as determining the response attitude feature based on the voice content or determining the response attitude feature based on the tone feature corresponding to the audio data.

[0253] In another embodiment, the fifth user behavior data includes user video data, and the evaluation result of the provoking time is determined based on the response attitude feature of the user video data. That is, the response attitude feature is determined by the gestures, expressions (such as disgust expression, complaint expression, happy expression), body movements, and the like of the user.

[0254] In some specific embodiments, the response attitude feature is evaluated to obtain an instigation timing evaluation result based on a mapping relationship between the response attitude feature in the labeled data and the labeled instigation timing evaluation result of the sample. The specific obtaining process of the mapping relationship can refer to the obtaining process of other mapping relationships, which will not be described here.

[0255] It should be noted that the node evaluation result of the device instigation node can also be determined based on other types of evaluation, which will not be described here.

[0256] The experience evaluation method of the voice interaction device provided by the embodiment of the application determines the instigation timing evaluation result based on the fifth user behavior data in the user behavior data, so as to determine the node evaluation result of the device instigation node based on the instigation timing evaluation result, and the fifth user behavior data is the data corresponding to the device instigation node, so that the instigation timing evaluation result of the device in the device instigation node in the interaction process is evaluated, and the node evaluation of the device instigation node is performed, thereby improving the experience evaluation accuracy of the voice interaction device, and the voice interaction device is improved based on the instigation timing evaluation result, the instigation timing suitability of the voice interaction device is improved, and the user experience of the voice interaction device is improved.

[0257] The experience evaluation device of the voice interaction device provided by the application will be described below. The experience evaluation device of the voice interaction device described below can be correspondingly referred to the experience evaluation method of the voice interaction device described above.

[0258] Figure 4 The structural diagram of the experience evaluation device of the voice interaction device provided by the application is shown in Figure 4 The experience evaluation device of the voice interaction device includes:

[0259] The determination module 410 is configured to determine a target voice interaction device to be evaluated and interaction data corresponding to the target voice interaction device.

[0260] The evaluation module 420 is configured to determine a plurality of experience evaluation results of the target voice interaction device in an interaction process based on the interaction data, the plurality of experience evaluation results including node evaluation results of each interaction node in the interaction process and / or an overall evaluation result of the interaction process, and the interaction process including an interaction instigation node, a user response node, a device response node and a device guidance node.

[0261] The evaluation module 430 is configured to determine the experience evaluation result of the target voice interaction device based on the plurality of experience evaluation results.

[0262] Figure 5 An example of an entity structure diagram of an electronic device is shown inFigure 5 As shown, the electronic device can include a processor 510, a communications interface 520, a memory 530, and a communications bus 540, wherein the processor 510, the communications interface 520, and the memory 530 complete mutual communication through the communications bus 540. The processor 510 can invoke a logical instruction in the memory 530 to execute an experience evaluation method of a voice interaction device, the method including determining a target voice interaction device to be evaluated and interaction data corresponding to the target voice interaction device; determining, based on the interaction data, a plurality of experience evaluation results of the target voice interaction device in an interaction flow, the plurality of experience evaluation results including node evaluation results of each interaction node in the interaction flow and / or an overall evaluation result of the interaction flow, the interaction flow including an interaction initiation node, a user response node, a device response node, and a device guidance node; and determining, based on the plurality of experience evaluation results, an experience evaluation result of the target voice interaction device.

[0263] In addition, the logical instruction in the memory 530 described above can be implemented in the form of a software functional unit and sold or used as an independent product, and can be stored in a computer-readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the methods described in the embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various program code storage media.

[0264] In yet another aspect, the present application also provides a non-transitory computer readable storage medium having stored thereon a computer program, which, when executed by a processor, implements the experience evaluation method of the voice interaction device provided by each of the above methods, and the method comprises: determining a target voice interaction device to be evaluated, and interaction data corresponding to the target voice interaction device; determining a plurality of experience evaluation results of the target voice interaction device in an interaction process based on the interaction data, the plurality of experience evaluation results comprising node evaluation results of each interaction node in the interaction process and / or an overall evaluation result of the interaction process, and the interaction process comprising an interaction initiation node, a user response node, a device response node, and a device guidance node; and determining an experience evaluation result of the target voice interaction device based on the plurality of experience evaluation results.

[0265] The device embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the embodiment scheme according to actual needs. Those skilled in the art can understand and implement without creative labor.

[0266] From the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be realized by means of software and the necessary general hardware platform, and of course can also be realized by hardware. Based on such understanding, the above technical solutions can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in each embodiment or some parts of the embodiments.

[0267] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement to some technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for evaluating the experience of a voice interaction device, characterized in that, The method comprises: determining a target voice interaction device to be evaluated, and interaction data corresponding to the target voice interaction device; based on the interaction data, determining a plurality of experience evaluation results of the target voice interaction device in an interaction process, the plurality of experience evaluation results comprising node evaluation results of each interaction node in the interaction process and / or an overall evaluation result of the interaction process, the interaction process comprising an interaction initiation node, a user response node, a device response node and a device guidance node; based on the plurality of experience evaluation results, determining an experience evaluation result of the target voice interaction device; the interaction initiation node comprises a device initiation node, and the node evaluation result of the device initiation node comprises an initiation content evaluation result; the interaction data comprises user behavior data of a user interacting with the target voice interaction device, and the initiation content evaluation result is determined based on the following steps: based on a second emotional feature of fourth user behavior data in the user behavior data, determining the initiation content evaluation result, the fourth user behavior data being data corresponding to the device initiation node, the fourth user behavior data comprising third user voice data and / or user video data; or, the interaction data comprises interaction guidance data of the target voice interaction device, and the initiation content evaluation result is determined based on the following steps: based on environmental information in which the target voice interaction device is located, and a first initiation content of second interaction guidance data in the interaction guidance data, determining a first matching degree evaluation result of the environmental information and the first initiation content, the second interaction guidance data being data corresponding to the device initiation node; based on the first matching degree evaluation result, determining the initiation content evaluation result; or, the interaction data comprises interaction guidance data of the target voice interaction device, and the initiation content evaluation result is determined based on the following steps: based on user personal information of a user interacting with the target voice interaction device, and a second initiation content of third interaction guidance data in the interaction guidance data, determining a second matching degree evaluation result of the user personal information and the second initiation content, the third interaction guidance data being data corresponding to the device initiation node; based on the second matching degree evaluation result, determining the initiation content evaluation result. 2.The method of claim 1, wherein, The interaction data comprises voice data of the target voice interaction device, and the overall evaluation result comprises a voice content consistency evaluation result; the voice content consistency evaluation result is determined based on the following steps: determining a first voice content of first voice data in the voice data and a second voice content of second voice data in the voice data, the first voice data and the second voice data being data corresponding to the same interaction process; based on a content consistency evaluation result of the first voice content and the second voice content, determining the voice content consistency evaluation result. 3.The method of claim 1, wherein, The interaction data comprises voice data of the target voice interaction device, and the overall evaluation result comprises a voice form consistency evaluation result; The voice form consistency evaluation result is determined based on the following steps: Determine the third voice content and the tone feature of the third voice data in the voice data; Determine the voice form consistency evaluation result based on the form consistency evaluation result of the third voice content and the tone feature; Alternatively, the interaction data further includes appearance form data of the target voice interaction device, and the voice form consistency evaluation result is determined based on the following steps: Determine the fourth voice content of the fourth voice data in the voice data, and the first emotional feature of the first appearance form data in the appearance form data, the fourth voice data and the first appearance form data being data corresponding to the same interaction node; Determine the voice form consistency evaluation result based on the form consistency evaluation result of the fourth voice content and the first emotional feature. 4.The method of claim 1, wherein, The interaction data includes voice data of the target voice interaction device and interaction guide data of the target voice interaction device, and the overall evaluation result includes a voice guide consistency evaluation result; The voice guide consistency evaluation result is determined based on the following steps: Determine the fifth voice content of the fifth voice data in the voice data, and the first guide content of the first interaction guide data in the interaction guide data, the interaction node corresponding to the first interaction guide data being the next interaction node corresponding to the fifth voice data; Determine the voice guide consistency evaluation result based on the guide consistency evaluation result of the fifth voice content and the first guide content. 5.The method of claim 1, wherein, The interaction data includes user behavior data interacting with the target voice interaction device; The node evaluation result of the user response node is determined based on the following steps: Determine the understanding degree evaluation result of the user response node based on the first response content of the first user behavior data in the user behavior data, the first user behavior data being data corresponding to the user response node; Determine the node evaluation result of the user response node based on the understanding degree evaluation result. 6.The method of claim 5, wherein, The first user behavior data includes first user voice data, and the determination of the understanding degree evaluation result of the user response node based on the first response content of the first user behavior data in the user behavior data includes: Determine the understanding degree evaluation result of the user response node based on the voice content of the first user voice data. 7.The method of claim 1, wherein, The interaction data includes user behavior data interacting with the target voice interaction device and voice data of the target voice interaction device; The node evaluation result of the device response node is determined based on the following steps: Determine the second response content of the second user behavior data in the user behavior data, and the sixth voice content of the sixth voice data in the voice data, the interaction node corresponding to the sixth voice data being the next interaction node corresponding to the second user behavior data; Determine the response correlation degree evaluation result of the device response node based on the correlation degree of the second response content and the sixth voice content; Based on the response correlation degree evaluation result, a node evaluation result of the device response node is determined. 8.The method of claim 1, wherein, The interaction data includes user behavior data of interaction with the target voice interaction device; The node evaluation result of the device guiding node is determined based on the following steps: Based on a degree of interest feature of third user behavior data in the user behavior data, a guiding reasonableness evaluation result of the device guiding node is determined, the third user behavior data being data corresponding to the device guiding node, and the third user behavior data including second user voice data and / or user video data; Based on the guiding reasonableness evaluation result, the node evaluation result of the device guiding node is determined. 9.The method of claim 1, wherein, The interaction data includes user behavior data of interaction with the target voice interaction device, and the interaction initiation node includes a device initiation node, and the node evaluation result of the device initiation node includes an initiation timing evaluation result; The initiation timing evaluation result is determined based on the following steps: Based on fifth user behavior data in the user behavior data, the initiation timing evaluation result is determined; The fifth user behavior data is data corresponding to the device initiation node, and the fifth user behavior data includes fourth user voice data and / or user video data.

10. An experience evaluation apparatus of a voice interaction apparatus, characterized by, Comprise: A determination module is configured to determine a target voice interaction device to be evaluated and interaction data corresponding to the target voice interaction device; An evaluation module is configured to determine, based on the interaction data, a plurality of experience evaluation results of the target voice interaction device in an interaction process, the plurality of experience evaluation results including node evaluation results of each interaction node in the interaction process and / or an overall evaluation result of the interaction process, and the interaction process including an interaction initiation node, a user response node, a device response node, and a device guiding node; An evaluation module is configured to determine, based on the plurality of experience evaluation results, an experience evaluation result of the target voice interaction device; The interaction initiation node includes a device initiation node, and the node evaluation result of the device initiation node includes an initiation content evaluation result; The interaction data includes user behavior data of interaction with the target voice interaction device, and the initiation content evaluation result is determined based on the following manner: Based on a second emotion feature of fourth user behavior data in the user behavior data, the initiation content evaluation result is determined, the fourth user behavior data being data corresponding to the device initiation node, and the fourth user behavior data including third user voice data and / or user video data; Alternatively, the interaction data includes interaction guiding data of the target voice interaction device, and the initiation content evaluation result is determined based on the following manner: Based on environmental information in which the target voice interaction device is located and a first initiation content of second interaction guiding data in the interaction guiding data, a first matching degree evaluation result of the environmental information and the first initiation content is determined, the second interaction guiding data being data corresponding to the device initiation node; Based on the first matching degree evaluation result, the initiation content evaluation result is determined. Or, the interaction data includes interaction guide data of the target voice interaction device, and the provoking content evaluation result is determined based on the following manner: determine a second matching degree evaluation result of the user personal information and second provoking content based on user personal information of a user interacting with the target voice interaction device and the third interaction guide data in the interaction guide data, the third interaction guide data being data corresponding to the device provoking node; determine the provoking content evaluation result based on the second matching degree evaluation result.

11. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor implements the experience evaluation method of the voice interaction device according to any one of claims 1 to 9 when executing the program. 12.A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the experience evaluation method of the voice interaction device according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Speech interaction satisfaction determination method and device

    CN108388926A

  • Speech recognition evaluation method and device, storage medium and equipment

    CN111681642A