Intelligent voice quality evaluation method and device for multi-modal intelligent terminal

By analyzing the Mel-frequency cepstral coefficients and text content of input and output speech in multimodal intelligent terminals, and combining clustering algorithms to evaluate emotion and text suitability, the subjective bias problem in the evaluation of intelligent speech quality in multimodal intelligent terminals is solved, and more accurate emotion matching and consistent interaction are achieved.

CN120913599BActive Publication Date: 2026-01-27ROPEOK TECHNOLOGY GROUP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511430876.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-09
Publication Date
2026-01-27
Estimated Expiration
2045-10-09

AI Technical Summary

Technical Problem

Existing methods for evaluating the quality of intelligent voice in multimodal intelligent terminals rely on human ratings, which are subject to subjective bias, leading to inconsistent evaluation results and making it difficult to accurately determine the emotional matching degree between the output voice and the user's input voice.

Method used

By acquiring the input and output speech of the same user from the database of multimodal smart terminals, the richness and similarity of emotional expression are analyzed using Mel frequency cepstral coefficients. Combined with clustering algorithms and speech-text content analysis, the emotional adaptability and text content adaptability of the output speech are evaluated, and a smart speech quality evaluation value is generated.

Benefits of technology

It enables accurate evaluation of the intelligent voice quality of multimodal intelligent terminals, ensuring that the output voice matches the emotion of the user's input voice, improving the naturalness and consistency of interaction, and providing a stable and predictable user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913599B_ABST
    Figure CN120913599B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of speech analysis, and specifically relates to a multi-modal intelligent terminal intelligent speech quality evaluation method and device, comprising: obtaining all input speech of a same user and output speech of a multi-modal intelligent terminal corresponding to each input speech in a database of the multi-modal intelligent terminal; obtaining output speech emotion richness adaptability according to a comparison of emotion expression richness in the input speech and the output speech; obtaining an output speech emotion expression ability value in combination with emotion expression similarity in the output speech corresponding to the input speech with similar emotion expression; and obtaining text content dissimilarity in the output speech corresponding to the input speech with different emotion expression of similar text content, to obtain a multi-modal intelligent terminal intelligent speech quality evaluation value, so as to judge the intelligent speech quality of the multi-modal intelligent terminal. The present application can obtain an accurate multi-modal intelligent terminal intelligent speech quality evaluation result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speech analysis technology, specifically to a method and apparatus for evaluating the intelligent speech quality of multimodal intelligent terminals. Background Technology

[0002] Multimodal smart terminals refer to intelligent devices that can communicate with users through multiple interaction methods (such as voice, touch, and vision). These devices typically integrate various sensors and input / output interfaces to allow users to interact in different ways. For example, smartphones, smartwatches, and smart home devices all fall under the category of multimodal smart terminals.

[0003] Ensuring that the emotional tone of the output voice of a multimodal smart terminal matches the emotional tone of the user's input voice, thereby enhancing the naturalness and friendliness of the interaction, is one of the important indicators for evaluating the intelligent voice quality of multimodal smart terminals. Matching the user's emotions can enable the output voice of multimodal smart terminals to provide a more humanized, efficient, and pleasant interactive experience, thereby enhancing users' trust and satisfaction with the smart terminal.

[0004] Existing problems: Different individuals may perceive and express the same emotion differently. Existing subjective evaluation methods often rely on the judgment of human raters, which may be subject to subjective bias, leading to inconsistent evaluation results. Furthermore, human emotional expression is diverse; the same emotion may be expressed in different ways. This increases the difficulty of judging whether the emotion of the output speech of a multimodal intelligent terminal matches the emotion of the user's input speech, easily leading to inaccurate evaluation results of the intelligent speech quality of multimodal intelligent terminals. Summary of the Invention

[0005] This invention provides a method and apparatus for evaluating the intelligent voice quality of multimodal intelligent terminals to solve existing problems.

[0006] The intelligent voice quality evaluation method and apparatus for multimodal intelligent terminals of the present invention adopts the following technical solution:

[0007] One embodiment of the present invention provides a method for evaluating the intelligent voice quality of a multimodal intelligent terminal, the method comprising the following steps:

[0008] In the database of multimodal smart terminals, obtain all input voices of the same user, as well as the output voices of the multimodal smart terminals corresponding to each input voice.

[0009] The emotional richness adaptability of the output speech is obtained by comparing the richness of emotional expression in the input speech and the output speech.

[0010] Based on the similarity of emotional expression in the output speech corresponding to the input speech with similar emotional expression, and combined with the emotional richness and adaptability of the output speech, the emotional expression capability value of the output speech is obtained.

[0011] Based on the dissimilarity of text content in the output speech corresponding to input speech with different emotional expressions of similar text content, and combined with the emotional expression capability value of the output speech, an intelligent speech quality evaluation value for the multimodal intelligent terminal is obtained; the intelligent speech quality of the multimodal intelligent terminal is judged based on the magnitude of the intelligent speech quality evaluation value.

[0012] Furthermore, the specific steps for obtaining the emotional richness and adaptability of the output speech are as follows:

[0013] The sum of the Euclidean distances between the Mel-frequency cepstral coefficients of each input speech and the Mel-frequency cepstral coefficients of all other input speech is denoted as the emotional extreme of each input speech.

[0014] Arrange all input speech in ascending order of emotional extremes to obtain the input speech sequence;

[0015] Preset quantity threshold In the input speech sequence, the preceding... The first set of input speech is composed of several input voices. The first set of input speech constitutes the second set of input speech. Each input speech is used to form a third set of input speech, and so on, until all input speech is used to form a final set of input speech, resulting in a sequence of input speech sets.

[0016] In the input speech set sequence, based on the Euclidean distance between the Mel frequency cepstral coefficients of the input speech in each input speech set and the Euclidean distance between the Mel frequency cepstral coefficients of the corresponding output speech, the input speech emotion expression richness sequence and the output speech emotion expression richness sequence are obtained.

[0017] Based on the input speech emotion expression richness sequence and the output speech emotion expression richness sequence, obtain the output speech emotion richness adaptability.

[0018] Furthermore, the specific steps for obtaining the input speech emotion expression richness sequence and the output speech emotion expression richness sequence are as follows:

[0019] Obtain the output speech set consisting of the output speech corresponding to all input speech in each input speech set, and then sequentially construct the output speech set sequence by combining the output speech sets corresponding to all input speech sets in the input speech set sequence.

[0020] In the sequence of input speech sets, the mean of the Euclidean distance of the cepstral coefficients of all any two input speech sets is obtained and denoted as the richness of emotion expression of each input speech set. The richness of emotion expression of all input speech sets are then used to form a sequence of richness of emotion expression of input speech.

[0021] In the output speech set sequence, the mean of the Euclidean distance of the Mel frequency cepstral coefficients of any two output speech in each output speech set is obtained and denoted as the emotional expression richness of each output speech set. The emotional expression richness of all output speech sets are then used to construct the output speech emotional expression richness sequence.

[0022] Furthermore, the specific steps for obtaining the output speech emotion richness adaptability based on the input speech emotion richness sequence and the output speech emotion richness sequence are as follows:

[0023] Obtain the normalized value of the difference between the last element of the output speech emotion expression richness sequence and the last element of the input speech emotion expression richness sequence, and denote it as the first difference value;

[0024] Obtain the normalized value of the Pearson correlation coefficient between the input speech emotion expression richness sequence and the output speech emotion expression richness sequence, and denote it as the first similarity value;

[0025] The mean of the first difference value and the first similarity value is denoted as the output speech emotion richness and adaptability.

[0026] Furthermore, the specific steps for obtaining the output speech emotion expression ability value are as follows:

[0027] The Euclidean distance between the Mel-frequency cepstral coefficients of any two input speech samples is used as the clustering distance. Clustering is performed on all input speech samples to obtain several clusters, and the contour coefficients of each cluster are obtained.

[0028] In each cluster, the output emotion consistency corresponding to each cluster is determined based on the Euclidean distance of the Mel frequency cepstral coefficients of the output speech in all the input speech.

[0029] The ratio of the silhouette coefficient of each cluster to the sum of the silhouette coefficients of all clusters is obtained and recorded as the weight of each cluster. The output emotion consistency corresponding to all clusters is weighted and summed to obtain the expression consistency of output speech emotion under the same input speech emotion.

[0030] The mean value of the richness and adaptability of the output speech emotion and the consistency of the output speech emotion expression under the same input speech emotion is recorded as the output speech emotion expression ability value.

[0031] Furthermore, the specific steps for determining the output sentiment consistency corresponding to each cluster are as follows:

[0032] In each cluster, among all the output speech corresponding to all input speech, the Euclidean distance between any two output speech frequencies is used as the clustering distance. Clustering is performed on all output speech to obtain several new clusters. The ratio of the number of output speech in the new cluster with the most output speech to the total number of output speech is obtained and recorded as the first ratio. The product of the reciprocal of the number of new clusters and the first ratio is normalized and recorded as the first output emotion uniformity. The inverse normalized value of the mean of the Euclidean distance between any two output speech frequencies is obtained and recorded as the second output emotion uniformity.

[0033] The mean of the first output sentiment consistency and the second output sentiment consistency is denoted as the output sentiment consistency corresponding to each cluster.

[0034] Furthermore, the specific steps for obtaining the intelligent voice quality evaluation value of the multimodal intelligent terminal are as follows:

[0035] The inversely proportional normalized value of the similarity between the text content of any two input speech samples is used as the clustering distance. Clustering operations are performed on all input speech samples to obtain several text clusters.

[0036] In each text cluster, the maximum value of the Euclidean distance between the Mel frequency cepstral coefficients of all two arbitrary input speech samples is obtained and recorded as the emotion fluctuation value of each text cluster. The two input speech samples corresponding to the maximum value are recorded as the target input speech.

[0037] Text clusters whose normalized emotional fluctuation values ​​are greater than a preset fluctuation threshold are denoted as target text clusters;

[0038] Within each target text cluster, the differences in output speech text representation for each target text cluster are determined based on the dissimilarity of the text content of the output speech corresponding to the two target input speech and the Euclidean distance of the Mel frequency cepstral coefficients.

[0039] The mean value of the output speech-text expression differences of all target text clusters is denoted as the output speech-text content fit.

[0040] The average of the output speech emotion expression ability value and the adaptability of the output speech text content is obtained and recorded as the intelligent speech quality evaluation value of the multimodal intelligent terminal.

[0041] Furthermore, the specific steps for determining the output speech-text representation differences of each target text cluster are as follows:

[0042] In each target text cluster, the inversely proportional normalized value of the similarity between the text content of the output speech corresponding to two target input speech is obtained, which is denoted as the first dissimilarity. The normalized value of the Euclidean distance between the Mel frequency cepstral coefficients of the output speech corresponding to two target input speech is obtained, which is denoted as the second dissimilarity. The mean of the first dissimilarity and the second dissimilarity is denoted as the difference in output speech text expression for each target text cluster.

[0043] Furthermore, the specific steps for judging the intelligent voice quality of the multimodal intelligent terminal are as follows:

[0044] When the intelligent voice quality evaluation value of the multimodal intelligent terminal is greater than the preset evaluation threshold, the intelligent voice quality of the multimodal intelligent terminal is judged to be good.

[0045] When the intelligent voice quality evaluation value of the multimodal intelligent terminal is less than or equal to the preset evaluation threshold, the intelligent voice quality of the multimodal intelligent terminal is judged to be average.

[0046] The present invention also proposes an intelligent voice quality evaluation device for a multimodal intelligent terminal, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor executes the computer program stored in the memory to implement the steps of the aforementioned intelligent voice quality evaluation method for a multimodal intelligent terminal.

[0047] The beneficial effects of the technical solution of the present invention are:

[0048] In this embodiment of the invention, all input voices of the same user and the output voices of the multimodal smart terminal corresponding to each input voice are obtained from the database of the multimodal smart terminal. Based on a comparison of the richness of emotional expression in the input and output voices, the emotional richness adaptability of the output voice is obtained. This involves analyzing whether the emotional changes in the voice generated by the smart terminal become more diverse as the emotional expression in the user's voice gradually becomes richer, making the interaction of the smart terminal more natural and human-like, and better adapting to the user's emotional changes, demonstrating the good quality of the intelligent voice of the multimodal smart terminal. Combining the similarity of emotional expression in the output voices corresponding to input voices with similar emotional expressions, the emotional expression capability value of the output voice is obtained. This ensures the consistency and predictability of the voice interaction of the multimodal smart terminal by analyzing the emotional similarity of the output voices of different input voices with similar emotions, allowing users to obtain a stable and expected interactive experience, thus demonstrating the good quality of the intelligent voice of the multimodal smart terminal. By combining the dissimilarity of text content in the output speech corresponding to input speech with different emotional expressions of similar text content, an intelligent speech quality evaluation value for the multimodal intelligent terminal is obtained. This value is used to judge the intelligent speech quality of the multimodal intelligent terminal. Furthermore, by analyzing the diversity of responses generated by the multimodal intelligent terminal that match the user's emotions when the terminal receives the same text content but user speech under different emotions, the high quality of the multimodal intelligent terminal is further demonstrated. Thus, this invention can obtain accurate intelligent speech quality evaluation results for multimodal intelligent terminals. Attached Figure Description

[0049] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0050] Figure 1 This is a flowchart illustrating the steps of the intelligent voice quality evaluation method for a multimodal intelligent terminal according to the present invention.

[0051] Figure 2 This is a schematic diagram of the time-domain waveform of a speech signal. Detailed Implementation

[0052] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effects of the intelligent voice quality evaluation method and apparatus for multimodal intelligent terminals proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.

[0053] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0054] The following description, in conjunction with the accompanying drawings, details the specific scheme of the intelligent voice quality evaluation method and device for multimodal intelligent terminals provided by the present invention.

[0055] Please see Figure 1 The diagram illustrates a flowchart of a method for evaluating the intelligent voice quality of a multimodal intelligent terminal according to an embodiment of the present invention. The method includes the following steps:

[0056] Step S001: In the database of the multimodal smart terminal, obtain all input voices of the same user, as well as the output voices of the multimodal smart terminal corresponding to each input voice.

[0057] It should be noted that for user input voice under different emotions, the synthesized voice output by the multimodal smart terminal should match or be appropriately adjusted to the emotion input by the user to provide a more human and appropriate interactive experience. For example, if the user expresses a happy emotion, the smart terminal's response should also have a positive and cheerful tone. When the user is more agitated or negative, the output voice should be appropriately soothed to avoid exacerbating the user's emotional fluctuations and to help the user calm down.

[0058] In the database of any multimodal smart terminal, retrieve all input voices of the same user, as well as the output voices of the multimodal smart terminal corresponding to each input voice.

[0059] It should be noted that the multimodal intelligent terminal in this embodiment has a speaker recognition function. This is a technology that can identify and distinguish different speakers by analyzing the acoustic features of the speech signal, such as pitch, timbre, and intonation, to identify the speaker's identity. Therefore, all input speech from the same user can be obtained. In this embodiment, the multimodal intelligent terminal is a smartphone, and the user is the owner. This is used as an example for description. By analyzing the emotional matching between the smartphone's output speech and the owner's input speech, the intelligent speech quality evaluation result of the multimodal intelligent terminal is obtained. The input speech undergoes wavelet transform denoising, which is a well-known technique, and the specific method will not be described here. A schematic diagram of the time-domain waveform of the speech signal is shown below. Figure 2 As shown, Figure 2 The horizontal axis represents time, in seconds (s), and the vertical axis represents amplitude.

[0060] Step S002: Based on the comparison of the richness of emotional expression in the input speech and the output speech, obtain the emotional richness adaptability of the output speech.

[0061] Preferably, in one embodiment of the present invention, the method for obtaining the emotional richness and adaptability of the output voice includes:

[0062] Obtain the Mel-frequency cepstral coefficients for each input speech, and then obtain the Mel-frequency cepstral coefficients for each output speech.

[0063] It should be noted that obtaining Mel-frequency cepstral coefficients (MFCCs) is a well-known technique, and the specific methods will not be described here. MFCCs can reflect emotional information in speech because they capture the spectral characteristics of the speech signal, which are related to emotional state. Emotion typically affects the pitch, volume, rhythm, and timbre of speech, all of which are features that MFCCs can reflect. MFCCs are a set of data values ​​that together constitute a feature vector.

[0064] Obtain the Euclidean distance between the Mel-frequency cepstral coefficients of any two input speech samples. Sum the Euclidean distances between each input speech sample and the Mel-frequency cepstral coefficients of all other input speech samples, and denote the emotional extreme of each input speech sample.

[0065] Arrange all input speech in ascending order of emotional extremes to obtain the input speech sequence.

[0066] It should be noted that the greater the Euclidean distance between the Mel frequency cepstral coefficients of different speech, the greater the difference in emotional expression between the speech. During the interaction between the user and the multimodal intelligent terminal, the emotional expression of most input speech should be stable and similar. When the emotion of a certain input speech differs greatly from that of other input speech, it may indicate that the input speech expresses a special or abnormal emotional state. For example: (1) Emotional fluctuation: Individuals may experience emotional fluctuations at different times or in different situations, and the emotion at a certain moment may be different from the usual emotional state. (2) Sudden event: A specific event may trigger a strong emotional reaction, causing the emotion of the speech to be different from the usual one. (3) Stress: When facing stress or challenges, individuals may show emotions different from the usual one, such as anxiety, anger or fear. (4) Stress response: Sudden stress events may cause a rapid change in the individual's emotional state, which is reflected in the speech.

[0067] Preset quantity threshold Let's take 10 as an example.

[0068] In the input speech sequence, the first The first set of input speech is composed of several input voices. The first set of input speech constitutes the second set of input speech. Each input speech is used to form a third set of input speech, and so on, until all input speech is combined to form a final set of input speech, resulting in a sequence of input speech sets.

[0069] Specifically, as the number of input speech samples with greater emotional extremes gradually increases, the emotional richness of the speech samples in the input speech set should gradually increase.

[0070] In the sequence of input speech sets, obtain the output speech set consisting of all the output speech corresponding to each input speech set, and then sequentially form the output speech set sequence of all the input speech sets.

[0071] In the sequence of input speech sets, the mean of the Euclidean distance of the cepstral coefficients of all any two input speech sets is obtained and denoted as the richness of emotion expression for each input speech set. The richness of emotion expression for all input speech sets is then used to construct a sequence of richness of emotion expression for input speech.

[0072] In the output speech set sequence, the mean of the Euclidean distance of the Mel frequency cepstral coefficients of any two output speech in each output speech set is obtained and denoted as the emotional expression richness of each output speech set. The emotional expression richness of all output speech sets are then used to construct the output speech emotional expression richness sequence.

[0073] The greater the Euclidean distance between the Mel-frequency cepstral coefficients of different speech sounds in the speech set, the richer the emotional expression of the speech sounds in the speech set.

[0074] Obtain the difference between the last element of the output speech emotion richness sequence and the last element of the input speech emotion richness sequence. The normalized value is denoted as the first difference value. The Pearson correlation coefficient between the input speech emotion richness sequence and the output speech emotion richness sequence is obtained. The normalized value is denoted as the first similarity value. The mean of the first difference value and the first similarity value is denoted as the output speech emotion richness and adaptability.

[0075] It should be noted that the Pearson correlation coefficient is a well-known technique, and the specific method will not be described here. The larger the Pearson correlation coefficient, the closer the changes in the data in the two sequences are to a positive correlation. and As respectively and The normalized value, where, This is a linear normalization function used to normalize data values ​​to between 0 and 1. During interaction between the user and the multimodal intelligent terminal, as the emotional expression in the user's voice becomes richer, the emotional changes in the voice generated by the intelligent terminal should also become more diverse. This design makes the interaction of the intelligent terminal more natural and human-like, better adapting to the user's emotional changes, thereby improving the user experience. Therefore, the richness of the output voice's emotional expression should increase with the increase of the richness of the input voice's emotional expression; that is, the larger the first similarity value, the more the output voice of the multimodal intelligent terminal is adapted to the user's increasingly rich emotional expression. Furthermore, when the richness of the output voice's emotional expression approaches or even exceeds the richness of the input voice's emotional expression, it indicates that the output voice of the multimodal intelligent terminal can better adapt to the user's emotional changes. Therefore, the larger the first difference value and the first similarity value, the better the emotional richness and adaptability of the output voice.

[0076] Step S003: Based on the similarity of emotional expression in the output speech corresponding to the input speech with similar emotional expression, and combined with the emotional richness and adaptability of the output speech, obtain the emotional expression ability value of the output speech.

[0077] It should be noted that the above analysis considered whether all output voices should also be emotionally rich when all input voices are emotionally rich, in order to better adapt to the user's emotional changes. That is, assuming the output voices are emotionally rich and well-adapted, it is further necessary to analyze whether the pairing of the emotion of each input voice with the emotion of its corresponding output voice is appropriate. For different input voices with similar emotions, the emotions of the output voices of the multimodal smart terminal should be similar. This is because when designing the voice interaction system of a multimodal smart terminal, consistency and predictability are usually pursued to ensure that users obtain a stable and expected interactive experience. If users receive similar responses under similar emotional states, this helps to establish user expectations and trust in the device's behavior. Users can more easily understand and predict the device's reactions, thereby improving overall user satisfaction. Furthermore, consistent responses help maintain the stability and reliability of the system, reducing confusion and frustration caused by emotion recognition errors or inconsistent responses.

[0078] Preferably, in one embodiment of the present invention, the method for obtaining the output voice emotion expression ability value includes:

[0079] Using the Euclidean distance between the Mel-frequency cepstral coefficients of any two input speech samples as the clustering distance, the K-means clustering algorithm is used to cluster all input speech samples, resulting in several clusters, and the silhouette coefficient of each cluster is obtained.

[0080] The silhouette coefficient is a metric used to evaluate clustering effectiveness. Its value ranges from -1 to 1, with a higher value indicating better clustering. The elbow method is used to obtain the K-value for the K-means clustering algorithm. The elbow method, the K-means clustering algorithm, and the acquisition of the silhouette coefficient are all well-known techniques, and their specific methods will not be described here. If a cluster contains only one output speech, the following analysis will not be performed.

[0081] In each cluster, among all the output speech corresponding to all input speech, the Euclidean distance between any two output speech's Mel-frequency cepstral coefficients is used as the clustering distance. The K-means clustering algorithm is then used to cluster all output speech, resulting in several new clusters. The ratio of the number of output speech in the new cluster with the largest number of output speech to the total number of output speech (the number of output speech corresponding to all input speech in each cluster) is recorded as the first ratio. The product of the reciprocal of the number of new clusters and the first ratio is then calculated. The normalized value is denoted as the first output emotional uniformity.

[0082] It should be noted that: with As The normalized value is calculated as follows: Each new cluster represents a class of output speech emotions. The emotions should be consistent across all output speech corresponding to each cluster. Therefore, the more new clusters there are, the less consistent the emotions of the output speech. The new cluster with the most output speech can represent the output speech emotion corresponding to the input speech emotion of that cluster. Therefore, the more output speech there is in the new cluster with the most output speech, the less consistent the emotions of the output speech. Thus, the larger the ratio of the reciprocal of the number of new clusters to the first value, the more consistent the output emotions.

[0083] In each cluster, among all the output speech corresponding to all input speech, obtain the mean of the Euclidean distance of the Mel-frequency cepstral coefficients of all any two output speech samples. The inversely proportional normalized value, the second output emotion uniformity.

[0084] Among them, with As The inverse proportional normalized value. That is... The smaller the value, the more consistent the emotional expression of all output voices corresponding to each cluster.

[0085] The mean of the first output sentiment consistency and the second output sentiment consistency is denoted as the output sentiment consistency corresponding to each cluster.

[0086] The sum of the silhouette coefficients of all clusters is obtained. The ratio of the silhouette coefficient of each cluster to the sum of the silhouette coefficients of all clusters is recorded as the weight of each cluster. The output emotion consistency corresponding to all clusters is weighted and summed to obtain the expression consistency of output speech emotion under the same input speech emotion.

[0087] It should be noted that: the larger the silhouette coefficient of each cluster, the more consistent the emotion of the input speech within the cluster, and therefore the more consistent the emotion of the output speech corresponding to that cluster should be. Therefore, the silhouette coefficient is used as the weight to perform a weighted summation of the output emotion consistency, yielding the consistency of the output speech emotion expression under the same input speech emotion. Since the weight of all clusters is 1, and the value of the output emotion consistency corresponding to each cluster is between 0 and 1, the value of the output speech emotion expression consistency under the same input speech emotion is also between 0 and 1.

[0088] The mean value of the richness and adaptability of the output speech emotion and the consistency of the output speech emotion expression under the same input speech emotion is recorded as the output speech emotion expression ability value.

[0089] Step S004: Based on the dissimilarity of text content in the output speech corresponding to input speech with different emotional expressions of similar text content, and combined with the emotional expression ability value of the output speech, obtain the intelligent speech quality evaluation value of the multimodal intelligent terminal; judge the intelligent speech quality of the multimodal intelligent terminal based on the magnitude of the intelligent speech quality evaluation value of the multimodal intelligent terminal.

[0090] It should be noted that the above analysis focused on the emotional expression capabilities of the output voice of multimodal intelligent terminals. When a multimodal intelligent terminal receives the same text content but user voice messages with different emotions, it should be able to recognize these emotional differences and generate a response that matches the user's emotion. For example, if a user's voice message is "I received a promotion notification today," in an excited tone, the multimodal intelligent terminal's output voice message might be "Congratulations! This is great news; your hard work and talent have been recognized." However, in a tense tone, the output voice message might be "A promotion brings new responsibilities, which can be nerve-wracking. But remember, this is also an opportunity to showcase your abilities. Are you ready?" Therefore, further analysis combining the voice message with the text content is needed to determine the intelligent voice quality evaluation results of the multimodal intelligent terminal.

[0091] Preferably, in one embodiment of the present invention, the method for obtaining the intelligent voice quality evaluation value of a multimodal intelligent terminal includes:

[0092] Using speech recognition technology, the text content of each input speech and the text content of each output speech are obtained.

[0093] Among them, speech recognition technology is a well-known technology, and the specific methods will not be introduced here.

[0094] Using the inversely proportional normalized value of the similarity between the text content of any two input speech samples as the clustering distance, the K-means clustering algorithm is applied to cluster all input speech samples, resulting in several text clusters. Within each text cluster, the text content of the input speech samples is similar.

[0095] It should be noted that in this embodiment, the text content is first converted into TF-IDF vectors, and then the cosine similarity between the two vectors is calculated as the similarity between the two text contents. This is a well-known technique, and the specific method will not be described here. Since the cosine similarity ranges from -1 to 1, the difference in similarity between the two text contents is subtracted from 1 and then divided by 2 to obtain the inversely proportional normalized value of the similarity between the two text contents. If there is only one input speech in a certain text cluster, the following analysis will not be performed.

[0096] Within each text cluster, the maximum value of the Euclidean distance between the Mel frequency cepstral coefficients of any two input speech samples is obtained and recorded as the emotion fluctuation value of each text cluster. The two input speech samples corresponding to the maximum value are both recorded as the target input speech.

[0097] The preset fluctuation threshold is 0.8, and we will use this as an example for explanation.

[0098] Obtain the sentiment fluctuation value for each text cluster. The normalized value of the emotional fluctuation value is used to identify the text clusters whose normalized emotional fluctuation value is greater than the preset fluctuation threshold. These clusters are then denoted as the target text clusters.

[0099] Among them, with As The normalized value. The target text cluster contains input speech with similar text content but significant differences in emotion.

[0100] Within each target text cluster, the inversely proportional normalized value of the similarity between the text content of the output speech corresponding to two target input speech samples is obtained, denoted as the first dissimilarity. The Euclidean distance of the Mel-frequency cepstral coefficients of the output speech corresponding to two target input speech samples is also obtained. The normalized value is denoted as the second dissimilarity. The mean of the first dissimilarity and the second dissimilarity is denoted as the output speech-text expression difference of each target text cluster.

[0101] It should be noted that: with As The normalized value indicates that when the text content and emotion of the output speech corresponding to two target input speech in the target text cluster are less similar, it means that for user speech with the same text content but different emotions, the multimodal intelligent terminal can generate response content that matches the user's emotion.

[0102] The mean of the differences in output speech-text expression among all target text clusters is obtained and denoted as the output speech-text content fit.

[0103] The average of the output speech emotion expression ability value and the adaptability of the output speech text content is obtained and recorded as the intelligent speech quality evaluation value of the multimodal intelligent terminal.

[0104] It should be noted that the stronger the ability of the output voice to express emotions and the stronger the adaptability of the output voice to the text content, the better the intelligent voice quality of the multimodal intelligent terminal.

[0105] The preset evaluation threshold is 0.6, and this will be used as an example for explanation.

[0106] When the intelligent voice quality evaluation value of the multimodal intelligent terminal is greater than the preset evaluation threshold, the intelligent voice quality of the multimodal intelligent terminal is judged to be good.

[0107] When the intelligent voice quality evaluation value of the multimodal intelligent terminal is less than or equal to the preset evaluation threshold, the intelligent voice quality of the multimodal intelligent terminal is judged to be average.

[0108] The present invention also provides an intelligent voice quality evaluation device for a multimodal intelligent terminal, including a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor executes the computer program stored in the memory to implement the steps of the aforementioned intelligent voice quality evaluation method for a multimodal intelligent terminal.

[0109] This invention is now complete.

[0110] In summary, in this embodiment of the invention, all input voices of the same user and the output voices of the multimodal intelligent terminal corresponding to each input voice are obtained from the database of the multimodal intelligent terminal. Based on the comparison of the richness of emotional expression in the input and output voices, the emotional richness adaptability of the output voice is obtained. Combined with the similarity of emotional expression in the output voices corresponding to input voices with similar emotional expressions, the emotional expression capability value of the output voice is obtained. Furthermore, combined with the dissimilarity of text content in the output voices corresponding to input voices with different emotional expressions of similar text content, the intelligent voice quality evaluation value of the multimodal intelligent terminal is obtained, which is used to judge the intelligent voice quality of the multimodal intelligent terminal. This invention can obtain accurate intelligent voice quality evaluation results for multimodal intelligent terminals.

[0111] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for evaluating the intelligent voice quality of a multimodal intelligent terminal, characterized in that, The method includes the following steps: In the database of multimodal smart terminals, obtain all input voices of the same user, as well as the output voices of the multimodal smart terminals corresponding to each input voice. The emotional richness adaptability of the output speech is obtained by comparing the richness of emotional expression in the input speech and the output speech. Based on the similarity of emotional expression in the output speech corresponding to the input speech with similar emotional expression, and combined with the emotional richness and adaptability of the output speech, the emotional expression capability value of the output speech is obtained. Based on the dissimilarity of text content in the output speech corresponding to input speech with different emotional expressions of similar text content, and combined with the emotional expression capability value of the output speech, an intelligent speech quality evaluation value of the multimodal intelligent terminal is obtained; the intelligent speech quality of the multimodal intelligent terminal is judged based on the magnitude of the intelligent speech quality evaluation value. The specific steps involved in obtaining the emotional richness and adaptability of the output speech are as follows: The sum of the Euclidean distances between the Mel-frequency cepstral coefficients of each input speech and the Mel-frequency cepstral coefficients of all other input speech is denoted as the emotional extreme of each input speech. Arrange all input speech in ascending order of emotional extremes to obtain the input speech sequence; Preset quantity threshold In the input speech sequence, the preceding... The first set of input speech is composed of several input voices. The first set of input speech constitutes the second set of input speech. Each input speech is used to form a third set of input speech, and so on, until all input speech is used to form the last set of input speech, resulting in a sequence of input speech sets; In the input speech set sequence, based on the Euclidean distance between the Mel frequency cepstral coefficients of the input speech in each input speech set and the Euclidean distance between the Mel frequency cepstral coefficients of the corresponding output speech, the input speech emotion expression richness sequence and the output speech emotion expression richness sequence are obtained. Based on the input speech emotion expression richness sequence and the output speech emotion expression richness sequence, obtain the output speech emotion richness adaptability.

2. The intelligent voice quality evaluation method for a multimodal intelligent terminal according to claim 1, characterized in that, The specific steps for obtaining the input speech emotion expression richness sequence and the output speech emotion expression richness sequence are as follows: Obtain the output speech set consisting of the output speech corresponding to all input speech in each input speech set, and then sequentially construct the output speech set sequence by combining the output speech sets corresponding to all input speech sets in the input speech set sequence. In the sequence of input speech sets, the mean of the Euclidean distance of the cepstral coefficients of all any two input speech sets is obtained and denoted as the richness of emotion expression of each input speech set. The richness of emotion expression of all input speech sets are then used to form a sequence of richness of emotion expression of input speech. In the output speech set sequence, the mean of the Euclidean distance of the Mel frequency cepstral coefficients of any two output speech in each output speech set is obtained and denoted as the emotional expression richness of each output speech set. The emotional expression richness of all output speech sets are then used to construct the output speech emotional expression richness sequence.

3. The intelligent voice quality evaluation method for a multimodal intelligent terminal according to claim 1, characterized in that, The specific steps for obtaining the output speech emotion richness adaptability based on the input speech emotion richness sequence and the output speech emotion richness sequence are as follows: Obtain the normalized value of the difference between the last element of the output speech emotion expression richness sequence and the last element of the input speech emotion expression richness sequence, and denote it as the first difference value; Obtain the normalized value of the Pearson correlation coefficient between the input speech emotion expression richness sequence and the output speech emotion expression richness sequence, and denote it as the first similarity value; The mean of the first difference value and the first similarity value is denoted as the output speech emotion richness and adaptability.

4. The intelligent voice quality evaluation method for a multimodal intelligent terminal according to claim 1, characterized in that, The specific steps for obtaining the output speech emotion expression ability value are as follows: The Euclidean distance between the Mel-frequency cepstral coefficients of any two input speech samples is used as the clustering distance. Clustering is performed on all input speech samples to obtain several clusters, and the contour coefficients of each cluster are obtained. In each cluster, the output emotion consistency corresponding to each cluster is determined based on the Euclidean distance of the Mel frequency cepstral coefficients of the output speech in all the input speech. The ratio of the silhouette coefficient of each cluster to the sum of the silhouette coefficients of all clusters is obtained and recorded as the weight of each cluster. The output emotion consistency corresponding to all clusters is weighted and summed to obtain the expression consistency of output speech emotion under the same input speech emotion. The mean value of the richness and adaptability of the output speech emotion and the consistency of the output speech emotion expression under the same input speech emotion is recorded as the output speech emotion expression ability value.

5. The intelligent voice quality evaluation method for a multimodal intelligent terminal according to claim 4, characterized in that, The specific steps involved in determining the output sentiment consistency for each cluster are as follows: In each cluster, among all the output speech corresponding to all input speech, the Euclidean distance between any two output speech at the Mel frequency cepstral coefficients is used as the clustering distance. Clustering is performed on all output speech to obtain several new clusters. The ratio of the number of output speech in the new cluster with the most output speech to the total number of output speech is obtained and denoted as the first ratio. The product of the reciprocal of the number of new clusters and the first ratio is normalized and denoted as the first output emotion uniformity. The inverse proportional normalized value of the mean of the Euclidean distance between any two output speech at the Mel frequency cepstral coefficients is obtained and denoted as the second output emotion uniformity. The mean of the first output sentiment consistency and the second output sentiment consistency is denoted as the output sentiment consistency corresponding to each cluster.

6. The intelligent voice quality evaluation method for a multimodal intelligent terminal according to claim 1, characterized in that, The specific steps for obtaining the intelligent voice quality evaluation value of the multimodal intelligent terminal are as follows: The inversely proportional normalized value of the similarity between the text content of any two input speech samples is used as the clustering distance. Clustering operations are performed on all input speech samples to obtain several text clusters. In each text cluster, the maximum value of the Euclidean distance between the Mel frequency cepstral coefficients of all two arbitrary input speech samples is obtained and recorded as the emotion fluctuation value of each text cluster. The two input speech samples corresponding to the maximum value are recorded as the target input speech. Text clusters whose normalized emotional fluctuation values ​​are greater than a preset fluctuation threshold are denoted as target text clusters. Within each target text cluster, the differences in output speech text representation for each target text cluster are determined based on the dissimilarity of the text content of the output speech corresponding to the two target input speech and the Euclidean distance of the Mel frequency cepstral coefficients. The mean value of the output speech-text expression differences of all target text clusters is denoted as the output speech-text content fit. The average of the output speech emotion expression ability value and the adaptability of the output speech text content is obtained and recorded as the intelligent speech quality evaluation value of the multimodal intelligent terminal.

7. The intelligent voice quality evaluation method for a multimodal intelligent terminal according to claim 6, characterized in that, The specific steps involved in determining the output speech-text representation differences for each target text cluster are as follows: In each target text cluster, the inversely proportional normalized value of the similarity between the text content of the output speech corresponding to two target input speech is obtained, which is denoted as the first dissimilarity. The normalized value of the Euclidean distance between the Mel frequency cepstral coefficients of the output speech corresponding to two target input speech is obtained, which is denoted as the second dissimilarity. The mean of the first dissimilarity and the second dissimilarity is denoted as the difference in output speech text expression for each target text cluster.

8. The intelligent voice quality evaluation method for a multimodal intelligent terminal according to claim 1, characterized in that, The specific steps for judging the intelligent voice quality of multimodal intelligent terminals are as follows: When the intelligent voice quality evaluation value of the multimodal intelligent terminal is greater than the preset evaluation threshold, the intelligent voice quality of the multimodal intelligent terminal is judged to be good. When the intelligent voice quality evaluation value of the multimodal intelligent terminal is less than or equal to the preset evaluation threshold, the intelligent voice quality of the multimodal intelligent terminal is judged to be average.

9. A smart voice quality evaluation device for a multimodal intelligent terminal, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the computer program is executed by the processor, it implements the steps of the intelligent voice quality evaluation method for a multimodal intelligent terminal as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Service evaluation method and device based on artificial intelligence, equipment and storage medium

    CN111311327A

  • Customer satisfaction analysis method and device based on voice emotion recognition

    CN116959486A