Multi-dimensional AI platform intelligent voice response system using voice synthesis technology

By selecting and constructing personalized speech training sets, the problem of insufficient emotional expression in speech synthesis technology is solved, achieving more natural and fluent speech synthesis, and improving user experience and model performance.

CN120808785APending Publication Date: 2025-10-17ROPEOK TECHNOLOGY GROUP CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202511292834.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-11
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing speech synthesis technologies lack personalized information and contain redundant data when training acoustic models, resulting in synthesized speech that is insufficient, monotonous, or mismatched in emotional expression, thus affecting user experience.

Method used

The system acquires current user question voices and historical user question voices through a voice data acquisition module. It then uses text matching and voice matching modules to filter out human responses with high matching scores, constructs a voice training set, trains an acoustic model, and combines user scores to determine the final matching score and weights, generating natural and fluent response voices.

Benefits of technology

It improves the generalization ability and user experience of speech synthesis systems, generates natural and fluent speech responses that meet user expectations, and shortens the time for acoustic models to reach the expected performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808785A_ABST
    Figure CN120808785A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of voice synthesis, in particular to a multi-dimensional AI platform intelligent voice response system using the voice synthesis technology, which screens out matched historical user question voices according to text semantic similarity corresponding to current user question voices and historical user question voices, and sends the matched historical user question voices to a user terminal. Obtaining a voice training set and the weight of each element in the voice training set according to the voice feature similarity between the question voice of the current user and the question voice of the matched historical user in combination with the user score value of the manual reply voice, training an acoustic model, obtaining a reply text corresponding to the question voice of the current user, and obtaining the question voice of the current user; and inputting into the trained acoustic model, generating a reply voice signal, and then outputting to the current user. According to the method, the voice training set is screened and constructed from historical manual reply voices, so that an acoustic model can learn a more natural and smooth voice synthesis mode, and the reply voice signal contains rich voice features and expression modes.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of speech synthesis, in particular to a multi-dimensional AI platform intelligent speech response system using speech synthesis technology. BACKGROUND

[0002] The multi-dimensional AI platform intelligent speech response system using speech synthesis technology is a system integrated with multiple artificial intelligence technologies, which can interact with users through voice and provide various services and information. After the user puts forward a question or request to the system through voice, the system first converts the user's voice into text using speech recognition technology, then understands the user's intention and context through natural language processing technology, and retrieves relevant information from the database or the Internet according to the user's intention to generate a piece of text as a response to the user's request, and finally converts the text into natural and fluent voice using speech synthesis technology, and transmits the voice result to the user through a loudspeaker or other output device.

[0003] The existing problem is that in the process of converting text into natural and fluent voice using speech synthesis technology, acoustic model training is required first. If the speech data in the training set used in training the acoustic model does not reflect the personalized information of the current user and there are useless redundant speech data, it is easy to cause the synthesized speech obtained by the trained acoustic model to have the problems of insufficient emotional expression, single or mismatched emotional expression, and emotional loss, so that the system cannot realize efficient, natural and intelligent voice interaction, and the user experience is reduced. SUMMARY

[0004] The present application provides a multi-dimensional AI platform intelligent speech response system using speech synthesis technology to solve the existing problems.

[0005] The multi-dimensional AI platform intelligent speech response system using speech synthesis technology of the present application adopts the following technical solutions: An embodiment of the present application provides a multi-dimensional AI platform intelligent speech response system using speech synthesis technology, which comprises the following modules: A voice data acquisition module is used to acquire current user question voice and a plurality of historical user question voices, and an artificial reply voice corresponding to each historical user question voice; wherein each artificial reply voice corresponds to a user score value; A text matching module is used to filter out a plurality of matching historical user question voices according to the text semantic similarity between the current user question voice and each historical user question voice; The voice matching module is configured to determine the final matching degree of the current user question voice and each matching historical user question voice according to the voice feature similarity between the current user question voice and each matching historical user question voice, and the user score of the artificial reply voice; The reply text voice generation module is configured to train an acoustic model according to the size of the final matching degree of the current user question voice and each matching historical user question voice, obtain a voice training set and a weight of each element in the voice training set, train the acoustic model to obtain a trained acoustic model, input the reply text corresponding to the current user question voice into the trained acoustic model to generate a reply voice signal, and output the reply voice signal to the current user.

[0006] Further, the screening of the matching historical user question voices includes: obtaining the text of the current user question voice and the text of each historical user question voice; obtaining the semantic similarity between the text of the current user question voice and the text of each historical user question voice; In the normalized value of the semantic similarity between the text of the current user question voice and the text of all historical user question voices, the historical user question voice greater than a preset similarity threshold is recorded as a matching historical user question voice.

[0007] Further, the determination of the final matching degree of the current user question voice and each matching historical user question voice includes: obtaining a plurality of voice components of the current user question voice and a feature value of each voice component, recording the voice component corresponding to the maximum feature value as a main voice component, and recording the voice component other than the main voice component as a reference voice component; obtaining the reference voice component of each matching historical user question voice according to the obtaining manner of the reference voice component of the current user question voice; determining the environmental interference similarity of the current user question voice and each matching historical user question voice according to the difference between the current user question voice and the main voice component of the current user question voice, and the matching condition of each reference voice component of the current user question voice and each reference voice component of each matching historical user question voice; determining the emotion similarity of the current user question voice and each matching historical user question voice according to the energy feature difference between the current user question voice and each matching historical user question voice; determining the final matching degree of the current user question voice and each matching historical user question voice according to the environmental interference similarity and the emotion similarity, and the user score of the artificial reply voice.

[0008] Further, the determining the environmental interference similarity of the current user question voice and each matched historical user question voice comprises: determining a noisy environment interference degree of the current user question voice according to a difference between the current user question voice and a main voice component of the current user question voice; obtaining a noisy environment interference degree of each matched historical user question voice according to an obtaining manner of the noisy environment interference degree of the current user question voice; determining an environmental type mismatch degree of the current user question voice and each matched historical user question voice according to a matching condition of each reference voice component of the current user question voice and each reference voice component of each matched historical user question voice; obtaining a normalized value of an absolute value difference of the noisy environment interference degree of the current user question voice and the kth matched historical user question voice, denoted as a first difference value, and obtaining an inverse proportional value of a mean value of the environmental type mismatch degree and the first difference value, denoted as an environmental interference similarity of the current user question voice and the kth matched historical user question voice.

[0009] Further, the determining the noisy environment interference degree of the current user question voice comprises: obtaining a normalized value of an absolute value difference of amplitudes at the same time in the current user question voice and the main voice component of the current user question voice, denoted as an interference deviation value, and obtaining a time sequence of the interference deviation value in a time sequence; in the time sequence of the interference deviation value, recording a time at which the interference deviation value is greater than a preset deviation threshold value as a high interference time, and constructing a high interference period by using adjacent high interference times; obtaining a ratio of a total length of all high interference periods to a length of the current user question voice, denoted as a first ratio, obtaining a ratio of a maximum value in lengths of all high interference periods to the length of the current user question voice, denoted as a second ratio, obtaining an inverse proportional normalized value of a sum value of time intervals between all adjacent high interference periods, denoted as a time proximity value, and obtaining a mean value of the first ratio, the second ratio and the time proximity value, denoted as an interference persistence of the current user question voice; obtaining a sum value of the interference deviation values of all times in the time sequence of the interference deviation value, denoted as a first sum value, and obtaining a product of the interference persistence of the current user question voice and the first sum value, denoted as the noisy environment interference degree of the current user question voice.

[0010] Further, the determining the environmental type mismatch degree of the current user question voice and each matched historical user question voice comprises: obtaining a minimum value in DTW distances between each reference speech component of the current user question speech and all reference speech components of the kth matching historical user question speech, denoted as an environmental dissimilarity of each reference speech component of the current user question speech, and a reference speech component of the kth matching historical user question speech corresponding to the minimum value is denoted as a matching historical speech component of each reference speech component of the current user question speech; obtaining a number of different matching historical speech components in the matching historical speech components of all reference speech components of the current user question speech, denoted as a first number value, obtaining a number of reference speech components of the kth matching historical user question speech, denoted as a second number value, obtaining a sum value of the environmental dissimilarities of all reference speech components of the current user question speech, denoted as a second sum value, and obtaining a normalized value of a product of an inverse value of a ratio of the first number value and the second number value and the second sum value, denoted as an environmental type mismatch degree of the current user question speech and the kth matching historical user question speech.

[0011] Further, the determining of the emotional similarity of the current user question speech and each matching historical user question speech comprises: obtaining a difference value of a maximum amplitude minus a minimum amplitude in amplitudes of all time points in the subject speech component of the current user question speech as an amplitude fluctuation size of the current user question speech; obtaining a variance of energy values at all frequencies in the subject speech component of the current user question speech, denoted as an energy distribution unevenness of the current user question speech; a preset number threshold H, for the subject speech component of the current user question speech, the time length is equally divided into H speech segments; obtaining a sum value of energy values at all frequencies in each speech segment, denoted as a total energy value of each speech segment; determining an emotional instability of the current user question speech according to differences between total energy values of adjacent speech segments, in combination with the amplitude fluctuation size and the energy distribution unevenness; obtaining an emotional instability of each matching historical user question speech according to the obtaining manner of the emotional instability of the current user question speech; obtaining an inverse value of an absolute value of a difference between the emotional instability of the current user question speech and the emotional instability of the kth matching historical user question speech, denoted as an emotional similarity of the current user question speech and the kth matching historical user question speech.

[0012] Further, the determining of the emotional instability of the current user question speech comprises: in a time sequence, total energy values of all speech segments are used to form a total energy value sequence; In the total energy value sequence, a normalized value of a sum value of absolute values of differences of all adjacent total energy values is obtained, and is recorded as an emotional fluctuation degree of the current user question voice; A normalized value of a product of the emotional fluctuation degree of the current user question voice and the energy distribution unevenness is obtained, and is recorded as a first product.

[0013] Further, the determining the final matching degree of the current user question voice and each matching historical user question voice according to the environmental interference similarity, the emotional similarity, and a user score value of the artificial reply voice comprises: A mean value of the emotional similarity and the environmental interference similarity of the current user question voice and the kth matching historical user question voice is obtained, and is recorded as a matching degree of the current user question voice and the kth matching historical user question voice. A mean value of the matching degree of the current user question voice and the kth matching historical user question voice and a user score value of the artificial reply voice corresponding to the kth matching historical user question voice is obtained, and is recorded as a final matching degree of the current user question voice and the kth matching historical user question voice.

[0014] Further, the obtaining the voice training set and the weight of each element in the voice training set comprises: In the final matching degrees of the current user question voice and all the matching historical user question voices, the matching historical user question voice corresponding to the final matching degree greater than a preset matching threshold value is recorded as a final matching historical user question voice. The final matching degree of the current user question voice and each final matching historical user question voice is taken as a weight of the artificial reply voice corresponding to each final matching historical user question voice. The artificial reply voices corresponding to all the final matching historical user question voices constitute a voice training set.

[0015] The technical scheme of the present application has the following beneficial effects: In the embodiment of the present application, according to the text semantic similarity between the current user question voice and each historical user question voice, a plurality of matching historical user question voices are screened out, so that when the artificial reply voice corresponding to the historical user question voice similar to the text of the current user question voice is used as the training set, the acoustic model can learn how to generate a suitable voice reply in different situations, thereby improving its generalization ability in actual application, and the artificial reply voice contains rich voice features and expression methods, which helps the acoustic model to learn a more natural and fluent voice synthesis method, that is, by imitating the artificial reply voice, the voice synthesis system can generate a voice reply that meets the user's expectations, thereby improving the user experience. According to the voice feature similarity between the current user question voice and each matching historical user question voice, and in combination with the user score of the artificial reply voice, the final matching degree of the current user question voice and each matching historical user question voice is determined, so as to obtain the voice training set and the weight of each element in the voice training set, train the acoustic model, and obtain the trained acoustic model, so that the historical artificial reply voice that meets the noise environment and emotional state of the current user and makes the user satisfied is further screened out as the training set, which can speed up the acoustic model training speed, shorten the time for the acoustic model to reach the expected performance, and by giving each element in the training set a weight, the acoustic model training can pay more attention to important samples, thereby improving the performance and generalization ability of the model. The reply text corresponding to the current user question voice is input into the trained acoustic model to generate a reply voice signal, which is then output to the current user. Thus, the present application screens and constructs a voice training set from historical artificial reply voices, which helps the acoustic model to learn a more natural and fluent voice synthesis method, so that the reply voice signal contains rich voice features and expression methods. BRIEF DESCRIPTION OF DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor.

[0017] Figure 1 Module flowchart of the multi-dimensional AI platform intelligent voice response system using voice synthesis technology of the present application; Figure 2 Process schematic diagram of the multi-dimensional AI platform intelligent voice response using voice synthesis technology. DETAILED DESCRIPTION

[0018] In order to further illustrate the technical means and effects taken by the present application to achieve the predetermined inventive purpose, the following describes in detail the specific implementation, structure, features and effects of the multi-dimensional AI platform intelligent voice response system using voice synthesis technology according to the present application, in combination with the accompanying drawings and preferred embodiments. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments can be combined in any suitable form.

[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs.

[0020] The specific scheme of the multi-dimensional AI platform intelligent voice response system using voice synthesis technology provided by the present application is described in detail below in combination with the accompanying drawings.

[0021] Please refer to Figure 1 which shows the module flowchart of the multi-dimensional AI platform intelligent voice response system using voice synthesis technology provided by one embodiment of the present application, which includes the following modules: Module 101: voice data acquisition module.

[0022] This module is used to acquire the current user question voice and a plurality of historical user question voices, and the artificial reply voice corresponding to each historical user question voice; wherein each artificial reply voice corresponds to a user score value.

[0023] It should be noted that the workflow of the multi-dimensional AI platform intelligent voice response using voice synthesis technology is as follows: (1) voice recognition: the system first captures the user's voice input through the voice recognition engine and converts it into text. (2) natural language understanding: the system uses the natural language understanding module to analyze the text and understand the user's intent and context. (3) intent recognition and processing: the system calls the corresponding service or performs the corresponding operation according to the user's intent. (4) generate response: the system generates text information to respond to the user. (5) voice synthesis: the text information is converted into natural voice by voice synthesis technology and output to the user. In the process of converting text information into natural voice by voice synthesis technology, a large amount of voice data needs to be used to train an acoustic model to learn the mapping relationship between linguistic features and acoustic features, and then the text information is input into the trained acoustic model to generate the corresponding sound signal.

[0024] First, the current user question voice is collected, and then a plurality of historical user question voices and the artificial reply voice corresponding to each historical user question voice are collected in the database of the multi-dimensional AI platform intelligent voice response system. Each artificial reply voice corresponds to a user score value.

[0025] It should be noted that in the embodiment, the user score value set by the system ranges from 0 to 1, in the process of artificial reply of the system, the voice of each question of the user and the corresponding artificial reply voice are collected, after the question of the user is answered, the system invites the user to score the reply, if the user does not score, the default user score value is 0.8, which is taken as an example for description. When there are multiple questions and answers in one call of the user, the user score value of this time is respectively given to the artificial reply voice corresponding to the multiple question voices.

[0026] Module 102: text matching module.

[0027] The module is used for screening a plurality of matching historical user question voices according to the text semantic similarity between the current user question voice and the text corresponding to each historical user question voice.

[0028] Preferably, in an embodiment of the application, the method for obtaining the matching historical user question voice comprises: The text of the current user question voice and the text of each historical user question voice are obtained by using a voice recognition technology.

[0029] The semantic similarity between the text of the current user question voice and the text of each historical user question voice is obtained by using a deep learning model.

[0030] It should be noted that the deep learning model used in the embodiment is RNN (recurrent neural network), in the text semantic similarity task, RNN can be used to encode two text sequences, and then the semantic similarity between them is evaluated by comparing their hidden states. Among them, the voice recognition technology and RNN are both known technologies, and the specific method is not introduced here.

[0031] The preset similarity threshold is 0.7, which is taken as an example for description.

[0032] The normalized value of the semantic similarity between the text of the current user question voice and the text of each historical user question voice is obtained, among the normalized values of the semantic similarity between the text of the current user question voice and the text of all historical user question voices, the historical user question voice greater than the preset similarity threshold is recorded as the matching historical user question voice.

[0033] It is required to be explained: in the embodiment, the semantic similarity between the text of the current user question voice and the text of each historical user question voice is normalized to 0 to 1 by using the maximum minimum normalization method. The maximum minimum normalization method is a known technology, and the specific method is not introduced here. Because the artificial customer service can flexibly adjust its tone, diction and expression according to the specific situation to adapt to the different emotional needs of the user. For example, when the user is in a low mood, the customer service can use more comforting and encouraging language; when the user is excited, more positive and enthusiastic responses can be used. Therefore, if the artificial reply voice corresponding to the historical user question voice similar to the text of the current user question voice is used as the training set, the acoustic model can learn how to generate appropriate voice replies in different situations, thereby improving its generalization ability in actual application, and the artificial reply voice contains rich voice features and expression methods, which helps the acoustic model to learn more natural and fluent voice synthesis methods, that is, by imitating the artificial reply voice, the voice synthesis system can generate voice replies that meet the user's expectations, thereby improving the user experience.

[0034] Module 103: voice matching module.

[0035] The module is used to determine the final matching degree of the current user question voice and each matching historical user question voice according to the voice feature similarity between the current user question voice and each matching historical user question voice, and the user score of the artificial reply voice.

[0036] It is required to be explained: in the artificial reply process, the customer service personnel will consider the environmental noise of the user and the user's emotional tone. For example, when the user is in a noisy environment, the customer service personnel will reduce the speaking speed and increase the volume to a certain extent when replying, so that the user can clearly and clearly hear the reply; when the user's emotional tone is irritable, the customer service will use a softer and soothing tone to avoid sharp or harsh tone, so as to reduce the user's nervousness, and the customer service will add more pauses between sentences to give the user time to digest the information. Therefore, further, it is necessary to compare and analyze the current user question voice and the matching historical user question voice, and screen out the matching historical user question voice similar to the environmental noise and the emotional tone of the current user.

[0037] Preferably, in an embodiment of the present application, the method for obtaining the final matching degree of the current user question voice and each matching historical user question voice comprises: The current user question voice is decomposed into a plurality of voice components by using principal component analysis method, and the characteristic value of each voice component is obtained. The voice component corresponding to the maximum characteristic value is recorded as the main voice component, and the voice component other than the main voice component is recorded as the reference voice component.

[0038] It should be noted that the principal component analysis method is a known technology, and the specific method is not introduced here. The components obtained by principal component analysis are equal in length to the original signal. The larger the eigenvalue of the component in principal component analysis, the more it usually indicates that the speech component contains the key features of the speech signal in speech signal processing, which is crucial for subsequent speech recognition. Therefore, in this embodiment, the main speech component is used as the user speech component, and the reference speech component is used as the user environmental noise interference component. Therefore, the deviation between the current user question speech and the main speech component of the current user question speech is the interference deviation of the environmental noise.

[0039] The difference between the amplitude of the current user question speech and the amplitude of the main speech component of the current user question speech at the same time is obtained. The normalized value of the difference is recorded as the interference deviation value. The time sequence of the interference deviation value is obtained in time sequence.

[0040] It should be noted that in this embodiment, the horizontal axis of the speech signal is time, and the vertical axis is amplitude. The normalized value of the difference is recorded as the interference deviation value. The time sequence of the interference deviation value is obtained in time sequence. is a linear normalization function for normalizing data values to between 0 and 1.

[0041] The preset deviation threshold is 0.65, and this example is described.

[0042] In the time sequence of the interference deviation value, the time corresponding to the interference deviation value greater than the preset deviation threshold is recorded as the high interference time. The adjacent high interference times form a high interference period.

[0043] The length of each high interference period is obtained. The ratio of the total length of all high interference periods to the length of the current user question speech is recorded as the first ratio. The ratio of the maximum value in the length of all high interference periods to the length of the current user question speech is recorded as the second ratio. The sum of the time intervals between all adjacent high interference periods is obtained. The inverse proportional normalized value of the sum of the time intervals between all adjacent high interference periods is recorded as the time proximity value. The mean of the first ratio, the second ratio, and the time proximity value is recorded as the interference persistence of the current user question speech.

[0044] It should be noted that the normalized value of the difference between the amplitude of the current user question speech and the amplitude of the main speech component of the current user question speech at the same time is recorded as the interference deviation value. The normalized value of the difference is recorded as the interference deviation value. The time sequence of the interference deviation value is obtained in time sequence. ​The higher the proportion of the maximum duration of the high interference period in the current user question voice duration, the greater the proportion of the total duration of the high interference period in the current user question voice duration, and the smaller the time interval between adjacent high interference periods, the more continuous the high interference period in the current user question voice. The environmental noise interference of the duration will continuously accumulate in the entire voice signal, which may cause the overall quality of the voice signal to decline, thereby affecting the accuracy of semantic analysis, and the continuous environmental noise may mask the key features in the voice signal.

[0045] The sum value of the interference bias values of all times in the interference bias value time sequence is obtained, denoted as a first sum value, and the product of the interference continuity of the current user question voice and the first sum value is denoted as the noisy environment interference degree of the current user question voice.

[0046] According to the acquisition manner of the subject voice component and the reference voice component of the current user question voice and the noisy environment interference degree of the current user question voice, the subject voice component and the reference voice component of each matched historical user question voice and the noisy environment interference degree of each matched historical user question voice are obtained.

[0047] It should be noted that: factory workshops, construction sites, restaurants, bars and the like are noisy environments. For industrial noise interference in factory workshops and construction sites, customer service personnel need to increase the voice volume, use simple and clear language, avoid using complex terms or long sentences, and ensure clear information transmission. For social noise in restaurants and bars, customer service personnel need to use a more friendly and patient tone and appropriately slow down the speech rate to ensure that the user can keep up with the pace of the conversation. Therefore, different types of environments in the same noisy environment require customer service personnel to make corresponding adjustments in answering voice to ensure that the user can clearly hear and understand the provided information, so it is necessary to analyze the differences between the current user and the matched historical user in the environment type.

[0048] The DTW algorithm is used to obtain the DTW distance between each reference speech component of the current user question speech and each reference speech component of the kth matching historical user question speech, to obtain the minimum value among the DTW distances between each reference speech component of the current user question speech and all reference speech components of the kth matching historical user question speech, and to record the minimum value as the environmental dissimilarity of each reference speech component of the current user question speech. The reference speech component of the kth matching historical user question speech corresponding to the minimum value is recorded as the matching historical speech component of each reference speech component of the current user question speech. The number of different matching historical speech components among the matching historical speech components of all reference speech components of the current user question speech is recorded as a first number value, the number of reference speech components of the kth matching historical user question speech is recorded as a second number value, the sum value of the environmental dissimilarities of all reference speech components of the current user question speech is recorded as a second sum value, and the product of the inverse value of the ratio of the first number value to the second number value and the second sum value is recorded as the normalized value of the environmental type mismatch degree of the current user question speech and the kth matching historical user question speech.

[0049] It should be noted that, in the embodiment, the normalized value of the environmental type mismatch degree is taken as The DTW algorithm (Dynamic Time Warping) is a known technology, and the specific method is not described herein. The smaller the DTW distance is, the more similar the two speech signals are, so the reference speech component of the kth matching historical user question speech corresponding to the minimum value is taken as the matching historical speech component of each reference speech component of the current user question speech. For example, the machine running noise (such as motor sound, gear meshing sound, and metal impact sound) of a factory is usually similar in spectral characteristics. The greater the minimum value is, the more dissimilar each reference speech component of the current user question speech is to the reference speech component of the kth matching historical user question speech, that is, the more different the environmental type in which the current user and the kth matching historical user are located is. Since the matching historical speech components of different reference speech components of the current user question speech can be the same, the number of different matching historical speech components among the matching historical speech components of all reference speech components of the current user question speech is taken, the difference between 1 and the ratio of the first number value to the second number value is taken as the inverse value of the ratio of the first number value to the second number value, that is, the smaller the ratio of the first number value to the second number value is, the fewer the matching historical speech components in the reference speech component of the kth matching historical user question speech are, and the more different the environmental type in which the current user and the kth matching historical user are located is.

[0050] ​​obtaining a normalized value of an absolute value of a difference between the noisy environment interference degree of the current user question voice and the noisy environment interference degree of the kth matching historical user question voice, denoted as a first difference value, and obtaining a mean value of the first difference value and the inverse proportional value of the environment type mismatch degree between the current user question voice and the kth matching historical user question voice , denoted as an environment interference similarity between the current user question voice and the kth matching historical user question voice.

[0051] It should be noted that in the embodiment, the inverse proportional value of the environment type mismatch degree between the current user question voice and the kth matching historical user question voice is used as the environment interference similarity between the current user question voice and the kth matching historical user question voice. If the noisy environment interference degrees of the current user question voice and the kth matching historical user question voice are consistent, and the environment types are more similar, the environment interference of the current user question voice and the kth matching historical user question voice is more similar.

[0052] It should be noted that further, the emotion similarity between the current user question voice and the matching historical user question voice needs to be analyzed. When the user's emotion is stable, the amplitude fluctuation of the user's voice signal is small, and the overall is relatively smooth, without obvious mutation or violent fluctuation, that is, the energy distribution of the voice signal is relatively uniform, without obvious energy concentration or dispersion phenomenon. When the user's emotion is unstable, the amplitude fluctuation of the user's voice signal is large, and sudden peak or valley value may appear, that is, the energy of the voice signal is significantly increased in some period, and the energy is low in other period, reflecting the ups and downs of the emotion.

[0053] Obtaining a difference between the maximum amplitude and the minimum amplitude in the amplitudes of all time points in the main voice component of the current user question voice as the amplitude fluctuation size of the current user question voice.

[0054] Performing short-time Fourier transform on the main voice component of the current user question voice to obtain energy values at different frequencies.

[0055] The short-time Fourier transform is a known technology, and the specific method is not introduced here.

[0056] Obtaining a variance of the energy values at all frequencies in the main voice component of the current user question voice as the energy distribution unevenness of the current user question voice.

[0057] The preset number threshold H is 30, and this is used as an example for description.

[0058] For the main voice component of the current user question voice, the time length is equally divided into H voice segments, that is, the time lengths of all voice segments are equal.

[0059] Performing short-time Fourier transform on each voice segment to obtain energy values at different frequencies.

[0060] ​The sum of energy values of all frequencies in each speech segment is obtained, and is recorded as the total energy value of each speech segment.

[0061] In the subject speech component of the current user question speech, the total energy values of all speech segments are arranged in time sequence to form a total energy value sequence.

[0062] In the total energy value sequence, the absolute value of the difference between adjacent total energy values is obtained, and the sum of the absolute values of the differences between all adjacent total energy values is obtained , which is recorded as the emotional fluctuation degree of the current user question speech.

[0063] The greater the difference between the energy values of the speech signal in adjacent time periods, the greater the fluctuation of the user's emotions. The is taken as the normalized value.

[0064] The product of the emotional fluctuation degree of the current user question speech and the energy distribution unevenness of the current user question speech is obtained , which is recorded as the first product. The product of the first product and the amplitude fluctuation size of the current user question speech is obtained , which is recorded as the emotional instability of the current user question speech.

[0065] It should be noted that in the embodiment, the and the are taken as the normalized values of the and the respectively. The greater the energy distribution unevenness of the current user question speech, the more unstable the user's emotions in the current user question speech, so the emotional fluctuation degree is taken as the adjustment value of the energy distribution unevenness to obtain the first product. The greater the amplitude fluctuation size of the current user question speech, the more likely it is that the user's emotions are unstable in the current user question speech. Therefore, the product of the first product and the amplitude fluctuation size is used to reflect the emotional instability of the current user question speech.

[0066] According to the manner of obtaining the emotional instability of the current user question speech, the emotional instability of each matching historical user question speech is obtained.

[0067] The inverse proportional value of the absolute value of the difference between the emotional instability of the current user question speech and the emotional instability of the kth matching historical user question speech is obtained, which is recorded as the emotional similarity between the current user question speech and the kth matching historical user question speech.

[0068] ​Wherein, the difference between 1 and the absolute value of the difference between the emotional instability of the current user question voice and the emotional instability of the kth matching historical user question voice is taken as the inverse proportional value of the absolute value of the difference between the emotional instability of the current user question voice and the emotional instability of the kth matching historical user question voice.

[0069] The mean value of the emotional similarity between the current user question voice and the kth matching historical user question voice and the environmental interference similarity between the current user question voice and the kth matching historical user question voice is taken as the matching degree between the current user question voice and the kth matching historical user question voice.

[0070] It is required to be explained that, in the artificial reply process, due to the working ability of different customer service personnel, the reply effect of the user question is different, therefore, when constructing the training set, the artificial reply that makes the user unsatisfied needs to be eliminated, and the artificial reply that makes the user satisfied needs to be reserved, so as to improve the user experience.

[0071] The mean value of the matching degree between the current user question voice and the kth matching historical user question voice and the user scoring value of the artificial reply voice corresponding to the kth matching historical user question voice is taken as the final matching degree between the current user question voice and the kth matching historical user question voice.

[0072] Module 104: reply text voice generation module.

[0073] The module is used for acquiring the voice training set and the weight of each element in the voice training set according to the size of the final matching degree between the current user question voice and each matching historical user question voice, training the acoustic model, acquiring the trained acoustic model, acquiring the reply text corresponding to the current user question voice, inputting into the trained acoustic model, generating the reply voice signal, and outputting the reply voice signal to the current user.

[0074] Preferably, in an embodiment of the present application, the method for acquiring the reply voice signal comprises: The preset matching threshold is 0.75, and this is taken as an example for description.

[0075] Among the final matching degrees between the current user question voice and all matching historical user question voices, the matching historical user question voice corresponding to the final matching degree greater than the preset matching threshold is taken as the final matching historical user question voice.

[0076] The artificial reply voice corresponding to all final matching historical user question voices is taken to constitute the voice training set.

[0077] The final matching degree of the current user question voice to each final matching historical user question voice is taken as the weight of the artificial reply voice corresponding to each final matching historical user question voice, that is, the weight of each element in the voice training set.

[0078] According to the voice training set and the weight of each element in the voice training set, training of the acoustic model is performed to obtain the trained acoustic model.

[0079] It should be noted that in the acoustic model training, the composition of the training set is crucial because it directly affects the performance and generalization ability of the model. The method of assigning a weight to each element in the training set is called weighted training samples in machine learning. By assigning a weight to each element in the training set, the acoustic model training can pay more attention to important samples, thereby improving the performance and generalization ability of the model. Therefore, in the embodiment, the historical artificial reply voice that is more consistent with the noise environment and emotional state of the current user and that makes the user satisfied is used as the training set, and the final matching degree is used as the weight, which can accelerate the acoustic model training speed and shorten the time for the acoustic model to reach the expected performance. The acoustic model used in the embodiment is a DNN, wherein the DNN (Deep Neural Network) is an artificial neural network model with multiple hidden layers and can learn and extract complex features in data. This is a known technology, and the specific method is not described here.

[0080] According to the workflow of the multi-dimensional AI platform intelligent voice response of the voice synthesis technology, the reply text corresponding to the current user question voice is obtained.

[0081] The reply text corresponding to the current user question voice is input into the trained acoustic model to generate a reply voice signal, and the reply voice signal is output to the current user. The workflow of the multi-dimensional AI platform intelligent voice response of the voice synthesis technology is shown in Figure 2 .

[0082] Thus, the present application is completed.

[0083] To sum up, in the embodiment of the present application, according to the text semantic similarity between the current user question voice and each historical user question voice, a plurality of matching historical user question voices are screened out, according to the voice feature similarity between the current user question voice and each matching historical user question voice, and in combination with the user score value of the artificial reply voice, the final matching degree of the current user question voice and each matching historical user question voice is determined, thereby obtaining the voice training set and the weight of each element in the voice training set, training the acoustic model, obtaining the trained acoustic model, obtaining the reply text corresponding to the current user question voice, inputting into the trained acoustic model, generating the reply voice signal, and then outputting to the current user. The present application constructs the voice training set by screening from the historical artificial reply voice, which is helpful for the acoustic model to learn a more natural and fluent voice synthesis mode, so that the reply voice signal contains rich voice features and expression modes.

[0084] The above merely describes the preferred embodiments of the present application and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. within the principles of the present application shall be included in the protection scope of the present application.

Claims

1. A multi-dimensional AI platform intelligent voice response system using speech synthesis technology, characterized by: The system includes the following modules: Voice data collection module: used to obtain the current user's question voice and several historical user question voices, as well as the manual reply voice corresponding to each historical user question voice; each manual reply voice corresponds to a user score value; Text matching module: used to screen out several matching historical user question voices based on the semantic similarity of the text corresponding to the current user question voice and each historical user question voice; Voice matching module: used to determine the final matching degree between the current user's question voice and each matching historical user's question voice based on the similarity of voice features between the current user's question voice and each matching historical user's question voice, combined with the user score of the manual reply voice; Reply text speech generation module: used to obtain the speech training set and the weight of each element in the speech training set according to the final matching degree between the current user's question speech and each matching historical user's question speech, train the acoustic model, and obtain the trained acoustic model; obtain the reply text corresponding to the current user's question speech, input it into the trained acoustic model, generate a reply speech signal, and output the reply speech signal to the current user.

2. The multi-dimensional AI platform intelligent voice response system using speech synthesis technology according to claim 1 is characterized in that: The filtering out of a plurality of matching historical user question voices includes: Get the text of the current user's question voice and the text of each historical user's question voice; Obtain the semantic similarity between the text of the current user's question speech and the text of each historical user's question speech; In the normalized value of the semantic similarity between the text of the current user question speech and the text of all historical user question speech, the historical user question speech that is greater than a preset similarity threshold is recorded as a matching historical user question speech.

3. The multi-dimensional AI platform intelligent voice response system using speech synthesis technology according to claim 1 is characterized in that: Determining the final matching degree between the current user question voice and each matching history user question voice includes: Obtain several speech components of the current user's question speech and the eigenvalues ​​of each speech component, record the speech component corresponding to the maximum eigenvalue as the main speech component, and record the speech components other than the main speech component as reference speech components; Obtain the reference speech component of each matching historical user question speech according to the method for obtaining the reference speech component of the current user question speech; Determining the similarity of environmental interference between the current user question voice and each matching historical user question voice based on a difference between the current user question voice and a main voice component of the current user question voice, and a matching condition between each reference voice component of the current user question voice and each reference voice component of each matching historical user question voice; Determine the emotional similarity between the current user's question voice and each matching historical user's question voice based on the energy feature difference between the current user's question voice and each matching historical user's question voice; The final matching degree between the current user question voice and each matching historical user question voice is determined based on the environmental interference similarity and the emotion similarity, as well as the user score value of the manual reply voice.

4. The multi-dimensional AI platform intelligent voice response system using speech synthesis technology according to claim 3 is characterized in that: Determining the similarity of environmental interference between the current user's question speech and each matching historical user's question speech includes: determining a noisy environment interference degree of the current user's question voice based on a difference between the current user's question voice and a main voice component of the current user's question voice; Obtain the noisy environment interference level of each matching historical user's question voice according to the method for obtaining the noisy environment interference level of the current user's question voice; Determining the degree of mismatch between the current user's question voice and the environment type of each matching historical user's question voice based on a matching condition between each reference voice component of the current user's question voice and each reference voice component of each matching historical user's question voice; Obtain the normalized value of the absolute value of the difference between the noisy environment interference degree of the current user's question voice and the k-th matching historical user's question voice, record it as the first difference value, and take the inverse proportional value of the environment type mismatch degree and the mean of the first difference value as the environmental interference similarity between the current user's question voice and the k-th matching historical user's question voice.

5. The multi-dimensional AI platform intelligent voice response system using speech synthesis technology according to claim 4 is characterized in that: Determining the noisy environment interference level of the current user's question speech includes: Obtain a normalized value of the absolute value of the difference in amplitude between the current user's question voice and the main voice component of the current user's question voice at the same time, record it as an interference deviation value, and obtain a time series of interference deviation values ​​in chronological order; In the interference deviation value time series, the moment corresponding to the interference deviation value greater than the preset deviation threshold is recorded as the high interference moment, and the adjacent high interference moments constitute the high interference period; Obtain the ratio of the total duration of all high-interference periods to the duration of the current user's problem voice, recorded as the first ratio; obtain the ratio of the maximum duration of all high-interference periods to the duration of the current user's problem voice, recorded as the second ratio; obtain the inversely proportional normalized value of the sum of the time intervals between all adjacent high-interference periods, recorded as the time proximity value; obtain the average of the first ratio, the second ratio, and the time proximity value, recorded as the duration of interference to the current user's problem voice; The sum of the interference deviation values ​​at all moments in the interference deviation value time series is obtained and recorded as the first sum value. The product of the interference persistence of the current user's problem speech and the first sum value is recorded as the noisy environment interference degree of the current user's problem speech.

6. The multi-dimensional AI platform intelligent voice response system using speech synthesis technology according to claim 4 is characterized in that: Determining the mismatch between the current user's question voice and the environment type of each matching historical user's question voice includes: Obtain the minimum value of the DTW distance between each reference speech component of the current user's question speech and all reference speech components of the k-th matching historical user's question speech, and record it as the environmental dissimilarity of each reference speech component of the current user's question speech. Record the reference speech component of the k-th matching historical user's question speech corresponding to the minimum value as the matching historical speech component of each reference speech component of the current user's question speech. Obtain the number of different matching historical voice components in the matching historical voice components of all reference voice components of the current user's question voice, recorded as the first quantity value, obtain the number of reference voice components of the k-th matching historical user's question voice, recorded as the second quantity value, obtain the sum of the environmental dissimilarity of all reference voice components of the current user's question voice, recorded as the second sum value, and normalize the product of the inverse proportional value of the ratio of the first quantity value to the second data quantity value and the second sum value, recorded as the environmental type mismatch degree between the current user's question voice and the k-th matching historical user's question voice.

7. The multi-dimensional AI platform intelligent voice response system using speech synthesis technology according to claim 3 is characterized in that: Determining the emotional similarity between the current user's question voice and each matching historical user's question voice includes: Obtain the difference between the maximum amplitude and the minimum amplitude of the main voice component of the current user's question voice at all times as the amplitude fluctuation of the current user's question voice; Obtain the variance of the energy values ​​at all frequencies in the main speech component of the current user's question speech, which is recorded as the energy distribution unevenness of the current user's question speech; A preset number threshold H is set, and the main speech component of the current user's question speech is divided into H speech segments according to the duration; Obtain the sum of the energy values ​​at all frequencies in each speech segment, and record it as the total energy value of each speech segment; Determining the emotional instability of the current user's question speech based on the difference in total energy values ​​of adjacent speech segments, combined with the amplitude fluctuation and the energy distribution unevenness; Obtain the emotional instability of each matching historical user's question speech according to the method for obtaining the emotional instability of the current user's question speech; Obtain an inversely proportional value of the absolute value of the difference in emotional instability between the current user's question voice and the k-th matching historical user's question voice, and record it as the emotional similarity between the current user's question voice and the k-th matching historical user's question voice.

8. The multi-dimensional AI platform intelligent voice response system using speech synthesis technology according to claim 7 is characterized in that: Determining the emotional instability of the current user's question speech includes: In chronological order, the total energy values ​​of all speech segments are used to form a total energy value sequence; In the total energy value sequence, obtain the normalized value of the sum of the absolute values ​​of the differences between all adjacent total energy values, and record it as the emotional fluctuation level of the current user's question voice; Obtain the normalized value of the product of the emotional fluctuation degree of the current user's question voice and the energy distribution unevenness, which is recorded as the first product. The normalized value of the product of the first product and the amplitude fluctuation size of the current user's question voice is recorded as the emotional instability of the current user's question voice.

9. The multi-dimensional AI platform intelligent voice response system using speech synthesis technology according to claim 3 is characterized in that: Determining the final matching degree between the current user question voice and each matching historical user question voice based on the environmental interference similarity and the emotion similarity, and the user score value of the manual reply voice includes: Obtain the average of the emotional similarity and environmental interference similarity between the current user's question voice and the k-th matching historical user's question voice, and record it as the matching degree between the current user's question voice and the k-th matching historical user's question voice; Get the average of the matching degree between the current user's question voice and the k-th matching historical user's question voice and the user's scoring value of the manual reply voice corresponding to the k-th matching historical user's question voice, and record it as the final matching degree between the current user's question voice and the k-th matching historical user's question voice.

10. The multi-dimensional AI platform intelligent voice response system using speech synthesis technology according to claim 1, characterized in that: The obtaining of the speech training set and the weight of each element in the speech training set includes: Among the final matching degrees between the current user's question voice and all matching historical user's question voices, the matching historical user's question voice corresponding to the final matching degree greater than the preset matching threshold is recorded as the final matching historical user's question voice; The final matching degree between the current user's question voice and each final matching historical user's question voice is used as the weight of the manual reply voice corresponding to each final matching historical user's question voice; All the human response voices that finally match the historical user question voices constitute the voice training set.

Citation Information

Patent Citations

  • Voice answering method for combining intelligent answer with artificial answer

    CN107315766A

  • Voice response optimization method, system and equipment based on intelligent comment and medium

    CN113327612A

  • Intelligent voice question answering method and system based on deep learning and emotion recognition

    CN114203177A

  • Online customer service platform based on AI voice interaction

    CN118173092A

  • Intelligent evaluation method and system for call voice, and medium

    CN118982978A