Multi-round voice reply generation method and system based on multi-modal feature optimization
By combining the text features and speech time-frequency characteristics of the big model, accurate mood recognition is achieved, and GANs are used to optimize the efficiency of chat robots, the problems of insufficient quality and difficult semantic understanding when digital employees respond to speech are solved, and the efficiency and quality of multiple rounds of conversations are improved.
Patent Information
- Application Number
- CN202510235915.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-28
AI Technical Summary
Existing digital employees have insufficient quality when replying to voice and it is difficult to accurately understand the semantics of speech, which leads to the inconsistency of the reply content with the user's intention, which increases the difficulty of model feedback.
Multi-round speech reply generation method based on multimodal feature optimization is adopted, and precise emotions are realized by combining the text features of the big model and the time-frequency features contained in the timing chart of speech extracted, and GANs are used to optimize the efficiency of chat robots.
The emotional perception ability in multiple rounds of conversations has been improved, the method of classifying emotions by pronunciation words for opportunity has been improved, the quality of generated voice data has been optimized, the load of overall services has been reduced, and the efficiency and quality of multi-round conversations of digital employees has been improved.
Smart Images

Figure CN120144713A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent AC optimization methods, and particularly relates to a multi-round voice response generation method based on multi-modal feature optimization. Background Art
[0002] Traditional digital employees are good at processing a large amount of structured text data and automatically executing repetitive business processes according to prefabricated rules and processes. They can only handle basic operations such as simple emails and texts to simulate manual work. As the application scenarios of intelligent customer service become more extensive, multi-round voice responses and real-time conversations have gradually become a necessity.
[0003] When digital employees based on AI agents reply to voices, since existing large models are mainly applied to reply to texts or pictures, there are deficiencies in the quality of voice replies. Moreover, such digital employees have certain difficulties in understanding the semantics expressed by voices. Even if the voice is accurately recognized as text, since the same voice segment may have different meanings in different contexts, digital employees may not be able to accurately judge its exact semantics, resulting in the reply content not matching the user's intention. In addition, multiple calls to the large model by AI agents will cause some users to interact with the large model using multi-modal data such as voices and pictures, increasing the difficulty of model feedback. Therefore, for the text input into the large model after voice conversion, the large model is not accurate enough in extracting the semantics of voices and images and then optimizing the data fed back to users based on semantic relevance. Summary of the Invention
[0004] The purpose of the present invention is to provide a multi-round voice response generation method based on multi-modal feature optimization to solve the problems raised in the above background art.
[0005] To achieve the above purpose, the present invention provides the following technical solution: A multi-round voice response generation method based on multi-modal feature optimization, including the following processes:
[0006] Process 1: Classify the text features of the large model and the time-frequency features included in the time series diagram extracted from the voice to achieve accurate emotion recognition of multi-round conversations:
[0007] Step 1: Collect some labeled sample pairs, including two formats: voice and text;
[0008] Step 2: Use voice-to-text, the interface large model to extract the semantics of the voice, use Clip to extract text features, pre-train the image encoder and text encoder to predict which images in the dataset are paired with which texts, convert all classes of the dataset into texts, and predict the best pairing of the caption class with the given image;
[0009] Step 3: Combine existing tags, such as prompt words with different weights like actively purchase, inclined, interested, etc., for pre-training. Converting speech into a matrix consists of two steps. First, convert the signal from space (time amplitude) to polar coordinates through the following operations:
[0010]
[0011] Among them, L is the total length of the speech in the prompt words. This conversion is objective and preserves the absolute time relationship. First, it is necessary to scale the data to obtain the normalized form of the input X as follows:
[0012]
[0013] Among them, UB is the upper bound of the parameter, LB is the lower bound of the parameter, and the calculation of the Gramian matrix is as follows:
[0014]
[0015] Local expressions may not be able to capture the overall tendency of emotional words involving real emotions. The combination method of context relationships in emotional words has a great impact on the final classification performance. Convert the speech signal to the Gramian field and combine the deep neural convolutional network to optimize the extraction of high-dimensional features;
[0016] Step 4: Convert the signal into a time-frequency diagram, fine-tune the time-frequency features extracted from the speech based on spatio-temporal attention, and adopt a channel spatio-temporal attention mechanism to optimize the expression of emotional features in the frequency domain:
[0017] M c (x) = ρ(mlp(avgpool(x)) + mlp(maxpool(x)))
[0018] M s (x) = ρ(con 7×7 ([avgpool(x)); maxpool(x)]))
[0019] M t (x) = ρ(con 7×7 (avgpool(x)) + conv(maxpool(x)))
[0020] Among them, M c (x), M s (x), M t (x) represent the weights of channel attention, spatial attention, and temporal attention respectively. ρ is the activation function, and mlp is the multi-layer perceptron, f 7×7It is a one-dimensional convolution operation that simultaneously considers information in three dimensions: channels, space, and time, improving the model's ability to represent spatio-temporal data and adaptively capturing the dynamic associations between nodes in the spatial dimension;
[0021] Step 5: Combine some of the labeled samples to train the emotion classification model. The loss function consists of two parts:
[0022]
[0023] The first part, \(l_s\), is the loss function for the sound feature category:
[0024]
[0025] where \(B\) is the number of labeled full waveforms, \(p\) b is the emotion category label, and \(p\) m \((y|a(x b ))\) is the probability distribution of the emotion category after the sample is predicted by the model. The unlabeled data undergoes two rounds of data augmentation, strong and weak, and is propagated forward respectively. The data after strong augmentation is input into the network \(G(\sigma s_t-1 )\) to calculate the predicted value, and the other branch inputs the weakly augmented data into the network \(G(\sigma_{\varsigma\_t - 1})\) to calculate the predicted value. The loss \(l\) of the unlabeled data is calculated as: u It is expressed as:
[0026] The second part is the loss function for the large model text features:
[0027]
[0028] where \(q\) b is the probability distribution of the model after weak augmentation of the unlabeled data. \(I(\max(q b ))\geq\Gamma\) is the threshold screening indicator function. If the prediction result after weak augmentation is greater than the threshold \(\eta\), it is considered that the true label is the same as the prediction result. These labels are retained and allowed to participate in the calculation of the loss function for the unlabeled waveforms, and the prediction result is recorded as a pseudo-label. If the prediction result does not exceed the threshold, the sample is discarded and not allowed to participate in the current training;
[0029] \(p\) m \((y|A(u b ))\) is the probability distribution of the model after strong augmentation of the full waveform \(u\) b . is the classification prediction result. Cross-entropy is used to calculate the distance between the predicted value of the strongly augmented data and the pseudo-label, and \(\eta\) u is used to specify the proportion of the unsupervised loss in the total loss function;
[0030] Step 6: Integrate the two features, optimize the results of emotion recognition using a fully connected layer, and based on the emotion recognition results, further optimize the efficiency of the chatbot using GANs, a generator network, and a discriminator network. Here, GANs are used for the classification network, the discriminator network is used to distinguish between real and fake images, and the generator network learns how to generate images according to text descriptions. Over time, the generator's ability to generate realistic images will be enhanced through the adversarial training process;
[0031] Process 2: Optimize the conversation logic of the chatbot's multi-turn conversation based on the emotional state in the current conversation:
[0032] Step 1: Based on Process 1, identify the emotional state of the user's current conversation and grade the state. It allows for the recognition, transformation, and analysis of input queries and metadata to generate appropriate responses. To this end, the input data is aggregated, transformed into structured data, and analyzed to plan and execute the necessary operations;
[0033] Step 2: If the user's emotion is more negative, the robot's performance should be more positive. Combine some threshold hint words, input the current information and hint words into the large model. Subsequently, through the MLP network, these inputs are complexly mapped to the probability space of emotion types:
[0034] Given a conversation U = {u 1 ,...,u t}, at each time t ∈ [1, t], identify y t , that is, the emotion of u t , which depends on the previous t - 1 predicted emotional states Y 1:t-1 = {y 1 ,...,y t-1}:
[0035]
[0036] where, M = {m 1 ,...,m t}, m t ∈ {0, 1} is an indicator random variable, m t = 1 indicates that the speaker's emotion is different from the previous one, m t = 0; otherwise, m 1 is defined as 1;
[0037] Step 3: In a multi-turn conversation, the user's emotion may change due to certain factors, such as seeing a more preferred product or having a little hesitation about the price, resulting in small fluctuations in emotion. The chatbot should be sensitive to such changes and adjust the introduction strategy in a timely manner;
[0038] In addition, identify the t-th emotional state y tThe probability can be factorized as
[0039]
[0040] where p(y t | x t ), p(y t | z t , Y 1:t-1 ), p(z t | Z 1:t-1 ) respectively simulate the emotional distribution of the current discourse, the probability of emotional assignment of the discourse predicted based on the past, and the probability of cross-time transition of the emotional state;
[0041] Step 4: The customer's emotion is easily affected by external factors, such as the evaluations of surrounding customers or small details of product displays. The chatbot needs to closely monitor the changes in the customer's emotion. When the customer hears a negative evaluation of the product from someone beside, the robot should promptly refute it and use facts and data to prove the quality and value of the product; if the customer is not satisfied with the details of the product display, the robot should quickly adjust the display method or provide a better display sample;
[0042] Step 5: In different application scenarios, the user's emotion grading methods are different. For example, in the application of intelligent grid workers, it can also be measured from dimensions such as the intensity, positivity, negativity, and complexity of emotion;
[0043] Step 6: Through multiple rounds of iteration, if the customer takes the initiative to purchase, the conversation is terminated, or if the customer fails to respond within the time limit, the conversation is terminated.
[0044] Preferably, the sample pair described in Step 1 of Process 1 has a regular update function, and the interval of the regular update is one day.
[0045] Preferably, the prompt words described in Step 3 of Process 1 can be manually added, and the weights of the prompt words can be manually adjusted.
[0046] Preferably, t described in Step 2 of Process 2 is set to ten seconds, and the setting of t can be manually changed.
[0047] Preferably, the negative evaluations described in Step 4 of Process 2 include keywords such as low cost performance and expensive, and the keywords of the negative evaluations can be manually added.
[0048] Preferably, the application scenarios described in Step 5 of Process 2 include holidays, long vacations, and online consumption festivals, etc., and the application scenarios can be manually added and set.
[0049] Preferably, the timeout time described in Step 6 of Process 2 is set to three hundred seconds, and the timeout time can be adjusted.
[0050] A multi-round voice reply generation system based on multi-modal feature optimization, comprising:
[0051] A voice reply production system: used to accurately identify the user's emotion and make corresponding replies;
[0052] Among them, the voice reply production system includes:
[0053] A voice processing module: collects and processes the user's voice information;
[0054] A model establishment module: generates the user's voice time-frequency diagram;
[0055] An emotion analysis module: analyzes the user's emotion based on the time-frequency diagram;
[0056] An output module: outputs corresponding answers according to the emotion data.
[0057] Preferably, the voice processing module includes: a voice collection node, a noise reduction processing node, and a translation input node; the emotion analysis module includes: an emotion analysis node and an emotion correction node.
[0058] The beneficial effects of the present invention are as follows:
[0059] In this paper, by converting voice into a time-frequency diagram and combining a mature classification architecture in the image field, a time-frequency classification method based on spatio-temporal attention is proposed to optimize the emotion perception ability in multi-round conversations, assign heavier weights to emotional features with obvious classification features, and dynamically adjust the multi-round conversation ability of the AI agent based on the weights of different emotions, improving the method of classifying emotions based on voice prompt words, optimizing the quality of generated voice data by combining the semantic information provided by the large model, and also improving the logic of applying chatbots to scenarios such as intelligent grid workers and shopping guides. At the same time, the efficiency and quality of multi-round conversations of digital employees are greatly optimized, thereby reducing the load of the overall service. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Figure 1 It is a schematic flowchart of emotion recognition of the present invention;
[0061] Figure 2 It is a schematic flowchart of optimizing the chatbot logic of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0062] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0063] As shown in Figures 1 to 2 the following, an embodiment of the present invention provides a multi-round speech response generation method optimized based on multi-modal features, including the following processes:
[0064] Process 1: Classify the text features of the large model and the time-frequency features included in the time series diagram extracted from the speech to achieve accurate emotion recognition in multi-round conversations:
[0065] Step 1: Collect partially labeled sample pairs, including two formats: speech and text;
[0066] Step 2: Use speech-to-text, interface the large model to extract the semantics of the speech, use Clip to extract text features, pre-train the image encoder and text encoder to predict which images in the dataset are paired with which texts, convert all classes of the dataset into texts, and predict the best pairing of the caption class with the given image;
[0067] Step 3: Combine existing labels, such as different-weighted prompt words like positive purchase, tendency, interest, etc. for pre-training. Converting the speech into a matrix consists of two steps. First, convert the signal from space (time amplitude) to polar coordinates through the following operations:
[0068]
[0069] Among them, L is the total length of the speech in the prompt word. This conversion is objective and preserves the absolute time relationship. First, it is necessary to scale the data to obtain the normalized form of the input X as follows:
[0070]
[0071] Among them, UB is the upper bound of the parameter, LB is the lower bound of the parameter, and the calculation of the Gramian matrix is as follows:
[0072]
[0073] Local expressions may not be able to grasp the overall tendency of emotion words involving real emotions. The combination method of context relationships in emotion words has a great impact on the final classification performance. Convert the speech signal to the Gramian field and combine the deep neural convolutional network to optimize the extraction of high-dimensional features;
[0074] Step 4: Convert the signal into a time-frequency diagram, fine-tune the time-frequency features extracted from the speech based on spatio-temporal attention, and adopt a channel spatio-temporal attention mechanism to optimize the expression of emotion features in the frequency domain:
[0075] M c (x) = ρ(mlp(avgpool(x)) + mlp(maxpool(x)))
[0076] M s(x) = ρ(con 7×7 ([avgpool(x)); maxpool(x)]))
[0077] M t (x) = ρ(con 7×7 (avgpool(x)) + conv(maxpool(x)))
[0078] Among them, M c (x), M s (x), M t (x) represent the weights of channel attention, spatial attention, and temporal attention respectively. ρ is the activation function, mlp is the multi-layer perceptron, f 7×7 is a one-dimensional convolution operation, considering the information of three dimensions: channel, spatial, and temporal simultaneously, improving the model's representation ability for spatio-temporal data, and adaptively capturing the dynamic associations between nodes in the spatial dimension;
[0079] Step Five: Train the emotion classification model by combining some of the labeled samples. The loss function consists of two parts:
[0080]
[0081] The first part, l_s, is the loss function for the sound feature category:
[0082]
[0083] Among them, B is the number of labeled full waveforms, p b is the emotion category label, p m (y|a(x b )) is the probability distribution of the emotion category after the sample is predicted by the model. The unlabeled data is forward-propagated through two data augmentations of strong and weak respectively. The data after strong augmentation is input into the network G(σ s_t-1 ) to calculate the predicted value, and the other branch inputs the weakly augmented data into the network G(\sigma_{\varsigma\_t - 1}) to calculate the predicted value. The loss l of the unlabeled data is calculated as: u It is expressed as:
[0084] The second part is the loss function for the large model text features:
[0085]
[0086] Among them, q b is the probability distribution of the model after the weak augmentation of the unlabeled data, I(max(q b)) ≥ Γ is the threshold screening indicator function. If the prediction result after weak enhancement is greater than the threshold η, it is considered that the true label is the same as the prediction result. These labels are retained and allowed to participate in the calculation of the loss function for unlabeled waveforms, and the prediction result is recorded as a pseudo-label. If the prediction result does not exceed the threshold, the sample is discarded and not allowed to participate in the current training;
[0087] p m (y|A(u b )) is the probability distribution of the full waveform u b after strong enhancement by the model, is the classification prediction result. Cross-entropy is used to calculate the distance between the predicted value of the data after strong enhancement and the pseudo-label, and η u is used to specify the proportion of the unsupervised loss in the total loss function;
[0088] Step Six: Fuse the two features, optimize the result of emotion recognition using a fully connected layer, and then optimize the efficiency of the chatbot based on the emotion recognition result using GANs, a generator network, and a discriminator network. Among them, GANs are used for the classification network, the discriminator network is used to distinguish between real and fake images, and the generator network learns how to generate images according to text descriptions. Over time, the ability of the generator to generate realistic images will be enhanced through the adversarial training process;
[0089] Process Two: Optimize the conversation logic of the chatbot's multi-turn conversation based on the emotional state in the current conversation:
[0090] Step One: Based on Process One, identify the emotional state of the user's current conversation and grade the state. It allows for the recognition, transformation, and analysis of input queries and metadata to generate appropriate responses. To this end, the input data is aggregated, transformed into structured data, and analyzed to plan and execute the necessary operations;
[0091] Step Two: If the user's emotion is more negative, the robot's performance should be more positive. Combine some threshold hint words, input the current information and hint words into the large model, and then, through the MLP network, these inputs are complexly mapped to the probability space of emotion types:
[0092] Given the conversation U = {u 1 ,..., u t}, at each time t ∈ [1, t], identify y t , that is, the emotion of u t , which depends on the previous t - 1 predicted emotional states Y 1:t-1 = {y 1 ,..., y t-1}:
[0093]
[0094] Among them, M = {m 1 ,..., m t}, where m t ∈ {0, 1} is an index random variable. When m t = 1, it means that the speaker's emotion is different from the previous one, and m t = 0; otherwise, m 1 is defined as 1;
[0095] Step 3: In a multi-round conversation, the user's emotion may change due to certain factors. For example, when seeing a more preferred product or having a little hesitation about the price, there may be a small fluctuation in emotion. The chatbot should keenly perceive this change and adjust the introduction strategy in a timely manner;
[0096] In addition, the probability of identifying the t-th emotional state y t can be factorized as
[0097]
[0098] Here, p(y t |x t ), p(y t |z t , Y 1:t-1 ), and p(z t |Z 1:t-1 ) simulate the emotional distribution of the current discourse, the probability of emotional assignment based on the past predicted discourse, and the probability of the transition of emotional states over time respectively;
[0099] Step 4: The customer's emotion is easily affected by external factors, such as the evaluations of surrounding customers or small details in product displays. The chatbot needs to closely monitor the changes in the customer's emotion. When the customer hears someone making negative comments about the product beside them, the robot should promptly refute with facts and data to prove the quality and value of the product; if the customer is not satisfied with the details of the product display, the robot should quickly adjust the display method or provide a better display sample;
[0100] Step 5: In different application scenarios, the way of grading the user's emotion is different. For example, in the application of intelligent grid workers, it can also be measured from dimensions such as the intensity, positivity, negativity, and complexity of emotion;
[0101] Step 6: Through multiple rounds of iteration, if the customer takes the initiative to purchase, the conversation is terminated, or if the customer does not respond after the timeout, the conversation is terminated.
[0102] Using speech-to-text, the interface large model extracts the semantics of speech, uses Clip to extract text features, pre-trains image encoders and text encoders to predict which images in the dataset are paired with which texts, converts all classes of the dataset into text, and predicts the best pairing of the title with a given image, thereby improving the quality of the chatbot's responses.
[0103] The sample pair in step 1 of process 1 has a regular update function, and the regular update interval is one day.
[0104] The regularly updated function can continuously improve the big model's ability to recognize emotions.
[0105] Among them, the prompt words in step three of process one can be manually added, and the weights of the prompt words can be manually adjusted.
[0106] Since different users are in different shopping circles, there are many different words to express their preferences and desire to buy. The manually added design allows operators to add Internet hot words to the prompt words in a timely manner.
[0107] Among them, the t in step 2 of process 2 is set to ten seconds, and the setting of t can be changed manually.
[0108] The ten-second setting of t can not only ensure that there will be no large number of emotional misjudgments when communicating with users, but also reduce the occupancy rate of the large model system when the chatbot is in use.
[0109] Among them, the negative comments in step four of process two include keywords such as low cost performance and expensive, and the keywords of negative comments can be added manually.
[0110] The keyword design of negative reviews enables the chatbot to identify negative reviews around the user.
[0111] Among them, the application scenarios of step five in process two include holidays, long vacations, and online consumption festivals, etc. The application scenarios can be manually added and set.
[0112] Different application scenarios are closely related to users' shopping intentions. The addable design enables operators to add new application scenarios to the large model in a timely manner.
[0113] Among them, the timeout time of step six in process two is set to 300 seconds, and the timeout time can be adjusted.
[0114] The timeout setting can prevent the chatbot's system resources from being wasted.
[0115] A multi-turn speech response generation system based on multimodal feature optimization, comprising:
[0116] Voice Reply Production System: Used to accurately identify the user's emotions and make corresponding replies;
[0117] Among them, the voice reply production system includes:
[0118] Voice Processing Module: Collects and processes the user's voice information;
[0119] Model Establishment Module: Generates the user's voice time-frequency diagram;
[0120] Emotion Analysis Module: Analyzes the user's emotions based on the time-frequency diagram;
[0121] Output Module: Outputs corresponding answers according to the emotion data.
[0122] The voice processing module first processes the voice information input by the user into text information that can be recognized by the large model and sends it to the model establishment module. The model establishment module processes the information and generates a time-frequency diagram. Then, the emotion analysis module analyzes the emotion data of the user in the current state based on the time-frequency diagram. Finally, the emotion data is input into the output module and corresponding automatic reply content is generated therefrom.
[0123] Among them, the voice processing module includes: a voice collection node, a noise reduction processing node, and a translation input node; the emotion analysis module includes: an emotion analysis node and an emotion correction node.
[0124] The noise reduction processing node can eliminate the invalid content contained in the user's input voice, so that this invalid content will not have a negative impact on the automatic reply of the final chatbot.
[0125] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.
[0126] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A multi-round voice response generation method based on multimodal feature optimization, characterized in that: The following processes are included: Process 1: Combine the text features of the large model and the time-frequency features contained in the time series graph of speech extraction for classification, and achieve accurate emotion recognition in multiple rounds of conversations: Step 1: Collect some labeled sample pairs, including speech and text formats; Step 2: Use speech-to-text conversion to extract the semantics of speech through the interface model, use Clip to extract text features, pre-train image encoders and text encoders to predict which images in the dataset are paired with which texts, convert all classes in the dataset into text, and predict the best pairing of the title class with a given image; Step 3: Combine existing labels, such as active purchase, tendency, interest, and other prompt words with different weights for pre-training. Converting speech into a matrix consists of two steps. First, convert the signal from space (time amplitude) to polar coordinates through the following operations: Among them, L is the total length of the speech in the prompt word. This conversion is objective and retains the absolute time relationship. First, the data needs to be scaled to obtain the normalized form of the input X as follows: \begin{equation}X_{norm,\i} =\\frac{x_{i}-UB+(x_{i}-LB)}{UB-LB}, \end{equation} Among them, UB is the upper bound of the parameter, LB is the lower bound of the parameter, and the calculation of the Gramian matrix is as follows: Local expressions may not be able to grasp the overall tendency of sentiment words involving real emotions. The combination of sentiment words including contextual relationships has a great impact on the final classification performance. The speech signal is converted into a Gramian field and combined with a deep neural convolutional network to optimize the extraction of high-dimensional features. Step 4: Convert the signal into a time-frequency graph, fine-tune the time-frequency features of speech extraction based on spatiotemporal attention, and use the channel spatiotemporal attention mechanism to optimize the expression of emotional features in the frequency domain: M c (x)=ρ(mlp(avgpool(x))+mlp(maxpool(x))) M s (x)=ρ(con 7×7 ([avgpool(x));maxpool(x)])) M t (x0=ρ(con 7×7 (avgpool(x))+conv(maxpool(x))) Among them, M c (x), M s (x), M t (x) represents the weights of channel attention, spatial attention, and temporal attention, respectively. ρ is the activation function, mlp is the multi-layer perceptron, and f 7×7 It is a one-dimensional convolution operation that considers information in three dimensions: channel, space, and time. It improves the model’s ability to represent spatiotemporal data and adaptively captures the dynamic associations between nodes in the spatial dimension. Step 5: Combine some labeled samples to train the emotion classification model. The loss function consists of two parts: \begin{equation}loss=l_s+\eta_ul_u\end{equation} The first part l_s is the sound feature category loss function: \begin{equation}l_s=\frac{1}{B}\sum_{b =1}^{B}H(p_b,p_m(y|a(x_b)))\end{equation} Where B is the number of full waveforms with labels, p b is the emotion category label, p m (y|a(x b )) is the probability distribution of emotion categories after the sample is predicted by the model. The unlabeled data is forward propagated after two strong and weak data enhancements respectively. The data after strong enhancement is input into the network G(σ s_t-1 ) calculates the predicted value, and the other branch inputs the weakly enhanced data into the network G(\sigma_{\varsigma\_t-1}) to calculate the predicted value, and the unlabeled data calculates the loss l u It is expressed as: The second part is the large model text feature loss function: \begin{equation}l_u =\frac{1}{\mu B}\sum_{1}^{\mu B}I(max(q_b)\ge \Gamma)H(\hat{q_b},p_m(y|A(u_b)))\end{equation} where q b is the probability distribution of the model after weak enhancement of unlabeled data, I(max(q b ))≥Γ is a threshold screening indicative function. If the prediction result after weak enhancement is greater than the threshold η, it is considered that the true label is the same as the prediction result. These labels are retained and allowed to participate in the loss function calculation of the unlabeled waveform. At the same time, the prediction result is recorded as a pseudo label. If the prediction result does not exceed the threshold, the sample is discarded and is not allowed to participate in the current training. p m (y|A(u b )) is the full waveform u b The probability distribution of the model after strong enhancement, is the classification prediction result, and the cross entropy is used to calculate the distance between the predicted value of the data after strong enhancement and the pseudo label, and through η u To specify the proportion of unsupervised loss in the total loss function; Step 6: Fusion of the two features, using full connectivity to optimize the results of emotion recognition, and then using GANs, the generator network, and the discriminator network to optimize the efficiency of the chatbot based on the emotion recognition results. GANs are used for the classification network, the discriminator network is used to distinguish between real and fake images, and the generator network learns how to generate images based on text descriptions. Over time, the generator's ability to generate realistic images will be enhanced through the adversarial training process. Process 2: Optimize the conversation logic of the chatbot's multi-round conversation based on the emotional state in the current conversation: Step 1: Identify the emotional state of the user in the current session based on process one and classify the state. It allows to identify, transform and analyze the input query and metadata to generate appropriate responses. To this end, the input data is aggregated, transformed into structured data and analyzed to plan and execute necessary actions. Step 2: If the user's emotion is more negative, the robot's performance should be more positive. Combined with some threshold prompt words, the current information and prompt words are input into the large model. Then, through the MLP network, these inputs are complexly mapped to the probability space of emotion types: Given a dialogue U = {u1,...,u t }, identify y at each time t∈[1,t] t , that is u t The emotion of , which depends on the previous t-1 predicted emotional states Y in U 1:t-1 ={y1, ..., y t-1 }: Where M = {m1, ..., m t },m t ∈{0, 1} is the indicator random variable, m t =1 means the speaker’s emotion is different from the previous one, m t =0; otherwise, m1 is defined as 1; Step 3: During multiple rounds of conversations, the user's emotions may change due to certain factors, such as seeing a more preferred product or being a little hesitant about the price. The chatbot should be sensitive to such changes and adjust the introduction strategy in time. In addition, identifying the tth emotional state y t The probability of can be factorized into Here p(y t |x t ),p(y t |z t , Y 1:t-1 ),p(z t |Z 1:t-1 ) respectively simulate the emotion distribution of the current text, the probability of emotion distribution based on the past predicted text, and the probability of emotional state transition across time; Step 4: Customer emotions are easily affected by external factors, such as the comments of surrounding customers or small details of product display. Chatbots need to pay close attention to changes in customer emotions. When customers hear someone nearby making negative comments about a product, the robot should refute it in a timely manner and use facts and data to prove the quality and value of the product. If the customer is not satisfied with the details of the product display, the robot should quickly adjust the display method or provide better display samples. Step 5: In different application scenarios, the user's emotions are graded differently. For example, in the smart grid worker application, the emotions can also be measured from the dimensions of intensity, positivity, negativity, and complexity. Step 6: After multiple rounds of iterations, if the customer actively purchases, the session is terminated; or if the customer does not respond within the time limit, the session is terminated.
2. The method for generating multi-round voice responses based on multimodal feature optimization according to claim 1, characterized in that: The sample pair described in step 1 of process 1 has a periodic update function, and the periodic update interval is one day.
3. The method for generating multi-round voice responses based on multimodal feature optimization according to claim 1, characterized in that: The prompt words described in step 3 of process 1 can be manually added, and the weights of the prompt words can be manually adjusted.
4. The method for generating multi-round voice responses based on multimodal feature optimization according to claim 1, characterized in that: In the step of process 2, the t is set to ten seconds, and the setting of t can be manually changed.
5. The method for generating multi-round voice responses based on multimodal feature optimization according to claim 1, characterized in that: The negative evaluation described in step 4 of process 2 includes keywords such as low cost performance and expensive, and the keywords of the negative evaluation can be added manually.
6. The method for generating multi-round voice responses based on multimodal feature optimization according to claim 1, characterized in that: The application scenarios described in step 5 of process 2 include holidays, long vacations, and online consumption festivals, etc. The application scenarios can be manually added and set.
7. The method for generating multi-round voice responses based on multimodal feature optimization according to claim 1, characterized in that: The timeout period in step 6 of process 2 is set to 300 seconds, and the timeout period can be adjusted.
8. A multi-round voice response generation system based on multimodal feature optimization, characterized in that: include: Voice response production system: used to accurately identify user emotions and make corresponding responses; Among them, the voice response production system includes: Voice processing module: collects and processes user voice information; Model building module: Generate user speech time-frequency graph; Sentiment analysis module: analyzes user emotions based on time-frequency graphs; Output module: Output corresponding answers based on emotional data.
9. The multi-round voice response generation system based on multimodal feature optimization according to claim 8, characterized in that: The speech processing module includes: a speech collection node, a noise reduction processing node, and a translation input node; the emotion analysis module includes: an emotion analysis node and an emotion correction node.
Citation Information
Patent Citations
Multi-mode based emotion recognition method
CN108805089A
Multi-round dialogue semantic comprehension subsystem based on multi-modal emotion identification system
CN108877801A
Voice signal analysis sub-system based on multi-modal emotion identification system
CN108899050A
Intelligent customer service voice rhythm adjusting method and device, equipment and storage medium
CN114842880A
Intelligent voice dialogue scene verbal skill intervention method and system based on customer portrait
CN116049360A