Intelligent voice interaction method and system
By processing the signal-to-noise ratio and semantic importance of user-questioned speech in voice interaction, combined with the prediction probability and similarity of the prediction model, the problem of inaccurate response text in voice interaction under the influence of background noise is solved, and the user experience is improved.
Patent Information
- Application Number
- CN202510661257.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-22
AI Technical Summary
In voice interaction, noisy background sounds will cause distortion of voice information and the inaccurate identification of questions raised by users, resulting in inaccurate response text and reduced user experience.
By obtaining the signal-to-noise ratio in the voice of the user's question, deleting the time frame with the signal-to-noise ratio smaller than the preset value, obtaining the voice signal and converting it into voice text. Calculate the semantic importance of words in pronunciation text, use semantic importance weighted word vectors to obtain problem semantics, and combine the prediction model trained by historical interaction information to calculate the prediction probability and similarity of standard problems, and determine the target probability to accurately identify the question.
It realizes accurate identification of user questions in a noisy context, obtains accurate response text, and improves the user experience of voice interaction.
Smart Images

Figure CN120183407A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of voice interaction technologies, and particularly to an intelligent voice interaction method and system. Background Art
[0002] Voice interaction involves voice recognition technology. By using voice recognition technology, the questions raised by users are recognized, the response content is determined according to the questions raised, and the response content is converted into voice information, so as to play the voice information to the users.
[0003] Currently, the patent application document with the publication number of CN117555916A discloses a voice interaction method and system based on natural language processing. The method therein includes: obtaining the voice information input by the user, and converting the voice information into natural language text information based on voice recognition technology; performing text preprocessing on the natural language text information to obtain the natural language text information after text preprocessing; converting the natural language text information after text preprocessing into structured query language based on the NLP semantic parsing model; obtaining the corresponding target data based on the structured query language; inputting the corresponding target data into the optimized natural language generation model, and converting the corresponding target data into a response text based on the optimized natural language generation model; converting the response text into a response voice based on voice synthesis technology, and outputting the response voice.
[0004] The above method directly converts the voice information input by the user into structured query language. After inputting the target data corresponding to the structured query language into the optimized natural language generation model, a response text is obtained. However, during the process of collecting the voice information input by the user, there may be noisy background sounds, which will cause the voice information to be distorted and unable to accurately recognize the questions raised by the user, resulting in inaccurate response texts and poor user experience during voice interaction. Summary of the Invention
[0005] In order to solve the technical problem of inaccurate response texts, this application provides an intelligent voice interaction method and system, which can accurately recognize the questions raised by users, and then obtain accurate response texts, improving the user experience during voice interaction.
[0006] In the first aspect of the present application, an intelligent voice interaction method is provided. The interaction method includes: obtaining the signal-to-noise ratio of each time frame in the user's question voice in the current round, deleting the time frames with a signal-to-noise ratio less than a preset value to obtain a voice signal, and converting the voice signal into a voice text; calculating the semantic importance of each word in the voice text, and weighted summing the word vectors of each word according to the semantic importance to obtain the question semantics, where the semantic importance is positively correlated with the TF-IDF value of the word in the standard question set and the energy value of the corresponding voice segment; training a prediction model based on historical interaction information, where the prediction model is used to obtain the prediction probability of each standard question in the current round according to the standard questions and response texts of each round before the current round of the user; calculating the similarity between the question semantics and each standard question, taking the weighted sum of the similarity and the prediction probability as the target probability of each standard question, and selecting the standard question with the maximum target probability as the question to be asked in the current round; inputting the question to be asked and the standard questions and response texts of each round before the current round into a text generation large model to obtain the response text of the current round, and converting the response text into a broadcast voice.
[0007] Collect the user's question voice in the current round, and obtain the signal-to-noise ratio of each time frame in the question voice. The smaller the signal-to-noise ratio, the greater the background noise of the corresponding time frame. After deleting the time frames with a signal-to-noise ratio less than a preset value, a voice signal is obtained, and the voice signal is converted into a voice text. Different words in the voice text have different abilities to represent semantic information. Therefore, the semantic importance of each word is accurately measured according to the TF-IDF value of the word in the standard question set and the energy value of the corresponding voice segment, and the word vectors of each word are weighted and summed using the normalized semantic importance to accurately obtain the question semantics of the question voice. According to the similarity between the question semantics and each standard question, it can be judged which standard question the user's question voice in the current round belongs to. At the same time, the time series of the standard questions and response texts of each round before the current round of the user is input into the prediction model to obtain the prediction probability of each standard question in the current round. The target probability of each standard question is determined by comprehensively considering the similarity and the prediction probability, and the standard question with the maximum target probability is selected as the question to be asked in the current round to achieve accurate identification of the question to be asked in the current round. Furthermore, the question to be asked is input into the text generation large model to obtain an accurate response text, improving the user experience during voice interaction.
[0008] Preferably, before calculating the semantic importance of each word in the voice text, the interaction method further includes: in response to the proportion of the deleted time frames in the total duration of the question voice being greater than a preset proportion, re-collecting the user's question voice in the current round.
[0009] When the proportion of the deleted time frame in the total duration of the question voice is greater than the preset proportion, it indicates that there is more background noise in the question voice, and the user is reminded to re-ask the question to re-collect the user's question voice in the current round to ensure the validity of the question voice.
[0010] Preferably, calculating the semantic importance of each word in the speech text includes: calculating the TF-IDF value of the word in the standard question set; performing forced alignment on the speech text and the speech information to obtain the word corresponding speech segment; using the ratio of the average short-time energy of the speech segment to the average short-time energy of the speech signal as the energy value of the word ; the semantic importance of the word is the product of the TF-IDF value and the energy value.
[0011] The semantic importance is comprehensively considered from two parts: the ability to distinguish standard questions and the degree of attention when the user expresses, so as to accurately quantify the semantic importance of each word in the speech text.
[0012] Preferably, the Aeneas tool or the Gentle tool is used to achieve forced alignment.
[0013] Preferably, obtaining the question semantics includes: using the ratio of the semantic importance of any word to the sum of the semantic importance of all words as the normalized semantic importance; weighted summing the word vectors of each word according to the normalized semantic importance to obtain the question semantics.
[0014] Preferably, the historical interaction information includes the standard questions and response texts of all rounds in the historical interaction process; training the prediction model based on the historical interaction information includes: using the Q&A records before any round in the historical interaction information as training samples, and using the standard question of the any round as the class label; inputting the training samples into the prediction model to obtain the output result, and performing gradient descent on the prediction model according to the cross-entropy loss between the output result and the class label to update the prediction model until the cross-entropy loss is less than the preset loss or the number of updates is greater than the preset number, then the training is completed.
[0015] Training the prediction model based on the Q&A records of all rounds in the historical interaction information, the trained prediction model can accurately obtain the prediction probability of each standard question in the current round according to the standard questions and response texts of each round before the current round of the user.
[0016] Preferably, calculating the similarity between the question semantics and each standard question includes: obtaining the standard semantics of each standard question, where the standard semantics is the result of weighted summing the word vectors of each word in the standard question according to the normalized TF-IDF value; calculating the similarity between the question semantics and the standard semantics of each standard question.
[0017] Preferably, the standard problem target probability of is: ; is the similarity between the problem semantics and the standard problem , is the predicted probability of the standard problem , is the adjustment coefficient.
[0018] The similarity can be regarded as the probability value of the standard problem obtained from the voice information of the question voice in the current round; while the predicted probability can be regarded as the probability value of the standard problem obtained from the time series of the standard problems and response texts in each round before the current round; comprehensively considering the similarity and the predicted probability to accurately obtain the target probability of each standard problem Preferably, the adjustment coefficient is negatively correlated with the number of rounds in the current round and negatively correlated with the proportion of the deleted time frames in the total duration of the question voice.
[0019] The adjustment coefficient is used to control the influence degree of the similarity between the problem semantics and the standard problem on the target probability; the larger the adjustment coefficient, the more the target probability depends on the question voice in the current round; the adjustment coefficient is determined according to the proportion of the deleted time frames in the total duration of the question voice and the number of rounds in the current round to ensure the accurate determination of the target probability.
[0020] In the second aspect of the present application, an intelligent voice interaction system is further provided, including a processor and a memory, where the memory stores computer program instructions, and when the computer program instructions are executed by the processor, an intelligent voice interaction method according to the first aspect of the present application is implemented.
[0021] The technical solution of the present application has the following beneficial technical effects: Collect the user's question voice in the current round, and obtain the signal-to-noise ratio of each time frame in the question voice. The smaller the signal-to-noise ratio, the greater the background noise of the corresponding time frame. After deleting the time frames with a signal-to-noise ratio less than the preset value, a voice signal is obtained, and the voice signal is converted into a voice text. Different words in the voice text have different abilities to represent semantic information. Therefore, the semantic importance of each word is accurately measured according to the TF-IDF value of the word in the standard question set and the energy value of the corresponding voice segment, and the word vectors of each word are weighted and summed using the normalized semantic importance to accurately obtain the question semantics of the question voice. According to the similarity between the question semantics and each standard question, it can be judged which standard question the user's question voice in the current round belongs to. At the same time, the time series of the standard questions and response texts in each round before the current round of the user are input into the prediction model to obtain the prediction probability of each standard question in the current round. The target probability of each standard question is determined by integrating the similarity and the prediction probability, and the standard question with the maximum target probability is selected as the question of the current round, realizing the accurate identification of the question of the current round. Furthermore, the question is input into the text generation large model to obtain an accurate response text, improving the user experience during voice interaction. Description of the Drawings
[0022] Figure 1 is a flowchart of an intelligent voice interaction method according to an embodiment of the present application.
[0023] Figure 2 is a structural block diagram of an intelligent voice interaction system according to an embodiment of the present application. Detailed Embodiments
[0024] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0025] According to the first aspect of the present application, the present application provides an intelligent voice interaction method. In the scenario of intelligent online customer service, users will ask questions, and the intelligent customer service system generates a response text according to the questions asked and converts the response text into a broadcast voice to complete an interactive question and answer in one round. Users and the intelligent customer service system will conduct multiple rounds of interactive question and answers until the user no longer asks new questions, ending the service process of the online customer service.
[0026] Figure 1 is a flowchart of an intelligent voice interaction method according to an embodiment of the present application. As Figure 1 shown, the intelligent voice interaction method includes steps S101 to S105, which will be described in detail below.
[0027] S101. Obtain the signal-to-noise ratio of each time frame in the user's question voice in the current round, delete the time frames with a signal-to-noise ratio less than a preset value to obtain a voice signal, and convert the voice signal into a voice text.
[0028] In one embodiment, during the voice interaction process, multiple interactive Q&A sessions will be carried out. Each interactive Q&A session includes two links: user question and system response. During the voice interaction process, collect the user's question voice in the current round. Since the process of collecting the question voice will be affected by the ambient sound, in a noisy background, some question voices may not be accurately collected. Therefore, it is necessary to obtain the signal-to-noise ratio of each time frame in the user's question voice in the current round.
[0029] Among them, the question voice is segmented into multiple time frames. Generally, the length of the time frame is any value from 20 milliseconds to 40 milliseconds. The signal-to-noise ratio is an index to measure the signal quality and is used to evaluate the quality of each time frame of the voice signal relative to the background noise. The larger the signal-to-noise ratio, the higher the voice quality of the corresponding time frame.
[0030] The preset value is 10. If the signal-to-noise ratio of any time frame is less than the preset value, it means that the time frame contains a large amount of background noise and the question voice cannot be accurately collected. Delete the voice information of this time frame to obtain a voice signal, and convert this voice signal into a voice text.
[0031] S102. Calculate the semantic importance of each word in the voice text, and weighted sum the word vectors of each word according to the semantic importance to obtain the question semantics. The semantic importance is positively correlated with the TF-IDF value of the word in the standard question set and the energy value of the corresponding voice segment.
[0032] In one embodiment, before calculating the semantic importance of each word in the voice text, the interactive method further includes: in response to the proportion of the deleted time frames in the total duration of the question voice being greater than a preset proportion, re-collect the user's question voice in the current round.
[0033] It can be understood that the preset proportion is 0.5. When the proportion of the deleted time frames in the total duration of the question voice is greater than the preset proportion, it means that there is more background noise in the question voice, reminding the user to re-ask the question to re-collect the user's question voice in the current round.
[0034] In one embodiment, calculating the semantic importance of each word in the voice text includes: calculating the TF-IDF value of the word in the standard question set; performing forced alignment on the voice text and the voice information to obtain the voice segment corresponding to the word ; taking the ratio of the average short-time energy of the voice segment to the average short-time energy of the voice signal as the word ; The energy value; the word The semantic importance of is the product of the TF-IDF value and the energy value.
[0035] Among them, tools such as Aeneas tool, Gentle tool or other alignment tools can be used to achieve forced alignment of speech text and speech information, and this application does not make restrictions.
[0036] The word The corresponding speech segment includes at least one time frame, and the short-time energy of each time frame can be calculated. Since the stressed frame usually has a higher short-time energy, the ratio of the average short-time energy of the speech segment to the average short-time energy of the speech signal can be used as the energy value. If the energy value of a word is larger, it means that the user has a higher pitch when pronouncing the word, attaches more importance to the pronunciation of the word, and the word is more important in the speech text.
[0037] The calculation method of the TF-IDF value is well-known to those skilled in the art and will be briefly introduced here. The word The TF-IDF value of is the product of the word frequency and the inverse document frequency; the word frequency is the frequency of the word appearing in the speech text, which reflects the importance of the word in the speech text; the inverse document frequency IDF value can measure the universality of the word . The fewer standard questions a word appears in, the higher the IDF value, indicating that its ability to distinguish standard questions is stronger. Therefore, the TF-IDF value can reflect the ability to distinguish the word to distinguish standard questions.
[0038] In this way, the semantic importance is comprehensively considered from two parts: the ability to distinguish standard questions and the degree of attention when the user expresses, so as to accurately quantify the semantic importance of each word in the speech text.
[0039] In one embodiment, after obtaining the semantic importance of each word, the obtaining of the question semantics includes: taking the ratio of the semantic importance of any word to the sum of the semantic importance of all words as the normalized semantic importance; weighted summing the word vectors of each word according to the normalized semantic importance to obtain the question semantics.
[0040] Among them, the word vector is obtained through the Word2Vec model or the Bert model.
[0041] S103. Train a prediction model according to historical interaction information. The prediction model is used to obtain the prediction probability of each standard question in the current round according to the standard questions and response texts of each round before the current round of the user.
[0042] In one embodiment, the prediction model includes a temporal feature sub-model and a classification sub-model. The temporal feature sub-model is used to extract temporal features from the Q&A records of each round before any round, obtain the interaction vector of any round, and input the interaction vector into the classification sub-model to obtain the prediction probabilities of each standard question in any round; a Q&A record includes a standard question and a response text corresponding to the standard question.
[0043] Among them, the temporal feature sub-model is an LSTM or GRU model. Input the interaction vector of any round into the classification sub-model to obtain the prediction probabilities of each standard question, and the sum of the prediction probabilities of each standard question is 1.
[0044] In one embodiment, the historical interaction information includes the Q&A records of all rounds in the historical interaction process, and the Q&A record includes a standard question and a response text; training the prediction model based on the historical interaction information includes: using the Q&A records before any round in the historical interaction information as training samples, and using the standard question of any round as the class label; inputting the training samples into the prediction model to obtain an output result, and performing gradient descent on the prediction model according to the cross-entropy loss between the output result and the class label to update the prediction model until the cross-entropy loss is less than a preset loss or the number of updates is greater than a preset number of times, and then the training is completed.
[0045] Among them, the preset loss is 0.01 and the preset number of times is 300.
[0046] In this way, the prediction model is trained based on the Q&A records of all rounds in the historical interaction information. The trained prediction model can accurately obtain the prediction probabilities of each standard question in the current round based on the standard questions and response texts of each round before the current round of the user.
[0047] S104, calculate the similarity between the question semantics and each standard question, take the weighted sum of the similarity and the prediction probability as the target probability of each standard question, and select the standard question with the maximum target probability as the question to be asked in the current round.
[0048] In one embodiment, calculating the similarity between the question semantics and each standard question includes: obtaining the standard semantics of each standard question, where the standard semantics is the result of weighted summing the word vectors of each word in the standard question based on the normalized TF-IDF values; calculating the similarity between the question semantics and the standard semantics of each standard question.
[0049] Understandably, the greater the similarity of the standard question, the greater the probability that the question voice in the current round belongs to the standard question. The similarity can be regarded as the probability value of the standard question obtained from the voice information of the question voice in the current round; while the prediction probability can be regarded as the probability value of the standard question obtained from the time series of the standard questions and response texts in the previous rounds before the current round; the target probability of each standard question is obtained by combining the similarity and the prediction probability.
[0050] Specifically, for the standard question the target probability is: ; is the similarity between the question semantics and the standard question , is the prediction probability of the standard question , is the adjustment coefficient.
[0051] Among them, the adjustment coefficient is used to control the influence degree of the similarity between the question semantics and the standard question on the target probability; the larger the adjustment coefficient, the more the target probability depends on the question voice in the current round. To ensure the accurate determination of the target probability, the value of the adjustment coefficient can be determined according to the proportion of the deleted time frames in the total duration of the question voice. The smaller the proportion of the deleted time frames in the total duration of the question voice, the more accurate the question voice can be collected, and the larger the adjustment coefficient should be; at the same time, the adjustment coefficient is also related to the round number of the current round. The larger the round number of the current round, the more standard questions and response texts the prediction model can collect, and the accuracy of the prediction probability will also increase accordingly, then the adjustment coefficient should be smaller; therefore, the adjustment coefficient is negatively correlated with the round number of the current round and negatively correlated with the proportion of the deleted time frames in the total duration of the question voice.
[0052] Specifically, the value of the adjustment coefficient is from 0 to 1; the adjustment coefficient is: , is the proportion of the deleted time frames in the total duration of the question voice, is the round number of the current round.
[0053] In one embodiment, after obtaining the target probability of each standard question, the standard question with the maximum target probability is selected as the question of the current round, which can accurately obtain the question of the current round in the case of large background noise, and this question is any one in the standard question set.
[0054] S105, Input the question of the current round and the standard questions and response texts of the previous rounds before the current round into the text generation large model to obtain the response text of the current round, and convert the response text into a broadcast voice.
[0055] In one embodiment, after obtaining the user's question in the current round, the question and the standard questions and response texts of each previous round are input into a text generation large model to obtain the response text of the current round, and then the response text of the current round is converted into a voice for broadcasting to achieve intelligent voice interaction. Among them, the text generation large model can be an existing large model such as DeepSeek or GPT.
[0056] According to the second aspect of the present application, the present application also provides an intelligent voice interaction system. Figure 2 It is a structural block diagram of an intelligent voice interaction system according to an embodiment of the present application. As Figure 2 shown, the system 50 includes a processor and a memory. The memory stores computer program instructions, and when the computer program instructions are executed by the processor, an intelligent voice interaction method according to the first aspect of the present application is implemented. The system also includes other components well known to those skilled in the art such as a communication bus and a communication interface, and their settings and functions are known in the art, so they will not be described in detail here.
[0057] It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can be made, and these all belong to the protection scope of the present application.
Claims
1. An intelligent voice interaction method, characterized in that, The interactive method includes: Obtain the signal-to-noise ratio of each time frame in the user's question voice in the current round, delete the time frames with a signal-to-noise ratio less than a preset value to obtain a voice signal, and convert the voice signal into a voice text; Calculate the semantic importance of each word in the voice text, and weighted sum the word vectors of each word according to the semantic importance to obtain the question semantics. The semantic importance is positively correlated with the TF-IDF value of the word in the standard question set and the energy value of the corresponding voice segment; Train a prediction model based on historical interaction information. The prediction model is used to obtain the prediction probability of each standard question in the current round according to the standard questions and response texts in each round before the current round of the user; Calculate the similarity between the question semantics and each standard question, take the weighted sum of the similarity and the prediction probability as the target probability of each standard question, and select the standard question with the maximum target probability as the question for the current round; Input the question for the current round and the standard questions and response texts in each round before the current round into a large text generation model to obtain the response text for the current round, and convert the response text into a voice for broadcasting.
2. The intelligent voice interaction method according to claim 1, characterized in that, Before calculating the semantic importance of each word in the voice text, the interactive method further includes: In response to the proportion of the deleted time frames in the total duration of the question voice being greater than a preset proportion, re-collect the user's question voice in the current round.
3. The intelligent voice interaction method according to claim 1, characterized in that, Calculating the semantic importance of each word in the voice text includes: Calculate the TF-IDF value of words in the speech text in the standard question set; Perform forced alignment on the speech text and speech information to obtain words corresponding speech segments; The ratio of the average short-time energy of the speech segment to the average short-time energy of the speech signal is used as the energy value of the word ; the semantic importance of the word is the product of the TF-IDF value and the energy value.
4. The intelligent voice interaction method according to claim 3, characterized in that, Implement forced alignment using the Aeneas tool or the Gentle tool.
5. The intelligent voice interaction method according to claim 1, characterized in that, The obtaining of the question semantics includes: Take the ratio of the semantic importance of any word to the sum of the semantic importance of all words as the normalized semantic importance; Weighted sum the word vectors of each word according to the normalized semantic importance to obtain the question semantics.
6. The intelligent voice interaction method according to claim 1, characterized in that, The historical interaction information includes the standard questions and response texts in all rounds during the historical interaction process; Training the prediction model based on historical interaction information includes: Use the Q&A records before any round in the historical interaction information as training samples, and use the standard question of the any round as the class label; Input the training samples into the prediction model to obtain an output result, and perform gradient descent on the prediction model according to the cross-entropy loss between the output result and the class label to update the prediction model until the cross-entropy loss is less than a preset loss or the number of updates is greater than a preset number of times, then the training is completed.
7. The intelligent voice interaction method according to claim 1, characterized in that, Calculating the similarity between the question semantics and each standard question includes: Obtain the standard semantics of each standard question. The standard semantics is the result of weighted summing the word vectors in the standard question according to the normalized TF-IDF values; Calculate the similarity between the question semantics and the standard semantics of each standard question.
8. The intelligent voice interaction method according to claim 1, characterized in that, Standard issue Target probability is as follows: ; is the similarity between the problem semantics and the standard problem , is the prediction probability of the standard problem , is the adjustment coefficient.
9. The intelligent voice interaction method according to claim 8, characterized in that, The adjustment coefficient is negatively correlated with the number of rounds in the current round and negatively correlated with the proportion of the deleted time frames in the total duration of the question voice.
10. An intelligent voice interaction system, characterized in that, It includes a processor and a memory. The memory stores computer program instructions, and when the computer program instructions are executed by the processor, it implements an intelligent voice interaction method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Voice interaction method and system based on natural language processing
CN117555916A
Question and answer feedback method and device based on deep learning, equipment and storage medium
CN109829038A
Intelligent question and answer method and system
CN114637760A
Question and answer system response method and device, equipment, medium and product
CN115146124A
Conference summary generation method based on AI identification
CN117151047A
Cited By
Multi-dimensional AI platform intelligent voice response system using voice synthesis technology
CN120808785A