Intelligent customer service voice interaction method and system based on voice recognition

By using hidden Markov model and semantic correlation calculation in speech recognition technology, the problem of low efficiency and accuracy of speech signal recognition in the prior art is solved, and more accurate and natural voice interaction is achieved, and user experience is improved.

CN120199247AActive Publication Date: 2025-06-24HUAZE ZHONGXI (BEIJING) TECH DEV CO LTD

Patent Information

Application Number
CN202510678456.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-06-24
Estimated Expiration
2045-05-26

AI Technical Summary

Technical Problem

The existing speech recognition technology is low in the recognition efficiency and accuracy of speech signals, mainly because the n-gram model cannot understand the actual meaning or grammatical structure between words in text, and ignores the deep semantics between words and longer range of context information.

Method used

By collecting voice signals and historical text data in voice interaction scenarios in real time, dividing the voice signals into multiple signal segments, and using the hidden Markov model to obtain phoneme sequences and their corresponding candidate texts. Compute the matching degree and semantic correlation degree of vocabulary in each candidate text, combine the correction of candidate probability, identify the speech signal and perform voice interaction.

Benefits of technology

It improves the accuracy of voice recognition, enables intelligent customer service to understand and respond to user voice commands more accurately, improves the accuracy, coherence and naturalness of voice interactions, and enhances user experience satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120199247A_ABST
    Figure CN120199247A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of voice recognition, in particular to an intelligent customer service voice interaction method and system based on voice recognition, and the method comprises the steps: collecting voice signals of a user in a voice instruction issuing process in a voice interaction scene in real time, and text data generated when the user interacts in a historical period, and forming a historical text set; dividing the voice signal into a plurality of signal segments; obtaining each instruction signal segment; obtaining a phoneme sequence corresponding to each instruction signal segment and all candidate texts corresponding to the phoneme sequence; calculating a matching degree and a semantic association degree of each vocabulary in each candidate text to obtain a candidate probability of each candidate text; and correcting the candidate probability to obtain a corrected candidate probability corresponding to each candidate text under each instruction signal segment, identifying a voice signal and performing voice interaction. According to the invention, the accuracy of voice recognition is improved, so that the intelligent customer service can more accurately understand and respond to the voice instruction of the user, and the accuracy of voice interaction is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of speech recognition technology, and in particular to an intelligent customer service speech interaction method and system based on speech recognition. Background Art

[0002] IoT voice recognition technology is a technical system that converts voice input into recognizable data or commands and responds through IoT devices. In the IoT environment, voice recognition not only relies on traditional voice signal processing and natural language processing technologies, but also combines the interconnection of IoT devices to achieve intelligent interaction between voice and devices, applications and services. Through voice recognition, users can directly use voice commands to interact with intelligent customer service, which simplifies the operation process and improves the user experience.

[0003] In the process of voice interaction, the hidden Markov model is usually used to screen the candidate texts corresponding to the voice signal, and the n-gram model is usually used to sort the candidate texts. However, the n-gram model is based on statistics and does not understand the actual meaning or grammatical structure between words in the text. The n-gram model can only capture the co-occurrence relationship between words, ignoring the deep semantics between words, and does not capture the longer-range context or global information in the sentence, which leads to low efficiency and accuracy in the recognition of voice signals. Summary of the invention

[0004] In order to solve the above technical problems, an intelligent customer service voice interaction method and system based on voice recognition are provided to solve the existing problems.

[0005] The solution to the technical problem of this application is to provide an intelligent customer service voice interaction method and system based on voice recognition, including the following steps: In a first aspect, an embodiment of the present application provides an intelligent customer service voice interaction method based on voice recognition, the method comprising the following steps: Real-time collection of voice signals from users issuing voice commands in voice interaction scenarios, as well as text data generated by users interacting in historical periods, to form a historical text set; Based on the signal strength in the voice signal, the voice signal is divided into a plurality of signal segments; and all the signal segments are screened to obtain each command signal segment; Using a hidden Markov model, the phoneme sequence corresponding to each instruction signal segment and all the corresponding candidate texts are obtained; Segment each candidate text, analyze the number of words represented by the combination of multiple phonemes corresponding to each word in each candidate text in the phoneme sequence, and calculate the matching degree of each word in each candidate text under each instruction signal segment; Determine the semantic correlation degree of each word in each candidate text under each instruction signal segment based on the distance between words and the number of co-occurrences in the text data where each word in each candidate text co-occurs with its adjacent words in the historical text set, and combine with the matching degree to obtain the candidate probability of each candidate text under each instruction signal segment; Correct the candidate probability based on the time length between each instruction signal segment and different previous instruction signal segments and the similarity of their candidate texts, to obtain the corrected candidate probability corresponding to each candidate text under each instruction signal segment, identify the voice signal and perform voice interaction.

[0006] Preferably, the dividing the voice signal into multiple signal segments includes: marking the moment when the signal intensity in the voice signal is 0 as the silent moment; if the signal intensities of two moments adjacent to the silent moment in the voice signal are not all 0, then marking the silent moment as the zero point; using the zero point as the segmentation point to divide the voice signal into multiple signal segments.

[0007] Preferably, the further obtaining method of each instruction signal segment is: calculating the average value of all signal intensities in each signal segment, clustering the average values of all signal segments to obtain two clustering clusters; marking all signal segments within the clustering cluster to which the maximum value of the average values of all signal segments belongs as each instruction signal segment.

[0008] Preferably, the obtaining the phoneme sequence corresponding to each instruction signal segment and all candidate texts corresponding thereto includes: Preprocessing each instruction signal segment, extracting the mel-frequency cepstral coefficients of each preprocessed instruction signal segment as the observation sequence; based on the observation sequence, obtaining the hidden state sequence through a hidden Markov model; mapping each hidden state in the hidden state sequence to the corresponding phoneme in the phoneme dictionary to obtain the phoneme sequence; Recording all texts corresponding to the phoneme sequence through the phoneme dictionary as each candidate text.

[0009] Preferably, the calculating the matching degree of each word in each candidate text under each instruction signal segment includes: Combining the phonemes corresponding to each word in each candidate text in the phoneme sequence to form a local phoneme sequence; Combining each phoneme in the local phoneme sequence with all the phonemes before it to form each subsequence; wherein, the subsequence formed by the last phoneme in the local phoneme sequence and all the phonemes before it is the local phoneme sequence; Combining all the phonemes in the local phoneme sequence with the next phoneme corresponding to the last phoneme in the phoneme sequence to form an extended sequence; marking all the subsequences except the local phoneme sequence and the extended sequence in all the subsequences as each characteristic phoneme sequence; Count the number of all words corresponding to the local phoneme sequence in the phoneme dictionary; count the number of all words corresponding to each characteristic phoneme sequence in the phoneme dictionary, which is denoted as the vocabulary size. The matching degree is the reciprocal of the product of the sum of the vocabulary sizes of all characteristic phoneme sequences and the number.

[0010] Preferably, determining the semantic association degree of each word in each candidate text under each instruction signal segment includes: Denote the words adjacent to any word in each candidate text as neighboring words; count the number of times the any word and its neighboring words co-occur in all text data in the historical text set. Calculate the cumulative sum of the distances between the any word and its neighboring words in all text data where they co-occur in the historical text set. Take the product of the number of times and the cumulative sum as the correlation degree between the any word and its neighboring words. The semantic association degree is the sum of the correlation degrees between the any word and all its neighboring words in each candidate text.

[0011] Preferably, the candidate probability of each candidate text under each instruction signal segment is the normalized result of the sum of the products of the matching degrees and the semantic association degrees of all words in each candidate text.

[0012] Preferably, obtaining the corrected candidate probability corresponding to each candidate text under each instruction signal segment includes: Obtain the word vector of each candidate text. The corrected candidate probability corresponding to the th candidate text under the th instruction signal segment is calculated as: , where is the similarity degree between the word vector of the candidate text corresponding to the maximum value of all corrected candidate probabilities under the th instruction signal segment before the th instruction signal segment and the word vector of the th candidate text under the th instruction signal segment; is the maximum value of all corrected candidate probabilities of the candidate texts under the th instruction signal segment before the th instruction signal segment, is the time interval between the th instruction signal segment and the th instruction signal segment before it, is the The number of all instruction signal segments before a certain instruction signal segment, where the correction candidate probability corresponding to each candidate text under the first instruction signal segment is the said candidate probability.

[0013] Preferably, the identifying the voice signal and performing voice interaction includes: selecting, as the result of voice recognition, the candidate text corresponding to the maximum value of all the said correction candidate probabilities under each instruction signal segment; and performing voice interaction according to the result of voice recognition by finding the corresponding answer in a preset knowledge base.

[0014] In a second aspect, an embodiment of the present application further provides an intelligent customer service voice interaction system based on voice recognition, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the steps of any one of the above-mentioned intelligent customer service voice interaction methods based on voice recognition are implemented.

[0015] The present application has at least the following beneficial effects: The present application divides the voice signal and screens all signal segments to obtain each instruction signal segment. The beneficial effect is that it screens the signals corresponding to each instruction issued by the user each time, so as to eliminate the signals at the silent moments, which helps to remove irrelevant voices or noise segments. Furthermore, through the hidden Markov model, the phoneme sequence corresponding to each instruction signal segment and all its corresponding candidate texts are obtained. The beneficial effect is that all texts matching the phonemes corresponding to the voice signal in the publicly available phoneme dictionary are corresponding, so as to screen out more matching texts in the follow-up; calculate the matching degree of each word in each candidate text under each instruction signal segment. The beneficial effect is the number of words matched in the publicly available phoneme dictionary through different combination methods among all the phonemes corresponding to the word, to illustrate the possibility that the phonemes corresponding to each word are misclassified, and to reflect the matching situation between each word and the word represented by the corresponding phoneme; determine the semantic correlation degree of each word in each candidate text. The beneficial effect is to consider the co-occurrence situation of each word in each candidate text and its adjacent words in the historical text, to reflect the semantic relationship between the words; obtain the candidate probability of each candidate text under each instruction signal segment, correct the candidate probability, obtain the corrected candidate probability corresponding to each candidate text under each instruction signal segment, identify the voice signal and perform voice interaction. The beneficial effect is to consider the similarity situation and time interval situation between each instruction signal segment and the text after the previous instruction signal segment, correct the candidate probability, select the candidate text corresponding to the maximum corrected candidate probability as the voice recognition result of each instruction signal segment, and then perform voice interaction, improving the accuracy of voice recognition, enabling the intelligent customer service to more accurately understand and respond to the user's voice instructions, enhancing the accuracy, coherence and naturalness of voice interaction, and enhancing the user experience satisfaction. Description of the Drawings

[0016] The following further elaborates in detail a voice interaction method of an intelligent customer service based on speech recognition according to the present application with reference to the accompanying drawings.

[0017] Figure 1 It is a flowchart of the steps of a voice interaction method of an intelligent customer service based on speech recognition provided by an embodiment of the present application; Figure 2 It is a flowchart of the steps of a method for obtaining the matching degree of each vocabulary in each candidate text provided by an embodiment of the present application; Figure 3 It is a flowchart of the steps of a method for obtaining the semantic correlation degree of each vocabulary in each candidate text provided by an embodiment of the present application. Specific embodiments

[0018] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the following further elaborates in detail a voice interaction method and system of an intelligent customer service based on speech recognition proposed by the present application with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs.

[0020] Please refer to Figure 1 , which shows a flowchart of the steps of a voice interaction method of an intelligent customer service based on speech recognition provided by an embodiment of the present application. The method includes the following steps: Step 1, collect in real time the voice signals of the user during the process of issuing a voice command in a voice interaction scenario, and the text data generated by the user during the interaction in a historical period, and form a historical text set.

[0021] With the rapid development of artificial intelligence, human-computer interaction methods represented by voice answering are widely used in service consultation industries such as e-commerce, electricity, finance, and education. Before the intelligent customer service makes a voice reply, it is necessary to accurately grasp the intention of the user in the voice signal, and convert the voice signal into text to identify the intention of the user.

[0022] Therefore, collect the voice signals of the user in the voice interaction scenario with the intelligent customer service, and form a historical text set with all the text data generated by the user during the interaction in a historical period; In this embodiment, the text data generated by converting the text work orders, chat records, and text records of customer service phone recordings recorded by the customer service during the interaction of different users in a historical period is used to form a historical text set.

[0023] So far, the voice signal of the user and the historical text set during voice interaction are obtained.

[0024] Step 2: Based on the signal strength in the voice signal, divide the voice signal into multiple signal segments; screen all the signal segments to obtain each instruction signal segment; use the hidden Markov model to obtain the phoneme sequence corresponding to each instruction signal segment and all the corresponding candidate texts.

[0025] When the user and the intelligent customer service are having a voice interaction, after the user issues a voice instruction once, the intelligent customer service understands the voice signal and then feedbacks the user's question. After that, the user will issue subsequent voice instructions, thus forming an information interaction. Among them, the voice signal only contains the user's voice, and the reply result of the intelligent customer service is not in the voice signal. Therefore, there are natural pauses during the interaction between the user and the intelligent customer service and silent periods when the intelligent customer service feedbacks questions. Through voice activity detection, the voice signal can be divided into multiple signal segments. Specifically: Mark the moment when the signal strength in the voice signal is 0 as the silent moment; if the signal strengths of the two moments adjacent to the silent moment in the voice signal are not all 0, then mark the silent moment as the zero point; Taking the zero point as the segmentation point, divide the voice signal into multiple signal segments; Secondly, the signal segments include the signals when the user issues an instruction and the signals during silence. Distinguish the signal segments through the voice intensity of the signal segments to obtain each instruction signal segment. Specifically: Calculate the average value of all the signal strengths in each signal segment, cluster the average values of all the signal segments to obtain two clustering clusters; In this embodiment, the k-means clustering algorithm is used for clustering. Among them, the k-means clustering algorithm is a well-known technology and will not be elaborated here. As other implementation manners, implementers can use other methods of the existing technology, such as the DBSCAN clustering algorithm, etc. This embodiment does not make special restrictions on this.

[0026] Mark all the signal segments within the clustering cluster to which the maximum value of the average values of all the signal segments belongs as each instruction signal segment; It should be noted that the larger the average value, the weaker the signal of the signal segment, which may be the signal during silence or when the user's expression has a long pause. Therefore, the signal segments within the clustering cluster with the largest average value are more likely to be the signal segments corresponding to when the user issues an instruction.

[0027] Furthermore, the Hidden Markov Model (HMM) is the most commonly used mathematical model in speech recognition. Generally, each individually recognizable word, phrase, or short sentence can be simulated by an HMM. The speech signals of each instruction signal segment are text-mapped through the Hidden Markov Model, specifically as follows: Preprocess each instruction signal segment, and extract the Mel Frequency Cepstral Coefficients (MFCCs) of each preprocessed instruction signal segment as the observation sequence; In this embodiment, the extraction of Mel Frequency Cepstral Coefficients (MFCC coefficients) is a well-known technique and will not be elaborated here. Secondly, the preprocessing process includes pre-emphasis, framing, windowing, FFT, MEL filter bank, DCT, etc. Among them, the preprocessing process is a well-known technique and will not be elaborated here. By extracting the Mel Frequency Cepstral Coefficients of each frame in each instruction signal segment, the observation sequence is formed.

[0028] It should be noted that the observation sequence , where is the MFCC coefficient extracted at time t.

[0029] Based on the observation sequence, through the Hidden Markov Model, obtain the hidden state sequence; map each hidden state in the hidden state sequence to the corresponding phoneme in the phoneme dictionary to obtain the phoneme sequence; It should be noted that the Hidden Markov Model is a well-known technique and will not be elaborated here. Secondly, the hidden state sequence , where represents the hidden state at time , and each hidden state corresponds to a specific phoneme; Secondly, the process of mapping the hidden state to the corresponding phoneme is generally implemented through the phoneme dictionary or the state-phoneme mapping relationship in the HMM model. Among them, the phoneme dictionary is generally created through the pronunciation rules of speech and a large-scale language dataset. The specific process is a well-known technique and will not be elaborated here. In this embodiment, the phoneme dictionary of Chinese entries uses the Chinese phoneme set lexcion in the Tsinghua University corpus.

[0030] It should be noted that the Hidden Markov Model needs to be trained. The specific training process is as follows: By collecting a large amount of labeled speech data, these speech data should contain the speech signals of different users and the corresponding text annotations, forming a dataset. Preprocess the speech signals in the dataset, extract the Mel Frequency Cepstral Coefficients (MFCC) as the observation sequence, and train the Hidden Markov Model; Through the observation sequence corresponding to each instruction signal segment, combined with the trained Hidden Markov Model, obtain the hidden state sequence.

[0031] Secondly, a phoneme is the smallest speech unit in linguistics. It is the smallest sound unit that distinguishes different words or semantics. Each phoneme represents an independent speech feature, usually a sound emitted through organs such as the oral cavity and larynx. Different languages have different phoneme systems, and the combination of phonemes constitutes the speech forms such as words and sentences in a language. Therefore, in the process of speech recognition, by mapping the phoneme sequence to the text, the user's intention can be grasped and understood through the speech signal. However, due to the existing semantic ambiguity problem, the same phoneme may correspond to multiple words, resulting in a phoneme sequence corresponding to multiple texts. For example, taking an English phoneme sequence [ / k / / a / / t / ] as an example, it may correspond to "cat" or "cot". Therefore, to obtain all candidate texts corresponding to each instruction signal segment, specifically: Through the phoneme dictionary, all texts corresponding to the phoneme sequence are recorded as each candidate text; It should be noted that by matching the corresponding words in the phoneme dictionary through the phoneme sequence corresponding to each instruction signal segment to form texts, multiple texts can be obtained.

[0032] So far, all candidate texts corresponding to each instruction signal segment are obtained.

[0033] Step 3: Segment each candidate text, analyze the number of words represented by the combination of multiple phonemes corresponding to each word in each candidate text within the phoneme sequence, and calculate the matching degree of each word in each candidate text under each instruction signal segment; determine the semantic association degree of each word in each candidate text under each instruction signal segment through the distance situation and the number of co-occurrences between the words in the text data where each word in each candidate text and its adjacent words co-occur in the historical text set, and combine the matching degree to obtain the candidate probability of each candidate text under each instruction signal segment.

[0034] Furthermore, in the traditional algorithm, the n-gram model is used to sort the candidate texts to obtain the candidate probability corresponding to the candidate text, and the candidate text with the largest candidate probability is selected as the final speech recognition result. However, when the n-gram model re-sorts the candidate texts, since the n-gram model only depends on the context of the previous n words for prediction, lacking in-depth understanding of the semantic relationship between words and the context understanding of the sentence, the matching degree between the selected candidate text and the actual speech signal is relatively low, thus affecting the accuracy and efficiency of the intelligent customer service speech interaction.

[0035] Based on the above analysis, first, by segmenting the candidate text, analyzing the number of words matched by the phonemes corresponding to the segmented words in the phoneme dictionary, and calculating the matching degree to evaluate the accurate situation of the phoneme division when obtaining the candidate text, the step flow chart of the method for obtaining the matching degree of each word in each candidate text provided by the embodiment of the present application is asFigure 2 As shown in the figure, specifically including: Perform word segmentation on each candidate text corresponding to the phoneme sequence of each instruction signal segment to obtain each vocabulary of each candidate text; In this embodiment, Jieba word segmentation is used for processing. Among them, Jieba word segmentation is a well-known technology and will not be elaborated here. As other implementation manners, implementers can adopt other methods of existing technologies, such as HanLP, SnowNLP, etc. This embodiment does not make special restrictions on this.

[0036] The phonemes corresponding to any vocabulary in each candidate text within the phoneme sequence are combined to form a local phoneme sequence; It should be noted that by splitting the candidate text into multiple words and phrases, denoted as each vocabulary. For example, for the sentence: "I have a problem", it can be split into "I", "have", "a", "problem". For each vocabulary in the candidate text, since the phoneme is the smallest sound unit of the vocabulary, the corresponding phoneme can be directly obtained through the pronunciation status of the vocabulary, so as to obtain the phoneme corresponding to any vocabulary within the phoneme sequence and form a local phoneme sequence.

[0037] Each phoneme within the local phoneme sequence and all the phonemes before it are combined to form each subsequence; All the phonemes within the local phoneme sequence and the next phoneme corresponding to the last phoneme within the phoneme sequence are combined to form an extended sequence; It should be noted that for easy understanding, assume that the phoneme sequence is [ / k / / a / / t / / n / / æ / / p / ], and the local phoneme sequence corresponding to a certain vocabulary is [ / k / / a / / t / ]. Each subsequence is [ / k / ], [ / k / / a / ], [ / k / / a / / t / ], where the subsequence [ / k / / a / / t / ] is the local phoneme sequence, and the next phoneme corresponding to the last phoneme / t / within the phoneme sequence is / n / . Combine / k / / a / / t / with / n / to form a sequence. Therefore, the extended sequence is [ / k / / a / / t / / n / ].

[0038] All the subsequences except the local phoneme sequence among all the subsequences and the extended sequence are denoted as each characteristic phoneme sequence; Count the number of all the vocabularies corresponding to the local phoneme sequence in the phoneme dictionary; Count the number of all the vocabularies corresponding to each characteristic phoneme sequence in the phoneme dictionary, denoted as the vocabulary amount. Take the reciprocal of the product of the sum value of the vocabulary amounts of all the characteristic phoneme sequences and the number as the matching degree of any vocabulary in each candidate text; It should be noted that the larger the number is, the richer the information contained in the local phoneme sequence is, and the more likely it corresponds to multiple meanings. If there are corresponding words in the phoneme dictionary for each characteristic phoneme sequence, and the more words there are, when obtaining the candidate text above, the more likely there will be deviations in the division of phonemes within the phoneme sequence, and the smaller the obtained matching degree. This indicates that the more words the local phoneme sequence represents, the less likely the word represented by the local phoneme sequence is the word in any of the candidate texts. The larger the matching degree, the fewer words the local phoneme sequence represents, and the more likely the word represented by the local phoneme sequence is the corresponding word in any of the candidate texts.

[0039] Secondly, the above-mentioned matching degree reflects the matching situation between phoneme features and words, and does not consider the relevance between words, which may cause the phenomenon of semantic confusion in the selected text. Therefore, by calculating the semantic relevance degree through the relevance between different words in any of the candidate texts, the step flow chart of the method for obtaining the semantic relevance degree of each word in each candidate text provided in the embodiments of the present application is as Figure 3 shown, and specifically includes: Mark the words adjacent to any of the words in each candidate text as neighboring words; Count the number of times any of the words and its neighboring words co-occur in all text data in the historical text set; Calculate the cumulative sum of the distances between any of the words and its neighboring words in all text data where they co-occur in the historical text set; In this embodiment, the distance is measured by calculating the character distance between any of the words and its neighboring words in all text data where they co-occur in the historical text set. Among them, the calculation of the character distance is a well-known technology and will not be elaborated here.

[0040] Take the product of the number of times and the cumulative sum as the relevance between any of the words and its neighboring words; Take the sum of the relevances between any of the words and all neighboring words in each candidate text as the semantic relevance degree of any of the words in each candidate text; In this embodiment, for each instruction signal segment, taking the k-th word in the m-th candidate text as an example, the calculation formula for its semantic relevance degree is: , where is the semantic relevance degree of the -th word in the -th candidate text, is the -th candidate text, the -th word and the The number of times that neighboring words co-occur in all text data within the historical text set, is the th candidate text, the th word, and the th neighboring word, the character distance between the two words in the th occurrence of text data within the historical text set, is the th candidate text, the number of all neighboring words of the th word; where is the said relevance.

[0041] It should be noted that the larger the number of times, the higher the frequency of co-occurrence of the th word and the th neighboring word in the texts of the historical period; the smaller the distance, the higher the relevance between the two words; the larger the obtained semantic relevance, the more relevant the semantic information between the two words in the intelligent customer service voice interaction scenario, and the smoother the overall sentence; conversely, the smaller the obtained semantic relevance, the less relevant the semantic information between the two words.

[0042] Furthermore, based on the matching degree and the semantic relevance, determine the candidate probability, specifically: Take the normalized result of the sum of the products of the matching degree and the semantic relevance of all words in each candidate text as the candidate probability of each candidate text; In this embodiment, the sigmoid function is used for normalization processing. The sigmoid function is a well-known technology and will not be elaborated here. As other implementation manners, implementers can use other methods of existing technologies, such as the softmax function, the tanh function, etc. This embodiment does not make special restrictions on this.

[0043] It should be noted that the larger the candidate probability, the more likely the text meaning represented by the candidate text is the true meaning expressed by each instruction signal segment; conversely, the smaller the candidate probability, the less likely the text meaning represented by the candidate text is the true meaning expressed by each instruction signal segment.

[0044] Thus, the candidate probability of each candidate text is obtained.

[0045] Step 4, based on the time length between each instruction signal segment and different previous instruction signal segments, and the similarity of their candidate texts, correct the candidate probability to obtain the corrected candidate probability corresponding to each candidate text under each instruction signal segment, identify the voice signal and perform voice interaction.

[0046] Further, due to the correlation between different instruction signal segments, based on the text corresponding to the initial instruction signal segment, the candidate probabilities of the candidate texts corresponding to the subsequent remaining instruction signal segments are corrected. Specifically: Obtain the word vectors of each candidate text; In this embodiment, the Word2Vec model is used to convert each candidate text into a word vector. Among them, the Word2Vec model is a well-known technology and will not be elaborated here. As other implementation manners, implementers can adopt other methods of existing technologies, such as BERT, etc. This embodiment does not make special restrictions on this.

[0047] The calculation formula for the corrected candidate probability corresponding to each candidate text under each instruction signal segment is: , where, is the corrected candidate probability corresponding to the th candidate text under the th instruction signal segment, is the similarity degree between the word vector of the candidate text corresponding to the maximum value of all corrected candidate probabilities under the th instruction signal segment before the th instruction signal segment and the word vector of the th candidate text under the th instruction signal segment; is the maximum value of the corrected candidate probabilities of all candidate texts under the th instruction signal segment before the th instruction signal segment, is the time interval between the th instruction signal segment and the th instruction signal segment before it, is the number of all instruction signal segments before the th instruction signal segment. Among them, the corrected candidate probability corresponding to each candidate text under the first instruction signal segment is the candidate probability.

[0048] In this embodiment, the similarity degree is measured by the cosine similarity between the word vector of the candidate text corresponding to the maximum value of all corrected candidate probabilities under the th instruction signal segment before the th instruction signal segment and the word vector of the th candidate text under the th instruction signal segment. Among them, the calculation of the cosine similarity is a well-known technology and will not be elaborated here; secondly, the time interval is calculated by the duration between the start time of the th instruction signal segment and the end time of the instruction signal segment before it.

[0049] It should be noted that for the first instruction signal segment, the candidate probability of each corresponding candidate text is used as the corrected candidate probability of each candidate text. The candidate text corresponding to the maximum value of all the corrected candidate probabilities under the first instruction signal segment is denoted as the matching text of the first instruction signal segment. Therefore, for the second instruction signal segment, the cosine similarity between the word vector of each candidate text under the second instruction signal segment and the word vector of the matching text corresponding to the previous instruction signal segment is calculated. Furthermore, the corrected candidate probability of each candidate text under the second instruction signal segment is obtained, and the candidate text corresponding to the maximum value of the corrected candidate probabilities under the second instruction signal segment is selected and denoted as the matching text corresponding to the second instruction signal segment. By analogy, the corrected candidate probability corresponding to each candidate text under each instruction signal segment can be obtained.

[0050] It should be noted that the greater the degree of similarity, the higher the relevance of the th candidate text under the th instruction signal segment to the text data corresponding to the previous instruction signal segment. The smaller the time interval, the closer the signal issuing times between the two instruction signal segments, and the higher the possible correlation between the two instruction signal segments. The greater the corrected candidate probability, the more likely the candidate text is the result of the speech recognition of the corresponding instruction signal segment.

[0051] Furthermore, the candidate text corresponding to the maximum value of all the corrected candidate probabilities under each instruction signal segment is selected as the result of the speech recognition. The key entities and intentions of the user are recognized through natural language processing algorithms. Based on the result of the user's speech recognition, the corresponding answer is found in the preset knowledge base to realize the voice interaction between the intelligent customer service and the user.

[0052] Based on the same inventive concept as the above method, an intelligent customer service voice interaction system based on speech recognition is further provided in an embodiment of the present application, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the steps of any one of the above methods for an intelligent customer service voice interaction method based on speech recognition are implemented.

[0053] It should be understood that although Figure 1 the steps in the flowchart are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, Figure 1At least a part of the steps may include multiple sub-steps or multiple stages, and these sub-steps or stages do not necessarily need to be completed at the same moment, but can be executed at different moments. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.

[0054] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0055] The above-described embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation to the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made. Therefore, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present application without departing from the content of the technical solution of the present application all belong to the protection scope of the technical solution of the present application.

Claims

1. An intelligent customer service voice interaction method based on speech recognition, characterized in that, The method includes the following steps: Collect in real time the voice signals of the user during the process of issuing voice commands in the voice interaction scenario, as well as the text data generated when the user interacts in the historical period, and form a historical text set; Based on the signal strength in the voice signal, divide the voice signal into multiple signal segments; and screen all the signal segments to obtain each command signal segment; Adopt a hidden Markov model to obtain the phoneme sequence corresponding to each command signal segment and all its corresponding candidate texts; Perform word segmentation on each candidate text, analyze the number of words represented by the combination of multiple phonemes corresponding to each word in the phoneme sequence in each candidate text, and calculate the matching degree of each word in each candidate text under each command signal segment; Determine the semantic association degree of each word in each candidate text under each command signal segment through the distance situation and the number of co-occurrences of the words in the text data co-occurring with each word and its adjacent words in the historical text set, and combine the matching degree to obtain the candidate probability of each candidate text under each command signal segment; Correct the candidate probability through the time length between each command signal segment and different command signal segments before it and the similarity of its candidate texts, obtain the corrected candidate probability corresponding to each candidate text under each command signal segment, and recognize the voice signal and perform voice interaction.

2. The intelligent customer service voice interaction method based on speech recognition according to claim 1, characterized in that, The dividing the voice signal into multiple signal segments includes: marking the moment when the signal strength in the voice signal is 0 as the silent moment; if the signal strengths of two moments adjacent to the silent moment in the voice signal are not all 0, then mark the silent moment as the zero point; and divide the voice signal into multiple signal segments with the zero point as the segmentation point.

3. The intelligent customer service voice interaction method based on speech recognition according to claim 1, wherein, The further method for obtaining each command signal segment is: calculate the average value of all signal strengths in each signal segment, perform clustering on the average values of all signal segments to obtain two clustering clusters; mark all the signal segments within the clustering cluster to which the maximum value of the average values of all signal segments belongs as each command signal segment.

4. The intelligent customer service voice interaction method based on speech recognition according to claim 1, characterized in that, The obtaining the phoneme sequence corresponding to each command signal segment and all its corresponding candidate texts includes: Perform preprocessing on each command signal segment, extract the mel-frequency cepstral coefficients of each command signal segment after preprocessing as the observation sequence; based on the observation sequence, obtain the hidden state sequence through a hidden Markov model; map each hidden state in the hidden state sequence to the corresponding phoneme in the phoneme dictionary to obtain the phoneme sequence; Through the phoneme dictionary, mark all the texts corresponding to the phoneme sequence as each candidate text.

5. The intelligent customer service voice interaction method based on speech recognition according to claim 1, characterized in that, The calculating the matching degree of each word in each candidate text under each command signal segment includes: Form a local phoneme sequence with the phonemes corresponding to each word in each candidate text in the phoneme sequence; Form each subsequence with each phoneme in the local phoneme sequence and all the phonemes before it; among them, the subsequence formed by the last phoneme in the local phoneme sequence and all the phonemes before it is the local phoneme sequence; Form an extended sequence by combining all the phonemes in the local phoneme sequence with the next phoneme corresponding to its last phoneme in the phoneme sequence; Denote all subsequences other than the local phoneme sequence and the extended sequence in all the subsequences as each characteristic phoneme sequence; Count the number of all the words corresponding to the local phoneme sequence in the phoneme dictionary; Count the quantity of all the words corresponding to each characteristic phoneme sequence in the phoneme dictionary, and denote it as the vocabulary size; The matching degree is the reciprocal of the product between the sum value of the vocabulary sizes of all the characteristic phoneme sequences and the number.

6. The intelligent customer service voice interaction method based on speech recognition according to claim 1, wherein, The determining the semantic association degree of each word in each candidate text under each instruction signal segment includes: Denote the words adjacent to any word in each candidate text as adjacent words; Count the number of times that the any word and its adjacent words co-occur in all the text data in the historical text set; Calculate the cumulative sum of the distances between the any word and its adjacent words in all the text data where they co-occur in the historical text set; Take the product of the number of times and the cumulative sum as the correlation degree between the any word and its adjacent words; The semantic association degree is the sum of the correlation degrees between the any word and all its adjacent words in each candidate text.

7. The intelligent customer service voice interaction method based on speech recognition according to claim 1, characterized in that, The candidate probability of each candidate text under each instruction signal segment is the normalized result of the sum of the products of the matching degrees and the semantic association degrees of all the words in each candidate text.

8. The intelligent customer service voice interaction method based on speech recognition according to claim 1, characterized in that, The obtaining the corrected candidate probability corresponding to each candidate text under each instruction signal segment includes: Obtain the word vector of each candidate text; The correction candidate probability corresponding to the th candidate text under the th instruction signal segment is calculated as follows: , where is the similarity between the word vector of the candidate text corresponding to the maximum value of all correction candidate probabilities under the th instruction signal segment before the th instruction signal segment and the word vector of the th candidate text under the th instruction signal segment; is the maximum value of all correction candidate probabilities of the candidate texts under the th instruction signal segment before the th instruction signal segment, is the time interval between the th instruction signal segment and the th instruction signal segment before it, is the number of all instruction signal segments before the th instruction signal segment. Among them, the correction candidate probability corresponding to each candidate text under the first instruction signal segment is the said candidate probability.

9. The intelligent customer service voice interaction method based on speech recognition according to claim 1, wherein The recognizing the speech signal and performing speech interaction includes: Select the candidate text corresponding to the maximum value of all the corrected candidate probabilities under each instruction signal segment as the result of speech recognition; According to the result of speech recognition, find the corresponding answer in the preset knowledge base to perform speech interaction.

10. An intelligent customer service voice interaction system based on speech recognition, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method for intelligent customer service speech interaction based on speech recognition according to any one of claims 1-9.

Citation Information

Patent Citations

  • Audio processing method and device, language model training method and device and computer equipment

    CN111933129A

  • Speech recognition method and device and storage medium

    CN114974249A

  • Human-robot voice interaction system based on end-to-end

    CN117542358A

  • Speech recognition method, system and device based on artificial intelligence and medium

    CN119943032A

  • Apparatus, method, and medium for generating grammar network for use in speech recognition and dialogue speech recognition

    US20060173686A1

Cited By

  • Voice instruction recognition method and device

    CN121096341A

  • Supplier voice communication auxiliary analysis system and method for logistics

    CN121281498A