Voice interaction method, system and device based on artificial intelligence and medium
Through attention mechanism and semantic analysis, the context dependence between dialogue rounds is extracted, the semantic consistency problem in complex dialogue scenarios is solved, and the logical coherence and consistency of multiple dialogues is achieved.
Patent Information
- Application Number
- CN202510429671.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-04-08
AI Technical Summary
The prior art is difficult to maintain semantic consistency in complex dialogue scenarios, resulting in logical confusion in the system during multiple rounds of dialogue.
The context dependencies between conversation rounds are extracted through the attention mechanism, the semantic loss and coherence are determined, intent classification and association are performed, and feedback control instructions are generated to maintain semantic integrity.
Enhance the consistency and consistency of dialogue, ensuring that the system accurately understands user intentions in multiple rounds of interactions, and maintains a logically coherent dialogue process.
Smart Images

Figure CN119964570A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of voice interaction technology, and more specifically, to a voice interaction method, system, device and medium based on artificial intelligence. Background Art
[0002] Voice interaction is a way of using artificial intelligence technology to achieve communication between people and devices. Users send commands to the system through voice input, and the system gives corresponding feedback after processing. Artificial intelligence plays a core role in voice interaction, enabling the system to understand voice content, identify user intentions, and respond intelligently based on the context. Widely used artificial intelligence technology improves the accuracy of voice recognition, makes interactions more natural and smooth, and supports multi-round conversations, emotion recognition and personalized recommendations. Voice interaction has been widely used in smart assistants, in-vehicle systems, smart homes and other fields, continuously optimizing user experience and improving the convenience and intelligence of interaction.
[0003] In the prior art, AI-based voice interaction is mainly implemented through speech recognition, natural language processing and generative models. Speech recognition converts voice signals into text, and then natural language processing analyzes the text semantics and understands the user's intention. Subsequently, the generative model generates appropriate replies based on the intention, and finally the reply is converted into speech output through speech synthesis technology. However, in multi-round conversations, especially in complex voice interaction scenarios, the user's intention and conversation context are dynamically changing. The voice interaction process may involve multiple topics, complex background information, and even emotional fluctuations, making it impossible for the system to continuously understand the user's intention in multi-round conversations, resulting in the system being unable to maintain semantic consistency in multi-round conversations, resulting in logical confusion and other problems. Therefore, how to achieve semantic consistency of contextual conversations in complex conversation scenarios has become a difficult problem faced by the industry. Summary of the invention
[0004] The present application provides a voice interaction method, system, device and medium based on artificial intelligence, which can maintain the semantic consistency of contextual dialogue in complex dialogue scenarios.
[0005] In a first aspect, the present application provides a voice interaction method based on artificial intelligence, comprising the following steps: Acquire the interactive voice of the target user in different conversation rounds through the voice interaction device; Extract the contextual dependencies between each dialogue turn from all interactive speech based on the attention mechanism; Determine the semantic loss amount of the interactive speech in each dialogue round according to the semantic features of the interactive speech in each dialogue round, and determine the dialogue coherence of the target user in each dialogue round through all the semantic loss amounts and the context dependency relationship between the respective dialogue rounds; Classify the target user's interactive speech in the current dialogue scenario, obtain the target user's different semantic intentions in the current dialogue scenario, and then determine the semantic association relationship between each semantic intention; The semantic completeness of the target user in the current dialogue scenario is determined by the semantic association relationship and the dialogue coherence of the interactive speech in each dialogue round, and then the feedback control instruction of the voice interaction device is generated according to the semantic completeness.
[0006] In some embodiments, extracting the context dependency between each dialogue turn from all interactive speech based on the attention mechanism specifically includes: Convert all interactive speech into multiple speech vectors; Construct an attention vector space for each speech vector based on the preset attention mechanism; Determine the semantic dependency of each conversation turn based on the attention vector space of each speech vector; The contextual dependencies between each dialogue turn are determined through the semantic dependencies of all dialogue turns.
[0007] In some embodiments, determining the semantic loss amount of the interactive speech in each dialogue turn according to the semantic features of the interactive speech in each dialogue turn specifically includes: Determine the semantic features of the interactive speech at each dialogue turn; Determine the semantic similarity between interactive speech in each dialogue round based on all semantic features; The semantic loss of the interactive speech in each dialogue round is determined by the semantic similarity between the interactive speech in each dialogue round.
[0008] In some embodiments, determining the dialogue coherence of the target user in each dialogue turn by using all semantic loss amounts and the context dependency relationship between the respective dialogue turns specifically includes: Determine the dialogue state quantity for each dialogue turn based on the preset sliding window and all semantic loss quantities; The dialogue coherence of the target user in each dialogue turn is determined according to the dialogue state quantity in each dialogue turn and the context dependency relationship between the dialogue turns.
[0009] In some embodiments, the target user's interactive speech in the current dialogue scene is classified into intents, and different semantic intents of the target user in the current dialogue scene are obtained, specifically including: Extract multiple conversation keywords from the target user's interactive speech in the current conversation scenario; Determine the semantic confidence of each conversation keyword based on the pre-trained large language model; All conversation keywords are feature classified according to their semantic confidence to obtain different semantic intentions of the target user in the current conversation scenario.
[0010] In some embodiments, determining the semantic association relationship between the semantic intents specifically includes: Get all conversation keywords; Perform topic analysis on all conversation keywords based on the pre-trained topic analysis model to obtain the topic features of the target user's interactive speech in the current conversation scenario; The semantic association relationship between various semantic intentions is determined according to the topic features of the target user's interactive speech in the current dialogue scenario and the different semantic intentions of the target user in the current dialogue scenario.
[0011] In some embodiments, determining the semantic integrity of the target user in the current dialogue scenario through the semantic association relationship and the dialogue coherence of the interactive speech in each dialogue turn specifically includes: Perform linear fitting on the dialogue coherence of the interactive speech in all dialogue rounds to obtain a fitting curve of the dialogue coherence; Get the semantic confidence of each conversation keyword; Determine a semantic association graph of the target user in the current conversation scenario according to all semantic confidences and the semantic association relationships; The semantic completeness of the target user in the current dialogue scenario is determined by the fitting curve of the semantic association graph and the dialogue coherence.
[0012] In a second aspect, the present application provides a voice interaction system based on artificial intelligence, comprising: An acquisition module, used to acquire the interactive voice of the target user in different conversation turns through a voice interaction device; A processing module that extracts contextual dependencies between conversation turns from all interactive speech based on an attention mechanism; The processing module is further used to determine the semantic loss amount of the interactive speech in each dialogue round according to the semantic features of the interactive speech in each dialogue round, and determine the dialogue coherence of the target user in each dialogue round through all the semantic loss amounts and the context dependency relationship between the various dialogue rounds; The processing module is further used to classify the target user's interactive speech in the current dialogue scene, obtain different semantic intentions of the target user in the current dialogue scene, and then determine the semantic association relationship between the various semantic intentions; An execution module is used to determine the semantic completeness of the target user in the current dialogue scenario through the semantic association relationship and the dialogue coherence of the interactive speech in each dialogue round, and then generate feedback control instructions for the voice interaction device according to the semantic completeness.
[0013] In a third aspect, the present application provides a computer device, comprising a memory and a processor, wherein the memory stores code, and the processor is configured to obtain the code and execute the above-mentioned artificial intelligence-based voice interaction method.
[0014] In a fourth aspect, the present application provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, it implements the above-mentioned artificial intelligence-based voice interaction method.
[0015] The technical solution provided by the embodiments disclosed in this application has the following beneficial effects: In the artificial intelligence-based voice interaction method, system, device and medium provided in the present application, the interactive speech of the target user in different dialogue rounds is obtained through the voice interaction device; the context dependency relationship between each dialogue round is extracted from all the interactive speech based on the attention mechanism; the semantic loss amount of the interactive speech in each dialogue round is determined according to the semantic features of the interactive speech in each dialogue round, and the dialogue coherence of the target user in each dialogue round is determined through all the semantic loss amounts and the context dependency relationship between the said each dialogue round; the interactive speech of the target user in the current dialogue scene is classified by intent to obtain different semantic intentions of the target user in the current dialogue scene, and then the semantic association relationship between the various semantic intentions is determined; the semantic completeness of the target user in the current dialogue scene is determined through the semantic association relationship and the dialogue coherence of the interactive speech in each dialogue round, and then the feedback control instructions of the voice interaction device are generated according to the semantic completeness.
[0016] It can be seen that in this application, the semantic completeness of the target user in the current dialogue scenario can be determined through the semantic association relationship and the dialogue coherence of the interactive speech in each dialogue round; wherein, firstly, the context dependency relationship between each dialogue round is extracted by using the attention mechanism, and the context dependency relationship enables the system to dynamically capture and associate the semantic context of the user in different rounds, and dynamically adjust the semantic weight, thereby enhancing the dialogue coherence; secondly, the semantic loss amount of the interactive speech in each dialogue round is determined, and the semantic loss amount can effectively improve the system's perception of changes in voice features, thereby reducing the semantic understanding errors caused by voice differences, and then determine the dialogue coherence of the target user in each dialogue round. The dialogue coherence can effectively measure the stability of the user's voice features and semantic information in multiple rounds of interaction, ensuring that the system can accurately understand the context and maintain logic. A coherent dialogue process; then, the interactive voice of the target user in the current dialogue scenario is classified by intent, and the different semantic intentions of the target user in the current dialogue scenario are obtained. The intent classification can effectively improve the system's ability to parse the user's dialogue content, ensure semantic consistency in multiple rounds of interaction, and then determine the semantic association relationship between each semantic intention, wherein the semantic association relationship can reflect the logical relationship at the semantic level, thereby maintaining the overall coherence of the dialogue; further, the semantic integrity of the target user in the current dialogue scenario is determined, and the semantic integrity can help the system evaluate the integrity of the dialogue information, ensuring that the contextual semantics are maintained throughout the interaction process; finally, the feedback control instructions of the voice interaction device are generated according to the semantic integrity; in summary, the solution of the present application can achieve the semantic consistency of the contextual dialogue in complex dialogue scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 is an exemplary flow chart of an artificial intelligence-based voice interaction method according to some embodiments of the present application; Figure 2 is a schematic diagram of a process for extracting context dependencies between various dialogue turns according to some embodiments of the present application; Figure 3 is a schematic diagram of a process for determining a semantic loss amount according to some embodiments of the present application; Figure 4 is a schematic diagram of the structure of a voice interaction system based on artificial intelligence according to some embodiments of the present application; Figure 5 It is a structural diagram of a computer device for implementing an artificial intelligence-based voice interaction method according to some embodiments of the present application. DETAILED DESCRIPTION
[0018] In order to better understand the technical solution of the present application, the technical solution of the present application will be described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0019] refer to Figure 1 , which is an exemplary flow chart of a voice interaction method based on artificial intelligence according to some embodiments of the present application. The voice interaction method 100 based on artificial intelligence mainly includes the following steps: In step 101, the interactive voice of the target user in different dialogue turns is obtained through a voice interaction device.
[0020] In specific implementation, obtaining the interactive voice of the target user in different conversation rounds through the voice interaction device can be achieved in the following manner, namely: obtaining the interactive voice of the target user in different conversation rounds through the cloud data transmission interface of the voice interaction device.
[0021] It should be noted that the interactive voice described in this application refers to the voice information input or output by the target user during the voice interaction process.
[0022] In step 102, context dependencies between each dialogue turn are extracted from all interactive speech based on the attention mechanism.
[0023] In some embodiments, reference Figure 2 As shown in FIG. 1 , this figure is a schematic diagram of a process for extracting context dependencies between various dialogue turns in some embodiments of the present application. In this embodiment, the context dependencies between various dialogue turns are extracted from all interactive voices based on the attention mechanism, which can be implemented by the following steps: First, in step 1021, all interactive speech is converted into multiple speech vectors; Secondly, in step 1022, an attention vector space of each speech vector is constructed based on a preset attention mechanism; Then, in step 1023, the semantic dependency of each dialogue turn is determined according to the attention vector space of each speech vector; Finally, in step 1024, the context dependency between each dialogue turn is determined through the semantic dependency of all dialogue turns.
[0024] In the specific implementation, all interactive voices can be converted into multiple voice vectors in the following manner, namely: a pre-trained voice model (such as Wav2Vec 2.0, Whisper, etc.) is used to convert each interactive voice into a voice vector, thereby obtaining multiple voice vectors, wherein the voice vector represents the numerical feature vector of the interactive voice. Other methods can also be used in other embodiments, which will not be repeated here.
[0025] In specific implementation, constructing the attention vector space of each speech vector based on the preset attention mechanism can be achieved in the following manner, namely: using the preset attention mechanism (such as the self-attention mechanism) to linearly transform each speech vector, and then obtain the query vector, key vector and value vector of each speech vector, so that the set consisting of the query vector, key vector and value vector of each speech vector is used as the attention vector space of each speech vector. Other methods can also be used in other embodiments, which are not limited here.
[0026] It should be noted that the attention vector space described in this application represents a vector set composed of query vectors, key vectors and value vectors after the speech vectors are converted into the speech vectors through the attention mechanism, which reflects the attention distribution of the speech vectors in different dialogue rounds.
[0027] In specific implementation, determining the semantic dependency of each dialogue turn based on the attention vector space of each voice vector can be implemented in the following manner, namely: first, select a voice vector as the selected voice vector, calculate the dot product of the query vector and the key vector in the attention vector space of the selected voice vector, and use the obtained dot product as the attention score of the selected voice vector, then calculate the cosine similarities between the query vector, key vector and value vector in the attention vector space of the selected voice vector and the query vector, key vector and value vector in the attention vector space of other voice vectors, then sum all the cosine similarities, multiply the summed value by the attention score of the selected voice vector, and use the multiplied value as the semantic dependency of the dialogue turn corresponding to the selected voice vector, and continue to determine the semantic dependencies of the dialogue turns corresponding to the remaining voice vectors. Other methods can also be used in other embodiments, which are not limited here.
[0028] It should be noted that the semantic dependency described in the present application represents the strength of association in semantic expression between each dialogue turn and the previous and next dialogue turns.
[0029] In specific implementation, determining the context dependency between each dialogue turn through the semantic dependencies of all dialogue turns can be implemented in the following manner, namely: first, removing the largest semantic dependency and the smallest semantic dependency from the semantic dependencies of all dialogue turns, then calculating the average of all remaining semantic dependencies after removing the largest semantic dependency and the smallest semantic dependency, and using the obtained average as the context dependency between each dialogue turn. In other embodiments, other methods can also be used for implementation, which will not be repeated here.
[0030] It should be noted that the context dependency relationship described in the present application represents the feature parameters that influence each other in semantic expression between each dialogue turn.
[0031] In step 103, the semantic loss amount of the interactive speech in each dialogue round is determined according to the semantic features of the interactive speech in each dialogue round, and the dialogue coherence of the target user in each dialogue round is determined through all the semantic loss amounts and the context dependencies between the respective dialogue rounds.
[0032] In some embodiments, reference Figure 3 As shown, this figure is a schematic diagram of a flow chart of determining the amount of semantic loss in some embodiments of the present application. In this embodiment, determining the amount of semantic loss of interactive speech in each dialogue round according to the semantic features of the interactive speech in each dialogue round can be implemented by the following steps: Determine the semantic features of the interactive speech at each dialogue turn; Determine the semantic similarity between interactive speech in each dialogue round based on all semantic features; The semantic loss of the interactive speech in each dialogue round is determined by the semantic similarity between the interactive speech in each dialogue round.
[0033] In specific implementation, determining the semantic features of the interactive speech in each dialogue round can be achieved in the following manner, namely: obtaining the speech vector of the interactive speech in each dialogue round, and then using the scikit-learn library in Python to perform principal component analysis on the speech vector of the interactive speech in each dialogue round, and using the vectors obtained after the principal component analysis as the semantic features of the interactive speech in the corresponding dialogue round, thereby obtaining the speech vector of the interactive speech in each dialogue round. In other embodiments, other methods can also be used for implementation, which will not be repeated here.
[0034] It should be noted that the semantic features described in this application represent digital vector representations of key semantic information of interactive speech in a dialogue turn.
[0035] In specific implementation, the semantic similarity between interactive speech in each dialogue round can be determined based on all semantic features in the following manner, namely: first, the cosine similarity is calculated for every two semantic features, then the mean of all cosine similarities is calculated, and the obtained mean is used as the global cosine similarity, and the maximum cosine similarity and the minimum cosine similarity are further selected from all cosine similarities, and then the difference between the maximum cosine similarity and the minimum cosine similarity is divided by the global cosine similarity, and the value obtained by the division is used as the semantic similarity between interactive speech in each dialogue round. In other embodiments, other methods can also be used for implementation, which is not limited here.
[0036] It should be noted that the semantic similarity described in the present application represents the degree of similarity of semantic information between interactive speech in each dialogue round. The greater the semantic similarity, the higher the degree of similarity of semantic information between interactive speech in each dialogue round.
[0037] In specific implementation, determining the semantic loss amount of interactive speech in each dialogue round through the semantic similarity between interactive speech in each dialogue round can be implemented in the following manner, namely: first, obtaining the cosine similarity between every two semantic features, then performing a differential operation on all cosine similarities, then summing all values obtained after the differential operation, and using the summed value as the similarity differential value, further, selecting a semantic feature as the selected semantic feature, and then selecting the largest cosine similarity from all cosine similarities between the selected semantic feature and other semantic features, dividing the largest cosine similarity by the similarity differential value, and then multiplying the value obtained by the division by the semantic similarity between interactive speech in each dialogue round, and using the multiplied value as the semantic loss amount of interactive speech in the dialogue round corresponding to the selected semantic feature, and continuing to determine the semantic loss amount of interactive speech in the dialogue round corresponding to the remaining semantic features. In other embodiments, other methods can also be used for implementation, which will not be repeated here.
[0038] It should be noted that the semantic loss amount described in the present application represents an indicator of the degree of deviation of semantic information between the interactive speech in the current dialogue round and the interactive speech in other dialogue rounds.
[0039] In some embodiments, determining the dialogue coherence of the target user in each dialogue turn by using all semantic loss amounts and the context dependency relationship between the dialogue turns can be implemented by the following steps: Determine the dialogue state quantity for each dialogue turn based on the preset sliding window and all semantic loss quantities; The dialogue coherence of the target user in each dialogue turn is determined according to the dialogue state quantity in each dialogue turn and the context dependency relationship between the dialogue turns.
[0040] It should be noted that the sliding window in the present application can be implemented according to the existing sliding window technology, wherein the size of the sliding window is set to the value obtained by taking the square root of the total number of all semantic loss amounts and rounding it down, and the sliding step size of the sliding window is set to 1.
[0041] In specific implementation, determining the dialogue state quantity under each dialogue turn based on the preset sliding window and all semantic loss quantities can be implemented in the following manner, namely: first, sorting all semantic loss quantities in ascending order, and taking the sequence obtained after sorting as the semantic loss quantity sequence, then aligning the first semantic loss quantity in the semantic loss quantity sequence using the preset sliding window, selecting the largest semantic loss quantity and the smallest semantic loss quantity from the sliding window, and taking the difference between the largest semantic loss quantity and the smallest semantic loss quantity as the dialogue state quantity under the dialogue turn corresponding to the first semantic loss quantity, moving the sliding window, aligning the sliding window with the second semantic loss quantity in the semantic loss quantity sequence, and then determining the dialogue state quantity under the dialogue turn corresponding to the second semantic loss quantity, and then, moving the sliding window in sequence until the sliding window aligns with the last semantic loss quantity in the semantic loss quantity sequence, thereby obtaining the dialogue state quantities under the dialogue turns corresponding to all the remaining semantic loss quantities in the semantic loss quantity sequence in sequence, and other methods can also be used in other embodiments, which are not limited here.
[0042] It should be noted that the dialogue state quantity described in this application represents a parameter that measures the stability of the interaction state of the current dialogue round compared to the previous dialogue round.
[0043] In specific implementation, determining the conversation coherence of the target user in each conversation round according to the conversation state quantity in each conversation round and the context dependency relationship between the conversation rounds can be implemented in the following manner, namely: multiplying the context dependency relationship between the conversation rounds by the conversation state quantity in each conversation round, and using the multiplied values as the conversation coherence of the target user in each conversation round. In other embodiments, other methods can also be used for implementation, which will not be repeated here.
[0044] It should be noted that the conversation coherence described in this application represents the logical coherence of the semantic expression of the target user in this conversation round. The greater the conversation coherence, the higher the logical coherence of the semantic expression of the target user in this conversation round, and vice versa.
[0045] In step 104, the interactive speech of the target user in the current dialogue scenario is classified for intent, and different semantic intents of the target user in the current dialogue scenario are obtained, and then the semantic association relationship between each semantic intent is determined.
[0046] In some embodiments, the following steps may be used to classify the target user's interactive speech in the current dialogue scenario and obtain different semantic intentions of the target user in the current dialogue scenario: Extract multiple conversation keywords from the target user's interactive speech in the current conversation scenario; Determine the semantic confidence of each conversation keyword based on the pre-trained large language model; All conversation keywords are feature classified according to their semantic confidence to obtain different semantic intentions of the target user in the current conversation scenario.
[0047] In specific implementation, extracting multiple conversation keywords from the interactive speech of the target user in the current conversation scenario can be achieved in the following manner, namely: first, using the Transformer model to convert the interactive speech of the target user in the current conversation scenario into text information, and then using a keyword extraction algorithm (such as the TextRank algorithm) to extract multiple keywords from the text information, and using the obtained keywords as conversation keywords. Other methods can also be used in other embodiments, which will not be repeated here.
[0048] It should be noted that the pre-trained large language model used in this application is the GPT-4 model. GPT-4 is a large multimodal language model developed by OpenAI, which performs well in language understanding. GPT-4 can handle complex sentence structures and detailed semantic understanding, and supports more natural conversations and text generation. GPT-4 evaluates the importance of each word in the sequence through the self-attention mechanism and assigns different weights according to their importance, so as to better understand language structure and patterns. In addition, GPT-4 can also understand the resolution of reference in multi-round conversations and more accurately handle complex semantic relationships such as context, metaphor, and humor in natural language.
[0049] In specific implementation, the semantic confidence of each dialogue keyword is determined based on the pre-trained large language model, which can be implemented in the following manner, namely: first, each dialogue keyword is input into the pre-trained large language model, and the output result of the large language model is used as the semantic information of each dialogue keyword; secondly, the Word2Vec model is used to convert each dialogue keyword and the semantic information of each dialogue keyword into a vector, thereby obtaining a vector representation of each dialogue keyword and a vector representation of the semantic information of each dialogue keyword; then, the dot product between the vector representation of each dialogue keyword and the vector representation of the semantic information of each dialogue keyword is calculated, and the obtained dot products are used as the semantic confidence of each dialogue keyword. Other methods can also be used in other embodiments, which are not limited here.
[0050] It should be noted that the semantic confidence described in this application represents the reliability of measuring the semantic understanding of conversation keywords.
[0051] In specific implementation, all conversation keywords are feature classified according to their semantic confidences, and different semantic intentions of the target user in the current conversation scenario can be obtained in the following manner, namely: first, a clustering algorithm (such as the Mean Shift algorithm) is used to cluster the semantic confidences of all conversation keywords, thereby obtaining multiple data clusters, wherein each data cluster is composed of semantic confidences, and then, a set of conversation keywords corresponding to each semantic confidence in the same data cluster is taken as a semantic intention of the target user in the current conversation scenario, thereby obtaining different semantic intentions of the target user in the current conversation scenario. Other methods may also be used in other embodiments, which will not be repeated here.
[0052] It should be noted that the semantic intent in this application represents the core demand characteristics of the same type expressed by the target user in the current dialogue scenario, where the semantic intent is a set of multiple dialogue keywords.
[0053] In some embodiments, determining the semantic association relationship between various semantic intents may be achieved by using the following steps: Get all conversation keywords; Perform topic analysis on all conversation keywords based on the pre-trained topic analysis model to obtain the topic features of the target user's interactive speech in the current conversation scenario; The semantic association relationship between various semantic intentions is determined according to the topic features of the target user's interactive speech in the current dialogue scenario and the different semantic intentions of the target user in the current dialogue scenario.
[0054] It should be noted that the pre-trained topic analysis model in this application is the BERTopic model. The BERTopic model is a powerful tool that combines the pre-trained model BERT and topic modeling. It captures the semantic information of all conversation keywords through the BERT model, and then uses topic modeling technology to cluster these semantic information to derive topics.
[0055] In specific implementation, topic analysis is performed on all conversation keywords based on a pre-trained topic analysis model to obtain the topic features of the target user's interactive speech in the current conversation scenario. This can be achieved in the following manner, namely: first, all conversation keywords are input into the pre-trained topic analysis model to output multiple topic candidate words, and then a set of all topic candidate words is used as the topic features of the target user's interactive speech in the current conversation scenario. Other methods may also be used in other embodiments, which are not limited here.
[0056] It should be noted that the topic features described in this application are a set consisting of multiple topic candidate words.
[0057] In specific implementation, the semantic association relationship between each semantic intention is determined according to the topic features of the target user's interactive speech in the current dialogue scenario and the different semantic intentions of the target user in the current dialogue scenario. This can be achieved in the following manner, namely: calculating the cosine similarity between each topic candidate word in the topic features of the target user's interactive speech in the current dialogue scenario and each dialogue keyword in each semantic intention of the target user in the current dialogue scenario, and then summing up all the cosine similarities, and using the summed value as the semantic association relationship between each semantic intention. Other methods may also be used in other embodiments, which will not be repeated here.
[0058] It should be noted that the semantic association relationship in this application represents the measurement parameters of the mutual influence of various semantic intentions at the semantic level.
[0059] In step 105, the semantic completeness of the target user in the current dialogue scenario is determined through the semantic association relationship and the dialogue coherence of the interactive speech in each dialogue round, and then the feedback control instruction of the voice interaction device is generated according to the semantic completeness.
[0060] In some embodiments, determining the semantic completeness of the target user in the current dialogue scenario through the semantic association relationship and the dialogue coherence of the interactive speech in each dialogue turn can be achieved by using the following steps: Perform linear fitting on the dialogue coherence of the interactive speech in all dialogue rounds to obtain a fitting curve of the dialogue coherence; Get the semantic confidence of each conversation keyword; Determine a semantic association graph of the target user in the current conversation scenario according to all semantic confidences and the semantic association relationships; The semantic completeness of the target user in the current dialogue scenario is determined by the fitting curve of the semantic association graph and the dialogue coherence.
[0061] In a specific implementation, linear fitting is performed on the dialogue coherence of the interactive speech in all dialogue turns to obtain a fitting curve of the dialogue coherence, which can be implemented in the following manner, namely: an existing linear fitting algorithm (such as a least squares support vector machine algorithm) is used to perform linear fitting on all dialogue coherences, and the curve obtained by fitting is used as the fitting curve of the dialogue coherence, wherein each value on the fitting curve is used as a dialogue coherence fitting value, and each dialogue coherence fitting value corresponds to a dialogue coherence. In other embodiments, other methods may also be used to implement the method, which will not be described in detail here.
[0062] In specific implementation, determining the semantic association graph of the target user in the current dialogue scenario based on all semantic confidences and the semantic association relationships can be implemented in the following manner, namely: based on graph theory, each semantic confidence is taken as a node, and the absolute value of the difference between each two semantic confidences is taken as the edge between each two semantic confidence corresponding nodes, thereby obtaining a graph structure, and further using the semantic association relationship to update the weights of all edges, and using the graph structure after updating the weights of all edges as the semantic association graph of the target user in the current dialogue scenario. Specifically, the updating method is to multiply the weight of each edge in the graph structure by the semantic association relationship, and replace the weight of each edge in the graph structure with all the multiplied values. Other methods may also be used in other embodiments, which are not limited here.
[0063] It should be noted that the semantic association graph described in the present application represents a graph structure composed of different semantic elements and their association relationships of the target user in the current dialogue scenario.
[0064] In specific implementation, determining the semantic completeness of the target user in the current dialogue scenario through the fitting curve of the semantic association graph and the dialogue coherence can be implemented in the following manner, namely: first, the minimum spanning tree of the semantic association graph is calculated using the Prim algorithm, the weights of all edges in the minimum spanning tree are summed, and the summed value is used as the semantic association weight value, then, the corresponding dialogue coherence is subtracted from each dialogue coherence fitting value on the fitting curve, and then the absolute values of all the subtracted values are calculated, and each absolute value is further multiplied by the semantic association weight value, so that all the multiplied values are normalized, and then the average of all the normalized values is calculated, and the average is used as the semantic completeness of the target user in the current dialogue scenario. In other embodiments, other methods can also be used for implementation, which are not limited here.
[0065] It should be noted that the semantic completeness described in this application represents the completeness of the semantic information of the target user in the current dialogue scenario. The larger the semantic completeness, the higher the completeness of the semantic information of the target user in the current dialogue scenario.
[0066] In specific implementation, the feedback control instructions of the voice interaction device are generated according to the semantic completeness in the following manner, namely: first, the semantic completeness is divided into different levels in advance, each level corresponds to a feedback control instruction, which can usually be divided into two levels, namely: high level and low level, and the level corresponding to the semantic completeness is judged. If the level corresponding to the semantic completeness is high level, the feedback control instruction corresponding to the voice content is directly output. If the level corresponding to the semantic completeness is low level, new voice content is supplemented first, and the corresponding feedback control instruction is generated according to the supplemented voice content. For example, if the semantic completeness is greater than 0.6, it is high level. When the semantic completeness is high level, the feedback control instruction is directly output according to the voice content. The corresponding feedback control instructions do not need to be supplemented with new voice content. When the semantic completeness is in the range of less than 0.6, the corresponding level of semantic completeness is low. When the semantic completeness is low, new voice content needs to be supplemented and then the corresponding feedback control instructions are output. In specific implementation, for example, when the voice content is "turn on the lights in the living room and adjust the brightness to 80%", the semantic completeness is judged to be high (i.e., greater than 0.6). At this time, the voice interaction device directly outputs the corresponding feedback control instructions, i.e., sends instructions to the smart home system to turn on the lights in the living room and adjust the brightness to 80%. When the voice content is "brighten the lights", the semantic completeness is judged to be low (i.e., less than 0.6). Because it is not clear which specific lamp it is, nor is it known to what degree it should be dimmed, new voice content needs to be supplemented first. For example, new voice content is supplemented to further ask the user "Do you mean the lamp in the living room or the lamp in the bedroom? To what brightness level do you need to adjust it?" etc. After the user supplements the answer, for example, the user says "It's the lamp in the living room, adjust it to the brightest", the corresponding feedback control instruction is generated according to the supplemented voice content, that is, an instruction is sent to the smart home system to adjust the lamp in the living room to the brightest. Other methods can also be used to achieve this in other embodiments, which will not be repeated here.
[0067] In addition, in another aspect of the present application, in some embodiments, the present application provides a voice interaction system based on artificial intelligence, referring to Figure 4 , which is a schematic diagram of the structure of a voice interaction system based on artificial intelligence according to some embodiments of the present application. The voice interaction system based on artificial intelligence 400 includes: an acquisition module 401, a processing module 402 and an execution module 403, which are described as follows: Acquisition module 401, in this application, acquisition module 401 is mainly used to acquire the interactive voice of the target user in different conversation rounds through the voice interaction device; Processing module 402, in the present application, is used to extract context dependencies between each dialogue turn from all interactive speech based on the attention mechanism; It should be noted that the processing module 402 in the present application is also used to determine the semantic loss amount of the interactive speech in each dialogue round according to the semantic features of the interactive speech in each dialogue round, and determine the dialogue coherence of the target user in each dialogue round through all the semantic loss amounts and the context dependency relationship between the dialogue rounds; In addition, the processing module 402 in the present application is also used to classify the target user's interactive speech in the current dialogue scene, obtain different semantic intentions of the target user in the current dialogue scene, and then determine the semantic association relationship between the various semantic intentions; Execution module 403. In the present application, execution module 403 is mainly used to determine the semantic completeness of the target user in the current dialogue scenario through the semantic association relationship and the dialogue coherence of the interactive voice in each dialogue round, and then generate feedback control instructions for the voice interaction device based on the semantic completeness.
[0068] In addition, the present application also provides a computer device, which includes a memory and a processor, the memory stores code, and the processor is configured to obtain the code and execute the above-mentioned artificial intelligence-based voice interaction method.
[0069] In some embodiments, reference Figure 5 , which is a schematic diagram of the structure of a computer device for implementing an artificial intelligence-based voice interaction method according to some embodiments of the present application. The artificial intelligence-based voice interaction method in the above embodiment can be Figure 5 The computer device 500 shown in the figure is implemented, and the computer device 500 includes at least one processor 501, a communication bus 502, a memory 503 and at least one communication interface 504.
[0070] Processor 501 can be a general-purpose central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more for controlling the execution of the artificial intelligence-based voice interaction method in the present application.
[0071] The communication bus 502 may be used to transmit information between the above-mentioned components.
[0072] The memory 503 may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, an optical disc storage (including a compressed optical disc, a laser disc, an optical disc, a digital versatile disc, a Blu-ray disc, etc.), a magnetic disk or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of an instruction or data structure and can be accessed by a computer, but is not limited thereto. The memory 503 may exist independently and be connected to the processor 501 via the communication bus 502. The memory 503 may also be integrated with the processor 501.
[0073] The memory 503 is used to store the program code for executing the solution of the present application, and the execution is controlled by the processor 501. The processor 501 is used to execute the program code stored in the memory 503. The program code may include one or more software modules. The method described in the above method embodiment can be implemented by the processor 501 and one or more software modules in the program code in the memory 503.
[0074] The communication interface 504 uses any transceiver or other device for communicating with other devices or communication networks, such as Ethernet, radio access network (RAN), wireless local area networks (WLAN), etc.
[0075] In a specific implementation, as an embodiment, a computer device may include multiple processors, each of which may be a single-CPU processor or a multi-CPU processor. The processor here may refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions).
[0076] The above-mentioned computer device may be a general-purpose computer device or a special-purpose computer device. In a specific implementation, the computer device may be a desktop computer, a portable computer, a network server, a personal digital assistant (PDA), a mobile phone, a tablet computer, a wireless terminal device, a communication device or an embedded device. The embodiment of the present application does not limit the type of computer device.
[0077] In addition, the present application also provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, it implements the above-mentioned artificial intelligence-based voice interaction method.
[0078] Although the preferred embodiments of the present application have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present application.
[0079] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.
Claims
1. A voice interaction method based on artificial intelligence, characterized in that: The steps include: Acquire the interactive voice of the target user in different conversation rounds through the voice interaction device; Extract the contextual dependencies between each dialogue turn from all interactive speech based on the attention mechanism; Determine the semantic loss amount of the interactive speech in each dialogue round according to the semantic features of the interactive speech in each dialogue round, and determine the dialogue coherence of the target user in each dialogue round through all the semantic loss amounts and the context dependency relationship between the respective dialogue rounds; Classify the target user's interactive speech in the current dialogue scenario, obtain the target user's different semantic intentions in the current dialogue scenario, and then determine the semantic association relationship between each semantic intention; The semantic completeness of the target user in the current dialogue scenario is determined by the semantic association relationship and the dialogue coherence of the interactive speech in each dialogue round, and then the feedback control instruction of the voice interaction device is generated according to the semantic completeness.
2. The method according to claim 1, characterized in that The contextual dependencies between each dialogue turn are extracted from all interactive speech based on the attention mechanism, including: Convert all interactive speech into multiple speech vectors; Construct an attention vector space for each speech vector based on a preset attention mechanism; Determine the semantic dependency of each conversation turn based on the attention vector space of each speech vector; The contextual dependencies between each dialogue turn are determined through the semantic dependencies of all dialogue turns.
3. The method according to claim 1, characterized in that Determining the semantic loss amount of the interactive speech in each dialogue round according to the semantic features of the interactive speech in each dialogue round specifically includes: Determine the semantic features of the interactive speech at each dialogue turn; Determine the semantic similarity between interactive speech in each dialogue round based on all semantic features; The semantic loss of the interactive speech in each dialogue round is determined by the semantic similarity between the interactive speech in each dialogue round.
4. The method according to claim 1, characterized in that Determining the dialogue coherence of the target user in each dialogue turn by using all semantic loss amounts and the context dependency relationship between the dialogue turns specifically includes: Determine the dialogue state quantity for each dialogue turn based on the preset sliding window and all semantic loss quantities; The dialogue coherence of the target user in each dialogue turn is determined according to the dialogue state quantity in each dialogue turn and the context dependency relationship between the dialogue turns.
5. The method according to claim 1, characterized in that The target user's interactive speech in the current dialogue scenario is classified into intents, and the different semantic intentions of the target user in the current dialogue scenario are obtained, including: Extract multiple conversation keywords from the target user's interactive speech in the current conversation scenario; Determine the semantic confidence of each conversation keyword based on the pre-trained large language model; All conversation keywords are feature classified according to their semantic confidence to obtain different semantic intentions of the target user in the current conversation scenario.
6. The method according to claim 1, characterized in that Determining the semantic association relationship between various semantic intents specifically includes: Get all conversation keywords; Perform topic analysis on all conversation keywords based on the pre-trained topic analysis model to obtain the topic features of the target user's interactive speech in the current conversation scenario; The semantic association relationship between various semantic intentions is determined according to the topic features of the target user's interactive speech in the current dialogue scenario and the different semantic intentions of the target user in the current dialogue scenario.
7. The method according to claim 1, characterized in that Determining the semantic integrity of the target user in the current dialogue scenario through the semantic association relationship and the dialogue coherence of the interactive speech in each dialogue turn specifically includes: Perform linear fitting on the dialogue coherence of the interactive speech in all dialogue rounds to obtain a fitting curve of the dialogue coherence; Get the semantic confidence of each conversation keyword; Determine a semantic association graph of the target user in the current conversation scenario according to all semantic confidences and the semantic association relationships; The semantic completeness of the target user in the current dialogue scenario is determined by the fitting curve of the semantic association graph and the dialogue coherence.
8. A voice interaction system based on artificial intelligence, characterized in that: include: An acquisition module, used to acquire the interactive voice of the target user in different conversation turns through a voice interaction device; A processing module that extracts contextual dependencies between conversation turns from all interactive speech based on an attention mechanism; The processing module is further used to determine the semantic loss amount of the interactive speech in each dialogue round according to the semantic features of the interactive speech in each dialogue round, and determine the dialogue coherence of the target user in each dialogue round through all the semantic loss amounts and the context dependency relationship between the various dialogue rounds; The processing module is further used to classify the target user's interactive speech in the current dialogue scene, obtain different semantic intentions of the target user in the current dialogue scene, and then determine the semantic association relationship between the various semantic intentions; An execution module is used to determine the semantic completeness of the target user in the current dialogue scenario through the semantic association relationship and the dialogue coherence of the interactive speech in each dialogue round, and then generate feedback control instructions for the voice interaction device according to the semantic completeness.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the artificial intelligence-based voice interaction method described in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the artificial intelligence-based voice interaction method as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Text filling method and device
CN112069810A
Vehicle man-machine voice interaction method and system and vehicle
CN118968992A
Method for correcting dialogue intention information of old people based on deep neural network
CN119577522A
AI intelligent customer service response method and system based on remote digital service
CN119719319A
Intent Identification for Agent Matching by Assistant Systems
US20230118962A1
Cited By
Interaction method and device for unmanned equipment, related equipment and program product
CN121260155A