Intention prediction method, device, equipment and medium for intelligent human-computer interaction

By integrating emotional fluctuations and interest focus data in real time in intelligent human-computer interaction, generating emotional state vectors and interest distribution vectors, and combining them with historical behavior databases for intention matching, the problem of lack of real-time perception capabilities in existing technologies is solved, and the smoothness of interaction and response accuracy are improved.

CN120523332BActive Publication Date: 2025-09-30GUANGZHOU HUAXIA VOCATIONAL COLLEGE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511020501.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-24
Publication Date
2025-09-30
Estimated Expiration
2045-07-24

AI Technical Summary

Technical Problem

Existing intelligent human-computer interaction technologies lack the ability to perceive real-time interaction dynamics and are unable to effectively quantify emotional fluctuations and focus of interest, resulting in delayed responses and difficulty in adapting to the flexible needs of dynamic scenarios.

Method used

The original emotional data matrix is ​​generated by extracting semantic emotional features from real-time user interaction data, and the emotional state vector is generated after dynamic calibration. The interest focus is modeled by combining the user topic interaction log, and the historical behavior database is used for bimodal intention matching to generate predicted user intentions.

Benefits of technology

It achieves real-time adaptation to the user's current psychological state, improves the smoothness of interaction and the accuracy of demand response, and solves the response lag problem caused by the traditional method's reliance on historical data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120523332B_ABST
    Figure CN120523332B_ABST
Patent Text Reader

Abstract

The present application provides an intention prediction method, device, equipment and medium for intelligent human-computer interaction, wherein the method includes: extracting semantic emotional features from the acquired real-time user interaction data to generate an original emotional data matrix, wherein the real-time interaction data includes a voice stream and a text input stream; dynamically calibrating the original emotional data matrix to obtain an emotional state vector; obtaining the user topic interaction log and combining the emotional state vector to perform interest focus modeling to generate a dynamic interest distribution vector, which is used to quantify the topic attention; performing bimodal intention matching processing based on the historical behavior database and the dynamic interest distribution vector to generate predicted user intentions. This method can be used to integrate emotional fluctuations and interest focus data in real time during the interaction process, thereby better fitting the user's current psychological state and improving the interaction fluency and demand response accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer technology, and in particular relates to an intention prediction method, device, equipment and medium for intelligent human-computer interaction. Background Art

[0002] With the evolution of intelligent human-computer interaction technology, intent prediction methods based on user behavior modeling are becoming mainstream. This type of technology builds statistical models by analyzing historical user interaction data (such as click patterns and dwell time) and uses machine learning algorithms to infer underlying intent. Its advantage lies in its ability to process structured behavioral data and achieve batch prediction. Traditional methods typically employ static feature extraction and offline training, such as collaborative filtering recommendations or hidden Markov models to capture behavioral patterns.

[0003] However, existing technologies rely too heavily on historical data and lack the ability to perceive the dynamics of real-time interactions. Specifically, emotional fluctuations (such as changes in voice tone and facial micro-expressions) are not effectively quantified, resulting in the system being unable to distinguish between a user's current state of excitement and frustration. Furthermore, shifts in focus (such as sudden shifts in attention to a topic) are often misled by historical behavior paths due to the lack of a dynamic monitoring mechanism. This results in traditional predictions being slow to respond in dynamic scenarios and unable to adapt to the flexible demands of real-time interactions. Summary of the Invention

[0004] Based on this, it is necessary to provide an intention prediction method, device, equipment and medium for intelligent human-computer interaction to address the above technical problems, which can integrate emotional fluctuations and interest focus data in real time during the interaction process, so as to better fit the user's current psychological state and improve the smoothness of interaction and the accuracy of demand response.

[0005] In a first aspect, the present application provides a method for predicting intentions in intelligent human-computer interaction, comprising:

[0006] Extracting semantic emotional features from the acquired real-time user interaction data to generate an original emotional data matrix, where the real-time interaction data includes voice streams and text input streams;

[0007] Dynamically calibrate the original emotion data matrix to obtain the emotion state vector;

[0008] Obtain user topic interaction logs and combine them with the emotional state vector to model interest focus and generate a dynamic interest distribution vector. The dynamic interest distribution vector is used to quantify topic attention.

[0009] Bimodal intent matching is performed based on the historical behavior database and dynamic interest distribution vector to generate predicted user intent.

[0010] In one embodiment, semantic emotion features are extracted from the acquired real-time user interaction data to generate an original emotion data matrix, including:

[0011] Perform multi-scale acoustic feature extraction and feature fusion on the speech stream to generate acoustic features;

[0012] Parse the text input stream into a six-dimensional emotion probability vector using the RoBERTa sentiment classifier, where the six-dimensional emotion probability vector includes joy, anger, sadness, fear, surprise, and neutral;

[0013] The acoustic features and the six-dimensional emotion probability vector are timestamp aligned, and the original emotion data matrix containing the time dimension is generated by feature concatenation.

[0014] In one embodiment, the original emotion data matrix is ​​dynamically calibrated to obtain the emotion state vector, including:

[0015] Obtain the current conversation history text and generate a contextual semantic vector through a bidirectional LSTM network. The contextual semantic vector is used to represent the semantic features of the conversation scene.

[0016] Calculate the attention weight distribution of the original sentiment data matrix based on the contextual semantic vector;

[0017] The original emotion data matrix is ​​weighted filtered according to the attention weight distribution and the emotion state vector is output.

[0018] In one embodiment, obtaining user topic interaction logs and combining them with the emotional state vector to perform interest focus modeling and generate a dynamic interest distribution vector includes:

[0019] Based on the user topic interaction log, the response frequency and duration of each topic within the preset time window are counted to generate the initial interest weight;

[0020] The sentiment state vector is input into the dilated convolutional neural network, and the sentiment-topic coupling coefficient is generated through time derivative analysis;

[0021] The initial interest weight is dynamically attenuated based on the emotion-topic coupling coefficient, and a dynamic interest distribution vector is generated through normalization.

[0022] In one embodiment, bimodal intent matching is performed based on a historical behavior database and a dynamic interest distribution vector to generate a predicted user intent, including:

[0023] Mining time series patterns on historical behavior databases to generate long-term intention preference vectors;

[0024] The dynamic time warping algorithm is used to calculate the similarity between the long-term intention preference vector and the dynamic interest distribution vector to generate the intention drift factor.

[0025] Adaptive fusion weights are constructed based on the intent drift factor to perform weighted fusion of the long-term intent preference vector and the dynamic interest distribution vector to generate a comprehensive intent feature vector.

[0026] The comprehensive intent feature vector is input into the pre-trained intent classifier for probability prediction to generate the predicted user intent.

[0027] In a second aspect, the present application also provides an intention prediction device for intelligent human-computer interaction, comprising:

[0028] An interaction feature extraction module is used to extract semantic emotional features from the acquired real-time user interaction data to generate an original emotional data matrix, where the real-time interaction data includes voice streams and text input streams;

[0029] Dynamic calibration module, used to dynamically calibrate the original emotion data matrix to obtain the emotion state vector;

[0030] The interest focus modeling module is used to obtain user topic interaction logs and combine them with the emotional state vector to perform interest focus modeling, generating a dynamic interest distribution vector. The dynamic interest distribution vector is used to quantify topic attention;

[0031] The bimodal intent matching module is used to perform bimodal intent matching processing based on the historical behavior database and dynamic interest distribution vector to generate predicted user intent.

[0032] In a third aspect, the present application also provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the above-mentioned intelligent human-computer interaction intention prediction method when executing the computer program.

[0033] In a fourth aspect, the present application also provides a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, the intention prediction method of the above-mentioned intelligent human-computer interaction is implemented.

[0034] The aforementioned intelligent human-computer interaction intention prediction method, device, equipment, and medium extract semantic emotional features from real-time user interaction data (covering voice and text input streams) to generate a raw emotional data matrix. This is then dynamically calibrated to obtain an emotional state vector. This is then combined with user topic interaction logs to model interest focus and generate a dynamic interest distribution vector. This dynamic interest distribution vector is then used to perform bimodal intent matching with a historical behavior database to predict user intent. By integrating emotional fluctuations and interest focus data in real time, this technical solution effectively addresses the overreliance on historical data and lack of real-time interaction dynamic perception capabilities of traditional methods. This approach achieves the technical effect of tailoring to the user's current state of mind, improving interaction fluency and the accuracy of demand response. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0036] Figure 1 A flowchart of an intention prediction method for intelligent human-computer interaction provided by an embodiment of the present invention;

[0037] Figure 2 A schematic diagram of the structure of an intelligent human-computer interaction intention prediction device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0038] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0039] In an exemplary embodiment, Figure 1 As shown, a method for intelligently predicting the audience flow of sports events is provided, comprising the following steps 101 to 104:

[0040] Step 101 : extracting semantic emotional features from acquired real-time user interaction data to generate an original emotional data matrix, wherein the real-time interaction data includes a voice stream and a text input stream.

[0041] Specifically, real-time interactive data covers two types of information sources: voice streams and text input streams. Voice streams can be processed using speech emotion recognition algorithms to analyze features such as intonation, speaking rate, and volume to extract emotional features. Text input streams, on the other hand, utilize natural language processing techniques to identify emotional vocabulary and semantic tendencies in text for feature extraction. Furthermore, in voice stream processing, after preprocessing the speech signal, emotion recognition models using recurrent neural networks (RNNs) or convolutional neural networks (CNNs) can be used to encode emotional features of speech frame sequences. When processing text input streams, pretrained language models such as BERT can be fine-tuned for text emotion classification, extracting emotion vectors from text segments, and integrating the emotional features in speech and text into a raw emotion data matrix, providing the basic data structure for subsequent emotional state analysis.

[0042] Step 102: Dynamically calibrate the original emotion data matrix to obtain an emotion state vector.

[0043] Specifically, considering the dynamic and volatile nature of user emotions during the interaction process, a dynamic calibration mechanism is adopted to improve the timeliness and accuracy of emotional data. For example, the Kalman filter algorithm can be used to treat the eigenvalues ​​in the original emotional data matrix as noisy measurement inputs, and through state prediction and update steps, a more accurate emotional state can be estimated; or the long short-term memory network (LSTM) can be used to process the emotional feature sequence with its memory and forgetting properties of time series information, so that the output emotional state vector can reflect the user's current emotional changes in real time. For example, when the user's voice tone suddenly rises during the interaction and the text input contains intense words, the dynamically calibrated vector can accurately identify the user's emotional state of "excitement" or "anger", and can better capture the instantaneous changes in emotion compared to static emotion analysis methods.

[0044] Step 103: Obtain user topic interaction logs and perform interest focus modeling in combination with the emotional state vector to generate a dynamic interest distribution vector. The dynamic interest distribution vector is used to quantify topic attention.

[0045] Specifically, based on the accumulation of user interaction logs to form topic interaction logs, the emotional state vector is incorporated as a weighting factor into the interest focus modeling process. For example, an attention mechanism can be employed, using the emotional state vector as the query vector and the characteristics of each topic in the topic interaction log as the key-value vector. The attention weight of each topic under the current emotional state is calculated, thereby generating a dynamic interest distribution vector. Furthermore, topic models such as Latent Dirichlet Allocation (LDA) can be used to mine topics in the interaction logs. The probability distribution of each topic is then adjusted based on the emotional vector. This allows the vector to not only reflect the user's long-term interests but also dynamically highlight immediate interests driven by specific emotions. For example, when a user is in a "curious" emotional state, the attention weight of emerging technology topics will be increased accordingly, achieving a deep integration of emotion and interest.

[0046] Step 104 : Perform bimodal intention matching processing based on the historical behavior database and the dynamic interest distribution vector to generate a predicted user intention.

[0047] Specifically, the historical behavior database stores a large number of user past behavior samples and their corresponding intention annotations, providing rich semantic and behavioral pattern references. During the matching process, on the one hand, the dynamic interest distribution vector is used to determine the topic range of the user's current interest tendency, and on the other hand, similar behavioral patterns are retrieved from the historical behavior database. For example, the cosine similarity between the dynamic interest vector and the feature vector of the historical behavior sample can be calculated to screen out a set of candidate intentions with high similarity; at the same time, the dual-tower model structure in deep learning is introduced to encode the dynamic interest vector and the historical behavior vector respectively, and the matching distance is optimized through metric learning to accurately predict the user's current intention. For example, when the dynamic interest points to "travel guide" and matches the "make travel plans" pattern in the historical behavior, it is predicted that the user's intention may be "query tourist attraction information" or "book hotel and air tickets", realizing efficient fusion prediction of dual-modal data.

[0048] The aforementioned intelligent human-computer interaction intention prediction method, device, equipment, and medium extract semantic emotional features from real-time user interaction data (covering voice and text input streams) to generate a raw emotional data matrix. This is then dynamically calibrated to obtain an emotional state vector. This is then combined with user topic interaction logs to model interest focus and generate a dynamic interest distribution vector. This dynamic interest distribution vector is then used to perform bimodal intent matching with a historical behavior database to predict user intent. By integrating emotional fluctuations and interest focus data in real time, this technical solution effectively addresses the overreliance on historical data and lack of real-time interaction dynamic perception capabilities of traditional methods. This approach achieves the technical effect of tailoring to the user's current state of mind, improving interaction fluency and the accuracy of demand response.

[0049] In one embodiment, semantic emotion features are extracted from the acquired real-time user interaction data to generate an original emotion data matrix, including:

[0050] Perform multi-scale acoustic feature extraction and feature fusion on the speech stream to generate acoustic features.

[0051] Specifically, multi-scale acoustic feature extraction technology is used for speech streams to capture emotional information from different time scales and frequency resolutions. For example, basic acoustic features such as Mel-frequency cepstral coefficients, intonation contours, and intensity changes are extracted. These multi-scale features are then integrated into a unified acoustic feature vector through a feature fusion algorithm. For example, a short-time Fourier transform combined with a wavelet transform can be used to first obtain the overall intonation trend at a coarser time scale, and then extract short emotional outburst features at a finer scale. Principal component analysis (PCA) or linear discriminant analysis (LDA) is then used to reduce the dimensionality of these features and fuse them together to generate compact and information-rich acoustic features.

[0052] The text input stream is parsed into a six-dimensional emotion probability vector through the RoBERTa sentiment classifier, where the six-dimensional emotion probability vector includes joy, anger, sadness, fear, surprise, and neutral.

[0053] Specifically, sentiment analysis is performed on the text input stream using a pre-trained RoBERTa sentiment classifier. The RoBERTa sentiment classifier is a pre-trained language model based on a modified Transformer architecture, optimized for text sentiment analysis. This model learns the deep semantic representation patterns of text through self-supervised pre-training on large-scale unlabeled corpora (e.g., masked language modeling tasks). Further fine-tuning on sentiment-labeled datasets (e.g., SST and EmoContext) enables it to accurately identify sentiment tendencies and parse input text sequences into six-dimensional sentiment probability vectors. Specifically, each sentence or phrase in the text input stream is encoded as a fixed-length vector. A multi-layer Transformer architecture captures the contextual semantics and sentiment tendencies of the text, outputting a probability distribution over the six basic emotions: joy, anger, sadness, fear, surprise, and neutrality. For example, when a user enters "Great! This feature exceeds my expectations!", the classifier outputs a sentiment vector with a high probability value for the joy dimension.

[0054] The acoustic features and the six-dimensional emotion probability vector are timestamp aligned, and the original emotion data matrix containing the time dimension is generated by feature concatenation.

[0055] Specifically, because the speech and text input streams may be out of sync in time, a timestamp correction algorithm, such as alignment based on dynamic time warping (DTW), is required to ensure their temporal consistency. Subsequently, through feature concatenation, the aligned acoustic feature vectors and the six-dimensional emotion probability vectors are concatenated along the time dimension into a comprehensive feature matrix, the original emotion data matrix. For example, for each segment of speech and corresponding text input of duration t, after timestamp alignment, their respective feature vectors are sequentially arranged along the time dimension to form an N × (M + 6) matrix (where N is the time step and M is the acoustic feature dimension). This resulting matrix not only contains the emotional features of the speech and text, but also preserves the dynamic information of emotion changes over time, laying the foundation for more accurate subsequent emotion analysis and intent prediction.

[0056] In one embodiment, the original emotion data matrix is ​​dynamically calibrated to obtain the emotion state vector, including:

[0057] The current conversation history text is obtained and a contextual semantic vector is generated through a bidirectional LSTM network. The contextual semantic vector is used to represent the semantic features of the conversation scene.

[0058] For example, the current conversation history is used as a foundation for semantic understanding. A bidirectional LSTM network is used to encode the continuous conversation sequence, and a gating mechanism is used to fuse forward and backward semantic features. Furthermore, a pooling layer is used to extract the core semantic representation of the conversation scene, generating a fixed-dimensional contextual semantic vector. This vector accurately represents the current interaction's topic, participant relationships, and context, providing a semantic benchmark for sentiment calibration.

[0059] Calculate the attention weight distribution of the original sentiment data matrix based on the contextual semantic vector.

[0060] Specifically, the contextual semantic vector is used as the query vector, and the emotional feature vectors in the original sentiment data matrix are used as key-value vectors. The correlation between them is calculated to obtain the attention weight distribution. For example, in an instant messaging application, for a captured user expression of sentiment words such as "very satisfied", the correlation between the sentiment feature vector containing this word and the contextual semantic vector is calculated, thereby obtaining the attention weight of this feature vector, highlighting key emotional features and providing a precise weight basis for subsequent emotional state calibration.

[0061] The original emotion data matrix is ​​weighted filtered according to the attention weight distribution and the emotion state vector is output.

[0062] Specifically, based on the calculated attention weight distribution, the original emotion data matrix is ​​weighted filtered to output an emotion state vector. For example, in the interactive phase of an online education platform, the original emotion data matrix containing the user's learning emotion fluctuations is weighted filtered based on the calculated attention weight distribution. Emotional features with high weights will have a greater impact on the final emotion state vector, while features with low weights will be appropriately suppressed. Through weighted filtering, the output emotion state vector can more accurately reflect the user's true emotional state in the current conversation scenario, avoid interference from noisy data, and provide more reliable emotion state information for subsequent emotion analysis and intent prediction.

[0063] In one embodiment, obtaining user topic interaction logs and combining them with the emotional state vector to perform interest focus modeling and generate a dynamic interest distribution vector includes:

[0064] Based on the user topic interaction log, the response frequency and duration of each topic within the preset time window are counted to generate the initial interest weight.

[0065] Specifically, based on user topic interaction logs, the response frequency and duration of each topic within a preset time window are counted to generate an initial interest weight. For example, on a social media platform, user interactions such as comments and likes on different topics within the last hour (preset time window) are collected. The number of responses to each topic (response frequency) and the length of time users interacted with the topic (duration) are counted. Then, using a specific weight calculation formula, such as assigning different weight coefficients to the response frequency and duration and then adding them together, an initial interest weight is generated to preliminarily quantify the user's level of interest in each topic.

[0066] The sentiment state vector is input into the atrous convolutional neural network, and the sentiment-topic coupling coefficient is generated through time derivative analysis.

[0067] For example, a dilated convolutional neural network (DCN) is a deep learning architecture that expands the receptive field by introducing a dilated sampling mechanism. For example, in online education platforms, students' emotional state vectors (e.g., focus, confusion, satisfaction, etc.) during the learning process are input into the network. Time derivative analysis is used to calculate the rate of change of emotional state, which in turn generates an emotion-topic coupling coefficient. This coefficient reflects the degree to which emotional fluctuations affect the attention paid to different topics. For example, confusion may reduce attention to the current learning topic, while satisfaction may increase it.

[0068] The initial interest weight is dynamically attenuated based on the emotion-topic coupling coefficient, and a dynamic interest distribution vector is generated through normalization.

[0069] Specifically, the initial interest weights are combined with the emotion-topic coupling coefficient to adjust the initial weights. For example, in an intelligent customer service system, if the emotion-topic coupling coefficient indicates that a user's interest in the current complaint topic has increased dramatically due to anger, the weight of that topic will be increased accordingly. Conversely, if a user's interest in a topic has decreased due to boredom, the weight will be decreased. The modified weights are then processed using a normalization function such as Softmax to generate a dynamic interest distribution vector. This ensures that the sum of the weights of each topic is 1, ensuring a reasonable and comparable distribution. This results in a dynamic vector that accurately reflects the user's real-time focus of interest.

[0070] In one embodiment, bimodal intent matching is performed based on a historical behavior database and a dynamic interest distribution vector to generate a predicted user intent, including:

[0071] Perform temporal pattern mining on the historical behavior database to generate long-term intention preference vectors.

[0072] For example, on an e-commerce platform, a historical behavior database contains records of users' past browsing, purchasing, and collection behaviors. Association rule mining techniques, such as the Apriori algorithm or the FP-Growth algorithm, are used to uncover frequent patterns and temporal regularities in user behavior, such as a user's interest in a certain type of product within a specific time period. This generates a long-term intention preference vector, which characterizes the user's long-term and stable interests and provides a historical basis for subsequent intention matching.

[0073] The dynamic time warping algorithm is used to calculate the similarity between the long-term intention preference vector and the dynamic interest distribution vector to generate the intention drift factor.

[0074] Specifically, the dynamic time warping algorithm can effectively handle the scaling and distortion of two time series on the time axis. For example, in an online video platform, the long-term intention preference vector (reflecting the user's long-term interest in a specific type of video) and the dynamic interest distribution vector (reflecting the user's current real-time focus of interest) are treated as time series inputs, and the similarity between them is calculated. By comparing the weight distribution of the two vectors in different dimensions (such as different video categories), the dynamic time warping algorithm outputs an intention drift factor, which quantifies the degree of deviation between the user's current intention and long-term intention. For example, when a user suddenly switches from a long-term focus on science and technology videos to entertainment videos, the intention drift factor will increase accordingly, providing a dynamic adjustment basis for subsequent intention fusion.

[0075] An adaptive fusion weight is constructed based on the intention drift factor, and the long-term intention preference vector and the dynamic interest distribution vector are weightedly fused to generate a comprehensive intention feature vector.

[0076] Specifically, the fusion weight is dynamically adjusted based on the intent drift factor. When the intent drift factor is large, the dynamic interest distribution vector is given greater weight, while when it is small, the long-term intent preference vector is given more weight. Through weighted fusion, long-term intent and real-time interests are organically combined to generate a comprehensive intent feature vector. This vector retains the user's historical behavior habits while reflecting the current immediate intent, providing a comprehensive feature representation for final intent prediction.

[0077] The comprehensive intent feature vector is input into the pre-trained intent classifier for probability prediction to generate the predicted user intent.

[0078] Specifically, pre-trained intent classifiers can use models based on the Transformer architecture, such as BERT, which are pre-trained and fine-tuned on large-scale user behavior data and intent annotation data. For example, in an intelligent customer service system, the comprehensive intent feature vector is input into the classifier, which outputs the probability distribution of the user's different intents (such as consultation, complaint, purchase, etc.). The user's intent is predicted based on the probability, and the intent with the highest probability is the predicted result. This allows for accurate prediction of the user's current interaction intent, improving the efficiency and accuracy of intent prediction responses.

[0079] In summary, the intention prediction method for intelligent human-computer interaction provided by this application captures the changes in the user's psychological state in real time during the interaction by fusing emotional fluctuations and focus of interest data, including extracting semantic emotional features from voice streams and text input streams to generate an original emotional data matrix, and dynamically calibrating to obtain an emotional state vector; statistically calculating the response frequency and duration within a preset time window to generate an initial interest weight, and combining the emotional state vector with a void convolutional neural network to calculate the emotion-topic coupling coefficient, dynamically correcting the initial interest weight to generate a dynamic interest distribution vector; mining the historical behavior database for time series patterns to generate a long-term intention preference vector, and calculating the similarity through a dynamic time warping algorithm to generate an intention drift factor, and then constructing an adaptive fusion weight, weighted fusion of the long-term intention preference vector and the dynamic interest distribution vector to generate a comprehensive intention feature vector, and using a pre-trained intention classifier for probabilistic prediction to generate predicted user intentions. The above technical solution constructs a dynamic intention prediction model by fusing emotional fluctuations and focus of interest data in real time, effectively solving the problem of delayed response in dynamic scenarios in existing technologies, and improving the fluency of the intelligent human-computer interaction system, the accuracy of demand response, and the adaptability to changes in user psychological state.

[0080] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0081] Based on the same inventive concept, the embodiments of the present application also provide an intelligent human-computer interaction intention prediction device 10 for implementing the above-mentioned intelligent human-computer interaction intention prediction method. The implementation solution provided by this device is similar to the implementation solution described in the above-mentioned method. Therefore, the specific limitations of one or more embodiments of the intelligent human-computer interaction intention prediction device 10 provided below can be referred to the limitations of the intelligent human-computer interaction intention prediction method above, and will not be repeated here.

[0082] In an exemplary embodiment, Figure 2 As shown, an intention prediction device 10 for intelligent human-computer interaction is provided, comprising:

[0083] The interaction feature extraction module 11 is used to extract semantic emotional features from the acquired real-time user interaction data to generate an original emotional data matrix, wherein the real-time interaction data includes a voice stream and a text input stream.

[0084] The dynamic calibration module 12 is used to dynamically calibrate the original emotion data matrix to obtain the emotion state vector.

[0085] The interest focus modeling module 13 is used to obtain user topic interaction logs and perform interest focus modeling in combination with the emotional state vector to generate a dynamic interest distribution vector, which is used to quantify topic attention.

[0086] The bimodal intention matching module 14 is used to perform bimodal intention matching processing based on the historical behavior database and the dynamic interest distribution vector to generate a predicted user intention.

[0087] In one embodiment, the interactive feature extraction module 11 includes:

[0088] Acoustic feature extraction unit, used to extract and fuse multi-scale acoustic features of speech streams to generate acoustic features;

[0089] Sentiment parsing unit, used to parse the text input stream into a six-dimensional sentiment probability vector using the RoBERTa sentiment classifier, where the six-dimensional sentiment probability vector includes joy, anger, sadness, fear, surprise, and neutral;

[0090] The feature fusion unit is used to align the timestamps of acoustic features and six-dimensional emotion probability vectors, and generate the original emotion data matrix containing the time dimension through feature splicing.

[0091] In one embodiment, the dynamic calibration module 12 includes:

[0092] The context encoding unit is used to obtain the current conversation history text and generate a context semantic vector through a bidirectional LSTM network. The context semantic vector is used to represent the semantic features of the conversation scene.

[0093] Attention calculation unit, used to calculate the attention weight distribution of the original sentiment data matrix based on the context semantic vector;

[0094] The filtering processing unit is used to perform weighted filtering on the original emotion data matrix according to the attention weight distribution and output the emotion state vector.

[0095] In one embodiment, the interest focus modeling module 13 includes:

[0096] The interest statistics unit is used to count the response frequency and duration of each topic within a preset time window based on the user topic interaction log and generate the initial interest weight;

[0097] The coupling analysis unit is used to input the emotional state vector into the dilated convolutional neural network and generate the emotion-topic coupling coefficient through time derivative analysis;

[0098] The dynamic correction unit is used to dynamically attenuate the initial interest weight based on the emotion-topic coupling coefficient and generate a dynamic interest distribution vector through normalization processing.

[0099] In one embodiment, the bimodal intent matching module 14 includes:

[0100] Pattern mining unit, used to mine time series patterns in the historical behavior database and generate long-term intention preference vectors;

[0101] a drift calculation unit, configured to calculate the similarity between the long-term intention preference vector and the dynamic interest distribution vector using a dynamic time warping algorithm to generate an intention drift factor;

[0102] The feature fusion unit is used to construct an adaptive fusion weight based on the intent drift factor, perform weighted fusion on the long-term intent preference vector and the dynamic interest distribution vector, and generate a comprehensive intent feature vector;

[0103] The intent prediction unit is used to input the comprehensive intent feature vector into the pre-trained intent classifier for probability prediction and generate predicted user intent.

[0104] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of the method for predicting intentions of intelligent human-computer interaction as described above are implemented.

[0105] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0106] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial description of the method embodiments. The device embodiments described above are merely illustrative, wherein the components described as separate parts may or may not be physically separated, and the parts displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the disclosed solution. A person of ordinary skill in the art can understand and implement it without expending creative work.

[0107] The above-described embodiments merely represent several implementation methods of the embodiments of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that a person skilled in the art may make various modifications and improvements without departing from the concept of the embodiments of the present application, and these modifications and improvements fall within the scope of protection of the embodiments of the present application.

Claims

1. A method for predicting intentions in intelligent human-computer interaction, characterized in that: The method comprises: Extracting semantic emotional features from the acquired real-time user interaction data to generate an original emotional data matrix, wherein the real-time interaction data includes a voice stream and a text input stream; Dynamically calibrating the original emotion data matrix to obtain an emotion state vector; Obtaining user topic interaction logs and combining them with the emotional state vector to perform interest focus modeling to generate a dynamic interest distribution vector, which is used to quantify topic attention; Performing bimodal intent matching processing based on the historical behavior database and the dynamic interest distribution vector to generate predicted user intent; The step of obtaining the user topic interaction log and combining the emotional state vector to perform interest focus modeling to generate a dynamic interest distribution vector includes: Based on the user topic interaction log, the response frequency and duration of each topic within a preset time window are counted to generate an initial interest weight; Inputting the emotional state vector into a dilated convolutional neural network, and generating an emotion-topic coupling coefficient through time derivative analysis; The initial interest weight is dynamically attenuated and corrected based on the emotion-topic coupling coefficient, and the dynamic interest distribution vector is generated through normalization processing.

2. The method according to claim 1, characterized in that The step of extracting semantic emotion features from the acquired real-time user interaction data to generate an original emotion data matrix includes: Performing multi-scale acoustic feature extraction and feature fusion on the speech stream to generate acoustic features; Parsing the text input stream into a six-dimensional emotion probability vector using a RoBERTa emotion classifier, wherein the six-dimensional emotion probability vector includes joy, anger, sadness, fear, surprise, and neutral; The acoustic features and the six-dimensional emotion probability vector are timestamp aligned, and the original emotion data matrix containing the time dimension is generated by feature splicing.

3. The method according to claim 1, characterized in that The dynamically calibrating the original emotion data matrix to obtain the emotion state vector includes: Obtain the current conversation history text and generate a contextual semantic vector through a bidirectional LSTM network, where the contextual semantic vector is used to represent the semantic features of the conversation scene; Calculating the attention weight distribution of the original sentiment data matrix based on the contextual semantic vector; The original emotion data matrix is ​​weighted filtered according to the attention weight distribution to output the emotion state vector.

4. The method according to claim 1, wherein The bimodal intention matching process based on the historical behavior database and the dynamic interest distribution vector is performed to generate a predicted user intention, including: Performing time series pattern mining on the historical behavior database to generate a long-term intention preference vector; Calculating the similarity between the long-term intention preference vector and the dynamic interest distribution vector by a dynamic time warping algorithm to generate an intention drift factor; An adaptive fusion weight is constructed based on the intention drift factor, and a weighted fusion is performed on the long-term intention preference vector and the dynamic interest distribution vector to generate a comprehensive intention feature vector; The comprehensive intention feature vector is input into a pre-trained intention classifier for probability prediction to generate the predicted user intention.

5. An intention prediction device for intelligent human-computer interaction, characterized in that: The device comprises: An interaction feature extraction module is used to extract semantic emotional features from the acquired real-time user interaction data to generate an original emotional data matrix, wherein the real-time interaction data includes a voice stream and a text input stream; A dynamic calibration module, configured to dynamically calibrate the original emotion data matrix to obtain an emotion state vector; An interest focus modeling module is used to obtain user topic interaction logs and perform interest focus modeling in combination with the emotional state vector to generate a dynamic interest distribution vector, which is used to quantify topic attention; A bimodal intent matching module, configured to perform bimodal intent matching processing based on a historical behavior database and the dynamic interest distribution vector to generate a predicted user intent; Wherein, the interest focus modeling module includes: An interest statistics unit, configured to count the response frequency and duration of each topic within a preset time window based on the user topic interaction log, and generate an initial interest weight; A coupling analysis unit, configured to input the emotion state vector into a dilated convolutional neural network and generate an emotion-topic coupling coefficient through time derivative analysis; A dynamic correction unit is used to perform dynamic attenuation correction on the initial interest weight based on the emotion-topic coupling coefficient, and generate the dynamic interest distribution vector through normalization processing.

6. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 4 is implemented.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Establishment method and device of application program prediction model, storage medium and terminal

    CN111258593A

  • Semantic response system based on emotional communication

    CN118132697A