A multimodal-based human-computer interaction method and system thereof
By comprehensively utilizing voice, environment and facial features, dynamically adjusting weights and weighting fusion, the problem that traditional interaction methods cannot capture user emotions and environmental changes is solved, and an efficient and humanized human-computer interaction experience is achieved.
Patent Information
- Application Number
- CN202510293143.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-03-13
Smart Images

Figure CN119806335B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of human-computer interaction, and particularly to a multi-modal based human-computer interaction method and system thereof. Background Art
[0002] With the rapid development of artificial intelligence technology, the human-computer interaction method is gradually changing from a single physical interaction such as a keyboard and a mouse to a more natural and diverse interaction mode. Among them, the multi-modal based human-computer interaction method has become a current research hotspot because it can comprehensively utilize various data sources such as sound, image, and environmental information to provide users with a richer and more accurate interaction experience.
[0003] Traditional human-computer interaction methods mainly rely on text input or physical button operations. This method is not only cumbersome to operate, but also unable to fully capture the user's emotional state and environmental changes, resulting in limited interaction experience.
[0004] Therefore, it is necessary to provide a multi-modal based human-computer interaction method and system to solve the above technical problems. Summary of the Invention
[0005] To solve the above technical problems, the present invention provides a multi-modal based human-computer interaction method and system. By comprehensively utilizing various data sources, it realizes a comprehensive understanding of the user's intention and emotional state, and at the same time can dynamically adjust the weights of each feature to generate a more intelligent and user-friendly interaction experience.
[0006] The present invention provides a multi-modal based human-computer interaction method, and the interaction method includes the following steps:
[0007] Receiving the user's voice input, collecting environmental information, and capturing the user's facial expression, and respectively obtaining voice features, environmental features, and facial features based on the voice input, environmental information, and facial expression;
[0008] Generating the user's emotional features based on the voice features and facial features;
[0009] Generating behavior features by using a pre-trained user behavior model according to the user's current operation information;
[0010] Dynamically adjusting the first weight and the second weight of each of the voice features, environmental features, emotional features, and behavior features based on the representations of the features, and performing feature weighted fusion based on the first weight and the second weight to obtain a comprehensive feature, where the first weight is a feature importance weight for reflecting the relative importance degree of different features in the interaction process, and the second weight is an attention weight for adjusting the influence degree of different features in the fusion process;
[0011] Generate multiple candidate responses based on the comprehensive features, and use a preset scoring mechanism to screen the multiple candidate responses to obtain and output the final response.
[0012] Preferably, receiving the user's voice input, collecting environmental information, and capturing the user's facial expression, and respectively obtaining voice features, environmental features, and facial features based on the voice input, environmental information, and facial expression, includes:
[0013] Receive the user's voice input through a voice receiving component, and extract voice features through a voice signal processing method, where the voice features include the pitch, volume, speech rate, and speech content of the voice;
[0014] Collect environmental information using environmental sensors, and extract environmental features through environmental data analysis methods;
[0015] Capture the user's facial expression through a face recognition camera component, and analyze the facial image using an image processing method to identify and extract facial features, where the facial features include the opening and closing state of the eyes, the shape of the eyebrows, the opening and closing degree of the mouth, and the overall facial expression type.
[0016] Preferably, generating the user's emotion features based on the voice features and facial features includes:
[0017] Input the extracted voice features and facial features into a pre-trained emotion recognition model, where the emotion recognition model is a convolutional neural network using a deep learning architecture for capturing emotion cues in voice and facial expressions;
[0018] Use the emotion recognition model to jointly analyze the input voice features and facial features, and calculate and output the probability distribution of the user's emotion categories through the trained weight and bias parameters, where the user's emotion categories include happy, sad, angry, surprised, fearful, and neutral;
[0019] According to the output probability distribution of the user's emotion categories, select the user's emotion category with the highest probability as the user's current emotion, and generate the user's emotion feature vector, where the emotion feature vector contains the emotion category.
[0020] Preferably, generating behavior features according to the user's current operation information using a pre-trained user behavior model includes:
[0021] Real-time capture and record the user's current operation information through a user operation monitoring component, and organize the user's current operation information into a user operation sequence in chronological order, where the current operation information includes the interaction history of the user with interactive elements on the interface;
[0022] Input the user operation sequence into a pre-trained user behavior model, where the user behavior model is trained based on historical user behavior data and has learned and understood the association between user operations and operation intentions;
[0023] Use the user behavior model to encode and extract features from the user operation sequence to generate a behavior feature vector, where the behavior feature vector contains feature elements for reflecting operation intentions.
[0024] Preferably, the adjustment of the first weight includes:
[0025] Real-time capture and parse the current representation data of the speech feature, environmental feature, emotional feature, and behavior feature, where the current representation data includes the quantization value, change rate of each feature, and the degree of association with the current interaction context;
[0026] Based on the current representation data of each feature, use statistical analysis methods to calculate the importance scores of each feature in the current interaction scenario;
[0027] According to the calculated importance scores, perform weight allocation on the speech feature, environmental feature, emotional feature, and behavior feature to obtain the first weight of each feature, where the importance score is positively correlated with the first weight.
[0028] Preferably, the adjustment of the second weight includes:
[0029] Real-time capture and record the reaction time of the user after receiving the voice command issued by the interaction device, where the reaction time is used to measure the response speed of the user to the current voice command;
[0030] Use the attention allocation model, combined with the reaction time, to calculate the attention allocation degree of the user to each feature in the current interaction scenario;
[0031] According to the attention allocation degree, dynamically adjust the second weight of the speech feature, environmental feature, emotional feature, and behavior feature, where the attention score is positively correlated with the second weight.
[0032] Preferably, the acquisition of the comprehensive feature includes:
[0033] According to the first weight and the second weight of each feature obtained by real-time adjustment, construct a feature weight matrix and perform normalization processing, where the rows of the feature weight matrix represent each feature, and the columns represent the first weight or the second weight, and the feature weight matrix is:
[0034]
[0035] Where, and The first weight and the second weight of the speech feature respectively, and The first weight and the second weight of the environmental feature respectively, and The first weight and the second weight of the emotion feature respectively, and The first weight and the second weight of the behavior feature respectively, is the feature weight matrix;
[0036] Based on the normalized feature weight matrix, each feature is weighted;
[0037] The weighted features are fused by the feature splicing method to generate a comprehensive feature vector.
[0038] Preferably, generating a plurality of candidate responses according to the comprehensive feature, and using a preset scoring mechanism to screen the plurality of candidate responses to obtain and output a final response, including:
[0039] Based on the comprehensive feature vector, a plurality of candidate responses are generated by using a pre-trained response generation model;
[0040] Using a pre-trained sentiment analysis model, calculate the response sentiment feature vector of each candidate response;
[0041] Perform sentiment similarity calculation on the response sentiment feature vector of each candidate response and the user's current sentiment feature vector, where the sentiment similarity is represented by cosine degree;
[0042] According to the calculated sentiment similarity, generate a sentiment matching degree score for each candidate response;
[0043] According to the sentiment matching degree score, select the candidate response with the highest score as the final response and output it.
[0044] The present invention also provides a multimodal human-computer interaction system for executing a multimodal human-computer interaction method, and the interaction system includes:
[0045] A feature acquisition module for receiving the user's voice input, collecting environmental information, and capturing the user's facial expression, and respectively obtaining a voice feature, an environmental feature, and a facial feature based on the voice input, environmental information, and facial expression;
[0046] An emotion feature generation module for generating the user's emotion feature based on the voice feature and the facial feature;
[0047] A behavior feature generation module for generating a behavior feature according to the user's current operation information by using a pre-trained user behavior model;
[0048] The comprehensive feature acquisition module is used to dynamically adjust the respective first weights and second weights based on the representations of the speech feature, environmental feature, emotional feature, and behavior feature, and perform feature weighted fusion based on the first weights and second weights to obtain comprehensive features, where the first weight is the feature importance weight, which is used to reflect the relative importance degree of different features in the interaction process, and the second weight is the attention weight, which is used to adjust the influence degree of different features in the fusion process;
[0049] The response generation module is used to generate multiple candidate responses according to the comprehensive features, and use a preset scoring mechanism to screen the multiple candidate responses to obtain the final response and output it.
[0050] Compared with the related technologies, a multi-modal based human-computer interaction method and system provided by the present invention have the following beneficial effects:
[0051] The present invention receives the user's voice input, collects environmental information, and captures the user's facial expression, and respectively obtains the speech feature, environmental feature, and facial feature based on these inputs; then, uses a pre-trained emotion recognition model to generate the user's emotional feature, and at the same time, according to the user's current operation information, uses a pre-trained user behavior model to generate the behavior feature; after obtaining these features, the present invention dynamically adjusts the respective first weights and second weights, and performs feature weighted fusion to obtain comprehensive features; finally, generates multiple candidate responses based on the comprehensive features, and uses a preset scoring mechanism to screen out the final response and output it. The present invention comprehensively utilizes multiple data sources to achieve a comprehensive understanding of the user's intention and emotional state, and at the same time can dynamically adjust the weights of each feature to generate a more intelligent and user-friendly interaction experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 It is a flowchart of a multi-modal based human-computer interaction method provided by the present invention;
[0053] Figure 2 It is a module diagram of a multi-modal based human-computer interaction system provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0054] The present invention will be further described in detail below with reference to the drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. In addition, it should be noted that only parts related to the present invention rather than all structures are shown in the drawings for the convenience of description. In addition, the embodiments in the present invention and the features in the embodiments can be combined with each other without conflict.
[0055] It should also be noted that, for the sake of convenience of description, only the parts related to the present invention rather than all the content are shown in the drawings. Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the operations (or steps) as sequential processes, many of the operations can be implemented in parallel, concurrently, or simultaneously. In addition, the order of the operations can be rearranged. The process can be terminated when its operations are completed, but there can also be additional steps not included in the drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, and so on.
[0056] Embodiment 1
[0057] The present invention provides a multimodal-based human-computer interaction method. Referring to Figure 1 as shown, the interaction method includes the following steps:
[0058] S1: Receive the user's voice input, collect environmental information, and capture the user's facial expressions, and respectively obtain voice features, environmental features, and facial features based on the voice input, environmental information, and facial expressions.
[0059] Specifically, step S1 includes the following steps:
[0060] S11: Receive the user's voice input through a voice receiving component, and extract voice features through a voice signal processing method, where the voice features include the pitch, volume, speech rate, and speech content of the voice.
[0061] In this embodiment, the voice receiving component can be, but is not limited to, a microphone array for capturing the user's voice input. The voice signal processing method mainly includes the following steps:
[0062] First, preprocess the captured voice signal, including denoising, noise reduction, and normalization processing to remove environmental noise and improve the signal quality.
[0063] Extract various features from the preprocessed voice signal, including pitch, volume, speech rate, and speech content. The pitch is extracted through spectral analysis, the volume is extracted through energy analysis, the speech rate is extracted through duration analysis, and the speech content is converted into text through automatic speech recognition (ASR) technology.
[0064] Combine the extracted various features into a feature vector for subsequent processing and analysis.
[0065] Through these steps, the voice features of users can be accurately captured and analyzed, providing reliable data support for subsequent emotion recognition and interaction. This method not only improves the accuracy of speech recognition but also captures the emotional changes of users, thereby enhancing the intelligence level of human-computer interaction.
[0066] S12: Use environmental sensors to collect environmental information and extract environmental features through environmental data analysis methods.
[0067] In this embodiment, the environmental sensors include temperature sensors, humidity sensors, light sensors, and sound sensors, which are used to collect environmental information around the user. The environmental data analysis method mainly includes the following steps:
[0068] Collect environmental data in real time through various environmental sensors, including temperature, humidity, light intensity, and background noise.
[0069] Preprocess the collected environmental data, including filtering, smoothing, and normalization processing, to remove noise and outliers and improve data quality.
[0070] Extract various features from the preprocessed environmental data, including environmental temperature, humidity, light intensity, and background noise level. These features can reflect the environmental conditions where the user is currently located.
[0071] Combine the extracted environmental features into a feature vector through a feature splicing method for subsequent processing and analysis.
[0072] Through these steps, the environmental conditions where the user is located can be comprehensively understood, providing important background information for subsequent behavior analysis and emotion recognition. This method not only improves the environmental perception ability of the system but also dynamically adjusts the interaction strategy according to environmental changes, enhancing the user experience.
[0073] S13: Capture the facial expressions of the user through the facial recognition camera component, and use image processing methods to analyze the facial image to identify and extract facial features, where the facial features include the opening and closing state of the eyes, the shape of the eyebrows, the opening and closing degree of the mouth, and the overall facial expression type.
[0074] In this embodiment, the facial recognition camera component can be, but is not limited to, a high-resolution camera, which is used to capture the facial image of the user. The image processing method mainly includes the following steps:
[0075] Collect the facial image of the user in real time through the camera to ensure that the image is clear and unobstructed.
[0076] Preprocess the collected facial image, including grayscale conversion, denoising, and normalization processing, to improve image quality and reduce noise interference.
[0077] Use a face detection algorithm (Haar cascade classifier or deep learning model) to detect and locate facial key points, including parts such as eyes, eyebrows, and mouth.
[0078] Extract various facial features from the detected key points, including the opening and closing state of the eyes, the shape of the eyebrows, the opening degree of the mouth, and the overall facial expression type. These features can be extracted through the analysis of geometric features and texture features.
[0079] Also combine the extracted various facial features into a feature vector through the feature splicing method for subsequent processing and analysis.
[0080] S2: Generate the emotional features of the user based on the speech features and facial features.
[0081] Specifically, step S2 includes the following steps:
[0082] S21: Input the extracted speech features and facial features into a pre-trained emotion recognition model, where the emotion recognition model is a convolutional neural network using a deep learning architecture, used to capture emotional cues in speech and facial expressions.
[0083] In this embodiment, the emotion recognition model is a convolutional neural network (CNN) based on deep learning, used to capture emotional cues in speech and facial expressions. The specific working principle and process are as follows:
[0084] The structure of the emotion recognition model includes:
[0085] Input layer: Receive the feature vector passed in from the speech and facial feature extraction module.
[0086] Convolutional layer: Perform convolution operations on the input features through multiple convolutional kernels to extract local features.
[0087] Pooling layer: Reduce the feature dimension through max pooling or average pooling operations and retain the main features.
[0088] Fully connected layer: Map the convolved and pooled feature maps to a fixed-length vector for classification tasks.
[0089] Output layer: Output the probability distribution of the user's emotion categories.
[0090] Input the features into the emotion recognition model, where the speech features include pitch, volume, speech rate, and speech content; the facial features include the opening and closing state of the eyes, the shape of the eyebrows, the opening degree of the mouth, and the overall facial expression type.
[0091] At the same time, splice the feature vectors of the speech features and the facial features into a single feature vector, and input it into the emotion recognition model. The model extracts the emotional cues in the speech and facial features through multiple layers of convolution and pooling operations.
[0092] S22: Use the emotion recognition model to jointly analyze the input speech features and facial features, and calculate and output the probability distribution of the user's emotion categories through the trained weight and bias parameters, where the user's emotion categories include happy, sad, angry, surprised, fearful, and neutral.
[0093] In this embodiment, the emotion recognition model jointly analyzes the input speech features and facial features through the trained weight and bias parameters, and calculates and outputs the probability distribution of the user's emotion categories. The specific working principle and process are as follows:
[0094] First, perform feature processing, specifically: splice the feature vectors of the speech features and the facial features into a single feature vector, and perform normalization processing on the feature vector to ensure that the feature values are within the same range.
[0095] Then, perform model inference, specifically: input the normalized feature vector into the emotion recognition model, extract high-level features through multiple layers of convolution and pooling operations, use the ReLU activation function to increase the non-linearity of the model, map the extracted high-level features to a fixed-length vector for classification tasks, and convert the output of the fully connected layer into a probability distribution through the softmax function, representing the probabilities of the user's emotion categories.
[0096] Finally, output the probability distribution of the user's emotion categories, including the probabilities of happy, sad, angry, surprised, fearful, and neutral.
[0097] S23: According to the output probability distribution of the user's emotion categories, select the user's emotion category with the highest probability as the user's current emotion, and generate the user's emotional feature vector, where the emotional feature vector contains the emotion category.
[0098] In this embodiment, according to the probability distribution of the user's emotion categories output by the emotion recognition model, select the user's emotion category with the highest probability as the user's current emotion, and generate the user's emotional feature vector. The specific working principle and process are as follows:
[0099] First, perform probability distribution analysis, specifically:
[0100] Obtain the probability distribution of the user's emotion categories from the output of the emotion recognition model, including the probability values of happy, sad, angry, surprised, fearful, and neutral, and select the user's emotion category with the highest probability as the user's current emotion.
[0101] Then, generate an emotional feature vector, specifically as follows:
[0102] Use the selected user emotion category as an emotional label to generate an emotional feature vector containing the emotional label.
[0103] S3: According to the user's current operation information, use a pre-trained user behavior model to generate behavior features.
[0104] Specifically, step S3 includes the following steps:
[0105] S31: The user operation monitoring component captures and records the user's current operation information in real time, and organizes the user's current operation information into a user operation sequence in chronological order, where the current operation information includes the interaction history of the user with the interaction elements on the interface.
[0106] In this embodiment, the user operation monitoring component is used to capture and record the user's current operation information in real time. The specific working principle and process are as follows:
[0107] First, the user operation monitoring component can be but is not limited to input devices such as touchscreens, mice, and keyboards, and is used to capture the user's operation behavior. Through these input devices, the user's operation information is captured in real time, including clicks, swipes, and text input.
[0108] Then, record the captured operation information, including the timestamp of the operation, the operation type (click, swipe), and the operation location, and organize the user's operation information into a user operation sequence in chronological order. Exemplarily, the operation sequence of a user on an interface may include: click button A → swipe the screen → input text → click button B.
[0109] Finally, the user operation sequence contains the interaction history of the user with the interaction elements on the interface, and this information can reflect the user's operation habits and intentions.
[0110] S32: Input the user operation sequence into the pre-trained user behavior model, where the user behavior model is trained based on historical user behavior data and has learned and understood the association between user operations and operation intentions.
[0111] In this embodiment, the user behavior model is a deep learning model trained based on historical user behavior data and is used to understand and predict user operation intentions. The specific working principle and process are as follows:
[0112] The structure of the user behavior model includes:
[0113] Input layer: Receive the user operation sequence as input.
[0114] Embedding layer: Convert information such as operation type and timestamp into fixed-length embedding vectors.
[0115] Recurrent Neural Network (RNN): Model the user operation sequence through RNN to capture the temporal dependencies between operations.
[0116] Fully connected layer: Map the output of RNN to a fixed-length vector for classification tasks.
[0117] Output layer: Output the probability distribution or feature vector of the user operation intention.
[0118] When building the user behavior model, collect a large amount of historical user behavior data, including the operation sequence of the user on the interface and its corresponding operation intention, and then clean and preprocess the historical data, including removing outliers, filling in missing values, train the user behavior model using the historical user behavior data, and optimize the model parameters through the backpropagation algorithm so that the model can accurately understand and predict the user operation intention.
[0119] Use the cross-entropy loss function to measure the difference between the model prediction value and the true value.
[0120] When processing the user behavior model, input the user operation sequence into the pre-trained user behavior model, and the user behavior model outputs the probability distribution or feature vector of the user operation intention, and this information can be used for subsequent behavior analysis and sentiment recognition.
[0121] S33: Use the user behavior model to encode and extract features from the user operation sequence to generate a behavior feature vector, where the behavior feature vector contains feature elements for reflecting the operation intention.
[0122] In this embodiment, the user behavior model encodes and extracts features from the user operation sequence to generate a behavior feature vector. The specific working principle and process are as follows:
[0123] Convert each operation type and timestamp in the user operation sequence into a fixed-length embedding vector.
[0124] Perform temporal modeling on the user operation sequence through RNN to capture the temporal dependencies between operations.
[0125] Map the output of RNN to a fixed-length feature vector, and this feature vector contains feature elements for reflecting the operation intention.
[0126] The generated behavior feature vector contains important features of the user operation sequence, and these features can reflect the user's operation intention and habits. Each element in the feature vector represents a specific feature, such as operation frequency, operation type distribution, operation time interval, etc.
[0127] S4: Dynamically adjust the respective first weights and second weights based on the representations of the speech features, environmental features, emotional features, and behavioral features, and perform feature weighted fusion based on the first weights and second weights to obtain comprehensive features, where the first weight is the feature importance weight, used to reflect the relative importance of different features during the interaction process, and the second weight is the attention weight, used to adjust the influence degree of different features during the fusion process.
[0128] In step S4, the adjustment of the first weight includes:
[0129] Capture and parse the current representation data of the speech features, environmental features, emotional features, and behavioral features in real time, where the current representation data includes the quantization values, change rates, and the degree of association with the current interaction context of each feature.
[0130] In this embodiment, capturing and parsing the current feature representation data of the user in real time is the basis for dynamically adjusting the feature weights. The specific working principle and process are as follows:
[0131] Through the voice receiving component, environmental sensor, facial recognition camera component, and user operation monitoring component, capture the user's speech features, environmental features, facial features, and behavioral features in real time, parse the captured data, and extract the quantization values of each feature. For example, the pitch, volume, speech rate, and speech content of the speech features; the temperature, humidity, light intensity, and background noise of the environmental features; the eye opening and closing state, eyebrow shape, mouth opening degree, and overall expression type of the facial features; the operation type, operation location, and timestamp of the behavioral features.
[0132] For the change rate, its calculation method is:
[0133] Set a time window and calculate the change rate of each feature within this time window. The change rate reflects the change trend of the feature over time.
[0134] Calculate the change rate of each feature through the differential calculation method, where the differential formula is:
[0135]
[0136] Where, is the change rate, is the feature value at the current time point is the feature value at the previous time point is the previous time point is the feature value, is the time interval;
[0137] Exemplarily, for the pitch change rate of the speech feature , the time point is , the time interval is 1 second, and its calculation formula is:
[0138]
[0139] Wherein, is the pitch at the current time point , is the pitch at the previous time point ;
[0140] For the temperature change rate of the environmental characteristics , the time point is , and the time interval is 1 second, and its calculation formula is:
[0141]
[0142] Wherein, is the temperature at the current time point , is the temperature at the previous time point ;
[0143] For the eye opening and closing change rate of the facial features , the time point is , and the time interval is 1 second, and its calculation formula is:
[0144]
[0145] Wherein, is the eye opening and closing value at the current time point , is the eye opening and closing value at the previous time point .
[0146] For the degree of association of the context, it can be obtained through the correlation analysis method, which can intuitively reflect the linear relationship between features. Specifically:
[0147] Organize the quantization values and change rates of the captured features into a feature vector, and define the information of the current interaction context, such as the user's current operation intention and emotional state.
[0148] Calculate the Pearson correlation coefficient between each feature and the current interaction context, where the Pearson correlation coefficient The calculation formula is:
[0149]
[0150] Wherein, is the feature value of each feature, is the context information value, and are the means of the eigenvalue and the context information value, respectively.
[0151] Convert the calculated Pearson correlation coefficient into an association degree score Q. The association degree score Q can be calculated using the following formula:
[0152]
[0153] where the range of the association degree score is from 0 to 1, and the larger the value, the higher the degree of association.
[0154] Based on the current representation data of each feature, use statistical analysis methods to calculate the importance scores of each feature in the current interaction context.
[0155] In this embodiment, based on the captured current feature representation data, use statistical analysis methods to calculate the importance scores of each feature. The specific working principle and process are as follows:
[0156] First, organize the quantization values, change rates, and association degrees of the captured features into a feature vector, and perform normalization processing on the feature vector to ensure that each feature value is within the same range and avoid the influence of some feature values being too large or too small on the results.
[0157] Then, extract the main components of each feature through the principal component analysis method. The principal component analysis can transform the high-dimensional feature space into a low-dimensional feature space while retaining most of the information of the original data. Then calculate the contribution degree of each feature in the main component. The contribution degree reflects the importance of the feature in the current interaction context, and the contribution degree can be evaluated by the variance explanation ratio of the main component.
[0158] Evaluate the influence degree of each feature on the current interaction context through a linear regression model. The linear regression model can be used to predict the user's emotional state. At the same time, by training the linear regression model, calculate the absolute value of the coefficient of each feature in the model, where the absolute value of the coefficient reflects the influence of the feature on the target variable.
[0159] Finally, combine the association degree score of the correlation analysis, the contribution degree of the principal component analysis, and the absolute value of the coefficient of the regression analysis, and calculate the importance scores of each feature by weighted calculation. At the same time, assign a weight to each analysis result, and these weights can be adjusted according to actual needs.
[0160] The importance scores reflect the relative importance degrees of each feature in the current interaction context. According to the importance scores, perform preliminary weight assignment to each feature. Features with high importance scores are assigned higher weights, and vice versa.
[0161] According to the calculated importance scores, weight distribution is performed on the voice features, environmental features, emotional features, and behavioral features to obtain the first weights of each feature, where the importance scores are positively correlated with the first weights.
[0162] In this embodiment, according to the calculated importance scores, weight distribution is performed on each feature to generate the first weights. The specific working principle and process are as follows:
[0163] Sort the calculated importance scores in descending order, and perform normalization processing on the importance scores to ensure that the sum of all scores is 1.
[0164] Linearly map the normalized importance scores to the weight interval (such as 0 to 1). Features with high importance scores are assigned higher first weights, and vice versa.
[0165] At the same time, according to the real-time feature representation data of the user, the first weights of the features are dynamically adjusted. When the feature representation data of the user changes, recalculate the importance scores and adjust the first weights.
[0166] In addition, in order to avoid drastic fluctuations in the first weights, the first weights can be smoothed, for example, using the Exponentially Weighted Moving Average (EWMA) method.
[0167] In step S4, the adjustment of the second weights includes:
[0168] Real-time capture and record the reaction time of the user after receiving the voice command issued by the interaction device, where the reaction time is used to measure the response speed of the user to the current voice command.
[0169] In this embodiment, real-time capture and record the reaction time of the user after receiving the voice command issued by the interaction device is the basis for dynamically adjusting the second weights. The specific working principle and process are as follows:
[0170] Send a voice command to the user through the interaction device, and record the time interval from when the user receives the voice command to when the user makes a response. The response can be a voice reply, a touch screen operation, or any other form of interaction behavior.
[0171] Record the timestamp when the voice command is issued and the timestamp when the user responds, and calculate the user's reaction time, that is, the time interval from when the voice command is issued to when the user responds. The reaction time reflects the response speed of the user to the current voice command.
[0172] Using the attention allocation model, combined with the reaction time, calculate the degree of attention allocation of the user to each feature in the current interaction context.
[0173] In this embodiment, an attention allocation model is used to calculate the attention allocation degree of the user to each feature in the current interaction scenario in combination with the user's response time. The specific working principle and process are as follows:
[0174] The attention allocation model is a deep learning-based model that can dynamically adjust the degree of attention to different features according to the user's response time.
[0175] The input layer of the attention allocation model receives the user's response time data and other feature data (such as voice features, environmental features, emotional features, and behavior features).
[0176] The attention layer of the attention allocation model calculates the attention weights of each feature through an attention mechanism. The attention weights reflect the degree of attention of the user to each feature in the current interaction scenario.
[0177] The user's response time and other feature data are organized into a feature vector, and the feature vector is normalized to ensure that each feature value is within the same range, avoiding the influence of some feature values being too large or too small on the results.
[0178] The attention weights of each feature are calculated through an attention mechanism. The attention weights reflect the degree of attention of the user to each feature.
[0179] According to the attention allocation degree, the second weights of the voice feature, environmental feature, emotional feature, and behavior feature are dynamically adjusted, where the attention score is positively correlated with the second weight.
[0180] In this embodiment, according to the calculated attention allocation degree, the second weights of each feature are dynamically adjusted. The specific working principle and process are as follows:
[0181] The calculated attention scores are sorted in descending order, and the attention scores are normalized to ensure that the sum of all scores is 1.
[0182] The normalized attention scores are linearly mapped to the weight interval (such as 0 to 1). Features with high attention scores are given higher second weights, and vice versa.
[0183] According to the user's real-time response time data, the second weights of the features are dynamically adjusted. When the user's response time data changes, the attention scores are recalculated and the second weights are adjusted.
[0184] In addition, in order to avoid drastic fluctuations in the second weights, the second weights can be smoothed, for example, using the exponential weighted moving average (EWMA) method.
[0185] In step S4, the acquisition of the comprehensive features includes:
[0186] Construct a feature weight matrix based on the first weight and the second weight of each feature obtained by real-time adjustment, and perform normalization processing, where the rows of the feature weight matrix represent each feature, and the columns represent the first weight or the second weight, and the feature weight matrix is:
[0187]
[0188] Wherein, and are the first weight and the second weight of the speech feature respectively, and are the first weight and the second weight of the environmental feature respectively, and are the first weight and the second weight of the emotional feature respectively, and are the first weight and the second weight of the behavior feature respectively, is the feature weight matrix.
[0189] Based on the normalized feature weight matrix, perform weighted processing on each feature.
[0190] Fuse the weighted features through a feature concatenation method to generate a comprehensive feature vector, where the concatenation method is simple vector concatenation.
[0191] S5: Generate multiple candidate responses according to the comprehensive feature, and use a preset scoring mechanism to screen the multiple candidate responses to obtain a final response and output it.
[0192] Specifically, step S5 includes the following steps:
[0193] S51: Based on the comprehensive feature vector, use a pre-trained response generation model to generate multiple candidate responses.
[0194] In this embodiment, the response generation model usually includes an input layer, an embedding layer, an encoding layer, a decoding layer, and an output layer. The input layer receives the comprehensive feature vector, the embedding layer converts the feature into an embedding vector, the encoding layer extracts high-level features through an encoder, the decoding layer generates multiple candidate responses, and the output layer outputs the generated response.
[0195] Collect a large amount of historical dialogue data, including the comprehensive feature vector of the user and the corresponding system response, and perform cleaning and preprocessing. Use the historical dialogue data to train the response generation model, and optimize the model parameters through the backpropagation algorithm to enable the model to generate high-quality responses. The loss function usually uses the cross-entropy loss function. Input the comprehensive feature vector into the pre-trained response generation model, and the model outputs multiple candidate responses, which can be in text form or other forms of feedback.
[0196] S52: Calculate the response sentiment feature vector for each candidate response using a pre-trained sentiment analysis model.
[0197] In this embodiment, the sentiment analysis model is a deep learning-based model for identifying sentiment features in text.
[0198] Collect historical text data with sentiment labels and perform cleaning and preprocessing.
[0199] Train the sentiment analysis model using the historical text data with sentiment labels, and optimize the model parameters through the backpropagation algorithm to enable the model to accurately identify sentiment features in text. The loss function usually uses the cross-entropy loss function.
[0200] Input each generated candidate response text into the pre-trained sentiment analysis model, and the model outputs the sentiment feature vector of each candidate response, including the probability distribution of sentiment categories.
[0201] S53: Calculate the sentiment similarity between the response sentiment feature vector of each candidate response and the user's current sentiment feature vector, where the sentiment similarity is represented by cosine similarity.
[0202] In this embodiment, the user sentiment feature vector is generated from step S23, and the candidate response sentiment feature vector is generated from step S52. Calculate the similarity between the two sentiment feature vectors using cosine similarity, and calculate the sentiment similarity score for each candidate response. The score range is from 0 to 1, and the larger the value, the higher the sentiment similarity.
[0203] S54: Generate a sentiment matching score for each candidate response based on the calculated sentiment similarity.
[0204] In this embodiment, normalize the sentiment similarity scores of each candidate response calculated in step S53 to ensure that the sum of all scores is 1. Use the normalized sentiment similarity score as the sentiment matching score, which reflects the matching degree between the candidate response and the user's current sentiment state.
[0205] S55: Select the candidate response with the highest score as the final response and output it according to the sentiment matching score.
[0206] In this embodiment, sort the sentiment matching scores of all candidate responses in descending order, select the candidate response with the highest sentiment matching score as the final response, generate the selected final response as the system output, and feedback the final response to the user in the form of voice, text, or other forms.
[0207] Embodiment II
[0208] The present invention also provides a multimodal-based human-computer interaction system for performing a multimodal-based human-computer interaction method, refer to Figure 2 As shown, the interaction system includes:
[0209] A feature acquisition module 100, configured to receive a user's voice input, collect environmental information, and capture the user's facial expressions, and respectively obtain voice features, environmental features, and facial features based on the voice input, environmental information, and facial expressions.
[0210] An emotion feature generation module 200, configured to generate the user's emotion features based on the voice features and facial features.
[0211] A behavior feature generation module 300, configured to generate behavior features according to the user's current operation information by using a pre-trained user behavior model.
[0212] A comprehensive feature acquisition module 400, configured to dynamically adjust the respective first weights and second weights based on the representations of the voice features, environmental features, emotion features, and behavior features, and perform feature weighted fusion based on the first weights and second weights to obtain comprehensive features, where the first weight is a feature importance weight for reflecting the relative importance of different features in the interaction process, and the second weight is an attention weight for adjusting the influence degree of different features in the fusion process.
[0213] A response generation module 500, configured to generate multiple candidate responses according to the comprehensive features, and use a preset scoring mechanism to screen the multiple candidate responses to obtain a final response and output it.
[0214] This application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of this application. It should be understood that each process and / or block in the flowcharts and / or block diagrams can be implemented by computer program instructions, and the combination of processes and / or blocks in the flowcharts and / or block diagrams can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in one Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0215] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium, which includes read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc memories, magnetic disc memories, tape memories, or any other medium that can be used to carry or store data and is computer-readable.
[0216] It should also be noted that the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such a process, method, commodity or device. Without more limitations, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, commodity or device including the element.
Claims
1. A multimodal human-computer interaction method, characterized in that: The interactive method comprises the following steps: Receiving a user's voice input, collecting environmental information, and capturing the user's facial expression, and obtaining voice features, environmental features, and facial features based on the voice input, environmental information, and facial expression, respectively; generating an emotion feature of the user based on the voice feature and the facial feature; Based on the user's current operation information, the pre-trained user behavior model is used to generate behavior features; Dynamically adjust the first weight and the second weight of each based on the representation of the speech feature, the environmental feature, the emotional feature and the behavioral feature, and perform weighted fusion of the features based on the first weight and the second weight to obtain a comprehensive feature, wherein the first weight is a feature importance weight, which is used to reflect the relative importance of different features in the interaction process, and the second weight is an attention weight, which is used to adjust the influence of different features in the fusion process; The adjustment of the first weight specifically includes: Based on the current representation data of each feature captured and parsed in real time, a statistical method is used to calculate the importance score of each feature in the current interaction context, and a weight is assigned to each feature according to the importance score to obtain a first weight of each feature that is positively correlated with the importance score, wherein the current representation data includes a quantitative value of each feature, a change rate, and a degree of association with the current interaction context; The adjustment of the second weight specifically includes: Based on real-time capture and recording of the user's reaction time after receiving the voice command issued by the interactive device, the attention allocation model is used to calculate the user's attention allocation degree to each feature in the current interactive situation, and the second weight of each feature that is positively correlated with it is dynamically adjusted according to the attention allocation degree; A plurality of candidate responses are generated according to the comprehensive features, and the plurality of candidate responses are screened using a preset scoring mechanism to obtain a final response and output it.
2. The multimodal human-computer interaction method according to claim 1, characterized in that: The receiving of the user's voice input, collecting environmental information, and capturing the user's facial expression, and obtaining voice features, environmental features, and facial features based on the voice input, environmental information, and facial expression, respectively, includes: Receive the user's voice input through the voice receiving component, and extract voice features through the voice signal processing method, wherein the voice features include the voice pitch, volume, speaking speed and voice content; Collect environmental information using environmental sensors and extract environmental features through environmental data analysis methods; The user's facial expression is captured by a facial recognition camera component, and the facial image is analyzed using an image processing method to identify and extract facial features, including the open and closed state of the eyes, the shape of the eyebrows, the degree of opening and closing of the mouth, and the overall facial expression type.
3. The multimodal human-computer interaction method according to claim 2, characterized in that: The generating the user's emotion feature based on the voice feature and the facial feature comprises: Inputting the extracted speech features and facial features into a pre-trained emotion recognition model, wherein the emotion recognition model is a convolutional neural network using a deep learning architecture to capture emotional clues in speech and facial expressions; The emotion recognition model is used to jointly analyze the input voice features and facial features, and the probability distribution of the user's emotion categories is calculated and output through the trained weight and bias parameters, wherein the user's emotion categories include happiness, sadness, anger, surprise, fear and neutrality; According to the output probability distribution of the user emotion categories, the user emotion category with the highest probability is selected as the user's current emotion, and an emotion feature vector of the user is generated, wherein the emotion feature vector contains the emotion category.
4. The multimodal human-computer interaction method according to claim 3, characterized in that: The generating behavior features by using a pre-trained user behavior model according to the user's current operation information includes: The user operation monitoring component is used to capture and record the user's current operation information in real time, and organize the user's current operation information into a user operation sequence in chronological order, wherein the current operation information includes the user's interaction history with the interactive elements on the interface; Inputting the user operation sequence into a pre-trained user behavior model, wherein the user behavior model is trained based on historical user behavior data and has learned and understood the relationship between user operations and operation intentions; The user operation sequence is encoded and feature extracted using the user behavior model to generate a behavior feature vector, wherein the behavior feature vector contains feature elements for reflecting operation intentions.
5. The multimodal human-computer interaction method according to claim 4, characterized in that: The acquisition of the comprehensive features includes: According to the first weight and the second weight of each feature obtained by real-time adjustment, a feature weight matrix is constructed and normalized, wherein the rows of the feature weight matrix represent each feature, and the columns represent the first weight or the second weight, wherein the feature weight matrix is: in, and are the first weight and the second weight of the speech feature, respectively. and are the first weight and the second weight of the environmental features, respectively. and are the first weight and the second weight of the sentiment feature, respectively. and are the first and second weights of the behavior characteristics, respectively. is the feature weight matrix; Based on the normalized feature weight matrix, each feature is weighted; The weighted features are fused through feature concatenation method to generate a comprehensive feature vector.
6. The multimodal human-computer interaction method according to claim 5, characterized in that: The generating of multiple candidate responses according to the comprehensive features, and screening the multiple candidate responses using a preset scoring mechanism to obtain and output a final response includes: Based on the comprehensive feature vector, generate a plurality of candidate responses using a pre-trained response generation model; Using the pre-trained sentiment analysis model, calculate the response sentiment feature vector of each candidate response; Calculate the emotional similarity between the response emotional feature vector of each candidate response and the user's current emotional feature vector, wherein the emotional similarity is expressed in cosine degree; Generate a sentiment matching score for each candidate response based on the calculated sentiment similarity; According to the sentiment matching score, the candidate response with the highest score is selected as the final response and output.
7. A multimodal human-computer interaction system, used to execute a multimodal human-computer interaction method according to any one of claims 1 to 6, characterized in that: The interactive system comprises: A feature acquisition module, used to receive a user's voice input, collect environmental information, and capture the user's facial expression, and acquire voice features, environmental features, and facial features based on the voice input, environmental information, and facial expression, respectively; An emotion feature generation module, used to generate the user's emotion features based on the voice features and facial features; A behavior feature generation module is used to generate behavior features based on the user's current operation information using a pre-trained user behavior model; A comprehensive feature acquisition module, used to dynamically adjust the first weight and the second weight of each of the voice feature, the environmental feature, the emotional feature and the behavioral feature based on the representation, and perform feature weighted fusion based on the first weight and the second weight to obtain a comprehensive feature, wherein the first weight is a feature importance weight, which is used to reflect the relative importance of different features in the interaction process, and the second weight is an attention weight, which is used to adjust the influence of different features in the fusion process; The response generation module is used to generate multiple candidate responses according to the comprehensive features, and screen the multiple candidate responses using a preset scoring mechanism to obtain a final response and output it.
Citation Information
Patent Citations
Visual interaction system based on multiple modes
CN118535023A
Emotion analysis method and system based on multi-modal fusion
CN119272224A
Cited By
Human-machine interaction oriented high-fidelity facial feature generation fusion method
CN122714922A