Intelligent robot interaction method

By constructing a modal quality assessment model and contextual semantic consistency detection, combined with dynamic weighting algorithms and user binding mechanisms, and analyzing speech features and micro-expressions, the problem of intelligent robots being unable to differentiate interactions and accurately identify intentions is solved, achieving efficient emotion recognition and improved interactive experience.

CN121838745APending Publication Date: 2026-04-10DIGITAL LITTLE NENGREN (JIANGSU) TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-28
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing intelligent robots cannot differentiate their interactions based on contextual information and user personality traits, and they cannot effectively integrate multiple input signals in multimodal scenarios to accurately identify the operational intentions of specific users, resulting in a lack of interactive experience and recognition accuracy.

Method used

A modal quality assessment model is constructed, and a contextual semantic consistency detection model is introduced to generate modal confidence vectors and semantic consistency results. A dynamic weighting algorithm is used to adjust the confidence, and a modality-user binding trusted channel mechanism is introduced. By combining Mel frequency cepstral coefficients and ActionUnits analysis, speech features and micro-expression details are extracted, user profiles are constructed, and emotion data is analyzed.

Benefits of technology

It improves the sensitivity and accuracy of emotion recognition, enhances the human-computer interaction experience, adapts to different user characteristics and contextual information, can accurately identify operation intentions in multimodal scenarios, solves the problem of unclear modality attribution, and improves the accuracy and trust of user recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121838745A_ABST
    Figure CN121838745A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent robot interaction, in particular to an intelligent robot interaction method which comprises the following steps: S1, constructing a modal quality evaluation model, introducing a context semantic consistency detection model, and generating a modal confidence vector and a semantic consistency result; s2, performing confidence adjustment on the modal confidence vector and the semantic consistency result by adopting a dynamic weighting algorithm, and outputting an intention judgment result; s3, introducing a mode-user binding credible channel mechanism, and capturing user data in real time, when the method is used, the MFCC is combined with the time sequence model to extract potential emotional fluctuations in voice, so that the sensitivity and accuracy of emotion recognition are improved, the problem of unclear mode attribution in a multi-user and multi-mode concurrent scene is solved, and the emotion recognition efficiency is improved. According to the method, the accuracy and credibility of user identification are enhanced, the method can adapt to individual characteristics and contextual information of different users, differentiated interaction is carried out, and the man-machine interaction experience feeling is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent robot interaction, and particularly relates to an intelligent robot interaction method. BACKGROUND

[0002] Intelligent robot interaction refers to the process of communication and cooperation between a robot equipped with artificial intelligence technology and a human or another machine; such interaction can be actual operation in the physical world or information exchange in a virtual environment.

[0003] Patent CN107203607A discloses an online creative interaction platform based on an intelligent robot, which comprises an intelligent robot and an online interaction software system. The intelligent robot is used for video acquisition and online communication and exchange. The online interaction software system comprises a creative submission module, a creative classification module, a creative automatic screening module and a video interaction module. The present application sets up an online interaction platform based on an intelligent robot. Users can submit creative works through the platform, which are then evaluated by service providers to realize creative exchange between users and children and develop children's intelligence. Although the above-mentioned technology integrates creative submission, classification, screening and interaction functions through an intelligent robot, and builds an efficient and interactive children's creative ecological system, which not only improves the use value of hardware but also promotes creativity cultivation and resource sharing, it still has the following problems: (1) the intelligent robot in the above-mentioned technology cannot interact differently according to context information and user individual characteristics, lacking interactive experience; (2) in addition, the intelligent robot in the above-mentioned technology cannot effectively fuse multiple input signals to accurately identify the operation intention of a specific user in a multi-modal scene.

[0004] In view of the above, it is still a key problem in the technical field of intelligent robot interaction to develop an intelligent robot interaction method. SUMMARY

[0005] The present application aims to solve the above-mentioned problems in the prior art that although the above-mentioned technology integrates creative submission, classification, screening and interaction functions through an intelligent robot, and builds an efficient and interactive children's creative ecological system, which not only improves the use value of hardware but also promotes creativity cultivation and resource sharing, it still has the following problems: (1) the intelligent robot in the above-mentioned technology cannot interact differently according to context information and user individual characteristics, lacking interactive experience; (2) in addition, the intelligent robot in the above-mentioned technology cannot effectively fuse multiple input signals to accurately identify the operation intention of a specific user in a multi-modal scene.

[0006] To achieve the above-mentioned purpose, the present application provides the following technical solutions. This invention provides an intelligent robot interaction method, comprising the following steps: S1, constructing a modal quality assessment model, introducing a contextual semantic consistency detection model, and generating a modal confidence vector and semantic consistency results; S2. The modality confidence vector and the semantic consistency result are adjusted using a dynamic weighting algorithm, and the intent judgment result is output. S3. Introduce a trusted channel mechanism of "modal-user binding" to build user profiles by capturing user data in real time and combining the results of intent judgment. S4. Extract speech features from the user profile using Mel frequency cepstral coefficients, and analyze user tone fluctuations using emotion vector models to identify emotion data. S5. Using the ActionUnits analysis method, detect micro-expression details based on the user data and emotion data to determine the user's emotional trend.

[0007] Furthermore, in step S1, the modal quality assessment model is constructed, and a contextual semantic consistency detection model is introduced. The method for generating the modal confidence vector and semantic consistency results is as follows: The modal quality assessment model receives the following multimodal inputs: ,in It is a collection of multimodal input signals, including but not limited to speech, image, text, signal strength, background noise, image sharpness, and user activity. Indicates the first The original input signal is processed by a modal feature encoder. High-dimensional embedding is performed on each mode to obtain the modality embedding, where It is the first Modality Embedded representation, It is the first Encoder function for each mode, It is the first The original input signal of each mode, This represents the embedding representation vector of the modality. The modality's feature dimension size is used to construct the modality quality evaluation function. Used to evaluate modal representation Information integrity, clarity, and noise level are expressed by a nonlinear quality mapping method based on residual attention and context-aware graph structure, using the following formula: ,in It is the first Quality assessment values ​​for each modality It is used to limit the result to between 0 and 1. It is the number of local residual mappings in the figure. It is the first a weight matrix of a local residual mapping, is used for processing a nonlinear transformation, is an embedding representation of the th modality, is a context transformation matrix based on a graph attention mechanism, is used for measuring the distance of a vector, is a balance coefficient, is used for measuring the probability distribution of the th modality between the difference between the uniform distribution , the modality confidence vector is obtained as: , wherein the modality confidence vector is represented by , the quality evaluation value of the th modality, represents the total number of modalities, represents that this vector belongs to a real number space, and the context semantic consistency detection model obtains the modality embedding, introduces a cross-modal semantic alignment module to determine whether the current modality expression is consistent with the global context semantics, and the cross-modal semantic alignment module is constructed based on a multi-modal graph convolution network and an attention alignment mechanism, and a semantic consistency vector is obtained as: wherein represents a semantic consistency result, represents the semantic consistency score of the th modality, represents the total number of modalities.

[0008] Further, in step S2, a dynamic weighting algorithm is used to adjust the confidence of the modality confidence vector and the semantic consistency result, and the method for obtaining the intention judgment result is: the modality confidence vector and the semantic consistency result are nonlinearly fused to generate normalized modality weights , and the expression formula is: wherein is generated by a softmax operation, the normalized modality weights, is the dynamic score of the th modality, is the th modality confidence vector, is the th semantic consistency result, is a very small positive number, represents the similarity between two vectors, is the hyperbolic tangent function, is the distance metric between modalities, is an adjustable hyperparameter.

[0009] Further, in step S2, a dynamic weighting algorithm is used to adjust the confidence of the modality confidence vector and the semantic consistency result, and the method for obtaining the intent judgment result is: using the modality weight weighting the modality embedding to obtain a unified fusion , the expression formula is: wherein is the representation after weighting and fusion of all modalities, is the total number of modalities, is the weight of the th modality, is the embedding representation of the th modality, represents the cosine similarity between the th and the th modality, represents the feature difference between the th and the th modality, is a balance factor, input the fused to the intent recognition function to obtain the intent judgment result , the expression formula is: wherein is the intent judgment result after softmax normalization, is the weight parameter of the output layer for linear transformation of the activation output of the previous layer, is the weight parameter of the hidden layer for is the representation after weighting and fusion of all modalities, is the bias term of the first layer for adjusting the translation amount of the neuron output, is the bias term of the second layer for bias adjustment of the output layer, is used to introduce a nonlinear representation that preserves positive values and suppresses negative values.

[0010] Further, in step S3, a trusted channel mechanism of "modality-user binding" is introduced, and the method for constructing a user portrait by capturing user data in real time and combining the intent judgment result is: The real-time capture of user data includes but is not limited to face recognition, voiceprint, operation habit, speaking speed, emotional stability, semantic tendency, and commonly used vocabulary, and there are users, modality channels, a modality-user binding matrix is constructed, and the expression formula is: wherein It is a modality-user binding matrix. This refers to the number of currently active users. It is the number of modes acquired simultaneously. Indicates the first Is the current modality related to the first modality? Individual user binding, for users The set of modes bound ,in It is a modality set; extract the input features of all modalities in the modality set. ,in It is the first The feature vector of each sample yes -D real space, construct weighted joint representation Formula: ,in It is a weighted joint input representation. It is a modal set. It is the first Feature representation of each modality It is the first Each modality for users Feature fusion contribution weights Hyperparameters are used to control the degree of influence of interaction terms between modes. Representing modes and modality The semantics between Representing modes and modality The differences in characteristics between them.

[0011] Furthermore, in step S3, a trusted channel mechanism of "modal-user binding" is introduced. The method for constructing a user profile by capturing user data in real time and combining it with the intent judgment results is as follows: Based on the intent judgment result Weighted joint input representation of the currently described user data Perform joint encoding to express the formula: ,in At the current moment For users The external emotion input feature vector, User Weighted joint input means, This indicates the outer product operation. To capture the moment Contextual drift introduced by changes in historical context It is a moment Based on the intent determination results, the user profile is constructed as follows: ,in User In time User profile at any time This refers to the space in which the vector resides, and the user profile update employs a nonlinear gating mechanism to fuse long-term interests and short-term intentions, expressed as: ,in User In time User profile at any time User In the previous moment The emotional state indicates, At the current moment For users The external emotion input feature vector, These are the linear transformation parameters for sentiment updates. and The accompanying bias term is used to adjust the translation offset of the linear transformation result. The sigmoid function is used to generate gated weights. This indicates that element-wise multiplication between corresponding elements of two vectors is used to achieve weighted fusion.

[0012] Furthermore, in step S4, the method for extracting speech features from the user profile using Mel frequency cepstral coefficients and analyzing user tone fluctuations using an emotion vector model to identify emotion data is as follows: The speech features are divided into frames of fixed length. Each frame is [length] Frame shift is , to obtain the frame sequence, for the first The frame-by-frame speech signal undergoes a Fourier transform, and then the logarithm of the Mel frequency cepstral coefficients (MFCCs) is taken followed by a discrete cosine transform to output an MFCC vector. The emotion vector model calculates the dynamic changes between consecutive frames to capture intonation fluctuations, thereby constructing a combined speech feature vector based on the user profile. It employs a cross-modal attention mechanism to fuse voice emotion and user personality bias, expressed as: ,in It is the first The emotion-perceived embedding vector of a frame. It is Hadamard's positional multiplication. It is the first Each input feature vector The bias term is used to adjust the output of the activation function. This indicates that the final output vector lies in a... - the real number space, is a nonlinear activation function, is a projection matrix, is a multilayer perceptron.

[0013] Further, in step S4, speech features are extracted from the user portrait using Mel frequency cepstral coefficients, and the user's intonation fluctuations are analyzed in combination with an emotion vector model to identify emotional data, and the method is as follows: Embedding vectors of emotional perception of all frames are input into a Bi-GRU+Attention neural network model combining bidirectional gated recurrent unit (Bi-GRU) and attention mechanism (Attention), to extract overall emotional representation of the user, and the expression formula is: wherein represents the hidden state calculated by the GRU (gated recurrent unit) network at the time, represents the hidden state of the network at time , is the emotional perception embedding vector of the frame, is the attention weight, is the weight vector used to calculate the correlation between the hidden state and the current hidden state, represents the hidden state calculated by the GRU (gated recurrent unit) network at the time, is the final emotional vector representation at time , and finally, the emotional vector is mapped to a dimensional value through a fully connected network, and the expression formula is: wherein is the emotional data output, is the final emotional vector representation at time , is the weight matrix used to convert the input to a dimension, is the bias term used to adjust the output.

[0014] Further, in step S5, the method for detecting micro-expression details and determining the emotional trend of the user according to the user data and emotional data using the ActionUnits analysis method is as follows: Set the sequence of continuous facial image frames of the user as: wherein is a set of continuous facial image data of the user, is the image at the time in the sequence, is the dimension space of images, representing the size of each image is height width and has 3 RGB channels, represents the total number of images in the image sequence, after face detection and 68 key point labeling and positioning of each frame of facial image of the user, the standard facial region is aligned through affine transformation, the AU vector of each frame of facial image of the user is extracted using the trained ActionUnits analysis method, and the expression formula is: wherein is the attention weight vector representing the attention weight at time . represents the degree of attention to the input part at time . is the dimension of the attention weight, representing the number of weight vector input parts, is the value range of the weight, represents is a dimensional real vector, user emotion data and user portrait are introduced to dynamically regulate and fuse the emotion priori of the AU vector.

[0015] Further, in step S5, the method for detecting micro-expression details and judging the emotional trend of the user according to the user data and emotion data by using the ActionUnits analysis method is: by constructing a time series activation state matrix, and then using a bidirectional LSTM network with residual connection to encode the AU time series, to capture the micro-expression change direction and rate , then calculating the first-order derivative vector of micro-expression change , constructing a micro-expression evolution tensor, and the expression formula is: wherein represents the feature vector at time . represents the hidden layer state at time . represents the difference between the hidden state at the current time and the previous time . represents the difference between the current time and the previous two times and . is the dimension of the hidden layer, is the dimension of the final feature vector, based on the micro-expression evolution tensor, the micro-expression details are detected, and the emotional trend of the user is predicted and judged, and the expression formula is: wherein is the emotion trend obtained after calculation, is the feature vector at time , is the dimension of the output vector, is the bias term, is the weight matrix used to convert the input vector to dimension.

[0016] Advantages Compared with the known prior art, the technical scheme provided by the present application has the following advantages: In use, the present application extracts potential emotional fluctuations in speech by combining MFCC with a time sequence model, which is beneficial to improving the sensitivity and accuracy of emotion recognition, solving the problem of unclear modality attribution in multi-user and multi-modal concurrent scenarios, enhancing the accuracy and trustworthiness of user recognition, adapting to different user individual characteristics and context information, and improving the human-computer interaction experience.

[0017] In use, the present application adjusts AU analysis by combining user portraits and emotion data, effectively addressing the real problem that the same expression has different meanings on different people, constructs a high-dimensional user portrait under dynamic context changes by fusing modal input, external emotion, and intention results, realizes the fusion memory of long-term interest and short-term intention by combining the current emotion and historical state through a nonlinear gating mechanism, and adapts to multi-modal scenarios to accurately identify the operation intention of specific different users by fusing multi-modal input signals. BRIEF DESCRIPTION OF DRAWINGS

[0018] Figure 1 is a flowchart of an intelligent robot interaction method of the present application. DETAILED DESCRIPTION

[0019] In order to better enable those skilled in the art to understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should fall within the scope of protection of the present application.

[0020] ​It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but includes other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0021] The present invention will now be described in further detail with reference to the accompanying drawings: Example: like Figure 1 As shown, the present invention provides an intelligent robot interaction method, including the following steps: S1, constructing a modal quality assessment model, introducing a contextual semantic consistency detection model, and generating modal confidence vectors and semantic consistency results; Furthermore, in step S1, the modality quality assessment model is constructed, and the contextual semantic consistency detection model is introduced. The method for generating the modality confidence vector and semantic consistency results is as follows: The modal quality assessment model receives the following multimodal inputs: ,in It is a collection of multimodal input signals, including but not limited to speech, image, text, signal strength, background noise, image sharpness, and user activity. Indicates the first The original input signal is processed by a modal feature encoder. High-dimensional embedding is performed on each mode to obtain the modality embedding, where It is the first Modality Embedded representation, It is the first Encoder function for each mode, It is the first The original input signal of each mode, This represents the embedding vector of the modality. The modality's feature dimension size is used to construct the modality quality evaluation function. Used to evaluate modal representation Information integrity, clarity, and noise level are expressed by a nonlinear quality mapping method based on residual attention and context-aware graph structure, using the following formula: ,in It is the first a quality evaluation value of the modal, is used to limit the result between 0 and 1, is the number of local residual mapping in the graph, is the weight matrix of the first local residual mapping, is used to process nonlinear transformation, is the embedding representation of the first modal, is the context transformation matrix based on the graph attention mechanism, is used to measure the distance of the vector, is a balance coefficient, is used to measure the difference between the probability distribution of the first modal and the uniform distribution, the modal confidence vector is obtained as: , wherein the modal confidence vector is represented by a quality evaluation value of the first modal, denotes the total number of modes, denotes that this vector belongs to the real number space, and the context semantic consistency detection model obtains the modal embedding, introduces a cross-modal semantic alignment module to determine whether the current modal expression is consistent with the global context semantics, and the cross-modal semantic alignment module is constructed based on a multi-modal graph convolutional network and an attention alignment mechanism, and a semantic consistency vector is obtained as: , denotes the semantic consistency result, denotes the semantic consistency score of the first modal, denotes the total number of modes. In the embodiment, the integrity and clarity of the modal information are automatically evaluated through the residual attention graph structure, which facilitates to improve the recognition ability of low-quality or noise interference input, and the semantic alignment mechanism is used to distinguish whether the modal expression deviates from the context, effectively eliminates the semantic irrelevant or conflicting modal, and is beneficial to improve the understanding ability of the method. At the same time, it has good expansibility and robustness, can adapt to different user states and environmental changes, and improves the intelligent level of human-computer interaction.

[0022] S2, a dynamic weighting algorithm is used to adjust the confidence of the modal confidence vector and the semantic consistency result, and an intention judgment result is output; Further, in step S2, a dynamic weighting algorithm is used to adjust the confidence of the modal confidence vector and the semantic consistency result, and an intention judgment result is output. the modal confidence vector ​and the semantic consistency results Perform nonlinear fusion to generate normalized mode weights. Formula: ,in The softmax operation is used to... Generate normalized modal weights. It is the first Dynamic scores for each modality It is the first A modal confidence vector It is the first A semantic consistency result, It is a very small positive number. This represents the similarity between two vectors. It is the hyperbolic tangent function. It is a distance metric between modes. It is an adjustable hyperparameter.

[0023] Further, in step S2, the method for adjusting the confidence of the modality confidence vector and the semantic consistency result using a dynamic weighting algorithm to obtain the intent judgment result is as follows: Using the modal weights The modal embeddings are weighted to obtain a unified fusion. Formula: ,in It is the representation after weighted fusion of all modalities. It is the total number of modes. It is the first Modal weights, It is the first Embedded representation of each modality Indicates the first The and the first Cosine similarity between modalities Indicates the first The and the first Feature differences between modes It is a balancing factor that will balance the fused components. The input is fed into the intent recognition function to obtain the intent determination result. Formula: ,in The intention judgment result is normalized by softmax. The weight parameters of the output layer are used to linearly transform the activation output of the previous layer. The weight parameters of the hidden layer are used for, It is the representation after weighted fusion of all modalities. The bias term in the first layer is used to adjust the shift of the neuron's output. The bias term of the second layer is used for bias adjustment of the output layer, For introducing the nonlinear representation to reserve positive values and suppress negative values; In this embodiment, through the dynamic weighting mechanism, the weight is automatically assigned in real time according to the modal confidence vector and the semantic consistency result, so as to overcome the information deviation caused by the failure or conflict of part of the modal, and the method not only considers the confidence level of the modal itself, but also introduces the cross-modal consistency judgment as a reference basis, so as to make the intention recognition more suitable for the real context.

[0024] S3, introduce a trusted channel mechanism of "modal-user binding", capture user data in real time, and construct a user portrait combined with the intention judgment result; Further, in step S3, a trusted channel mechanism of "modal-user binding" is introduced, user data is captured in real time, and a user portrait is constructed combined with the intention judgment result, and the method is: The real-time capture of user data includes but is not limited to face recognition, voiceprint, operation habit, speaking speed, emotional stability, semantic tendency and commonly used vocabulary, and it is set that there are users, modal channels, a modal-user binding matrix is constructed, and the expression formula is: Wherein is the modal-user binding matrix, is the number of currently active users, is the number of simultaneously collected modalities, indicates whether the th modal is currently bound to the th user, and the modal set bound to the user is Wherein is the modal set, and the input features of all modalities in the modal set are extracted Wherein is the feature vector of the th sample, is -dimensional real space, and a weighted joint representation is constructed, and the expression formula is: Wherein is the weighted joint input representation, is the modal set, is the feature representation of the th modality, is the feature fusion contribution weight of the th modality to the user , is a hyperparameter used to control the influence degree of the interaction term between modalities, indicates the modal and modalities between semantics, representing modalities and modalities between feature differences.

[0025] Further, in step S3, the trusted channel mechanism of "modal - user binding" is introduced, and the method of constructing user portrait by capturing user data in real time combined with the intention judgment result is: According to the intention judgment result , the weighted joint input representation of the current user data is jointly encoded, and the expression formula is: Wherein is the external emotional input feature vector of the user at the current time , is the weighted joint input representation of the user , represents the outer product operation, is the context drift term introduced by the historical context change at the capture time , is the intention judgment result at time , and the user portrait is constructed as: Wherein is the user portrait of the user at time , is the space where the vector is located, and the user portrait is updated using a nonlinear gating mechanism. The expression formula is: Wherein is the user portrait of the user at time , is the emotional state representation of the user at the last time , is the external emotional input feature vector of the user at the current time , is the linear transformation parameter of emotional update, and the bias term matched with it is used to adjust the translation offset of the linear transformation result, is the Sigmoid function used to generate the gating weight, denotes the bit-by-bit multiplication between the corresponding elements of the two vectors for realizing weighted fusion; In this embodiment, taking a smart home companion robot as an example, this smart robot can simultaneously perceive the activity status of multiple family members in the home. When the father enters the living room and says, "Schedule a meeting for me," this method confirms his identity through voiceprint, identifies his tense state through speech rate, and recognizes his frequent use of words such as "schedule" and "meeting." Combined with past operating habits, it determines that he is a high-frequency schedule management user. At this time, the smart robot binds this multimodal information to the "father" user channel and dynamically generates a personalized user profile. In future interactions, it continuously memorizes and optimizes response strategies, solving the problem of unclear modality attribution in multi-user, multimodal concurrent scenarios. This enhances the accuracy and trustworthiness of user identification. By fusing modal input, external emotions, and intention results, a high-dimensional user profile is constructed under dynamic context changes. The nonlinear gating mechanism combines current emotions and historical states to achieve the fusion of long-term interests and short-term intentions, adapting to the evolution of user needs.

[0026] S4. Extract speech features from the user profile using Mel frequency cepstral coefficients, and analyze user tone fluctuations using emotion vector models to identify emotion data. Furthermore, in step S4, the method for extracting speech features from the user profile using Mel frequency cepstral coefficients and analyzing user tone fluctuations using an emotion vector model to identify emotion data is as follows: The speech features are divided into frames of fixed length. Each frame is [length] Frame shift is , to obtain the frame sequence, for the , The frame-by-frame speech signal undergoes a Fourier transform, and then the logarithm of the Mel frequency cepstral coefficients (MFCCs) is taken followed by a discrete cosine transform to output an MFCC vector. The emotion vector model calculates the dynamic changes between consecutive frames to capture intonation fluctuations, thereby constructing a combined speech feature vector based on the user profile. It employs a cross-modal attention mechanism to fuse voice emotion and user personality bias, expressed as: ,in It is the first The emotion-perceived embedding vector of a frame. It is Hadamard's positional multiplication. It is the first Each input feature vector The bias term is used to adjust the output of the activation function. This indicates that the final output vector lies in a In a 3-dimensional real space, It is a non-linear activation function. It is a projection matrix. It is a multilayer perceptron.

[0027] Furthermore, in step S4, the method for extracting speech features from the user profile using Mel frequency cepstral coefficients and analyzing user tone fluctuations using an emotion vector model to identify emotion data is as follows: Embed the emotion perception of all frames into vectors The input is processed by Bi-GRU+Attention (a neural network model that combines a bidirectional gated recurrent unit (Bi-GRU) and an attention mechanism), which extracts the user's overall sentiment representation, expressed as: ,in Indicates the first The hidden state calculated by the GRU (Gated Recurrent Unit) network at time 1. Indicates at time The hidden state of the network at that time It is the first The emotion-perceived embedding vector of a frame. It is attention weight. The weight vector is used to calculate the correlation between the hidden state and the current hidden state. Indicates the first The hidden state calculated by the GRU (Gated Recurrent Unit) network at time 1. It is a moment The final emotion vector representation is achieved by mapping the emotion vector to dimension values ​​through a fully connected network, expressed as follows: ,in It is an output of sentiment data. It is a moment The final emotion vector representation, The weight matrix is ​​used to transform the input to a different dimension. The bias term is used to adjust the output; In this embodiment, for the intelligent robot designed for the elderly, this method can analyze changes in the elderly's tone of voice in real time. For example, when a user says, "I feel a little unwell today," the speech rate slows down and the tone becomes lower. The method analyzes the emotion labels of "anxiety" or "depression" through MFCC and emotion vector model. By combining MFCC with a temporal model, the potential emotional fluctuations in the speech are extracted, which helps to improve the sensitivity and accuracy of emotion recognition. By using Bi-GRU+Attention, the trend of tone changes over time is effectively captured, making emotion judgment more continuous and consistent with the context.

[0028] S5. Using the ActionUnits analysis method, detect micro-expression details based on the user data and emotion data to determine the user's emotional trend; Further, in step S5, the method for detecting micro-expression details and judging the emotional trend of the user according to the user data and the emotion data by using the ActionUnits analysis method is: Set the sequence of continuous facial image frames of the user as: Wherein is a set of continuous facial image data of the user, is the image at the moment in the sequence, is the dimensional space of the image, each image is represented by the size of height width and has 3 RGB channels, represents the total number of images in the image sequence, after face detection and 68 key point labeling and positioning of each frame of facial image of the user, the standard facial region is aligned by affine transformation, the ActionUnits analysis method trained is used to extract the AU vector of each frame of facial image of the user, and the expression formula is: Wherein is an attention weight vector representing the attention weight at the moment, represents the degree of attention to the part of the input at the moment, is the dimension of the attention weight, representing the number of weight vector input parts, is the value range of the weight, represents is a dimensional real vector, the user emotion data and the user portrait are introduced to dynamically regulate and fuse the emotion priori of the AU vector.

[0029] Further, in step S5, the method for detecting micro-expression details and judging the emotional trend of the user according to the user data and the emotion data by using the ActionUnits analysis method is: By constructing a time series activation state matrix, a bidirectional LSTM network with residual connection is used to encode the AU time series, so as to capture the change direction and rate of micro-expression , then calculate the first-order derivative vector of micro-expression change , construct a micro-expression evolution tensor, and the expression formula is: Wherein represents the feature vector at the moment, represents the hidden layer state at the moment, represents the current moment and the last moment the difference between the hidden states, denotes the current time and the previous two times the difference between , is the hidden layer dimension, is the dimension of the final feature vector, based on the micro-expression evolution tensor, detects micro-expression details, predicts the emotional trend of the user, and the expression formula is: wherein is the emotional trend obtained after calculation at time , denotes the feature vector at time , is the dimension of the output vector, is the bias term, is the weight matrix used to convert the input vector to dimension; In the embodiment, by modeling the AU and the time sequence, the method can distinguish different emotional expressions of the same facial action, such as “strong smiling” and “true happiness”, and by fusing the time evolution and the first-order derivative trend analysis, the method is convenient for predicting the emotional trend of the user in advance, rather than only identifying the current emotion, and by combining the user portrait and the emotional data to adjust the AU analysis, the method can effectively cope with the realistic problem that the same expression has different meanings on different people.

[0030] The above embodiments are only used to illustrate the technical solutions of the present application, but not limit it; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement to some technical features; and these modifications or replacements will not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for intelligent robot interaction, characterized in that, Includes the following steps: S1. Construct a modality quality assessment model, introduce a contextual semantic consistency detection model, and generate modality confidence vectors and semantic consistency results; S2. The modality confidence vector and the semantic consistency result are adjusted using a dynamic weighting algorithm, and the intent judgment result is output. S3. Introduce a trusted channel mechanism of "modal-user binding" to build user profiles by capturing user data in real time and combining the results of intent judgment. S4. Extract speech features from the user profile using Mel frequency cepstral coefficients, and analyze user tone fluctuations using emotion vector models to identify emotion data. S5. Using the ActionUnits analysis method, detect micro-expression details based on the user data and emotion data to determine the user's emotional trend.

2. The intelligent robot interaction method according to claim 1, characterized in that, In step S1, the modality quality assessment model is constructed, and the contextual semantic consistency detection model is introduced. The method for generating the modality confidence vector and semantic consistency results is as follows: The modal quality assessment model receives the following multimodal inputs: ,in It is a set of multimodal input signals, including but not limited to speech, image, text, signal strength, background noise, image sharpness, and user activity. Indicates the first The original input signal is processed by a modal feature encoder. High-dimensional embedding is performed on each mode to obtain the modality embedding, where It is the first Modality Embedded representation, It is the first Encoder function for each mode, It is the first The original input signal of each mode, This represents the embedding representation vector of the modality. The modality's feature dimension size is used to construct the modality quality evaluation function. Used to evaluate modal representation Information integrity, clarity, and noise level are expressed by a nonlinear quality mapping method based on residual attention and context-aware graph structure, using the following formula: ,in It is the first Quality assessment values ​​for each modality It is used to limit the result to between 0 and 1. It is the number of local residual mappings in the figure. It is the first The weight matrix of each local residual mapping. It is used to handle nonlinear transformations. It is the first Embedded representation of each modality It is a context transformation matrix based on graph attention mechanism. It is used to measure the distance between vectors. It is a balance coefficient. It is used to measure the first Probability distribution of each mode With uniform distribution The difference between them yields the modal confidence vector as follows: , Wherein represents the modal confidence vector. It is the first Quality assessment values ​​for each modality This represents the total number of modes. This indicates that the vector belongs to the real number space. The context semantic consistency detection model obtains the modal embedding and introduces a cross-modal semantic alignment module to determine whether the current modal representation is consistent with the global context semantics. The cross-modal semantic alignment module is constructed based on a multimodal graph convolutional network and an attention alignment mechanism, and the resulting semantic consistency vector is: ,in Indicates semantic consistency results. Indicates the first Semantic consistency score for each modality This represents the total number of modes.

3. The intelligent robot interaction method according to claim 2, characterized in that, In step S2, the method for adjusting the confidence of the modality confidence vector and the semantic consistency result using a dynamic weighting algorithm to obtain the intent judgment result is as follows: For the modal confidence vector and the semantic consistency results Perform nonlinear fusion to generate normalized mode weights. Formula: ,in The softmax operation is used to... Generate normalized modal weights. It is the first Dynamic scores for each modality It is the first A modal confidence vector It is the first A semantic consistency result, It is a very small positive number. This represents the similarity between two vectors. It is the hyperbolic tangent function. It is a distance metric between modes. It is an adjustable hyperparameter.

4. The intelligent robot interaction method according to claim 3, characterized in that, In step S2, the method for adjusting the confidence of the modality confidence vector and the semantic consistency result using a dynamic weighting algorithm to obtain the intent judgment result is as follows: Using the modal weights The modal embeddings are weighted to obtain a unified fusion. Formula: ,in It is the representation after weighted fusion of all modalities. It is the total number of modes. It is the first Modal weights, It is the first Embedded representation of each modality Indicates the first The and the first Cosine similarity between modalities Indicates the first The and the first Feature differences between modes It is a balancing factor that will balance the fused components. The input is fed into the intent recognition function to obtain the intent determination result. Formula: ,in The intention judgment result is normalized by softmax. The weight parameters of the output layer are used to linearly transform the activation output of the previous layer. The weight parameters of the hidden layer are used for, It is the representation after weighted fusion of all modalities. The bias term in the first layer is used to adjust the shift of the neuron's output. The bias term in the second layer is used for bias adjustment of the output layer. It is used to introduce nonlinear representations that retain positive values ​​and suppress negative values.

5. The intelligent robot interaction method according to claim 4, characterized in that, In step S3, a trusted channel mechanism of "modal-user binding" is introduced. The method for constructing a user profile by capturing user data in real time and combining it with the intent judgment results is as follows: The real-time capture of user data includes, but is not limited to, facial recognition, voiceprint, operating habits, speaking speed, emotional stability, semantic tendencies, and frequently used vocabulary. One user, For each modal channel, construct a modality-user binding matrix, expressed as follows: ,in It is a modality-user binding matrix. This refers to the number of currently active users. It is the number of modes acquired simultaneously. Indicates the first Is the current modality related to the first modality? Individual user binding, for users The set of modes bound ,in It is a modality set; extract the input features of all modalities in the modality set. ,in It is the first The feature vector of each sample yes -D real space, construct weighted joint representation Formula: ,in It is a weighted joint input representation. It is a modal set. It is the first Feature representation of each modality It is the first Each modality for users Feature fusion contribution weights Hyperparameters are used to control the degree of influence of interaction terms between modes. Representing modes and modality The semantics between Representing modes and modality The differences in characteristics between them.

6. The intelligent robot interaction method according to claim 5, characterized in that, In step S3, a trusted channel mechanism of "modal-user binding" is introduced. The method for constructing a user profile by capturing user data in real time and combining it with the intent judgment results is as follows: Based on the intent judgment result Weighted joint input representation of the currently described user data Perform joint encoding to express the formula: ,in At the current moment For users The external emotion input feature vector, User Weighted joint input means, This indicates the outer product operation. To capture the moment Contextual drift introduced by changes in historical context It is a moment Based on the intent determination results, the user profile is constructed as follows: ,in User In time User profile at any time This refers to the space in which the vector resides, and the user profile update employs a nonlinear gating mechanism to fuse long-term interests and short-term intentions, expressed as: ,in User In time User profile at any time User In the previous moment The emotional state indicates, At the current moment For users The external emotion input feature vector, These are the linear transformation parameters for mood updates. and The accompanying bias term is used to adjust the translation offset of the linear transformation result. The sigmoid function is used to generate gated weights. This indicates that element-wise multiplication between corresponding elements of two vectors is used to achieve weighted fusion.

7. The intelligent robot interaction method according to claim 6, characterized in that, In step S4, the method for extracting speech features from the user profile using Mel frequency cepstral coefficients and analyzing user tone fluctuations using an emotion vector model to identify emotion data is as follows: The speech features are divided into frames of fixed length. Each frame is [length] Frame shift is , to obtain the frame sequence, for the , The frame-by-frame speech signal undergoes a Fourier transform, and then the logarithm of the Mel frequency cepstral coefficients (MFCCs) is taken followed by a discrete cosine transform to output an MFCC vector. The emotion vector model calculates the dynamic changes between consecutive frames to capture intonation fluctuations, thereby constructing a combined speech feature vector based on the user profile. It employs a cross-modal attention mechanism to fuse voice emotion and user personality bias, expressed as: ,in It is the first The emotion-perceived embedding vector of a frame. It is Hadamard's positional multiplication. It is the first Each input feature vector The bias term is used to adjust the output of the activation function. This indicates that the final output vector lies in a In a 3-dimensional real space, It is a non-linear activation function. It is a projection matrix. It is a multilayer perceptron.

8. The intelligent robot interaction method according to claim 7, characterized in that, In step S4, the method for extracting speech features from the user profile using Mel frequency cepstral coefficients and analyzing user tone fluctuations using an emotion vector model to identify emotion data is as follows: Embed the emotion perception of all frames into vectors The input is processed by Bi-GRU+Attention (a neural network model that combines a bidirectional gated recurrent unit (Bi-GRU) and an attention mechanism), which extracts the user's overall sentiment representation, expressed as: ,in Indicates the first The hidden state calculated by the GRU (Gated Recurrent Unit) network at time 1. Indicates at time The hidden state of the network at that time It is the first The emotion-perceived embedding vector of a frame. It is attention weight. The weight vector is used to calculate the correlation between the hidden state and the current hidden state. Indicates the first The hidden state calculated by the GRU (Gated Recurrent Unit) network at time 1. It is a moment The final emotion vector representation is achieved by mapping the emotion vector to dimension values ​​through a fully connected network, expressed as follows: ,in It is an output of sentiment data. It is a moment The final emotion vector representation, The weight matrix is ​​used to transform the input to a different dimension. The bias term is used to adjust the output.

9. The intelligent robot interaction method according to claim 8, characterized in that, In step S5, the ActionUnits analysis method is used to detect micro-expression details and determine the user's emotional trend based on the user data and emotion data. The user's continuous facial image frame sequence is set as follows: ,in It is a set of continuous facial image data of users. It is the first in the sequence Images of moments The dimensional space of an image represents the size of each image, which is its height. width It also has 3 RGB channels. This represents the total number of images in the image sequence. After performing face detection and 68 keypoint annotation and localization on each frame of the user's facial image, the standard facial region is aligned through affine transformation. The trained ActionUnits analysis method is then used to extract the AU vector from each frame of the user's facial image, expressed by the formula: ,in The attention weight vector represents the value at time 1. Attention weights Indicates at time For the input of the first The level of attention, The dimension representing the attention weights is... The number of vector input parts, It refers to the range of values ​​for the weights. express It is 3D real vectors, incorporating user sentiment data and user profiles Dynamically adjust the AU vector and fuse it with emotion priors.

10. The intelligent robot interaction method according to claim 8, characterized in that, In step S5, the ActionUnits analysis method is used to detect micro-expression details and determine the user's emotional trend based on the user data and emotion data. By constructing a time-series activation state matrix and then encoding the AU time series using a bidirectional LSTM network with residual connections, the direction and rate of micro-expression changes can be captured. Next, the first-order steering variable of micro-expression changes was calculated. Construct a micro-expression evolution tensor, expressed by the formula: ,in Indicates time eigenvectors, Indicates time The hidden state, Indicates the current time and the previous moment The difference in hidden states between them Indicates the current time And the previous two moments and The differences between them It is the hidden layer dimension. This refers to the dimension of the final feature vector, based on the micro-expression evolution tensor, which detects micro-expression details and predicts the user's emotional trend. The formula is as follows: ,in At any moment The sentiment trend obtained after calculation Indicates time eigenvectors, It is the dimension of the output vector. It is a bias term. The weight matrix is ​​used to transform the input vector to... Dimension.

Citation Information

Patent Citations

  • Online creative interaction platform based on intelligent robot

    CN107203607A