Hybrid interaction method and device, electronic equipment and readable medium

By collecting and processing multimodal interaction data, using the interaction recognition model to determine the interaction intention and generate feedback, the problem of users' difficulty in interacting with the traditional graphical user interface in specific scenarios is solved, and a more convenient and natural interactive experience is achieved.

CN120029445APending Publication Date: 2025-05-23CHINA TELECOM ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411876200.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-18
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

In some scenarios, it is difficult for users to interact with traditional graphical user interfaces, resulting in inconvenient interaction process.

Method used

The hybrid interaction method is adopted to collect multimodal interaction data of users (including voice data, touch data, gesture data, eye movement data and facial expression data), and extract feature encoding information through the interaction recognition model, determine the interaction recognition information, and generate interactive feedback information.

Benefits of technology

Through the collection and processing of multimodal interactive data, users can choose the most appropriate interaction method in different situations, which improves the convenience and nature of interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029445A_ABST
    Figure CN120029445A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a hybrid interaction method and device, electronic equipment and a readable medium. The method comprises the steps of collecting multi-modal interaction data of a user, wherein the multi-modal interaction data comprises at least two kinds of interaction data of voice data, touch data, gesture data, eye movement data and facial expression data; extracting feature coding information in the multi-modal interaction data through the interaction recognition model, and determining interaction recognition information based on the feature coding information; and based on the interaction identification information, interaction feedback information is generated and fed back to a user. Therefore, interaction with the system can be carried out by collecting user inputs of various different modes, so that the interaction convenience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a hybrid interaction method, a hybrid interaction device, an electronic device and a computer-readable medium. Background Art

[0002] Human-Computer Interaction (HCI) refers to the design, evaluation and implementation of interactive computing systems to improve the interaction experience and efficiency between humans and computers. Generally speaking, human-computer interaction can usually be achieved by providing users with a graphical user interface (GUI), and users can interact with the system through graphical elements (such as buttons and menus).

[0003] However, in some specific scenarios, such as in a driving environment that requires the driver's full concentration, or when the user's hands are occupied and unable to operate traditional devices, it is difficult for the user to interact with the traditional graphical user interface, resulting in the interaction process being not very convenient. Summary of the invention

[0004] The embodiments of the present invention provide a hybrid interaction method, device, electronic device and computer-readable storage medium to solve the problem of difficulty in interaction between a user and a system.

[0005] An embodiment of the present invention discloses a hybrid interaction method, which includes:

[0006] Collecting multimodal interaction data of the user, wherein the multimodal interaction data includes at least two types of interaction data among voice data, touch data, gesture data, eye movement data, and facial expression data;

[0007] Extracting feature coding information from the multimodal interaction data through the interaction recognition model, and determining interaction recognition information based on the feature coding information;

[0008] Based on the interaction identification information, interaction feedback information is generated and fed back to the user.

[0009] Optionally, the step of extracting feature coding information from the multimodal interaction data by using the interaction recognition model, and determining interaction recognition information based on the feature coding information includes:

[0010] Extracting feature information corresponding to the interaction data respectively through the interaction recognition model;

[0011] Performing feature encoding processing on the feature information respectively to obtain at least two feature encoding information; fusing the at least two feature encodings based on a preset attention weight matrix to obtain fused feature information;

[0012] Based on the fused feature information, interaction identification information is determined.

[0013] Optionally, the feature information includes speech feature information corresponding to the language data;

[0014] The step of respectively extracting feature information corresponding to the interaction data through the interaction recognition model comprises:

[0015] Extract Mel-frequency cepstral coefficients from speech data;

[0016] Extracting convolution feature information from the Mel-frequency cepstral coefficients based on a preset convolutional neural network;

[0017] The convolution feature information is input into a preset recurrent neural network, and a speech recognition result output by the recurrent neural network is obtained as speech feature information.

[0018] Optionally, the feature information includes touch feature information corresponding to the touch data;

[0019] The step of respectively extracting feature information corresponding to the interaction data through the interaction recognition model comprises:

[0020] Smoothing the touch data to obtain smooth trajectory data;

[0021] At least one of a speed feature, a curvature feature, a direction feature, and a shape feature is extracted from the smooth trajectory data as touch feature information.

[0022] Optionally, the characteristic information includes gesture characteristic information corresponding to the gesture data;

[0023] The step of respectively extracting feature information corresponding to the interaction data through the interaction recognition model comprises:

[0024] A preset image recognition model is used to determine the gesture key points corresponding to the gesture data as gesture feature information.

[0025] Optionally, the feature information includes eye movement feature information corresponding to the eye movement data;

[0026] The step of respectively extracting feature information corresponding to the interaction data through the interaction recognition model comprises:

[0027] Based on the eye movement data, determining eyeball positions and pupil positions;

[0028] Based on the eyeball position and the pupil position, the user's sight direction is determined as eye movement feature information.

[0029] Optionally, the feature information includes expression feature information corresponding to the facial expression data:

[0030] The step of respectively extracting feature information corresponding to the interaction data through the interaction recognition model comprises:

[0031] Extracting key expression information from the facial expression data;

[0032] Based on the expression key information, an expression recognition result is determined as the expression feature information.

[0033] An embodiment of the present invention further provides a hybrid interaction device, the device comprising:

[0034] A data collection module, used to collect multimodal interaction data of the user, wherein the multimodal interaction data includes at least two types of interaction data among voice data, touch data, gesture data, eye movement data, and facial expression data;

[0035] An interaction identification module, configured to extract feature coding information from the multimodal interaction data through the interaction identification model, and determine interaction identification information based on the feature coding information;

[0036] The feedback module is used to generate interaction feedback information based on the interaction identification information and provide feedback to the user.

[0037] Optionally, the interaction identification module includes:

[0038] A feature information extraction submodule, used to extract feature information corresponding to the interaction data respectively through the interaction recognition model;

[0039] The encoding submodule is used to perform feature encoding processing on the feature information respectively to obtain at least two types of feature encoding information;

[0040] A fusion submodule, used for fusing the at least two feature codes based on a preset attention weight matrix to obtain fused feature information;

[0041] The interaction identification submodule is used to determine the interaction identification information based on the fused feature information.

[0042] Optionally, the feature information includes speech feature information corresponding to the language data;

[0043] The feature information extraction submodule includes:

[0044] A mel-frequency cepstral coefficient extraction unit, used for extracting mel-frequency cepstral coefficients in speech data;

[0045] A convolution feature extraction unit, used to extract convolution feature information from the Mel-frequency cepstral coefficients based on a preset convolutional neural network;

[0046] The speech feature extraction unit is used to input the convolution feature information into a preset recurrent neural network and obtain the speech recognition result output by the recurrent neural network as the speech feature information.

[0047] Optionally, the feature information includes touch feature information corresponding to the touch data;

[0048] The feature information extraction submodule includes:

[0049] A smooth trajectory acquisition unit, used for smoothing the touch data to obtain smooth trajectory data;

[0050] The touch feature extraction unit is used to extract at least one of a speed feature, a curvature feature, a direction feature, and a shape feature from the smooth trajectory data as touch feature information.

[0051] Optionally, the characteristic information includes gesture characteristic information corresponding to the gesture data;

[0052] The feature information extraction submodule includes:

[0053] The gesture feature extraction unit is used to use a preset image recognition model to determine the gesture key points corresponding to the gesture data as gesture feature information.

[0054] Optionally, the feature information includes eye movement feature information corresponding to the eye movement data;

[0055] The feature information extraction submodule includes:

[0056] An eye movement position determination unit, used to determine the eyeball position and the pupil position based on the eye movement data;

[0057] The eye movement feature extraction unit is used to determine the user's line of sight direction as eye movement feature information based on the eyeball position and the pupil position.

[0058] Optionally, the feature information includes expression feature information corresponding to the facial expression data:

[0059] The feature information extraction submodule includes:

[0060] An expression key information extraction unit, used to extract expression key information from the facial expression data;

[0061] The expression feature extraction unit is used to determine the expression recognition result as the expression feature information based on the expression key information.

[0062] The embodiment of the present invention further discloses an electronic device, comprising a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus;

[0063] The memory is used to store computer programs;

[0064] The processor is used to implement the method described in the embodiment of the present invention when executing the program stored in the memory.

[0065] The embodiment of the present invention further discloses one or more computer-readable media on which instructions are stored. When executed by one or more processors, the processors are enabled to execute the method described in the embodiment of the present invention.

[0066] The embodiments of the present invention include the following advantages:

[0067] Through the hybrid interaction method provided by the embodiment of the present invention, the multimodal interaction data of the user is collected, and the multimodal interaction data includes at least two types of interaction data among voice data, touch data, gesture data, eye movement data, and facial expression data; the feature coding information in the multimodal interaction data is extracted through the interaction recognition model, and the interaction recognition information is determined based on the feature coding information; based on the interaction recognition information, interaction feedback information is generated and fed back to the user. In this way, the user input of multiple different modes can be collected to interact with the system, thereby improving the convenience of interaction. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] Figure 1 is a flow chart of steps of a hybrid interaction method provided in an embodiment of the present invention;

[0069] Figure 2 is a structural block diagram of a hybrid interaction device provided in an embodiment of the present invention;

[0070] Figure 3 is a block diagram of an electronic device provided in an embodiment of the present invention;

[0071] Figure 4 is a schematic diagram of a computer-readable medium provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0072] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0073] Reference Figure 1 , shows a flow chart of the steps of a hybrid interaction method provided in an embodiment of the present invention, which may specifically include the following steps:

[0074] Step 101, collecting multimodal interaction data of a user, wherein the multimodal interaction data includes at least two types of interaction data among voice data, touch data, gesture data, eye movement data, and facial expression data;

[0075] In the embodiment of the present invention, the user can use multiple different modes to input to the system and interact with the system. The multimodal interaction data includes at least two types of interaction data among voice data, touch data, gesture data, eye movement data, and facial expression data.

[0076] In the specific implementation, voice data can be collected through a microphone array, touch data can be obtained through a capacitive sensor, gesture data, facial expression data, and eye movement data can be obtained through depth cameras, infrared cameras and other camera devices.

[0077] Step 102, extracting feature coding information from the multimodal interaction data through the interaction recognition model, and determining interaction recognition information based on the feature coding information;

[0078] In an embodiment of the present invention, an interaction recognition model can be set up to process and analyze the collected data, perform feature encoding on the multimodal interaction data, obtain key information in the multimodal interaction data as feature information, and then encode the feature information, and perform feature recognition through the interaction recognition model to determine the interaction recognition information to understand the user's current interaction intention.

[0079] Specifically, the interaction recognition model may include a variety of different sub-models, and different sub-models may have different functions. For example, the sub-model may be used to extract feature information from specific modal data, or may be used to encode feature information, or may be used to recognize the user's interaction intention and learn the user's current interaction intention.

[0080] Step 103: Generate interaction feedback information based on the interaction identification information and feed it back to the user.

[0081] After the system determines the user's current interaction intention, it can generate corresponding interaction feedback information based on the user's interaction intention and provide feedback to the user, so that the user can interact with the system in a variety of different modes, thereby improving the convenience of interaction.

[0082] Specifically, the interactive feedback information may be information associated with the interactive identification information and related to the functions that the system can provide. For example, the interactive identification information may indicate that the user's current intention is to check the weather, and the interactive feedback information may be feedback of text or voice information containing the current weather conditions. For another example, the interactive identification information may indicate that the user's current intention is to turn on the lights in the living room, and the interactive feedback information may be feedback of confirmation information to the user after the system turns on the lights in the living room. For another example, the interactive identification information may indicate that the user's current intention is to switch songs, and the interactive feedback information may be feedback of confirmation information to the user after the system switches songs.

[0083] Through the hybrid interaction method provided by the embodiment of the present invention, the multimodal interaction data of the user is collected, and the multimodal interaction data includes at least two types of interaction data among voice data, touch data, gesture data, eye movement data, and facial expression data; the feature coding information in the multimodal interaction data is extracted through the interaction recognition model, and the interaction recognition information is determined based on the feature coding information; based on the interaction recognition information, interaction feedback information is generated and fed back to the user. In this way, the user input of multiple different modes can be collected to interact with the system, thereby improving the convenience of interaction.

[0084] In an embodiment of the present invention, the step of extracting feature coding information from the multimodal interaction data by using the interaction recognition model, and determining interaction recognition information based on the feature coding information includes:

[0085] S11, extracting feature information corresponding to the interaction data respectively through the interaction recognition model;

[0086] In a specific implementation, the interaction recognition model can first extract the feature information corresponding to the interaction data. For example, after the voice data is collected by the microphone array, the fast Fourier transform (FFT) is used to extract the time-frequency domain features to obtain the feature information; the touch data is obtained by the capacitive sensor, and the Kalman filter is used to suppress noise and smooth the data to obtain the feature information; the gesture data is obtained by the depth camera, and the model is used to perform gesture recognition to obtain the feature information.

[0087] S12, performing feature coding processing on the feature information respectively to obtain at least two types of feature coding information;

[0088] In the embodiment of the present invention, feature encoding processing may be performed on the feature information respectively, and the deep representation of feature information of different modalities may be analyzed to obtain feature encoding information.

[0089] As a specific example of the present invention, a deep belief network (DBN) may be used to perform feature encoding processing on the feature information respectively to obtain at least two types of feature encoding information.

[0090] Specifically, the Deep Belief Network (DBN) is a generative model composed of multiple layers of restricted Boltzmann machines (RBMs). It trains each layer of RBM layer by layer through unsupervised learning to extract high-level features of the data. In the feature encoding layer, the input data is processed by stacked RBMs, and the data representation learned by each layer of RBM is used as the input of the next layer. The hidden layer uses the ReLU activation function to increase nonlinearity, which helps the model learn more complex features. Finally, the feature decoding layer is fine-tuned through the back-propagation algorithm to ensure that the network can accurately reconstruct the input data, thereby optimizing the feature representation.

[0091] S13, fusing the at least two feature codes based on a preset attention weight matrix to obtain fused feature information;

[0092] In a specific implementation, an attention weight matrix can be pre-set, and the attention weight matrix can distinguish the contribution of feature encoding information of different models to the intent recognition task. At least two feature encodings can be fused based on the attention weight matrix to obtain fused feature information.

[0093] S14: Determine interactive identification information based on the fused feature information.

[0094] Afterwards, the attention-weighted fusion feature information can be integrated and the final decision can be made. A feedforward neural network structure can be used to effectively integrate the features of different modalities. In the last layer of the model, the softmax function is used to convert the network output into a probability distribution to make a classification decision. The Softmax function can ensure that the sum of the output probabilities is one, so that the prediction results of the model have probabilistic meaning, which facilitates the selection of the most likely category as the final output to obtain interactive recognition information.

[0095] As a specific example of the present invention, a deep belief network (DBN) is a generative model composed of multiple layers of restricted Boltzmann machines (RBMs) stacked together to learn deep feature representations of data;

[0096] RBM training: Each layer of RBM learns the representation of data through comparison paging and reconstruction process;

[0097] Feature encoding: Use the hidden layer activation of RBM as the feature input of the previous layer;

[0098] Feature decoding: Use the RBM reconstruction process to decode high-level features back to low-level features;

[0099] Energy function of RBM:

[0100]

[0101] Among them, v is the visible layer, h is the hidden layer, w is the weight, and a and b are biases.

[0102] Specifically, the attention mechanism is used to dynamically adjust the importance of different modal data, allowing the model to focus on the most relevant information;

[0103] Weight calculation: Use a trainable weight matrix and softmax activation function to calculate the attention score of each modality;

[0104] Feature weighting: weight the features according to the attention scores to enhance the features of important modalities;

[0105] Calculation of attention score:

[0106]

[0107] Among them, e k is the unnormalized score of the kth modality, α k is the corresponding attention weight.

[0108] Specifically, the fusion decision layer uses a feedforward neural network to integrate features of different modalities and uses a softmax function to make classification decisions;

[0109] Feature integration: input the weighted features into the feedforward neural network;

[0110] Classification decision: Use the softmax function to output the final classification result;

[0111] Softmax function:

[0112]

[0113] Among them, z is the output of the network, p k is the probability of the kth category.

[0114] As a specific example of the present invention, all decisions obtained under each mode can be preset. Suppose there are n decisions in total, and the set of each decision is [d 1 ,d 2 ,…,d n}, the decision fusion module generates the following probability matrix according to the decisions and decision probabilities obtained under each mode:

[0115]

[0116] Among them, p 1j Make decisions based on the user’s voice modality j The probability, p 2j Indicates that the user makes a decision in gesture mode d j The probability, p3j Make decisions based on the user's eye movement modality j The probability, p 4j Represents the user's decision-making mode j The probability that for all j satisfies p ij ∈[0,1].

[0117] Optionally, the decision fusion module generates a weight matrix according to the probability matrix, and the weight matrix is ​​obtained as follows: calculate the average probability of the user making a decision on each modality:

[0118]

[0119] in, Help users make decisions on each modality j The average probability, p ij Make a decision d for the user in the i-th mode j The probability of; calculate:

[0120]

[0121] Among them, δ ij Represents the user making a decision d in the i-th mode j The larger the distance is, the smaller the correlation between the decision made under this mode and the average decision is.

[0122] Optionally, according to the δ ij Generate a weight matrix to assign weights to the decisions under each mode. The weight matrix is ​​as follows:

[0123]

[0124] Among them, w 1j Make decisions based on the user’s voice modality j The weight, w 2j Indicates that the decision is made in the user's voice mode d j The weight, w 3j Make decisions based on user gesture modality j The weight, w 4j Represents the user's decision-making under the eye movement mode j The weights of ij ∈[0,1].

[0125] Optionally, for each weight w in the weight matrix ij ,satisfy:

[0126]

[0127] Among them, max(δ) represents all spacing δ ij The maximum value in .

[0128] Optionally, the probability matrix P and the weight matrix W may be multiplied to generate a final decision matrix, and decisions corresponding to values ​​greater than a threshold value of 0.6 in the final decision matrix are extracted as interaction identification information.

[0129] In the specific implementation, the interactive recognition model can be trained in the following way: the Xavier method is used for parameter initialization, which helps to maintain the activation value and gradient size of neurons in each layer in the early stage of training to avoid the problem of gradient disappearance or explosion. The stochastic gradient descent (SGD) algorithm is used for model training to minimize the loss function, namely the cross entropy loss, by continuously updating the network weights. During the training process, the model is trained using the training data set, and the generalization ability of the model is evaluated through the verification data set, and the hyperparameters are adjusted to optimize the model performance.

[0130] In one embodiment of the present invention, the feature information includes speech feature information corresponding to the language data;

[0131] The step of respectively extracting feature information corresponding to the interaction data through the interaction recognition model comprises:

[0132] S21, extracting Mel frequency cepstral coefficients in the speech data;

[0133] S22, extracting convolution feature information from the Mel-frequency cepstral coefficients based on a preset convolutional neural network;

[0134] S23, inputting the convolution feature information into a preset recurrent neural network, and obtaining a speech recognition result output by the recurrent neural network as speech feature information.

[0135] In a specific implementation, the time domain signal can be first converted into a frequency domain signal by using a fast Fourier transform (FFT), and then the Mel frequency cepstral coefficients (MFCCs) are calculated. Mel is an important parameter for describing the characteristics of human voice, and the calculation formula is:

[0136] MFCCk=log(∑m=1MDm,k)

[0137] Thereafter, the Mel-frequency cepstral coefficients can be input into a preset convolutional neural network to extract convolution feature information in the Mel-frequency cepstral coefficients. The convolutional neural network can use convolution kernels of different sizes to capture convolution feature information of different scales.

[0138] The convolution feature information can be input into a preset recurrent neural network, and a speech recognition result output by the recurrent neural network is obtained as the speech feature information.

[0139] Specifically, the recurrent neural network may be a bidirectional LSTM network. The bidirectional LSTM network may include a forward LSTM and a backward LSTM. The forward LSTM is used to process a sequence from time step t=1 to t=T. The backward LSTM is used to process a sequence from time step t=T to t=1. The convolution feature information may be input into the bidirectional LSTM network as sequence information, and the outputs of the forward and backward LSTMs may be combined to form hidden state features.

[0140] Among them, the basic unit formula of LSTM can be expressed as:

[0141] f t =σ(W f ·[h t-1 , x t ]+b f )

[0142] i t =σ(W i ·[h t-1 , x t ]+b i )

[0143] o t =σ(W o ·[h t-1 , x t ]+b o )

[0144]

[0145] h t =o t *tanh(C t )

[0146] Among them, f t ,i t , o t Respectively represent the forget gate, input gate and output gate, C t and h t denote the cell state and hidden state respectively.

[0147] The LSTM network can further decode the hidden state features and finally output the speech recognition results as speech feature information.

[0148] In one embodiment of the present invention, the characteristic information includes touch characteristic information corresponding to the touch data;

[0149] The step of respectively extracting feature information corresponding to the interaction data through the interaction recognition model comprises:

[0150] S31, smoothing the touch data to obtain smooth trajectory data;

[0151] S32: extract at least one of a speed feature, a curvature feature, a direction feature, and a shape feature from the smooth trajectory data as touch feature information.

[0152] Specifically, a Kalman filter may be used to smooth the touch trajectory data to eliminate noise and irregularities and obtain smooth trajectory data.

[0153] The Kalman filter is a recursive filter that uses the dynamic model and measurement model of the system to estimate the state of the system. The basic equation of the Kalman filter includes two steps: prediction and update:

[0154] Prediction steps:

[0155]

[0156] in, is the state prediction at time k, F k is the state transfer matrix, B k and u k are the control input and control matrix, P k∣k-1 is the covariance of the forecast errors, Q k is the covariance of the prediction noise;

[0157] Update steps:

[0158]

[0159] P k =(IK k H k ) k∣k-1

[0160] Among them, K k is the Kalman gain, z k is the measured value at time k, H k is the measurement matrix, R k is the covariance of the measurement noise.

[0161] Thereafter, at least one of the speed feature, the curvature feature, the direction feature, and the shape feature may be further extracted from the smooth trajectory data as touch feature information.

[0162] Speed ​​feature: calculate the displacement change rate of the touch point between consecutive time points;

[0163]

[0164] where Δx and Δy are the changes in x and y coordinates between consecutive time points, respectively;

[0165] Direction feature: Calculate the direction change of the touch track:

[0166] θ=arctan2(Δy,Δx)

[0167] Curvature feature: calculates the local curvature of the touch trajectory to recognize complex gestures;

[0168]

[0169] As a specific example of the present invention, touch data is collected by a capacitive sensor, which can detect the contact point of the user's finger on the touch surface. The sensor outputs a series of time series data, including the coordinates of the touch point (x t ,y t ) and touch force p t , where t represents time;

[0170] Track direction analysis: Use the Histogram of Directed Gradients (HOG) descriptor to analyze the direction changes of the touch track. The HOG descriptor captures the characteristics of the shape by calculating the edge directional gradient histogram of the local area;

[0171] For each touch point, first calculate its gradient in the x and y directions:

[0172]

[0173] Where I is the image intensity, and the gradient can be calculated by the central difference method;

[0174] Use the Sobel operator to convolve the image to obtain the gradient amplitude and direction of each pixel; divide the touch track into several small units, and calculate the histogram of the gradient direction in each unit; normalize and connect the histogram to form a feature vector;

[0175] Trajectory velocity analysis: Analyze velocity characteristics by calculating the displacement change rate of the touch point between consecutive time points, velocity vector V t Reflects the speed at which the user's finger moves on the touch surface;

[0176] Velocity vector V t It can be calculated by the following formula:

[0177]

[0178] Where (x t ,y t ) and (x t+1 ,y t+1 ) are the coordinates of consecutive time points;

[0179] The coordinates of continuous touch points are differentiated to calculate the velocity vector; a series of velocity vectors are statistically analyzed, such as calculating the average velocity, standard deviation of velocity change, etc., to obtain velocity characteristics.

[0180] In a specific implementation, the trajectory shape analysis can be further performed, and the shape features of the touch trajectory can be extracted using a contour detection algorithm; these features are helpful in identifying specific gestures or symbols;

[0181] The Canny edge detection algorithm is used to detect the edge of the touch track; the geometric features of the extracted contour, such as length, area, convex hull, etc., are calculated to form shape features.

[0182] In one embodiment of the present invention, the characteristic information includes gesture characteristic information corresponding to the gesture data;

[0183] The step of respectively extracting feature information corresponding to the interaction data through the interaction recognition model comprises:

[0184] S41, using a preset image recognition model to determine gesture key points corresponding to the gesture data as gesture feature information.

[0185] Specifically, MaskRCNN can be used for skeleton segmentation first, and then YOLOv4 can be used for skeleton key point detection. Darknet53 is used as the feature extraction network, and multi-scale prediction is used to enhance the detection performance. A window is slid on the feature map to predict the position and confidence of the key point for each window.

[0186] The prediction formula for key point position:

[0187]

[0188] Among them, (b x ,b y ) is the center position of the anchor point, is the predicted keypoint offset and σ is the sigmoid activation function.

[0189] The key point detection algorithm can then be used to extract the fingertip and joint positions. The key point detection algorithm uses a convolutional neural network to extract features of the gesture image and predict the fingertip and joint positions through the trained model.

[0190] The calculation of the key point position can be expressed as:

[0191] P i =f(I;θ)

[0192] Among them, P i is the position of the i-th key point, I is the input image, θ is the model parameter, and f is the key point detection function.

[0193] The optical flow algorithm is used for key point tracking to track the dynamic changes of key points in continuous image frames to achieve accurate recognition of dynamic gestures.

[0194] Specifically, the Lucas-Kanade method can be used to calculate the optical flow between two frames, and the positions of the key points can be updated according to the optical flow information.

[0195] The basic equation of optical flow can be expressed as:

[0196]

[0197] Where I is the image intensity, is the gradient of the image, is the optical flow vector, I t is the time derivative of the image intensity.

[0198] In one embodiment of the present invention, the feature information includes eye movement feature information corresponding to the eye movement data;

[0199] The step of respectively extracting feature information corresponding to the interaction data through the interaction recognition model comprises:

[0200] S51, determining the eyeball position and the pupil position based on the eye movement data;

[0201] S52: Based on the eyeball position and the pupil position, determine the user's sight direction as eye movement feature information.

[0202] Specifically, the image is first grayscaled and Gaussian blurred to reduce noise and smooth the image, then the Canny algorithm is used to detect the edges in the image, and then the Hough transform is applied to the edge image to detect the circular structure, and then the identified circular area of ​​the eye is used to locate the eyeball and pupil;

[0203]

[0204] Among them, (x c ,y c ) are the coordinates of the circle center, f(x,y) is the edge strength function, and δ is the Dirac delta function.

[0205] Specifically, a geometric model can be used to calculate the line of sight direction. Based on the geometric structure of the eye, the line of sight direction can be estimated through the positional relationship between the pupil center and the corneal reflection point;

[0206] Corneal reflection point detection: Detect corneal reflection points in eye images, usually bright spots;

[0207] Sight vector calculation: Based on the position of the pupil center and the corneal reflection point, the sight direction is calculated using vector geometry;

[0208] The calculation formula of sight direction is:

[0209]

[0210] in, is the pupil center, It is the corneal reflex point. is the view direction vector.

[0211] Specifically, the line of sight direction is converted into spatial coordinates using a perspective projection model; the perspective projection model takes into account the distance from the eye to the observation plane and the angle between the line of sight and the observation plane;

[0212] Calculation of the intersection of line of sight and plane: Calculate the intersection of line of sight and plane according to the line of sight direction and the distance from the eye to the observation plane;

[0213] Space coordinate conversion: convert the screen coordinates of the intersection point into three-dimensional space coordinates;

[0214]

[0215] in, is the position of the eye, t is the distance from the intersection of the line of sight and the observation plane to the eye, is the sight direction vector, are the spatial coordinates of the focus.

[0216] In one embodiment of the present invention, the feature information includes expression feature information corresponding to the facial expression data:

[0217] The step of respectively extracting feature information corresponding to the interaction data through the interaction recognition model comprises:

[0218] S61, extracting key expression information from the facial expression data;

[0219] S62, based on the expression key information, determining an expression recognition result as the expression feature information.

[0220] Specifically, the facial feature detection algorithm processing unit is first used to locate the user's facial area through the face detection subunit. Then the expression key point extraction subunit identifies the position information of key parts such as eyebrows, eyes, nose, and mouth. The expression tracking subunit is responsible for tracking the changes of these key points in a series of continuous image frames in order to capture the changes in the user's facial expressions. This process can help the system capture the user's facial expression data in real time and convert it into data that can be used for multimodal fusion, which is then analyzed together with data from other modalities by the AI ​​large model processing module to understand the user's interaction intentions.

[0221] Furthermore, a Haar feature cascade classifier is used to locate the face; the expression key point extraction subunit uses the OpenPose framework to extract the position information of eyebrows, eyes, nose, mouth, etc.; the expression tracking subunit uses a Kalman filter to track the dynamic changes of expression key points in continuous image frames.

[0222] Specifically, the Haar feature cascade classifier is used, which is an effective face detection method based on image features. It can quickly locate the face position in the image. The expression key point extraction subunit uses the OpenPose framework, which is good at human posture estimation tasks and can accurately extract facial feature points such as eyebrows, eyes, nose, and mouth. The expression tracking subunit uses the Kalman filter to maintain stable tracking of facial feature points in continuous image frames, and can maintain high accuracy even when the user's head moves or the expression changes.

[0223] Furthermore, local binary patterns (LBP) can be used to extract expression features from expression key point information; the expression recognition subunit uses a support vector machine (SVM) to match predefined expression patterns according to the extracted expression features and output a recognition result.

[0224] Specifically, the local binary pattern (LBP) algorithm is used to extract expression features from the extracted expression key point information. LBP is an effective texture descriptor that can capture local details well. The extracted features are then sent to the expression recognition subunit, which uses a support vector machine (SVM) model for classification, matches the predefined expression patterns according to the extracted expression features, and outputs the recognition results to determine the user's current emotional state or intention.

[0225] The hybrid interaction method provided by the embodiment of the present invention collects the user's voice data through a voice recognition module, uses a convolutional neural network (CNN) to extract features, and then recognizes it through a recurrent neural network (RNN). Through a touch trajectory analysis algorithm processing unit, including trajectory smoothing processing, trajectory feature extraction and gesture recognition, the user's touch action is recognized. Through a skeleton recognition algorithm processing unit, including skeleton detection, key point extraction and key point tracking, the user's gesture action is recognized. Through a line of sight tracking algorithm processing unit, including eye feature detection, line of sight direction estimation and focus mapping, the user's line of sight focus is tracked. This multimodal fusion interaction method allows the user to choose the most appropriate interaction method in different situations, such as using voice or eye control in a driving environment, or using gesture control when both hands are occupied. This flexibility significantly improves the user experience and makes the interaction more natural and intuitive.

[0226] The hybrid interaction method provided in the embodiment of the present invention adopts a multimodal fusion algorithm to fuse the features of different modalities to obtain fused features, and then uses the fused features to identify interaction intentions and understand the scenes, and transmits the identification results to the interaction response module to implement corresponding functional operations. According to user feedback, the system parameters are adjusted to optimize the interaction experience. Through the intelligent decision-making ability of the interaction identification model, the system can more accurately understand the user's interaction intentions and needs, thereby providing a more accurate response. At the same time, the adaptive learning ability allows it to self-optimize according to the user's usage habits and preferences, so that the interaction experience continues to improve over time.

[0227] The hybrid interaction method provided by the embodiment of the present invention not only enriches the system's perception ability through the facial expression recognition function, but also enhances the understanding of the user's emotional state. The system can judge the user's emotional changes by detecting the user's facial expressions, such as smiling, frowning, etc., and adjust the interaction strategy accordingly, thereby providing a more humane and personalized interaction experience; for example, when the user expresses confusion or dissatisfaction, the system can timely adjust the interaction interface or provide additional help information to improve the user's experience. In addition, through the multimodal fusion algorithm, the system can flexibly select the most appropriate interaction method in different situations. For example, when the user cannot use gestures or voice, control is performed through eye movements or facial expressions, which significantly improves the naturalness and convenience of the interaction. This multi-dimensional interaction method not only improves the user experience, but is also particularly suitable for human-computer interaction in special groups or specific environments, making the technology more popular and inclusive.

[0228] It should be noted that, for the sake of simplicity, the method embodiments are described as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present invention are not limited by the order of the actions described, because according to the embodiments of the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present invention.

[0229] Reference Figure 2 , shows a structural block diagram of a hybrid interaction device provided in an embodiment of the present invention, which may specifically include the following modules:

[0230] The data collection module 201 is used to collect multimodal interaction data of the user, wherein the multimodal interaction data includes at least two types of interaction data selected from the group consisting of voice data, touch data, gesture data, eye movement data, and facial expression data;

[0231] An interaction identification module 202, configured to extract feature coding information from the multimodal interaction data through the interaction identification model, and determine interaction identification information based on the feature coding information;

[0232] The feedback module 203 is used to generate interaction feedback information based on the interaction identification information and feed it back to the user.

[0233] Optionally, the interaction identification module includes:

[0234] A feature information extraction submodule, used to extract feature information corresponding to the interaction data respectively through the interaction recognition model;

[0235] The encoding submodule is used to perform feature encoding processing on the feature information respectively to obtain at least two types of feature encoding information;

[0236] A fusion submodule, used for fusing the at least two feature codes based on a preset attention weight matrix to obtain fused feature information;

[0237] The interaction identification submodule is used to determine the interaction identification information based on the fused feature information.

[0238] Optionally, the feature information includes speech feature information corresponding to the language data;

[0239] The feature information extraction submodule includes:

[0240] A Mel-frequency cepstral coefficient extraction unit, used for extracting Mel-frequency cepstral coefficients in speech data;

[0241] A convolution feature extraction unit, used to extract convolution feature information from the Mel-frequency cepstral coefficients based on a preset convolutional neural network;

[0242] The speech feature extraction unit is used to input the convolution feature information into a preset recurrent neural network and obtain the speech recognition result output by the recurrent neural network as the speech feature information.

[0243] Optionally, the feature information includes touch feature information corresponding to the touch data;

[0244] The feature information extraction submodule includes:

[0245] A smooth trajectory acquisition unit, used for smoothing the touch data to obtain smooth trajectory data;

[0246] The touch feature extraction unit is used to extract at least one of a speed feature, a curvature feature, a direction feature, and a shape feature from the smooth trajectory data as touch feature information.

[0247] Optionally, the characteristic information includes gesture characteristic information corresponding to the gesture data;

[0248] The feature information extraction submodule includes:

[0249] The gesture feature extraction unit is used to use a preset image recognition model to determine the gesture key points corresponding to the gesture data as gesture feature information.

[0250] Optionally, the feature information includes eye movement feature information corresponding to the eye movement data;

[0251] The feature information extraction submodule includes:

[0252] An eye movement position determination unit, used to determine the eyeball position and the pupil position based on the eye movement data;

[0253] The eye movement feature extraction unit is used to determine the user's line of sight direction as eye movement feature information based on the eyeball position and the pupil position.

[0254] Optionally, the feature information includes expression feature information corresponding to the facial expression data:

[0255] The feature information extraction submodule includes:

[0256] An expression key information extraction unit, used to extract expression key information from the facial expression data;

[0257] The expression feature extraction unit is used to determine the expression recognition result as the expression feature information based on the expression key information.

[0258] In addition, an embodiment of the present invention further provides an electronic device, such as Figure 3 As shown, it includes a processor 301, a communication interface 302, a memory 303 and a communication bus 304, wherein the processor 301, the communication interface 302, and the memory 303 communicate with each other through the communication bus 304.

[0259] Memory 303, used for storing computer programs;

[0260] The processor 301 is used to execute the program stored in the memory 303 to implement the following steps:

[0261] Collecting multimodal interaction data of the user, wherein the multimodal interaction data includes at least two types of interaction data among voice data, touch data, gesture data, eye movement data, and facial expression data;

[0262] Extracting feature coding information from the multimodal interaction data through the interaction recognition model, and determining interaction recognition information based on the feature coding information;

[0263] Based on the interaction identification information, interaction feedback information is generated and fed back to the user.

[0264] Optionally, the step of extracting feature coding information from the multimodal interaction data by using the interaction recognition model, and determining interaction recognition information based on the feature coding information includes:

[0265] Extracting feature information corresponding to the interaction data respectively through the interaction recognition model;

[0266] Performing feature encoding processing on the feature information respectively to obtain at least two feature encoding information; fusing the at least two feature encodings based on a preset attention weight matrix to obtain fused feature information;

[0267] Based on the fused feature information, interaction identification information is determined.

[0268] Optionally, the feature information includes speech feature information corresponding to the language data;

[0269] The step of respectively extracting feature information corresponding to the interaction data through the interaction recognition model comprises:

[0270] Extract Mel-frequency cepstral coefficients from speech data;

[0271] Extracting convolution feature information from the Mel-frequency cepstral coefficients based on a preset convolutional neural network;

[0272] The convolution feature information is input into a preset recurrent neural network, and a speech recognition result output by the recurrent neural network is obtained as speech feature information.

[0273] Optionally, the feature information includes touch feature information corresponding to the touch data;

[0274] The step of respectively extracting feature information corresponding to the interaction data through the interaction recognition model comprises:

[0275] Smoothing the touch data to obtain smooth trajectory data;

[0276] At least one of a speed feature, a curvature feature, a direction feature, and a shape feature is extracted from the smooth trajectory data as touch feature information.

[0277] Optionally, the characteristic information includes gesture characteristic information corresponding to the gesture data;

[0278] The step of respectively extracting feature information corresponding to the interaction data through the interaction recognition model comprises:

[0279] A preset image recognition model is used to determine the gesture key points corresponding to the gesture data as gesture feature information.

[0280] Optionally, the feature information includes eye movement feature information corresponding to the eye movement data;

[0281] The step of respectively extracting feature information corresponding to the interaction data through the interaction recognition model comprises:

[0282] Based on the eye movement data, determining eyeball positions and pupil positions;

[0283] Based on the eyeball position and the pupil position, the user's sight direction is determined as eye movement feature information.

[0284] Optionally, the feature information includes expression feature information corresponding to the facial expression data:

[0285] The step of respectively extracting feature information corresponding to the interaction data through the interaction recognition model comprises:

[0286] Extracting key expression information from the facial expression data;

[0287] Based on the expression key information, an expression recognition result is determined as the expression feature information.

[0288] The communication bus mentioned in the above terminal can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0289] The communication interface is used for communication between the above terminal and other devices.

[0290] The memory may include a random access memory (RAM) or a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.

[0291] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0292] like Figure 4 As shown, in another embodiment provided by the present invention, a computer-readable storage medium 401 is also provided, in which instructions are stored. When the computer-readable storage medium is run on a computer, the computer executes the hybrid interaction method described in the above embodiment.

[0293] In another embodiment of the present invention, a computer program product including instructions is provided. When the computer program product is run on a computer, the computer executes the hybrid interaction method described in the above embodiment.

[0294] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented by software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website site, computer, server or data center to another website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk Solid State Disk (SSD)), etc.

[0295] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.

[0296] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and reference can be made to the corresponding part of the method embodiment for the relevant content.

[0297] The above are only the preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention are all included in the protection scope of the present invention.

Claims

1. A hybrid interaction method, characterized in that: The method comprises: Collecting multimodal interaction data of the user, wherein the multimodal interaction data includes at least two types of interaction data among voice data, touch data, gesture data, eye movement data, and facial expression data; Extracting feature coding information from the multimodal interaction data through the interaction recognition model, and determining interaction recognition information based on the feature coding information; Based on the interaction identification information, interaction feedback information is generated and fed back to the user.

2. The method according to claim 1, characterized in that The step of extracting feature coding information from the multimodal interaction data through the interaction recognition model and determining interaction recognition information based on the feature coding information includes: Extracting feature information corresponding to the interaction data respectively through the interaction recognition model; Performing feature coding processing on the feature information respectively to obtain at least two types of feature coding information; Fusing the at least two feature codes based on a preset attention weight matrix to obtain fused feature information; Based on the fused feature information, interaction identification information is determined.

3. The method according to claim 2, characterized in that The feature information includes speech feature information corresponding to the language data; The step of respectively extracting feature information corresponding to the interaction data through the interaction recognition model comprises: Extract Mel-frequency cepstral coefficients from speech data; Extracting convolution feature information from the Mel-frequency cepstral coefficients based on a preset convolutional neural network; The convolution feature information is input into a preset recurrent neural network, and a speech recognition result output by the recurrent neural network is obtained as speech feature information.

4. The method according to claim 2, characterized in that: The characteristic information includes touch characteristic information corresponding to the touch data; The step of respectively extracting feature information corresponding to the interaction data through the interaction recognition model comprises: Smoothing the touch data to obtain smooth trajectory data; At least one of a speed feature, a curvature feature, a direction feature, and a shape feature is extracted from the smooth trajectory data as touch feature information.

5. The method according to claim 2, characterized in that: The characteristic information includes gesture characteristic information corresponding to the gesture data; The step of respectively extracting feature information corresponding to the interaction data through the interaction recognition model comprises: A preset image recognition model is used to determine the gesture key points corresponding to the gesture data as gesture feature information.

6. The method according to claim 2, characterized in that The feature information includes eye movement feature information corresponding to the eye movement data; The step of respectively extracting feature information corresponding to the interaction data through the interaction recognition model comprises: Based on the eye movement data, determining eyeball positions and pupil positions; Based on the eyeball position and the pupil position, the user's sight direction is determined as eye movement feature information.

7. The method according to claim 2, characterized in that The feature information includes expression feature information corresponding to the facial expression data: The step of respectively extracting feature information corresponding to the interaction data through the interaction recognition model comprises: Extracting key expression information from the facial expression data; Based on the expression key information, an expression recognition result is determined as the expression feature information.

8. A hybrid interactive device, characterized in that: The device comprises: A data collection module, used to collect multimodal interaction data of the user, wherein the multimodal interaction data includes at least two types of interaction data among voice data, touch data, gesture data, eye movement data, and facial expression data; An interaction identification module, configured to extract feature coding information from the multimodal interaction data through the interaction identification model, and determine interaction identification information based on the feature coding information; The feedback module is used to generate interaction feedback information based on the interaction identification information and provide feedback to the user.

9. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; The memory is used to store computer programs; The processor is used to implement the method according to any one of claims 1 to 7 when executing the program stored in the memory.

10. A computer-readable medium having instructions stored thereon, which, when executed by one or more processors, cause the processors to perform the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • User desktop interaction control device and method based on AI

    CN120371677A