Multi-mode audio-visual media remote interaction system
Through the multimodal audio-visual media remote interaction system, the visual, audio and body movement data are integrated to generate the optimal interaction strategy, solving the one-sided user experience caused by single modal evaluation, and achieving high accuracy and intelligent interaction optimization.
Patent Information
- Application Number
- CN202510353481.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-03-25
AI Technical Summary
The existing remote interactive systems rely on single modal data to cause user experience evaluation to be one-sided, affecting the system optimization direction and effect.
A multimodal audio-visual media remote interaction system is adopted to integrate visual, audio and body movement data, generate optimal interaction strategies through deep learning models, and optimize them with user behavior pattern matrix and interaction effect evaluation values.
It realizes multi-dimensional quantitative analysis of user interaction behavior, improves the accuracy and reliability of evaluation results, dynamically adapts to user needs and scenario changes, and enhances user experience and system intelligence.
Smart Images

Figure CN120295463A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of interaction technologies, and particularly to a multi-modal audio-visual media remote interaction system. Background Art
[0002] With the rapid development of information technology, remote interaction systems have been widely applied in many fields such as education, medical care, entertainment, and office work. Traditional remote interaction systems usually rely only on single-modal data (such as audio or video), and this limitation makes it difficult for the interaction effect of the system to meet the diverse needs of users. For example, in a remote meeting, relying solely on voice communication may not accurately convey emotional information; while in remote teaching, the lack of capture of body movements may lead to insufficient interactivity.
[0003] In recent years, the rise of multi-modal technologies has brought new opportunities to remote interaction systems. A multi-modal audio-visual media remote interaction system can more comprehensively capture the interaction behaviors of users by integrating visual data (such as facial expressions, eye directions), audio data (such as speech intonation, speech rate), and body movement data (such as gestures, postures), thereby improving the intelligence level and user experience of the system.
[0004] It is difficult to comprehensively reflect the actual interaction experience of users based on a single index, mainly because the remote interaction process involves many complex factors, and a single index can only be evaluated from a certain specific dimension and cannot cover the comprehensive needs of users at the emotional, behavioral, and cognitive levels. In addition, the interaction experience of users is often jointly affected by multiple modal data. For example, in a remote meeting, clear speech may seem monotonous due to the lack of body movement feedback, and high-quality video may lose its attractiveness due to plain speech intonation. Therefore, relying solely on a single index for evaluation will lead to a one-sided understanding of the user experience, which in turn affects the optimization direction and final effect of the system. Summary of the Invention
[0005] The purpose of the present invention is to provide a multi-modal audio-visual media remote interaction system to solve the problem that relying solely on a single index for evaluation will lead to a one-sided understanding of the user experience, which in turn affects the optimization direction and final effect of the system.
[0006] To achieve the above purpose, the present invention provides the following technical solution: A multi-modal audio-visual media remote interaction system, including the following modules:
[0007] A visual data acquisition module: used to acquire the real-time visual data of the user;
[0008] An audio data acquisition module: used to acquire the real-time audio data of the user;
[0009] A body movement data acquisition module: used to acquire the real-time body movement data of the user;
[0010] Preprocessing module: Connected to the visual data acquisition module, audio data acquisition module, and limb movement data acquisition module, and used to preprocess real-time visual data, real-time audio data, and real-time limb movement data;
[0011] Initial interaction strategy generation module: Connected to the preprocessing module, and used to input the preprocessed real-time visual data, real-time audio data, and real-time limb movement data into a pre-trained model to generate an initial interaction strategy;
[0012] Optimized interaction strategy generation module: Used to obtain a user behavior pattern matrix and an interaction effect evaluation adjustment amount, calculate a comprehensive matrix based on the user behavior pattern matrix and the interaction effect evaluation adjustment amount, and optimize the initial interaction strategy according to the comprehensive matrix to generate an optimized interaction strategy.
[0013] Furthermore, the optimized interaction strategy generation module includes the following sub-modules:
[0014] User behavior pattern matrix construction sub-module: Used to construct a user behavior pattern matrix through the user's historical visual data, historical audio data, and historical limb movement data;
[0015] Interaction effect evaluation value prediction sub-module: Used to construct a deep learning model, and train the deep learning model to obtain the optimal model parameters; Obtain the real-time user behavior pattern matrix and input it into the deep learning model for prediction, and calculate the interaction effect evaluation value based on the optimal model parameters, the real-time facial expression feature value, the real-time speech intonation feature value, and the real-time limb movement feature value extracted from the real-time user behavior pattern matrix;
[0016] Preset sub-module: Used to preset an ideal interaction effect evaluation value and an optimization threshold;
[0017] Interaction effect judgment sub-module: Calculate the interaction effect evaluation difference according to the interaction effect evaluation value and the ideal interaction effect evaluation value; Compare the interaction effect evaluation difference with the optimization threshold. If the interaction effect evaluation value is less than the optimization threshold, generate an optimized interaction strategy according to the interaction effect evaluation value; If the interaction effect evaluation value is greater than or equal to the optimization threshold, transmit the interaction effect evaluation value to the interaction effect evaluation adjustment amount calculation sub-module;
[0018] Interaction effect evaluation adjustment amount calculation sub-module: Used to calculate the sensitivity values of the real-time facial expression feature value, the real-time speech intonation feature value, and the real-time limb movement feature value to the interaction effect evaluation value according to the partial derivative formula, and construct a sensitivity matrix, and calculate the interaction effect evaluation adjustment amount based on the sensitivity matrix and the interaction effect evaluation difference;
[0019] Comprehensive matrix construction sub-module: used to obtain a comprehensive matrix based on the real-time user behavior pattern matrix and the interaction effect evaluation adjustment amount, and input it into the deep learning model to predict the interaction effect evaluation value again; transmit the currently predicted interaction effect evaluation value to the interaction effect evaluation adjustment amount calculation sub-module until the currently predicted interaction effect evaluation value is less than the optimization threshold, then stop the prediction;
[0020] Optimal interaction strategy generation sub-module: used to receive the optimized interaction strategy and generate candidate interaction strategies, calculate the final scores for the optimized interaction strategy and the candidate interaction strategies, and screen out the optimal interaction strategy based on the final scores.
[0021] Further, the optimal interaction strategy generation sub-module includes the following units:
[0022] Candidate interaction strategy generation unit: used to calculate the cosine similarity between the comprehensive matrix or the real-time user behavior pattern matrix corresponding to the optimized interaction strategy and the user behavior pattern matrix; use the user behavior pattern matrix with a cosine similarity greater than k as the candidate matrix, where 0 < k < 1; calculate the feature mean based on the features of the candidate matrix and add the feature mean to the optimized interaction strategy to obtain the candidate interaction strategy;
[0023] Priority weight calculation unit; used to calculate the priority weight of the candidate interaction strategy according to the cosine similarity;
[0024] Final score calculation unit: used to calculate the final scores of the candidate interaction strategy and the optimized interaction strategy according to the cosine similarity and the priority weight;
[0025] Optimal interaction strategy screening unit: used to take the optimized interaction strategy or the candidate interaction strategy with the highest final score as the optimal interaction strategy.
[0026] Further, the user behavior pattern matrix includes a facial expression feature vector, a speech intonation feature vector, and a body movement feature vector, and the facial expression feature vector, the speech intonation feature vector, and the body movement feature vector form a 3x3 user behavior pattern matrix.
[0027] Further, construct a deep learning model and obtain the optimal model parameters through training of the deep learning model, including: using the user behavior pattern matrix as the training set of the deep learning model, using the historical interaction effect evaluation value of the user behavior pattern matrix as the prediction label, and the input of the deep learning model being the real-time user behavior pattern matrix; constructing the loss function of the deep learning model based on the absolute error and relative error between the historical interaction effect evaluation value and the predicted interaction effect evaluation value, and using the gradient descent method to minimize the loss function to obtain the optimal model parameters, thereby obtaining the trained deep learning model.
[0028] Further, the calculation formula for the interaction effect evaluation value is:
[0029]
[0030] Among them, E represents the interactive effect evaluation value, f face,i represents the i-th facial expression feature value in the real-time user behavior pattern matrix, f voice,i represents the i-th speech intonation feature value in the real-time user behavior pattern matrix, f gesture,i represents the i-th body movement feature value in the real-time user behavior pattern matrix, a, b, c, and d respectively represent the optimal parameters of the model, n represents the total number of feature values in the real-time user behavior pattern matrix, log represents the logarithmic function, and e represents the exponential decay function.
[0031] Furthermore, the calculation formula for the adjustment amount of the interactive effect evaluation is:
[0032] ΔM = γ·ΔE·F;
[0033] Among them, ΔM represents the adjustment amount of the interactive effect evaluation, γ represents the adjustment coefficient, ΔE represents the difference in the interactive effect evaluation, and F represents the sensitivity matrix.
[0034] Technical effects and advantages of the present invention:
[0035] The present invention gets rid of the way of realizing interaction by preset rules, but through integrating visual data, audio data and body movement data, realizes multi-dimensional quantitative analysis of user interaction behaviors, and generates an interactive effect evaluation value, thereby providing a scientific basis for the formulation of the optimal interaction strategy. The fusion of multi-modal data can comprehensively capture the interaction state of users, avoid information loss or deviation that may be brought by a single data source, and significantly improve the accuracy and reliability of the evaluation results; secondly, converting complex behavior data into a quantifiable interactive effect evaluation value is convenient for the system to quickly respond to and optimize the adjustment of user needs, improving the intelligence level of the interaction process; finally, the optimal interaction strategy generated based on the evaluation value can dynamically adapt to the needs and scenario changes of different users, enhance the user experience, and also provide strong support for the design of personalized services, ultimately promoting the performance of the human-computer interaction system to a higher level. Description of the Drawings
[0036] Figure 1 It is a schematic structural diagram of the multi-modal audio-visual media remote interaction system provided in this embodiment. Detailed Embodiments
[0037] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0038] This embodiment has a high degree of flexibility and scalability and can be applied to multiple fields: in the field of telemedicine, doctors can conduct remote video consultations with patients through this system and use various modal information for diagnosis; in the field of education, teachers can provide remote teaching and real-time interaction for students through this system to improve the teaching effect and learning experience, and it can also be used in remote meetings. In short, the multi-modal audio-visual media remote interaction system is becoming an important development direction in the future remote interaction field with its rich interaction methods, high degree of intelligence, and wide application fields.
[0039] This embodiment integrates visual data, audio data, and body movement data to achieve multi-dimensional quantitative analysis of user interaction behaviors and generate an interaction effect evaluation value, thereby providing a scientific basis for the formulation of the optimal interaction strategy. The beneficial effects of this method are as follows: First, the fusion of multi-modal data can comprehensively capture the user's interaction state, avoid information loss or deviation that may be brought by a single data source, and significantly improve the accuracy and reliability of the evaluation results; Second, converting complex behavior data into a quantifiable interaction effect evaluation value facilitates the system to quickly respond to and optimize user needs, improving the intelligence level of the interaction process; Finally, the optimal interaction strategy generated based on the evaluation value can dynamically adapt to the needs and scenario changes of different users, enhance the user experience, and also provide strong support for the design of personalized services, ultimately promoting the performance of the human-computer interaction system to a higher level.
[0040] As Figure 1 shown, this embodiment discloses a multi-modal audio-visual media remote interaction system, including the following modules:
[0041] Visual data acquisition module: used to acquire the user's real-time visual data.
[0042] Use a camera installed on the user side (such as a web camera, the front camera of a mobile phone) to capture the user's RGB image. The RGB image includes visual information such as facial expressions, eye movement changes, and head movements. The camera can be a high-resolution RGB camera or a depth camera (such as an infrared camera) to support more complex visual feature extraction.
[0043] The real-time visual data includes the coordinates of real-time acquired facial key points, head rotation angle, eye movement tracking data, facial expression classification labels, etc.
[0044] Among them, based on a pre-trained model, user facial key points are extracted (facial key points refer to the two-dimensional or three-dimensional coordinates of specific parts on a person's face (such as eyes, nose, mouth, eyebrows, etc.) in an image), so as to judge the user's expression state according to the relative position changes of the key points.
[0045] The head rotation angle describes the rotation of the user's head relative to a reference plane, and judges whether the user is looking directly at the system or deviating from the line of sight, so as to judge the user's concentration or emotional state. It usually includes angles in three dimensions:
[0046] Pitch angle: The angle of the head swinging up and down;
[0047] Yaw angle: The angle of the head turning left and right;
[0048] Roll angle: The angle of the head tilting.
[0049] Eye movement tracking data records the movement trajectory and fixation point position of the user's eyeballs, and is used to judge the user's attention distribution and points of interest, and evaluate the user's fatigue level through the blink frequency and pupil size. It mainly includes:
[0050] Fixation point coordinates: The (x, y) coordinates of the user's line of sight focus on the screen. Example: The fixation point is (300, 200), indicating that the user is paying attention to a certain area of the screen.
[0051] Pupil size: Reflects the user's attention concentration or emotional state. Example: The pupil diameter is 4mm.
[0052] Blink frequency: The number of blinks per unit time. Example: The blink frequency is 20 times per minute.
[0053] Audio data acquisition module: Used to acquire the user's real-time audio data.
[0054] The user's voice signal is captured in real time through the microphone on the user side, and the existing automatic speech recognition (ASR) technology is used to convert the audio signal into text content. Therefore, the real-time audio data includes the real-time obtained speech text, acoustic features, emotion labels, background noise information, etc.
[0055] Among them, signal processing techniques (such as short-time Fourier transform STFT, Mel frequency cepstral coefficients MFCC) are used to extract the acoustic features of the speech signal, the speech signal is segmented into short-time frames, and the speech rate, volume and pitch features of each frame are calculated.
[0056] Noise detection algorithms (such as Wiener filtering, deep learning denoising models) are used to analyze the type and intensity of background noise.
[0057] Limb movement data acquisition module: Used to acquire the user's real-time limb movement data.
[0058] Capture the three-dimensional spatial information of the user through a depth camera (such as Microsoft Kinect, Intel RealSense). Real-time limb movement data is obtained through accelerometers and gyro sensors worn on the user's body. The real-time limb movement data includes the coordinates of the skeletal key points obtained in real time, action labels, the movement trajectories of limb parts, the movement speeds and accelerations of limb parts, etc.
[0059] Among them, a deep learning-based human pose estimation algorithm (such as OpenPose, AlphaPose, HRNet) is used to extract the coordinates of the skeletal key points.
[0060] Use an action classification model based on time series analysis (such as LSTM, Transformer) to classify the user's limb movements and obtain the action categories.
[0061] Use the Kalman filter or particle filter algorithm to smooth and track the user's limb movement trajectories.
[0062] Use numerical differentiation methods to calculate speed and acceleration.
[0063] Preprocessing module: Connected to the visual data acquisition module, audio data acquisition module, and limb movement data acquisition module, and used to preprocess real-time visual data, real-time audio data, and real-time limb movement data.
[0064] Preprocessing process of real-time visual data:
[0065] Denoising process: Use median filtering or Gaussian filtering to remove random noise in the image.
[0066] Lighting correction: Adjust the image brightness and contrast to ensure clear images under different lighting conditions.
[0067] Size adjustment: Scale the image to a unified resolution to reduce the amount of calculation.
[0068] Depth data calibration: Calibrate the depth image to eliminate sensor errors.
[0069] Preprocessing process of real-time audio data:
[0070] Noise reduction process: Use Wiener filtering or a deep learning noise reduction model to remove background noise.
[0071] Volume normalization: Adjust the volume of the audio signal to a unified range.
[0072] The preprocessing process of real-time limb movement data is:
[0073] Coordinate data processing: Downsample and denoise the coordinate data.
[0074] Timestamp alignment: Align the acceleration and angular velocity data with the timestamp of the camera.
[0075] Motion trajectory smoothing: Use the Kalman filter or particle filter algorithm to smooth the motion trajectory of the limb parts.
[0076] Initial interaction strategy generation module: Connected to the preprocessing module, it is used to input the preprocessed real-time visual data, real-time audio data, and real-time limb movement data into the pre-trained model to generate an initial interaction strategy.
[0077] The pre-trained model is a multi-classification problem. Use historical visual data, historical audio data, and historical limb movement data as the training set of the pre-trained model to obtain a pre-trained model. The predicted labels are: manually calibrated facial expression types and key point coordinates (eye gaze horizontal coordinate and eye gaze vertical coordinate), manually calibrated emotion labels, speech rate, volume, and manually calibrated action labels, and the motion trajectory of the limb parts. Input the preprocessed real-time visual data, real-time audio data, and real-time limb movement data into the pre-trained model to obtain three prediction results and fuse them to generate an initial interaction strategy.
[0078] Optimized interaction strategy generation module: Used to obtain the user behavior pattern matrix and the interaction effect evaluation adjustment amount, calculate the comprehensive matrix based on the user behavior pattern matrix and the interaction effect evaluation adjustment amount, and optimize the initial interaction strategy according to the comprehensive matrix to generate an optimized interaction strategy.
[0079] The optimized interaction strategy generation module includes the following sub-modules:
[0080] User behavior pattern matrix construction sub-module: Used to construct the user behavior pattern matrix through the user's historical visual data, historical audio data, and historical limb movement data; the user behavior pattern matrix includes facial expression feature vectors, speech intonation feature vectors, and limb movement feature vectors, and the facial expression feature vectors, speech intonation feature vectors, and limb movement feature vectors form a 3x3 user behavior pattern matrix.
[0081] Interaction effect evaluation value prediction sub-module: Used to construct a deep learning model, train the optimal parameters of the deep learning model; obtain the real-time user behavior pattern matrix and input it into the deep learning model for prediction, and calculate the interaction effect evaluation value based on the optimal model parameters, real-time facial expression feature values, real-time speech intonation feature values, and real-time limb movement feature values extracted from the real-time user behavior pattern matrix;
[0082] Build a deep learning model and obtain the optimal model parameters through training, including: using the user behavior pattern matrix as the training set of the deep learning model, manually calibrating the historical interaction effect evaluation values of the user behavior pattern matrix as prediction labels, and using the real-time user behavior pattern matrix as the input of the deep learning model; constructing the loss function of the deep learning model based on the absolute error and relative error between the historical interaction effect evaluation value and the predicted interaction effect evaluation value, and using the gradient descent method to minimize the loss function to obtain the optimal model parameters, thus obtaining the trained deep learning model.
[0083] Construct the loss function for the number of samples in the real-time user behavior pattern matrix as follows:
[0084]
[0085] where L is the loss function, is the absolute error, ω1 is the weight coefficient of the absolute error, represents the relative error, ω2 is the weight coefficient of the relative error, y j is the predicted value of the j-th sample (i.e., the interaction effect evaluation value predicted by the deep learning model), is the true value of the j-th sample (i.e., the prediction label in the training set).
[0086] The absolute error ensures that the deep learning model pays attention to the prediction deviation of each sample, and the relative error further ensures that the model can maintain high accuracy even when predicting small values. This dual consideration enables the model to optimize the prediction performance more evenly during training. Especially when dealing with data with a large dynamic range, it can effectively avoid the model bias problem caused by a single error metric.
[0087] Based on this loss function, use the gradient descent method to minimize the loss function to obtain the optimal model parameters, and apply the optimal model parameters to the calculation process of the interaction effect evaluation value:
[0088]
[0089] where E represents the interaction effect evaluation value, f face,i represents the i-th facial expression feature value in the real-time user behavior pattern matrix, f voice,i represents the i-th voice intonation feature value in the real-time user behavior pattern matrix, f gesture,i represents the i-th body movement feature value in the real-time user behavior pattern matrix, a, b, c, and d respectively represent the optimal model parameters, n represents the total number of feature values in the real-time user behavior pattern matrix, log represents the logarithmic function, and e represents the exponential decay function.
[0090] Since the data in the above formula contains different units, before prediction, each data is standardized and normalized to obtain data with the same dimension, thus eliminating the influence of units.
[0091] Using the optimal parameters of the deep learning model, combined with the facial expression feature values, speech intonation feature values, and body movement feature values to calculate the interaction effect evaluation value, it is possible to achieve multi-dimensional quantification and accurate description of the user state. The optimal parameters ensure the rationality of the eigenvalue weight distribution, enabling different modalities of features to exert their due influence in the evaluation. By fusing and calculating multiple eigenvalue, the complexity and diversity of human interaction behaviors are fully considered, resulting in a more representative and comprehensive evaluation result. This method not only effectively reflects the subtle changes in the interaction process but also provides a reliable numerical basis for understanding user emotions and behaviors, significantly enhancing the practical application value of the evaluation value.
[0092] Preset sub-module: used to preset an ideal evaluation value of the interaction effect and an optimization threshold.
[0093] Collect the historical interaction data set, which includes several samples. Each sample includes a user behavior pattern matrix and the corresponding interaction effect evaluation value. Arrange the interaction effect evaluation values from low to high, and define the samples corresponding to the last 5% of the interaction effect evaluation values as high-score samples. Take the mean value of the interaction effect evaluation values of the high-score samples as the ideal evaluation value of the interaction effect. The optimization threshold is calculated based on the ideal evaluation value of the interaction effect:
[0094] P = ideal×(1 - c), where P is the optimization threshold, ideal is the ideal evaluation value of the interaction effect, and c is the tolerance coefficient, and c is preferably 0.1.
[0095] Interaction effect judgment sub-module: calculate the interaction effect evaluation difference |E - ideal| based on the interaction effect evaluation value and the ideal evaluation value of the interaction effect; compare the interaction effect evaluation difference with the optimization threshold. If the interaction effect evaluation value is less than the optimization threshold, generate an optimized interaction strategy based on the interaction effect evaluation value; if the interaction effect evaluation value is greater than or equal to the optimization threshold, transmit the interaction effect evaluation value to the interaction effect evaluation adjustment amount calculation sub-module.
[0096] Based on the interaction effect evaluation value, the ideal evaluation value of the interaction effect, and the optimization threshold, dynamically determine whether the current interaction state needs to be optimized and generate corresponding optimization strategies or adjustment amounts.
[0097] The optimization strategy is specific improvement measures generated based on the judgment of the optimization threshold on the basis of the initial strategy:
[0098] (1) Visual-related optimization
[0099] Facial expression adjustment: According to facial expression tags (such as smiling, frowning, etc.), prompt the speaker to adjust their emotional state.
[0100] Eye gaze direction guidance: Based on the horizontal and vertical coordinate data of the eyes, remind participants to keep their eyes focused on the screen or a specific speaker.
[0101] (2) Audio-related optimization
[0102] Speech speed and volume adjustment: According to the speech speed and volume data, suggest that the speaker adjust their expression style.
[0103] Emotion label feedback: Combine the manually calibrated emotion labels to provide emotion management suggestions for the speaker.
[0104] (3) Body movement optimization
[0105] Action label guidance: According to action labels (such as nodding, shaking the head, waving, etc.), encourage participants to enhance interaction through body language.
[0106] Motion trajectory analysis: Track the motion trajectory of body parts to ensure that the actions are natural and in line with the meeting context.
[0107] Interaction effect evaluation adjustment amount calculation sub-module: Used to calculate the sensitivity values of real-time facial expression feature values, real-time speech intonation feature values, and real-time body movement feature values to the interaction effect evaluation value according to the partial derivative formula, and construct a sensitivity matrix. Based on the sensitivity matrix and the interaction effect evaluation difference, calculate the interaction effect evaluation adjustment amount;
[0108] The sensitivity value is:
[0109]
[0110] Is the sensitivity value of the i-th facial expression feature value in the real-time user behavior pattern matrix to the interaction effect evaluation value, Is the sensitivity value of the i-th speech intonation feature value in the real-time user behavior pattern matrix to the interaction effect evaluation value, Is the sensitivity value of the i-th body movement feature value in the real-time user behavior pattern matrix to the interaction effect evaluation value, Represents the partial derivative.
[0111] Combine the sensitivity values of each feature value in the real-time user behavior pattern matrix to the interaction effect evaluation value into a sensitivity matrix.
[0112]
[0113] Where, ΔM represents the adjustment amount of the interaction effect evaluation, γ represents the adjustment coefficient, ΔE represents the difference of the interaction effect evaluation, F represents the sensitivity matrix, max represents the maximum operation, and max(F) represents the maximum eigenvalue in the sensitivity matrix.
[0114] By combining the difference of the interaction effect evaluation with the sensitivity matrix and adjusting it using the adjustment coefficient, the system can dynamically determine the optimal adjustment amount. It can not only quickly identify the gap between the current interaction state and the ideal state, but also perform targeted optimization according to the different sensitivities of each eigenvalue to the interaction effect. For example, when a certain eigenvalue has a greater impact on the interaction effect, its corresponding sensitivity value is higher, and the system will focus more on adjusting this feature to achieve a better interaction effect. This adjustment strategy based on sensitivity and difference enables the system to flexibly respond in a complex and changing interaction environment, ensuring that each adjustment is more accurate and effective, thereby greatly improving the overall interaction quality and user satisfaction.
[0115] Comprehensive matrix construction sub-module: used to obtain a comprehensive matrix based on the real-time user behavior pattern matrix and the adjustment amount of the interaction effect evaluation, and input it into the deep learning model to predict the interaction effect evaluation value again; transmit the currently predicted interaction effect evaluation value to the interaction effect evaluation adjustment amount calculation sub-module until the currently predicted interaction effect evaluation value is less than the optimization threshold, then stop the prediction.
[0116] Optimal interaction strategy generation sub-module: used to receive the optimized interaction strategy and generate candidate interaction strategies, calculate the final scores for the optimized interaction strategy and the candidate interaction strategies, and screen out the optimal interaction strategy based on the final scores.
[0117] The optimal interaction strategy generation sub-module includes the following units:
[0118] Candidate interaction strategy generation unit: used to calculate the cosine similarity between the comprehensive matrix corresponding to the optimized interaction strategy or the real-time user behavior pattern matrix and the user behavior pattern matrix; use the user behavior pattern matrix with a cosine similarity greater than k as the candidate matrix, 0 < k < 1, preferably, k is greater than or equal to 0.7; calculate the feature mean based on the features of the candidate matrix, and add the feature mean to the optimized interaction strategy to obtain the candidate interaction strategy.
[0119] Where, the mean of each dimension feature in the candidate matrix is added to the eigenvalue corresponding to the optimized strategy, thereby expanding the expression ability of the optimized strategy and making up for potential deficiencies. There are at least two candidate strategies. After the feature mean adjusts the optimized interaction strategy, multiple candidate interaction strategies are obtained.
[0120] Priority weight calculation unit; used to calculate the priority weights of the candidate interaction strategies according to the cosine similarity.
[0121] f(s) = s2 ;
[0122] Among them, f(s) represents the priority weight, and s represents the cosine similarity between the comprehensive matrix corresponding to the optimized interaction strategy or the real-time user behavior pattern matrix and the user behavior pattern matrix.
[0123] Final score calculation unit: used to calculate the final score of the candidate interaction strategy according to the cosine similarity and the priority weight.
[0124] S = w1×s + w2×f(s), where S is the final score, w1 represents the influence coefficient of the cosine similarity, w1 = 0.1s, w2 = 0.1f(s), and w2 represents the influence coefficient of the priority weight.
[0125] Optimal interaction strategy screening unit: used to take the optimized interaction strategy or candidate interaction strategy with the highest final score as the optimal interaction strategy.
[0126] During the interaction strategy optimization process, although the optimized interaction strategy is generated based on the predicted evaluation value, the cosine similarity between it and the target strategy is not necessarily high. This may be due to the fact that the optimized strategy fails to fully cover or match the feature requirements of the target strategy. Therefore, it is an effective means to supplement the optimized strategy by introducing the candidate matrix and its mean. Specifically, the feature mean of the candidate matrix can be added to the optimized strategy in a weighted fusion manner, and the mean of each dimension feature in the candidate matrix is proportionally superimposed on the feature value corresponding to the optimized strategy, so as to expand the expression ability of the optimized strategy and make up for potential deficiencies. When calculating the score, this fused strategy can more comprehensively reflect the importance of multi-dimensional features, thereby improving the consistency with the target strategy. The significance of finally selecting the optimal interaction strategy is that it not only combines the advantages of the optimized strategy, but also enhances the adaptability and robustness to the target requirements by introducing the candidate matrix mean, ensuring that the generated strategy can achieve a higher interaction effect evaluation value in practical applications and better meet the user needs and system goals.
[0127] As described above, it is only the preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent replacements or changes, and all should be covered by the protection scope of the present invention.
Claims
1. A multi-modal audio-visual media remote interaction system, characterized in that, It includes the following modules: Visual data acquisition module: used to acquire the user's real-time visual data; Audio data acquisition module: used to acquire the user's real-time audio data; Limb movement data acquisition module: used to acquire the user's real-time limb movement data; Preprocessing module: connected to the visual data acquisition module, audio data acquisition module, and limb movement data acquisition module, and used to preprocess the real-time visual data, real-time audio data, and real-time limb movement data; Initial interaction strategy generation module: connected to the preprocessing module, and used to input the preprocessed real-time visual data, real-time audio data, and real-time limb movement data into a pre-trained model to generate an initial interaction strategy; Optimized interaction strategy generation module: used to obtain the user behavior pattern matrix and the interaction effect evaluation adjustment amount, calculate the comprehensive matrix based on the user behavior pattern matrix and the interaction effect evaluation adjustment amount, and optimize the initial interaction strategy according to the comprehensive matrix to generate an optimized interaction strategy.
2. The multimodal audio-visual media remote interaction system according to claim 1, characterized in that: The optimized interaction strategy generation module includes the following sub-modules: User behavior pattern matrix construction sub-module: used to construct the user behavior pattern matrix through the user's historical visual data, historical audio data, and historical limb movement data; Interaction effect evaluation value prediction sub-module: used to construct a deep learning model, train the deep learning model to obtain the optimal model parameters; obtain the real-time user behavior pattern matrix and input it into the deep learning model for prediction, and calculate the interaction effect evaluation value based on the optimal model parameters, the real-time facial expression feature value, real-time speech intonation feature value, and real-time limb movement feature value extracted from the real-time user behavior pattern matrix; Preset sub-module: used to preset an ideal interaction effect evaluation value and an optimization threshold; Interaction effect judgment sub-module: calculate the interaction effect evaluation difference according to the interaction effect evaluation value and the ideal interaction effect evaluation value; compare the interaction effect evaluation difference with the optimization threshold. If the interaction effect evaluation value is less than the optimization threshold, generate an optimized interaction strategy according to the interaction effect evaluation value; if the interaction effect evaluation value is greater than or equal to the optimization threshold, transmit the interaction effect evaluation value to the interaction effect evaluation adjustment amount calculation sub-module; Interaction effect evaluation adjustment amount calculation sub-module: used to calculate the sensitivity values of the real-time facial expression feature value, real-time speech intonation feature value, and real-time limb movement feature value to the interaction effect evaluation value according to the partial derivative formula, construct a sensitivity matrix, and calculate the interaction effect evaluation adjustment amount based on the sensitivity matrix and the interaction effect evaluation difference; Comprehensive matrix construction sub-module: used to obtain the comprehensive matrix according to the real-time user behavior pattern matrix and the interaction effect evaluation adjustment amount, and input it into the deep learning model to predict the interaction effect evaluation value again; transmit the currently predicted interaction effect evaluation value to the interaction effect evaluation adjustment amount calculation sub-module until the currently predicted interaction effect evaluation value is less than the optimization threshold, and then stop the prediction; Optimal interaction strategy generation sub-module: used to receive the optimized interaction strategy and generate a candidate interaction strategy, calculate the final scores for the optimized interaction strategy and the candidate interaction strategy, and select the optimal interaction strategy based on the final scores.
3. The multimodal audiovisual media remote interaction system according to claim 2, characterized in that: The optimal interaction strategy generation sub-module includes the following units: Candidate interaction strategy generation unit: used to calculate the cosine similarity between the comprehensive matrix corresponding to the optimized interaction strategy or the real-time user behavior pattern matrix and the user behavior pattern matrix; Take the user behavior pattern matrix with a cosine similarity greater than k as the candidate matrix, where 0 < k < 1; Calculate the feature mean based on the features of the candidate matrix, and add the feature mean to the optimized interaction strategy to obtain the candidate interaction strategy; Priority weight calculation unit; Used to calculate the priority weight of the candidate interaction strategy according to the cosine similarity; Final score calculation unit: used to calculate the final scores of the candidate interaction strategy and the optimized interaction strategy according to the cosine similarity and the priority weight; Optimal interaction strategy screening unit: used to take the optimized interaction strategy or the candidate interaction strategy with the highest final score as the optimal interaction strategy.
4. The multimodal audiovisual media remote interaction system according to claim 3, characterized in that: The user behavior pattern matrix includes a facial expression feature vector, a speech intonation feature vector, and a body movement feature vector. The facial expression feature vector, the speech intonation feature vector, and the body movement feature vector form a 3x3 user behavior pattern matrix.
5. The multimodal audio-visual media remote interaction system according to claim 4, characterized in that: Build a deep learning model and obtain the optimal model parameters through training of the deep learning model, including: using the user behavior pattern matrix as the training set of the deep learning model, using the historical interaction effect evaluation value of the user behavior pattern matrix as the prediction label, and the input of the deep learning model being the real-time user behavior pattern matrix; constructing the loss function of the deep learning model based on the absolute error and relative error between the historical interaction effect evaluation value and the predicted interaction effect evaluation value, and using the gradient descent method to minimize the loss function to obtain the optimal model parameters, thereby obtaining the trained deep learning model.
6. The multimodal audiovisual media remote interaction system according to claim 5, characterized in that: The calculation formula for the interaction effect evaluation value is: Among them, E represents the interactive effect evaluation value, and f face,i represents the i-th facial expression eigenvalue in the real-time user behavior pattern matrix, and f voice,i represents the i-th speech intonation eigenvalue in the real-time user behavior pattern matrix, and f gesture,i represents the i-th body movement eigenvalue in the real-time user behavior pattern matrix. a, b, c, and d respectively represent the optimal parameters of the model, n represents the total number of eigenvalues in the real-time user behavior pattern matrix, log represents the logarithmic function, and e represents the exponential decay function.
7. The multimodal audiovisual media remote interaction system according to claim 6, characterized in that: The calculation formula for the interaction effect evaluation adjustment amount is: ΔM = γ·ΔE·F; Among them, ΔM represents the interaction effect evaluation adjustment amount, γ represents the adjustment coefficient, ΔE represents the interaction effect evaluation difference, and F represents the sensitivity matrix.
Citation Information
Patent Citations
Visual interaction system based on multiple modes
CN118535023A
Context aware data system using biometric and identifying data
US20250005966A1