Multi-modal human-computer interaction interface and behavior prediction system
Through the collaborative work of multimodal data acquisition and fusion, behavior prediction and system optimization modules, the problem of slow response of multimodal human-computer interaction systems in the existing technology is solved, and a personalized and flexible user experience is achieved.
Patent Information
- Application Number
- CN202510257066.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-05-30
AI Technical Summary
Due to the lack of efficient data fusion and accurate behavior prediction, existing multimodal human-computer interaction systems cannot achieve adaptive adjustments in real-time interaction, resulting in slow response and unable to provide a personalized and flexible user experience.
A multimodal human-computer interaction interface and behavior prediction system was designed. Through the coordinated work of multimodal data acquisition, data synchronization and preprocessing, data fusion, behavior prediction and system optimization modules, accurate prediction of user behavior and dynamic adjustment of interaction methods are achieved.
It realizes more accurate user behavior prediction and more flexible interactive experience, can quickly respond when user behavior changes, and provides a personalized and adaptive user experience.
Smart Images

Figure CN120066277A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence, specifically a multimodal human-computer interaction interface and behavior prediction system. Background Art
[0002] With the rapid development of artificial intelligence technology, human-computer interaction has become an important part of the design of intelligent devices and systems. Especially in application fields such as smart home, virtual reality, and augmented reality, how to achieve more intelligent and personalized interaction methods has become a hot issue in technical research and product development. Traditional human-computer interaction methods usually rely on a single input modality, such as voice, image, or tactile signals. This single-modal interaction method is often restricted by environmental interference or data acquisition accuracy in practical applications, and it is difficult to accurately capture the true intentions of users. Especially in complex interaction scenarios, it is difficult to achieve a flexible and adaptive interaction experience.
[0003] In existing multimodal human-computer interaction systems, the data collected by multiple sensors is usually processed and fused, but there are still some deficiencies in the multimodal data fusion method. Most technologies rely on relatively simple data fusion methods and do not fully consider the complex temporal relationships and non-linear characteristics between multimodal data. In addition, existing behavior prediction methods usually adopt linear models or algorithms based on fixed rules, and these methods have poor adaptability to the dynamic changes of user behavior and are difficult to handle the complexity and diversity of user behavior. Although some technologies attempt to achieve better adaptability through reinforcement learning or deep learning, due to the lack of accurate behavior prediction, these systems are still unable to make efficient adjustments under changing user behaviors.
[0004] There is a major problem in existing technologies regarding multimodal data fusion and behavior prediction: due to the lack of efficient data fusion and accurate behavior prediction, the system cannot achieve adaptive adjustment in real-time interaction. This results in the system being slow to respond when facing user preferences or behavior changes and being difficult to provide a personalized and flexible user experience. Summary of the Invention
[0005] Aiming at the deficiencies of the existing technology, the present invention provides a multimodal human-computer interaction interface and behavior prediction system, which solves the problem that in the existing technology, due to the lack of accurate behavior prediction and intelligent interaction adjustment, the system is slow to respond when the user behavior changes and cannot provide a personalized and flexible interaction experience.
[0006] To achieve the above objectives, the present invention is realized through the following technical solutions: A multimodal human-computer interaction interface and behavior prediction system, including: A multimodal data acquisition module, used to collect the visual, auditory, and tactile data of users in real-time; A data synchronization and preprocessing module, which is used to synchronize the time of the collected multi-modal data and denoise and correct the errors of the data; A data fusion module, which is used to fuse different modal data and extract the core features of the data; A behavior prediction module, which is used to predict the user's behavior based on the fused data; A system optimization module, which is used to dynamically adjust the interaction mode and functions of the system based on the behavior prediction results to achieve a personalized user experience.
[0007] Preferably, the multi-modal data acquisition module includes: An RGB-D camera, which is used to collect the visual data of the user and provide the three-dimensional image and depth information of the user; A microphone array, which is used to collect the voice data of the user, ensure a high-quality voice signal and perform noise reduction processing; A tactile sensor, which is used to collect the tactile data of the user and detect the contact pressure and displacement of the user in real time; Each sensor works together through a synchronous trigger mechanism to ensure that different modal data is collected at the same time point.
[0008] Preferably, the data synchronization and preprocessing module synchronizes the multi-modal data through an extended Kalman filter, specifically including: Align the time of visual, voice and tactile data; Correct the data error through an extended Kalman filter; Output the optimized synchronized data.
[0009] Preferably, the data synchronization and preprocessing module further optimizes the data through time series analysis, specifically including: Convert the multi-modal data into a time series format; Identify and smooth the outliers in the time series; Output the optimized time series data for subsequent processing.
[0010] Preferably, the data fusion module realizes data fusion through the following steps: Receive and extract the features of each modal data; Use the tensor decomposition method to reduce the dimension of the data and generate a low-rank matrix; Fuse the low-rank matrix to generate a comprehensive feature matrix for the behavior prediction module to use.
[0011] Preferably, the behavior prediction module realizes user behavior prediction through the following steps: Receive the feature matrix from the data fusion module; Use a non-linear model to predict the user behavior trend; Optimize the prediction results and output the prediction results for the system optimization module to adjust.
[0012] Preferably, the system optimization module performs optimization through the following steps: Receive the behavior prediction results; Evaluate the system interaction method based on the prediction results; Dynamically adjust the system parameters through reinforcement learning; Output the optimized system settings.
[0013] Preferably, the system optimization module dynamically adjusts the following interaction parameters according to the behavior prediction results: The density of interface information presentation; The intensity of tactile feedback; The voice response delay.
[0014] Preferably, the system dynamically adjusts at least one of the following interaction parameters or triggers a protection mechanism according to the behavior prediction results: The density of interface information presentation; The intensity of tactile feedback; The voice response delay; Initiate multi-modal authentication for abnormal behaviors; Insert a secondary confirmation process for high-risk operations; Initiate system locking for continuous incorrect behaviors.
[0015] Preferably, the system further includes: An adaptive security protection module for judging the user operation risk according to the behavior prediction results and triggering a protection mechanism: Initiate multi-modal authentication; Insert a secondary confirmation process; Initiate system locking.
[0016] The present invention provides a multi-modal human-computer interaction interface and a behavior prediction system. It has the following beneficial effects: 1. The present invention adopts a multi-modal data fusion technology to efficiently fuse the data from visual, auditory, and tactile sensors, extract core features, and achieve a more accurate user behavior prediction effect. Compared with the existing technologies that only rely on single-modal data, the present invention can solve the problem of insufficient prediction accuracy caused by a single data source and provide a more comprehensive and reliable user behavior analysis.
[0017] 2. By introducing a non - linear dynamics model and combining historical behavior data with real - time data for comprehensive prediction, the present invention realizes the dynamic modeling of user behavior, achieving higher prediction accuracy and adaptability. Compared with existing linear models, this technical solution can better capture the complex non - linear relationships in user behavior and overcome the defect that traditional methods cannot accurately predict user behavior in complex situations.
[0018] 3. By combining reinforcement learning technology, the present invention makes real - time optimization adjustments to system interaction parameters, enabling the system to make adaptive adjustments according to changes in user behavior, achieving the effect of enhancing interaction flexibility and personalization. Compared with the existing solutions for statically adjusting interaction methods, the present invention enables the system to be optimized after each user interaction through an intelligent learning process, solving the problem that traditional methods are slow to respond to changes in user preferences.
[0019] 4. Through the close collaborative work of the behavior prediction module and the system optimization module, the present invention forms a closed - loop feedback mechanism, ensuring that the system can self - adjust according to real - time data, achieving the effect of improving the system's adaptive ability and response speed. Compared with existing independent processing systems, the modular collaborative design of the present invention solves the problem that traditional solutions are insufficient in responding to changes in user needs, greatly improving the intelligence level of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 It is a schematic diagram of the multi - modal data acquisition module of the present invention; Figure 2 It is a schematic diagram of the data synchronization and pre - processing module of the present invention; Figure 3 It is a schematic diagram of the data fusion module of the present invention; Figure 4 It is a schematic diagram of the behavior prediction module of the present invention; Figure 5 It is a schematic diagram of the system optimization module of the present invention; Figure 6 It is a schematic diagram of the dynamic interaction adjustment module of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0021] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0022] Please refer to the attached Figure 1 - attached Figure 6, embodiments of the present invention provide a multimodal human-computer interaction interface and behavior prediction system, including: A multimodal data acquisition module for real-time acquisition of users' visual, auditory, and tactile data; A data synchronization and preprocessing module for time synchronization of the acquired multimodal data and denoising and error correction of the data; A data fusion module for fusing different modal data and extracting the core features of the data; A behavior prediction module for predicting users' behaviors based on the fused data; A system optimization module for dynamically adjusting the interaction mode and functions of the system based on the behavior prediction results to achieve a personalized user experience.
[0023] The multimodal data acquisition module includes: An RGB-D camera for acquiring users' visual data and providing three-dimensional images and depth information of users; A microphone array for acquiring users' voice data, ensuring high-quality voice signals and performing noise reduction processing; A tactile sensor for acquiring users' tactile data and real-time detecting the contact pressure and displacement of users; Each sensor works together through a synchronous triggering mechanism to ensure that data of different modalities are acquired at the same time point.
[0024] The data synchronization and preprocessing module synchronizes the multimodal data through an extended Kalman filter, specifically including: Performing time alignment on visual, voice, and tactile data; Correcting data errors through an extended Kalman filter; Outputting the optimized synchronized data.
[0025] The data synchronization and preprocessing module further optimizes the data through time series analysis, specifically including: Converting the multimodal data into a time series format; Identifying and smoothing outliers in the time series; Outputting the optimized time series data for subsequent processing.
[0026] The data fusion module realizes data fusion through the following steps: Receiving and extracting the features of each modal data; Using tensor decomposition method to reduce the dimension of the data and generate a low-rank matrix; Fusing the low-rank matrix to generate a comprehensive feature matrix for use by the behavior prediction module.
[0027] The behavior prediction module realizes user behavior prediction through the following steps: Receive the feature matrix from the data fusion module; Use a non-linear model to predict the user behavior trend; Optimize the prediction result and output the prediction result for the system optimization module to adjust.
[0028] The system optimization module performs optimization through the following steps: Receive the behavior prediction result; Evaluate the system interaction method based on the prediction result; Dynamically adjust the system parameters through reinforcement learning; Output the optimized system settings.
[0029] The system optimization module dynamically adjusts the following interaction parameters according to the behavior prediction result: The density of interface information presentation; The intensity of tactile feedback; The voice response delay.
[0030] The system dynamically adjusts at least one of the following interaction parameters or triggers a protection mechanism according to the behavior prediction result: The density of interface information presentation; The intensity of tactile feedback; The voice response delay; Initiate multi-modal authentication for abnormal behaviors; Insert a secondary confirmation process for high-risk operations; Initiate system locking for continuous incorrect behaviors.
[0031] The system further includes: An adaptive security protection module, which is used to judge the user operation risk according to the behavior prediction result and trigger a protection mechanism: Initiate multi-modal authentication; Insert a secondary confirmation process; Initiate system locking.
[0032] The multi-modal data acquisition module in this embodiment is used to collect the visual, auditory and tactile data of the user in real time, as the basis for subsequent data processing and analysis. This module consists of multiple sensors, which ensure that the user behavior information in different modalities can be obtained through collaborative work, and provide a rich data source for subsequent synchronization, preprocessing, fusion and prediction.
[0033] In the aforementioned system architecture, the data acquisition module is directly connected to the data synchronization and preprocessing module, and the acquired data is transmitted to this module in real time for further processing. In this embodiment, the data acquisition module includes an RGB-D camera, a microphone array, and a tactile sensor. Each sensor is responsible for acquiring data of a specific modality, and the timing consistency of the data is ensured through a synchronization mechanism. The design of this module aims to provide high-precision multi-modal data support for the system.
[0034] The RGB-D camera is used to acquire the visual data of the user, providing a three-dimensional image and depth information of the user. The camera can not only capture the user's facial expressions and body movements, but also sense the interaction between the user and the environment, thereby providing the system with the user's spatial positioning information.
[0035] In specific implementation, the RGB-D camera usually uses an infrared light source for depth perception and works in coordination with a traditional RGB camera. The output data generated by the camera includes a color image and a depth map. The color image provides the visual information of the user, while the depth map provides the distance information between the user and the object. Through precise depth sensors and image recognition technologies, this information can help the system identify the user's action trajectories, postures, and distances in three-dimensional space.
[0036] Specifically, the output data of the RGB-D camera can be expressed as: ; where represents the pixel value of the RGB image, is the pixel position in the image.
[0037] The depth information can be expressed as: ; where represents the depth value of each pixel in the depth map, is still the coordinate position in the image. Through the joint processing of the RGB and depth maps, the system can obtain more abundant user behavior data.
[0038] The microphone array is used to acquire the user's voice data, and can provide high-quality voice signals, and also has a noise reduction function to ensure that voice signals can be effectively extracted in a complex environment. In some cases, the microphone array can also achieve sound source localization, that is, determine the sound source by measuring the time difference of the sound signals received by multiple microphones.
[0039] In general, a microphone array consists of multiple microphone units, which are distributed in a plane or three-dimensional space. Each microphone can independently receive sound signals. To process the speech signals received by different microphones, the system adopts beamforming technology, which can screen and enhance sounds from different directions according to the characteristics of spatial signals.
[0040] Specifically, the beamforming algorithm can improve the signal quality through the following steps: Perform weighted processing on the input signals of each microphone; Adjust the weighting coefficients according to the direction of the sound source to enhance the signals from a specific direction; Superimpose and analyze the enhanced signals to extract the user's speech content.
[0041] In a certain implementation, the speech data of the microphone array is represented as: ; where, is the weighted speech signal, is the weighting coefficient of the th microphone, is the th microphone's received signal at time , is the number of microphones.
[0042] Tactile sensors are used to collect the user's tactile data and detect the contact pressure and displacement of the user in real time. In practical applications, tactile sensors can monitor information such as the pressure change, the movement of the touch point, and the intensity of the contact during the user's interaction with the device, and these information are crucial for accurately judging the user's intention.
[0043] Tactile sensors generally achieve the sensing function through capacitive, resistive, or piezoelectric technologies. The system can obtain the following data through these sensors: Contact pressure: The pressure generated when the user touches the device, expressed as ; Touch displacement: The change in displacement during the user's touch, expressed as .
[0044] The data of the tactile sensor can be modeled in the following way: , where, represents the pressure signal received at time , is the coordinate of the touch point, is the pressure function.
[0045] To ensure the synchronization of data from different modalities, multiple sensors in this embodiment work together through a synchronization triggering mechanism. The goal of the synchronization triggering mechanism is to ensure that data from different modalities can be collected at the same time point, avoiding data distortion or out-of-sync problems caused by time differences.
[0046] Specifically, when the system starts data collection, each sensor, through the cooperation of hardware and software, ensures that their collection processes are strictly synchronized. The synchronization signal is generated by the central control unit and is transmitted between the sensors to ensure that all sensors start collecting data at the same moment. This synchronization mechanism greatly improves the timeliness and accuracy of the data.
[0047] In a possible implementation, the synchronization of the sensors can be expressed by the following formula: ; where represents the time point of synchronous collection, and represent the collection times of the RGB camera, microphone array, and tactile sensor respectively.
[0048] The data synchronization and preprocessing module in this embodiment undertakes an important function. Its task is to perform time synchronization, denoising, and error correction on the visual, auditory, and tactile data collected by the multi-modal data collection module. Through effective synchronization mechanisms and preprocessing techniques, it ensures that subsequent data processing and analysis modules can perform further operations based on high-quality data. The data synchronization and preprocessing module is connected to the aforementioned multi-modal data collection module through a data interface to ensure the timeliness and accuracy of the data. Its effectiveness is crucial for the performance of the entire system.
[0049] In this embodiment, the data synchronization and preprocessing module adopts a combination of the extended Kalman filter (EKF) and time series analysis to ensure that data from each modality is not only accurately synchronized in time, but also, after denoising and error correction, can provide high-quality input for the data fusion module. Specifically, this module includes multiple processing steps to ensure that the multi-modal data received by the system from different sources is highly consistent in time and space, and the errors in the data are effectively corrected.
[0050] Data synchronization is the first task of this module. The goal is to align the sensor data from different modalities in time so that each data point corresponds to the corresponding visual, auditory, and tactile information within the same time window. In some embodiments, by using the extended Kalman filter (EKF) to model the data of each modality, non-linear synchronization of the data is achieved.
[0051] Specifically, assume that at the time point , the data collected by the RGB camera, microphone array, and tactile sensor are , and these data sources come from different sensors, and their respective acquisition time points are usually inconsistent. To solve this problem, the extended Kalman filter algorithm is applied to predict the state of each data source at a certain moment through a model and correct the prediction result based on the measurement data.
[0052] In a possible implementation, the prediction process of the extended Kalman filter can be represented by the following equation: ; where, is the predicted value of the state at time , is the state transition matrix, is the state estimate of the previous moment, is the control matrix, is the control input.
[0053] For the error correction of visual, speech, and tactile data, it can be achieved through the following Kalman gain: ; where, is the Kalman gain, is the predicted covariance matrix, is the observation matrix, is the observation noise covariance matrix.
[0054] Finally, the system corrects the predicted values of each modality through Kalman filtering, obtains the synchronized data, and outputs the denoised and error-corrected signals, providing high-quality input for the further processing of the data fusion module.
[0055] After data synchronization, the next step is to perform data preprocessing. Specifically, time series analysis is applied to optimize the collected multimodal data, remove noise, and correct potential errors. In some embodiments, the data of each modality is first converted into a time series format for further analysis. By analyzing the patterns and trends of these time series data, the system can identify potential outliers and smooth them to ensure the stability and consistency of the data in subsequent processing.
[0056] During the process of time series data processing, techniques such as extended Kalman filtering and moving average are used for smoothing. Specifically, in the processed time series data , the data after removing noise can be represented by the following formula: ; where, is the smoothed data, represents the original data at time is the size of the moving average window.
[0057] In addition, outlier detection is often applied to the optimization of time series. The system will mark the inconsistent or mutated data in the time series according to the set threshold, and correct it through interpolation or smoothing methods. Through these optimization steps, the final output time series data can be used in subsequent data fusion and behavior prediction modules.
[0058] After completing data synchronization and preprocessing, the output data of the data synchronization and preprocessing module will be passed to the data fusion module. Through the above processing process, the data has gone through multiple steps such as denoising, time synchronization, and outlier correction, ensuring the optimization of data quality and avoiding any impacts caused by noise, deviations in data from different time windows, and outliers. The output data form can be time series data for each modality, or converted into a unified feature vector format according to requirements.
[0059] Generally speaking, the work of the data synchronization and preprocessing module ensures the quality of the input data through various technical means, providing a reliable foundation for subsequent data fusion, behavior prediction, and system optimization. The system can perform flexible synchronization and preprocessing according to the characteristics of each data source, thus ensuring the efficiency and accuracy of the system when processing different modality data.
[0060] The data fusion module in this embodiment is responsible for integrating data from different modalities to extract useful core features and generate a comprehensive feature matrix for subsequent processing. The core task of this module is to construct a multi-dimensional feature space by efficiently fusing visual, auditory, and tactile data provided by different sensors through efficient data fusion methods, so as to provide accurate input for the behavior prediction module. The data fusion module works closely with the aforementioned data synchronization and preprocessing module to ensure that the data after time synchronization and error correction can be efficiently subjected to feature extraction and fusion.
[0061] Specifically, the data fusion module reduces the dimensionality of multi-modal data by using tensor decomposition methods, thereby reducing the complexity of the data while retaining the most critical information. Through this process, the system can fuse different data types of multiple modalities into a unified representation for subsequent behavior prediction. This module not only requires high-precision mathematical calculations, but also needs to ensure the unity of various data in the time series and feature space.
[0062] In this embodiment, the data fusion module converts different types of data (such as visual data, speech data, and tactile data) into a unified feature vector through feature extraction and dimensionality reduction processing of multimodal data. Specifically, first, the unique features of each modality's raw data are extracted, and then these features are dimensionally reduced through a high-order tensor decomposition method to achieve data fusion.
[0063] In some embodiments, feature extraction is performed through the following steps: Feature extraction of visual data: Key visual features, such as facial expressions and limb movements, are extracted from the images collected by an RGB-D camera. The image data is converted into a feature vector through image processing techniques (such as convolutional neural networks). Let the feature vector of visual data be Then its expression form is: ; Among them, is an RGB image, is an image processing function.
[0064] Feature extraction of speech data: The speech data collected by a microphone array is converted into a speech feature vector through speech signal processing (such as MFCC feature extraction) , that is: ; Among them, is the collected speech signal, is a speech signal processing function.
[0065] Feature extraction of tactile data: The contact pressure and displacement information collected by a tactile sensor is converted into a tactile feature vector through a tactile perception function , that is: ; Among them, is tactile data, is a tactile perception processing function.
[0066] To effectively fuse multimodal data, this embodiment adopts a tensor decomposition method to perform dimensionality reduction processing on each modality's data. The tensor decomposition method maps high-dimensional data to a low-dimensional space, thereby reducing the computational complexity and extracting key information.
[0067] Generally, the basic idea of tensor decomposition is to decompose multi-dimensional data (tensor) into the product of several matrices, and these matrices represent different modes of the data. Specifically, assume that we have a three-dimensional tensor It represents the data obtained from three modalities: visual, speech, and tactile, where each dimension corresponds to a different modality and time window.
[0068] In a possible implementation, the tensor decomposition is performed by the following formula ; where is the original three-dimensional tensor, are the weight coefficients of the tensor decomposition, respectively represent the low-rank matrices of visual, speech, and tactile data, represents the tensor product, is the rank of the decomposition.
[0069] Through this decomposition, the original multimodal data is converted into low-dimensional matrices that can represent the relevant information between different modalities. The data after dimensionality reduction will be further fused to generate a comprehensive feature matrix.
[0070] The low-rank matrices of different modalities are fused through the tensor decomposition method to finally generate a comprehensive feature matrix , which combines the data features of all modalities and can be used for the subsequent behavior prediction module. The form of this feature matrix can be expressed as: ; where is the fused feature matrix, which contains all the key information from visual, speech, and tactile data.
[0071] In some embodiments, to enhance the robustness and accuracy of the system, the fused feature matrix can also be optimized through further feature selection and weighting algorithms to highlight the most informative parts of each modality. The optimized feature matrix will be used by the behavior prediction module to ensure the accuracy of the prediction results.
[0072] The behavior prediction module in this embodiment receives the fused feature matrix from the data fusion module and predicts the user's behavior trend based on this data. The core task of this module is to model the user's behavior based on high-dimensional multimodal data and predict the user's future behavior performance through a nonlinear dynamics model. The behavior prediction module is a crucial part of the entire system, directly affecting the adjustment of the interaction method of the system optimization module and the provision of personalized services. Through accurate behavior prediction, the system can more intelligently adapt to the changing needs of users and achieve more user-friendly and personalized interactions.
[0073] The behavior prediction module is closely connected to the aforementioned data fusion module. Based on the fused feature matrix, the system can comprehensively capture the user's behavior information. These feature matrices include key information from visual, auditory, and tactile modalities, which already have sufficient representational ability after being processed by the data fusion module. Through in-depth analysis of this information, the behavior prediction module uses a nonlinear dynamics model for prediction, generates the trend of the user's behavior, and provides the prediction result to the system optimization module, thereby realizing dynamic adjustment of the interaction strategy.
[0074] In this embodiment, the behavior prediction module adopts a nonlinear dynamics model to simulate and predict the user's behavior. The nonlinear dynamics model can well capture the complexity and temporal dependence in the user's behavior. Compared with the linear model, it can more accurately reflect the nonlinear change trend of the behavior in reality.
[0075] Specifically, assume that the user's behavior state at time is represented by a set of variables The nonlinear dynamics model describes the evolution of the behavior through the following formula: ; where, represents the user's behavior state at time , is the nonlinear state transition function, is the state at the previous time, is the current control input (such as the user's real-time operation), is the parameter of the system.
[0076] This model can predict the user's future behavior based on the historical behavior state and the current user input. In particular, the system continuously optimizes the parameter through the training process, so that the model can better adapt to the behavior patterns of different users.
[0077] In this embodiment, the behavior prediction based on historical behavior data and real-time data In some embodiments, the behavior prediction module not only makes predictions based on the current input data, but also combines historical behavior data. This method of combining historical and real-time data enables the prediction to more accurately reflect the long-term behavior trend of the user.
[0078] Specifically, the behavior prediction module trains the nonlinear dynamics model through historical behavior data (such as the user's past operation records, preference settings, etc.), and then generates predictions of the user's future behavior. Assume is the historical behavior data set, and the model is trained through the following formula: ; where, is the future behavior predicted by the model, is a function for integrating historical and current data, is the current behavior state, are the model parameters obtained through learning.
[0079] By combining historical data and real-time data, the system can obtain more accurate predictions of user behavior, thus providing more precise feedback to the system optimization module. During the prediction process, the output of the model is usually a probability distribution or trend sequence of user behavior. In some embodiments, the prediction results are further optimized to ensure their accuracy. For example, the system can adjust the prediction results according to an optimization algorithm to eliminate possible errors and ensure that the predicted behavior meets the actual needs of the user.
[0080] The optimized prediction results can be represented by the following formula: ; where, represents the optimized predicted behavior, is the given fused feature matrix and the optimization parameter is the conditional probability, representing the most likely state of the user's behavior given the features and optimization parameters.
[0081] Through this optimization process, the system can eliminate prediction biases caused by incomplete input data or noise and output a result that best conforms to the user's behavior trend.
[0082] The output results of the prediction module directly affect the work of the system optimization module. Specifically, the system optimization module adjusts the interaction method according to the behavior prediction results, such as adjusting the interface layout, voice feedback delay, etc., to adapt to the changing needs of the user. The accuracy of behavior prediction directly determines the effect of system optimization, and accurate prediction can achieve a more personalized user experience.
[0083] For example, if the system predicts that the user may need more voice feedback at a certain moment, the system optimization module will adjust the delay threshold of the voice response to respond to the user's needs in a timely manner. The prediction trend generated by the behavior prediction module provides an effective basis for system optimization, enabling the system to automatically adjust its operation mode according to the user's behavior changes.
[0084] The system optimization module evaluates the current system interaction method and functional performance according to the prediction results generated by the behavior prediction module, and adjusts the system's response mode according to the evaluation results. The output of the behavior prediction module provides the behavior trend of the user in the next period of time, and the system optimization module makes corresponding decisions by analyzing these trends to ensure a smoother and more personalized interaction process.
[0085] In general, the behavior prediction results received by the system optimization module include information such as the trends, patterns, and change speeds of user behaviors. Specifically, the system optimization module processes these data in the following ways: Evaluate the system interaction method: The system optimization module first evaluates the current interaction method. These interaction methods include, but are not limited to, the display density of the visual interface, the intensity of tactile feedback, the latency of voice response, etc. The evaluation results are based on the user's historical behavior data and real-time prediction results.
[0086] Use reinforcement learning for adjustment: As an option, the system optimization module adopts reinforcement learning algorithms (such as Q-learning or deep Q-network) to dynamically adjust the interaction parameters of the system according to the evaluation results. The goal of reinforcement learning is to maximize the interaction quality between the system and the user through repeated learning.
[0087] The application of reinforcement learning in the system optimization module is mainly to optimize the interaction behavior of the system by learning the feedback information of the user. The basic idea of reinforcement learning is to learn the optimal decision-making strategy through experiments and feedback to achieve the goal of maximizing the long-term reward. In this case, the system optimization module regards each adjusted interaction behavior as an "action", the state of the system can be represented as the current user behavior pattern, and the interaction feedback between the system and the user is used as the "reward".
[0088] In a possible implementation, the system optimization module takes the following steps for reinforcement learning according to the predicted behavior trends and historical feedback: State representation: The state of the system consists of multiple factors, including the user's visual, auditory, and tactile data, etc. Each state represents the current behavior pattern of the user.
[0089] Action selection: The system optimization module selects an action based on the current state i.e., adjust a certain interaction parameter. The action may be to modify the latency of the voice response, adjust the density of the interface information presentation, or change the intensity of the tactile feedback, etc.
[0090] Reward function: The reward is calculated according to the user's feedback, reflecting the effect of the current interaction parameter adjustment. The reward function can be defined by the following formula: ; where, The reward function is evaluated based on the current state and the selected action. Specifically, the reward function may be related to factors such as the user's satisfaction with the interaction and the fluency of the operation.
[0091] Q-value update: Based on the user's feedback and historical experience, the system updates the Q-value to optimize the decision-making strategy. The Q-value update can be achieved through the following formula: ; where is the learning rate, is the discount factor, is the maximum Q-value in state .
[0092] In this way, the system continuously optimizes its interaction parameters and gradually improves the interaction experience with the user.
[0093] The optimized system settings include adjusted interaction methods and functional parameters, and these outputs will be fed back to the behavior prediction module for further adjustment. The adjustment results of the system optimization module provide new data for the behavior prediction module, thus forming a closed-loop feedback system. In this way, the system can achieve self-adjustment and continuously optimize as the user's behavior changes.
[0094] Specifically, the system optimization module works in cooperation with the behavior prediction module in the following ways: Dynamic adjustment of interaction methods: The system optimization module dynamically adjusts the interaction methods (such as interface display, voice feedback, etc.) according to the output results of the behavior prediction module. For example, when it is predicted that the user may need more voice interaction, the system can adjust the delay of the voice response in advance to optimize the fluency of the voice interaction.
[0095] Feedback update and prediction optimization: The feedback information of the system optimization module enters the behavior prediction module again as new input, helping it to more accurately adjust future behavior predictions. Through this collaborative work, the system can continuously optimize in each interaction to meet the user's needs.
[0096] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. Multimodal human-computer interaction interface and behavior prediction system, characterized in that: include: Multimodal data collection module, used to collect users' visual, auditory and tactile data in real time; The data synchronization and preprocessing module is used to synchronize the collected multimodal data, and to perform denoising and error correction on the data; Data fusion module, used to fuse data from different modalities and extract the core features of the data; Behavior prediction module, used to predict user behavior based on the fused data; The system optimization module is used to dynamically adjust the system's interaction mode and functions based on the behavior prediction results to achieve a personalized user experience.
2. The multimodal human-computer interaction interface and behavior prediction system according to claim 1, characterized in that: The multimodal data acquisition module comprises: RGB-D camera, used to collect the user's visual data and provide the user's three-dimensional image and depth information; Microphone array, used to collect user voice data, ensure high-quality voice signals and perform noise reduction processing; A tactile sensor is used to collect the user's tactile data and detect the user's contact pressure and displacement in real time; Each sensor works together through a synchronous trigger mechanism to ensure that data of different modes are collected at the same time point.
3. The multimodal human-computer interaction interface and behavior prediction system according to claim 1, characterized in that: The data synchronization and preprocessing module synchronizes the multimodal data through extended Kalman filtering, specifically including: Temporal alignment of visual, audio, and tactile data; Correct data errors through extended Kalman filter; Output optimized synchronization data.
4. The multimodal human-computer interaction interface and behavior prediction system according to claim 1, characterized in that: The data synchronization and preprocessing module further optimizes the data through time series analysis, specifically including: Convert multimodal data into time series format; Identify and smooth outliers in time series; Output optimized time series data for subsequent processing.
5. The multimodal human-computer interaction interface and behavior prediction system according to claim 1, characterized in that: The data fusion module realizes data fusion through the following steps: Receive and extract features of each modal data; Use tensor decomposition method to reduce the dimension of data and generate low-rank matrix; The low-rank matrices are fused to generate a comprehensive feature matrix for use in the behavior prediction module.
6. The multimodal human-computer interaction interface and behavior prediction system according to claim 1, characterized in that: The behavior prediction module implements user behavior prediction through the following steps: receiving a feature matrix from a data fusion module; Use nonlinear models to predict user behavior trends; The prediction results are optimized and output for adjustment by the system optimization module.
7. The multimodal human-computer interaction interface and behavior prediction system according to claim 1, characterized in that: The system optimization module is optimized by the following steps: receiving behavior prediction results; Evaluate how the system interacts based on the predictions; Dynamically adjust system parameters through reinforcement learning; Output optimized system settings.
8. The multimodal human-computer interaction interface and behavior prediction system according to claim 1, characterized in that: The system optimization module dynamically adjusts the following interaction parameters based on the behavior prediction results: Interface information presentation density; Haptic feedback intensity; Delayed voice response.
9. The multimodal human-computer interaction interface and behavior prediction system according to claim 1, characterized in that: The system dynamically adjusts at least one of the following interaction parameters or triggers a protection mechanism based on the behavior prediction results: Interface information presentation density; Haptic feedback intensity; Delayed speech response; Initiate multimodal authentication for abnormal behavior; Insert a secondary confirmation process for high-risk operations; Initiates a system lockout for continuous erroneous behavior.
10. The multimodal human-computer interaction interface and behavior prediction system according to claim 1, characterized in that: The system further comprises: Adaptive security protection module, used to determine user operation risks and trigger protection mechanisms based on behavior prediction results: Enable multimodal authentication; Insert a second confirmation process; Initiate system lock.
Citation Information
Patent Citations
Man-machine interaction system and man-machine interaction method
CN118151763A
Multi-mode interactive intelligent control system
CN118226967A
Intelligent access control management method and system based on multi-mode identification and Internet of Things technology
CN118968665A
Cited By
VR interaction behavior prediction method based on AI
CN120429585A
Intelligent motion capture system and method of multi-mode sensor
CN120447746A
Industrial humanoid robot teleoperation cooperation system and method based on multi-modal motion capture fusion
CN120839787A
Man-machine interaction behavior prediction method and system based on large language model
CN121502010A