Fusion multi-modal interaction method and device, and storage medium
By collecting and fusion of touch, voice and motion data, and combining cloud-based large models for multimodal interaction, the problem of single interaction mode of wearable devices is solved, natural user connection and personalized feedback are achieved, and interactive experience is improved.
Patent Information
- Application Number
- CN202510651367.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-08-15
AI Technical Summary
The interaction mode of existing wearable devices is relatively single, which leads to the overall interaction performance being unnatural enough and it is difficult to establish a stable user connection.
Touch data, voice data and motion data are collected, gesture recognition, voice recognition and motion recognition processing are combined with preset cloud models to perform multi-modal interaction, and the results of gesture, voice and motion recognition are integrated to generate operational recognition data for multi-modal interactions, and natural interaction is carried out through display, tactile and auditory feedback.
It realizes more natural multimodal interaction, enhances user connection, improves the intelligence and efficiency of interaction, can make reasonable responses in complex environments, and provides personalized feedback and emotional expression.
Smart Images

Figure CN120491828A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent wearable devices, and in particular to a fusion multimodal interaction method, device and storage medium. Background Art
[0002] Currently, there are a variety of wearable devices with interactive functions on the market, such as smart bracelets, smart brooches, and electronic pet machines. These devices typically integrate displays, microcontrollers, sensors, and voice modules to achieve functions such as exercise tracking, time display, and simple interaction.
[0003] These products can display basic visual information (such as icons and animations), provide simple touch responses, or provide notifications, but these features are mostly limited to "information output" and lack a true "interactive experience" or "emotional expression." Overall, their relatively simple interaction methods make the overall interaction unnatural and difficult to establish a stable user connection. Summary of the Invention
[0004] The main purpose of the present invention is to solve the technical problem that the existing technology is relatively single in interaction mode, resulting in the overall interaction performance being unnatural and difficulty in establishing a stable user connection feeling.
[0005] A first aspect of the present invention provides a fusion multimodal interaction method, the fusion multimodal interaction method comprising:
[0006] Collecting input signals, the input signals including: touch data, voice data, and motion data;
[0007] Performing gesture recognition processing on the touch data to obtain a gesture trajectory;
[0008] Comparing the gesture trajectory with a preset gesture model to obtain a gesture recognition result;
[0009] Performing speech recognition processing on the speech data to obtain a speech recognition result;
[0010] According to a preset motion feature model, the motion data is processed to obtain a motion recognition result;
[0011] According to the preset cloud-based large model, the gesture recognition results, the voice recognition results, and the motion recognition results are fused and analyzed to obtain multimodal interactive operation recognition data.
[0012] Optionally, in a first implementation of the first aspect of the present invention, the touch data includes: the number of contact points, the contact area, and the movement direction, and performing gesture recognition processing on the touch data to obtain the gesture trajectory includes:
[0013] Generating continuous coordinate data according to the number of contact points, the contact area, and the moving direction;
[0014] Based on the continuous coordinate data, a gesture trajectory is obtained.
[0015] Optionally, in a second implementation of the first aspect of the present invention, performing speech recognition processing on the speech data to obtain a speech recognition result includes:
[0016] Analyze the voice data based on a preset voice activity detection algorithm to obtain voice audio frames;
[0017] Triggering a voice wake-up instruction according to the voice audio frame;
[0018] Based on the voice wake-up instruction, the voice data is subjected to voice recognition processing to obtain a voice recognition result. Optionally, in a third implementation of the first aspect of the present invention, the motion data recognition operation is processed according to a preset motion feature model to obtain a motion recognition result, including:
[0019] Analyzing the acceleration, angular velocity, and movement direction of the motion data to obtain change data;
[0020] Performing integral calculation based on the change data to obtain the displacement of the fusion multimodal interaction device in various directions;
[0021] Based on a preset behavior recognition model, the displacement of the fusion multimodal interactive device in various directions is analyzed to obtain a motion recognition result.
[0022] Optionally, in a fourth implementation of the first aspect of the present invention, the input signal further includes image data, and collecting the input signal includes:
[0023] Performing recognition processing on the image data to obtain an image recognition result;
[0024] According to the preset cloud-based large model, the gesture recognition results, the voice recognition results, the motion recognition results, and the image recognition results are fused and analyzed to obtain high-end multimodal interactive operation recognition data.
[0025] Optionally, in a fifth implementation of the first aspect of the present invention, after fusing and analyzing the gesture recognition result, the voice recognition result, and the motion recognition result according to the preset cloud-based large model to obtain multimodal interaction operation recognition data, the method further includes:
[0026] According to the preset agent behavior model, the operation identification data and the current state of the fusion multimodal interaction device are dynamically selected and processed to obtain a response strategy corresponding to the operation identification data.
[0027] Optionally, in a sixth implementation of the first aspect of the present invention, after fusing and analyzing the gesture recognition result, the voice recognition result, and the motion recognition result according to the preset cloud-based large model to obtain multimodal interaction operation recognition data, the method further includes:
[0028] Based on the long-term memory mechanism of the cloud-based large model, N interaction data are recorded, where N is a positive integer;
[0029] Vectorizing the N interaction data to obtain N interaction memory matrices;
[0030] Based on the similarities between the N interaction memory matrices, calling historical data related to the operation identification data;
[0031] The historical data is matched with the current operation identification data and the behavioral decisions corresponding to the current interaction data to obtain an accurate state judgment result.
[0032] Optionally, in a seventh implementation of the first aspect of the present invention, after matching the historical data with the current operation identification data and the behavior decision corresponding to the current interaction data to obtain an accurate state determination result, the method further includes:
[0033] generating feedback data corresponding to the state judgment result by combining the recognition operation recognition data with the state judgment result;
[0034] The feedback data is output by the fusion multimodal interaction device.
[0035] A second aspect of the present invention provides a fusion multimodal interaction device, the fusion multimodal interaction device comprising:
[0036] A display module, configured to display the feedback data;
[0037] A touch control module, configured to collect the touch data, identify gestures based on the touch data, and obtain a gesture trajectory;
[0038] A sound module, used for collecting the voice data and playing sound feedback;
[0039] Motion sensing module, used to detect user's shaking, tapping, and wearing posture motion data;
[0040] Communication module, used for remote communication or memory content update;
[0041] The main control module is connected to the display module, the touch module, the sound module, the motion sensing module, and the communication module respectively, and is used to fuse and analyze the gesture recognition results, the voice recognition results, and the motion recognition results to obtain multimodal interactive operation recognition data.
[0042] Optionally, in a first implementation of the second aspect of the present invention, the fused multimodal interaction device further includes a camera module, which is connected to the main control module and is used to collect image data and perform user expression recognition, image recognition analysis, or image interaction processing.
[0043] Optionally, in a second implementation of the second aspect of the present invention, the communication module is a 4G communication module or an NB-IoT communication module, which is used to realize remote networking, content update or data interaction with cloud services in a non-Wi-Fi environment.
[0044] Optionally, in a third implementation of the second aspect of the present invention, the fusion multimodal interaction device further includes a vibration module, which is connected to the main control module and is used to provide tactile feedback including simulated emotional expression or execution event reminders.
[0045] Optionally, in a fourth implementation of the second aspect of the present invention, the communication module includes a Bluetooth module and a Wi-Fi module, and the Bluetooth module and the Wi-Fi module are respectively connected to the main control module. The Bluetooth communication module is used for device pairing, mobile terminal interaction, data synchronization and firmware upgrade, and the Wi-Fi module is used for device networking and cloud data communication.
[0046] Optionally, in a fifth implementation of the second aspect of the present invention, the fusion multimodal interaction device further includes: a memory and at least one processor, the memory storing instructions, and the memory and the at least one processor being interconnected via a line;
[0047] The at least one processor calls the instructions in the memory to enable the fusion multimodal interaction device to execute the above-mentioned fusion multimodal interaction method.
[0048] A third aspect of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores instructions that, when executed on a computer, enable the computer to execute the above-mentioned fusion multimodal interaction method.
[0049] In an embodiment of the present invention, touch data, voice data, and motion data are collected, and gesture recognition processing is performed on the touch data to obtain a gesture trajectory, the gesture trajectory is compared with a preset gesture model to obtain a gesture recognition result, the voice data is subjected to voice recognition processing to obtain a voice recognition result, the motion data is subjected to operation recognition processing according to a preset motion feature model to obtain a motion recognition result, and the gesture recognition result, the voice recognition result, and the motion recognition result are fused and analyzed according to a preset cloud-based large model to obtain multimodal interaction operation recognition data. The fused multimodal interaction device can perform multimodal interaction, making the overall interaction performance more natural to establish a stable sense of user connection. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 Schematic diagram of a first specific embodiment of the method for integrating multimodal interaction in an embodiment of the present invention;
[0051] Figure 2 This is a schematic diagram of a specific embodiment of step 104 in an embodiment of the present invention;
[0052] Figure 3 This is a schematic diagram of a specific embodiment of step 105 in an embodiment of the present invention;
[0053] Figure 4 Schematic diagram of a second specific embodiment of the method for integrating multimodal interaction in an embodiment of the present invention;
[0054] Figure 5 Schematic diagram of an embodiment of a fusion multimodal interaction device according to an embodiment of the present invention;
[0055] Figure 6 2 is a schematic diagram of a module of a fusion multimodal interaction device in an embodiment of the present invention. DETAILED DESCRIPTION
[0056] Embodiments of the present invention provide a fusion multimodal interaction method, device, and storage medium.
[0057] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0058] In the description of the embodiments disclosed herein, the term "including" and similar terms should be understood as open inclusion, i.e., "including but not limited to." The term "based on" should be understood as "based, at least in part, on." The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment." The terms "first," "second," etc. may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
[0059] For ease of understanding, the specific process of the embodiment of the present invention is described below. Figure 1 , Figure 1 FIG1 is a schematic diagram of a first specific embodiment of a fusion multimodal interaction method in an embodiment of the present invention. The fusion multimodal interaction method in an embodiment of the present invention is applied to a fusion multimodal interaction device. An embodiment of the fusion multimodal interaction method includes:
[0060] 101. Collect input signals, where the input signals include touch data, voice data, and motion data.
[0061] In this embodiment, the input signal may also include image data, which may enable enhanced interactions such as environment recognition, user expression collection, and image interaction.
[0062] The input signal also includes image data. In step 101, the following steps may be performed:
[0063] 1011. Perform recognition processing on the image data to obtain an image recognition result;
[0064] 1012. Based on a preset cloud-based large model, the gesture recognition result, the voice recognition result, the motion recognition result, and the image recognition result are integrated and analyzed to obtain high-configuration multimodal interactive operation recognition data.
[0065] In steps 1011-1012, by fusing gesture recognition results, voice recognition results, motion recognition results, and image recognition results, more comprehensive and accurate recognition results are provided across different perceptual dimensions. This multimodal data fusion can reduce errors in a single recognition mode and improve overall recognition accuracy. In multimodal interactions, the fusion analysis of different recognition results can handle more complex tasks and operations. For example, while the user is performing gesture operations, the system can combine voice commands and image analysis to perform multiple levels of operations, improving the intelligence and efficiency of the operations. By pre-setting large cloud models, the system can quickly adapt to different application scenarios and user needs, especially in complex or changing environments, and can make reasonable responses based on different input forms.
[0066] 102. Perform gesture recognition processing on the touch data to obtain a gesture trajectory;
[0067] 103. Compare the gesture trajectory with a preset gesture model to obtain a gesture recognition result;
[0068] In this embodiment, gesture recognition can be implemented using capacitive, resistive, infrared, or other touch technologies. It works by detecting the contact position and trajectory of a user's finger on the screen, generating continuous coordinate data that is then compared with built-in gesture templates to identify the specific user input. The gesture recognition methods described above can be flexibly selected based on the actual product structure.
[0069] In addition, the touch data includes: the number of contact points, the contact area, and the moving direction. In step 102, the following steps may be performed:
[0070] 1021. Generating continuous coordinate data according to the number of contact points, the contact area, and the moving direction;
[0071] 1022. Obtain a gesture trajectory based on the continuous coordinate data.
[0072] In steps 1021-1022, the finger contact trajectory on the screen surface of the fusion multimodal interactive device is collected in real time and analyzed based on the number of contact points, contact area, and movement direction. The fusion multimodal interactive device can recognize single-finger and multi-finger gestures, including clicks, slides, long presses, double clicks, and two-finger zooming.
[0073] 104. Performing speech recognition processing on the speech data to obtain a speech recognition result;
[0074] In this embodiment, the voice recognition processing of the voice data may be performed in a manner of actively triggering a voice wake-up instruction: the user may actively turn on the voice recognition function by clicking a physical button or a virtual control on the screen;
[0075] Natural language wake-up method: The system can monitor natural voice input and trigger voice wake-up commands through voice activity detection (VAD), voice feature analysis, or specific intonation / rhythm recognition;
[0076] No reliance on traditional "wake-up word" mechanisms: Although "voice wake-up words" are still compatible, wake-up words are not the only method. They are not the only triggering condition. When the fusion multimodal interaction device detects that the user makes a sound (such as calling, speaking, or making a certain tone), the voice wake-up command will be triggered. Specific methods include recognizing specific voice wake-up words (such as "Hello Badge"), or non-keyword wake-up based on sound characteristics (such as long tones, stresses, and voice activity).
[0077] like Figure 2 As shown, Figure 2This is a schematic diagram of a specific embodiment of step 104 in an embodiment of the present invention. In step 104, the following steps may be performed:
[0078] 1041. Analyze the voice data based on a preset voice activity detection algorithm to obtain a voice audio frame;
[0079] 1042. Trigger a voice wake-up instruction according to the voice audio frame;
[0080] 1043. Based on the voice wake-up instruction, perform voice recognition processing on the voice data to obtain a voice recognition result.
[0081] In steps 1041-1043, the preset voice activity detection algorithm accurately identifies valid audio frames in the voice data. This helps reduce interference from noise and invalid speech, improving the accuracy of subsequent processing. After detecting specific voice activity, the system can automatically trigger a voice wake-up command. This allows for a quick response to user commands without manual intervention, improving the user experience. After the voice wake-up command is triggered, voice recognition processing is performed on the voice data, ensuring that the recognition process is initiated at the correct time. This precise startup timing helps improve the success rate and accuracy of voice recognition and reduces the probability of misidentification.
[0082] In step 104, the following steps may also be performed:
[0083] 1044. Identify the voice wake-up word and detect the start and end words of the voice wake-up word;
[0084] 1045. Extracting audio features from the start and end words;
[0085] 1046. Compare the audio feature with a preset wake-up model to obtain an audio comparison result;
[0086] 1047. Trigger a voice wake-up command based on the audio comparison result.
[0087] In steps 1044-1047, detect the start and end words: After recognizing the voice wake-up word, it is necessary to determine the starting and ending positions of the wake-up word to ensure that subsequent processing is only for this part of the audio, and extract features from the audio clip of the voice wake-up word (such as the spectral characteristics of the audio signal, time domain characteristics, etc.). These features help to compare with the pre-trained wake-up model, and compare the extracted audio features with the preset wake-up model to verify whether the voice meets the wake-up conditions. If the audio features match the model, it means that the voice wake-up word has been successfully recognized. If the audio comparison result meets the preset conditions, the system will trigger the voice wake-up command, and then start the voice assistant or perform subsequent voice recognition and processing.
[0088] 105. Perform recognition operations on the motion data according to a preset motion feature model to obtain a motion recognition result;
[0089] In this embodiment, this method supports the perception of the user's movements and wearing status based on multiple posture detection modules, and determines whether the user has performed a specific operation by real-time acquisition of the acceleration, angular velocity and direction change signals generated by the device in space and combining the action feature model.
[0090] For example:
[0091] Shake detection: Identify medium-amplitude, continuous acceleration changes;
[0092] Knock recognition: Capture high acceleration signals generated by instantaneous impact;
[0093] Flip / take-off recognition: Identify the front and back of the device, and whether it is worn or not, through continuous changes in the direction of gravity;
[0094] Rotation detection and direction judgment: Determine whether turning or rolling occurs based on the gyroscope angular velocity signal.
[0095] In practical applications, multi-dimensional information fusion can be combined with Hall effect sensors, infrared / light sensors, capacitive proximity sensors, and more to enhance the accuracy and robustness of gesture recognition. Different product versions can utilize different combinations of gesture detection modules to achieve the desired functionality, depending on cost and usage scenarios.
[0096] like Figure 3 As shown, Figure 3 This is a schematic diagram of a specific embodiment of step 105 in an embodiment of the present invention. In step 105, the following steps may be performed:
[0097] 1051. Analyze the acceleration, angular velocity, and motion direction of the motion data to obtain change data;
[0098] 1052. Perform integral calculation based on the change data to obtain the displacement of the fusion multimodal interaction device in each direction;
[0099] 1053. Based on a preset behavior recognition model, analyze the displacement of the fusion multimodal interaction device in various directions to obtain a motion recognition result.
[0100] In steps 1051-1053, by collecting motion data such as acceleration and angular velocity and analyzing their changes, the dynamic characteristics of the motion are derived, providing the necessary data foundation for subsequent calculations. By integrating the changing data, the displacement information of the multimodal interactive device in various directions can be obtained. The role of the integral calculation here is to convert the instantaneous change into a cumulative amount, thereby obtaining the overall displacement of the device.
[0101] 106. Based on a preset cloud-based large model, the gesture recognition result, the voice recognition result, and the motion recognition result are integrated and analyzed to obtain multimodal interactive operation recognition data.
[0102] By receiving multimodal input signals, including voice, touch, and motion data, the system extracts key features such as keywords, intonation, operation rhythm, and movement amplitude. Combined with contextual information such as the current time and the state of multimodal interactive devices, the system uses a deep learning-based behavior recognition model to analyze these features and determine the user's operational intent.
[0103] To enhance the system's understanding, the model incorporates diverse user behavior data during training. This allows for generalization, enabling it to handle ambiguous or incomplete input and infer the user's likely intent. Furthermore, the system references the user's past behavior records to aid in its judgment, ensuring that recognition results are more consistent with individual habits.
[0104] like Figure 4 As shown, Figure 4 This is a schematic diagram of a second specific embodiment of the method for integrating multimodal interaction according to an embodiment of the present invention. After step 106, the following steps may be performed:
[0105] 107. According to the preset intelligent agent behavior model, a strategy dynamic selection process is performed on the operation identification data and the current state of the fusion multimodal interaction device to obtain a response strategy corresponding to the operation identification data.
[0106] In step 107, the system intelligently selects an appropriate response strategy based on the pre-set agent behavior model, rather than relying solely on a fixed response method. This means the system can dynamically adjust its response strategy based on different operation recognition data and the current state of the multimodal interaction device, improving the intelligence and personalization of the response.
[0107] After step 106, the following steps may be performed:
[0108] 108. Based on the long-term memory mechanism of the cloud-based large model, record N interaction data, where N is a positive integer;
[0109] 109. Vectorize the N interaction data to obtain N interaction memory matrices;
[0110] 110. Based on the similarities between the N interaction memory matrices, call historical data related to the operation identification data;
[0111] 111. Match the historical data with the current operation identification data and the behavioral decision corresponding to the current interaction data to obtain an accurate state judgment result.
[0112] In steps 108-111, a long-term memory mechanism based on a large cloud-based model continuously records user interactions and forms a personalized memory representation. Through vectorized storage and similarity retrieval, relevant historical data is retrieved in new interactions to assist in current intent recognition and behavioral decision-making. Memory content is dynamically updated based on time decay and importance ratings to ensure long-term retention of user preferences and habits.
[0113] After step 111, the following steps may be performed:
[0114] 112. Generate feedback data corresponding to the state judgment result by combining the recognition data with the state judgment result through the recognition operation;
[0115] 113. The fusion multimodal interaction device outputs the feedback data.
[0116] In steps 112-113, feedback data is presented in various ways such as visual animation, voice output, vibration prompts, etc., which are determined by the current interaction state, user historical behavior and emotional tendencies.
[0117] For example: when the user taps the fusion multimodal interactive device after a long period of inactivity, it can be recognized as a "wake-up" behavior, and the animated character will open his eyes and play the voice of "You are back"; if the system recognizes that the user's voice is depressed, it can play an expression comfort animation and say in a gentle voice "Do you want me to tell you a joke?"; if a user is accustomed to interacting before going to bed every night, the system can actively display the "Good night interaction" animation and enter night mode.
[0118] By integrating the built-in basic behavioral data of multimodal interactive devices (such as commonly used expression animations, prompt voice, and standard response strategies), core feedback functions can be executed offline to ensure a basic interactive experience. For feedback that requires advanced model judgment or personalized content generation, cloud services can be called for completion and enhancement while connected to the Internet, supporting a hybrid deployment model that combines local operation with cloud collaboration. Feedback behaviors are selected by the behavior library or dynamically generated according to rules to match the current context and user habits.
[0119] To achieve a more natural interaction process, the behavioral control logic of the AI Agent feature is introduced. The AI core operates in an event-driven manner, combining modal recognition results, interaction history, and current status to dynamically select response strategies from the "behavior library" to control the device output process. It can display an "active personality" when the user interacts frequently, and adopt "waiting" or "lost" states when the user is indifferent or has no action for a long time. This simulates the interaction process with emotional fluctuations and personality changes, further enhancing the sense of companionship.
[0120] Through the above mechanism, a scalable multimodal feedback generation process with continuous learning capabilities is realized, enhancing the intelligence and personalization of the interaction.
[0121] like Figure 6 As shown, Figure 6 This is a schematic diagram of a module of a fusion multimodal interaction device in an embodiment of the present invention. In a second aspect of the present invention, a fusion multimodal interaction device is provided. The fusion multimodal interaction device includes:
[0122] Display module 7, used to display the feedback data. Display module 7 can adopt various display technologies such as OLED, LCD (including TFT-LCD, IPS, etc.), electronic ink screen, MiniLED, MicroLED, dot matrix LED screen, etc. The screen size and shape (such as circular, rectangular, oval or special-shaped structure) can also be flexibly set according to product positioning and usage requirements;
[0123] A touch control module 6 is configured to collect the touch data, identify gestures based on the touch data, and obtain a gesture trajectory;
[0124] The sound module 4 is used to collect the voice data and play the voice feedback. The sound module 4 has a built-in low-power microphone, which can continuously monitor the ambient sound in standby mode. It can use a local voice recognition chip or an integrated microphone + external cloud recognition system. The voice output can also be achieved through a buzzer, a TTS chip or an audio decoding module;
[0125] Motion sensing module 1, used to detect the user's shaking, tapping, and wearing posture motion data. The motion sensing module 1 used may include but is not limited to: a three-axis acceleration sensor, a six-axis inertial sensor (including a gyroscope), a nine-axis posture fusion unit (including a magnetometer), an infrared distance sensor, a light sensor, etc.;
[0126] Communication module 2, used for remote communication or memory content update;
[0127] The main control module 3 is respectively connected to the display module 7, the touch module 6, the sound module 4, the motion sensing module 1, and the communication module 2, and is used to fuse and analyze the gesture recognition results, the voice recognition results, and the motion recognition results to obtain multimodal interactive operation recognition data. It can adopt a system-on-chip (SoC) or controller architecture with communication capabilities and edge intelligent processing capabilities, support local operation of embedded operating systems, and have voice recognition, image processing, neural network reasoning, multimodal fusion and other capabilities to support the device's perception, judgment and interactive logic processing. In lightweight application scenarios with limited resources or sensitive power consumption, a low-power microcontroller platform can also be used to achieve basic control, state management and local communication tasks. In addition to 4G, the communication part can also support Bluetooth, Wi-Fi, NB-IoT and other methods. The communication module 2 can be divided into a basic communication module and an extended communication module according to the device configuration. Among them, the basic communication module includes a Bluetooth communication module (supporting low-power Bluetooth protocol) and a Wi-Fi module, which are used to realize conventional communication needs such as pairing connection, data synchronization, remote control and content update between the device and the mobile terminal; the extended communication module includes a cellular communication module, such as a communication sub-module that supports 4G or NB-IoT protocol, which can be used to realize independent networking of the device, periodic data synchronization, device status reporting or cloud service access without Wi-Fi conditions.
[0128] The specific combination of the above-mentioned main control module and communication module can be flexibly selected and deployed according to product configuration, usage environment and system function requirements to meet the balanced needs of different types of equipment in terms of response speed, power consumption, communication mode and computing power.
[0129] When a user interacts with the device (e.g., tapping, speaking, shaking the device), the various sensor modules simultaneously collect input signals. The system integrates these signals and processes them to determine the user's intent based on the current interaction context. For example, continuous tapping of the device might be interpreted as a "call," while keywords in the tone of voice could trigger a "response" or "comfort" response.
[0130] After analyzing the input, the main control module 3 relies primarily on a cloud-based model to analyze and judge the current state of the device (e.g., "happy," "sleepy," or "waiting"). When a user initiates an interaction via voice, gesture, or touch, the system uploads the relevant data to the cloud. The cloud-based model then conducts a comprehensive analysis based on voice intonation, keywords, historical behavior patterns, and other information to infer the emotional tendency and context of the current interaction, and accordingly determines the emotional state the device should present.
[0131] Devices supporting multimodal interaction retain a simplified state marker mechanism to maintain minimal feedback capabilities in offline states. For example, it automatically enters a "waiting" state after a long period of inactivity and defaults to a "sleepy" state during nighttime. State variables have a lifecycle, with a configurable duration or trigger conditions, allowing states to automatically expire or be overwritten by new interactions.
[0132] When connected to the internet, cloud-based models are prioritized for more accurate and detailed state judgments, enabling a more personalized, context-consistent emotional feedback experience. Fusion multimodal interactive devices feature a "memory" mechanism that records and updates user behavior frequency and interaction preferences, personalizing device feedback and enhancing the sense of human-machine connection. Fusion multimodal interactive devices can respond quickly through local voice recognition modules or, when connected to the internet, utilize cloud-based recognition services to improve semantic understanding and recognition accuracy.
[0133] The fusion multimodal interaction device may further include an output control unit: generating visual, auditory, tactile and other feedback commands according to the behavioral decision, driving the display module 7 to display expression animations, triggering sound responses, vibration prompts, etc.;
[0134] Basic function service module: realizes basic smart terminal functions such as time display, alarm reminder, content broadcast, holiday animation, etc., ensuring practicality and fun.
[0135] To meet different usage requirements and cost control strategies, the present invention can provide multiple configuration versions. In the basic version, the device has complete visual, tactile, voice, and motion perception and feedback capabilities to meet daily interaction and emotional expression functions.
[0136] In the basic version, the device includes at least a Bluetooth module (supporting the BLE protocol) by default, which is used for pairing with mobile terminals, data synchronization, remote control, and content updates. Some basic versions can also be equipped with an optional Wi-Fi module to support direct networking and access to cloud services.
[0137] In the high-end version, the fusion multimodal interaction device also includes a camera module, which is connected to the main control module 3 and is used to collect image data and perform user expression recognition, image recognition analysis or image interaction processing. It is also used for enhanced interactive functions such as environment recognition, user expression collection, image interaction, etc. The camera module is a high-end optional module, and its resolution, viewing angle, and processing method can be flexibly selected according to the application scenario. The image recognition function can be supported by local or cloud algorithms.
[0138] The fusion multimodal interaction device further includes a vibration module, which is connected to the main control module 3 and is used to provide tactile feedback including simulating emotional expression or executing event reminders. It is also used to provide tactile feedback to achieve emotional prompts or event reminders. A vibration motor is used or replaced with a linear vibrator, piezoelectric element or small electromagnetic according to the device volume and power consumption requirements to achieve different forms of tactile feedback;
[0139] The communication module 2 includes a Bluetooth module and a Wi-Fi module. The Bluetooth module and the Wi-Fi module are respectively connected to the main control module 3. The Bluetooth communication module is used for device pairing, mobile terminal interaction, data synchronization and firmware upgrade, and the Wi-Fi module is used for device networking and cloud data communication.
[0140] In addition, the main control module 3 is configured as an MCU or SoC chip with local AI reasoning capabilities or image processing capabilities, which is used to improve the intelligence and real-time nature of interactive responses, improve storage and AI processing capabilities, and have stronger local reasoning and image processing capabilities. The fusion multimodal interactive device supports remote social interaction and behavior log management functions, including virtual event synchronization between multiple devices, user behavior data upload or personalized behavior recording.
[0141] Power module: built-in rechargeable lithium battery, combined with power management circuit to ensure stable operation of the equipment;
[0142] Storage and communication module: Flash storage and wireless communication module 2 (such as Wi-Fi, BLE) can be expanded as needed for remote communication or content updates.
[0143] Among them, the 4G communication module, camera module, vibration module, power module, storage and communication module are respectively connected to the main control module 3.
[0144] The addition of these modules can further enhance the "perceptual depth" and "expressive richness" of integrated multimodal interactive devices at the human-computer interaction level, and improve their functional breadth and application extension capabilities as intelligent emotional terminals.
[0145] The overall technical solution of the present invention integrates active interaction, emotional expression and basic practical functions on an extremely small wearable terminal - a fusion multimodal interactive device through the collaboration of software and hardware. It has significant technical integration and scalability, and is suitable for new smart device forms with emotional interaction needs.
[0146] The present invention adopts a unified input processing and behavior-driven architecture to systematically integrate perception modules such as touch, voice, and action with execution structures such as screen display, sound output, and vibration feedback, forming a continuous closed-loop "perception-judgment-feedback" mechanism, which is different from the processing method of separate responses of each module in the existing technology.
[0147] By building an internal behavioral model, the integrated multimodal interactive device can dynamically adjust the response content according to the user status and interaction history, realize the anthropomorphic behavioral expression of emotional changes and personality evolution, and form an active interactive structure that is different from the traditional rule-driven system. The core components such as the display module 7, touch module 6, sound module 4, motion sensing module 1, communication module 2, and main control module 3 are integrated in the wearable terminal. At the same time, the complete structural closure is achieved through a compact PCB layout, which improves the portability and functional density of the device. The high-end version is compatible with the system through structural reservation, and supports the addition of camera modules, 4G communication modules 2, etc., to form an enhanced smart terminal with remote social and visual recognition capabilities, reflecting the modular upgrade design concept at the structural level.
[0148] The present invention integrates a multimodal perception system and AI Agent control logic into a fusion multimodal interactive device to create an intelligent device with active interaction, emotional expression and personalized feedback capabilities. It not only achieves a breakthrough at the functional level, but also brings unique user value at the emotional connection level. By integrating multimodal interaction methods such as touch, voice, action, and image, the system can provide real-time feedback based on user behavior and status, enabling the device to respond to users in a more natural, human-like manner, significantly improving the immersion and emotional expressiveness of the interactive experience.
[0149] Relying on the behavioral logic built on the AIAgent architecture, the device can simulate emotions, adjust "personality", and produce anthropomorphic reactions to meet users' psychological needs for "response", "understanding", and "interactive feedback" in individual scenarios such as solitude, rest, and daily use, thereby increasing the sense of companionship and usage stickiness. Whether it is the emotional animation display in a static state, or the voice communication and action awakening initiated by the user, the device can combine the current interaction rhythm with historical behavioral data to generate a matching feedback strategy, forming a more coherent and memorable "relational interaction", breaking the cold interaction boundaries of traditional products. Achieving high integration in a compact form factor, through system-level software and hardware collaboration, the device not only has the practicality of a basic smart terminal (such as time display, reminders, etc.), but also has a "sense of warmth" and "emotional expression ability" that exceeds traditional products.
[0150] In the embodiment of the present invention, by integrating the multimodal interaction method, the integrated multimodal interaction device can perform multimodal interaction, making the overall interaction performance more natural to establish a stable user connection feeling.
[0151] Figure 5: This is a schematic diagram of the structure of a fusion multimodal interaction device provided by an embodiment of the present invention. The fusion multimodal interaction device 500 may have relatively large differences due to different configurations or performances, and may include one or more processors (central processing units, CPU) 510 (for example, one or more processors) and a memory 520, and one or more storage media 530 (for example, one or more mass storage devices) for storing application programs 533 or data 532. Among them, the memory 520 and the storage medium 530 can be temporary storage or permanent storage. The program stored in the storage medium 530 may include one or more modules (not shown in the figure), each module may include a series of instruction operations in the fusion multimodal interaction device 500. Furthermore, the processor 510 can be configured to communicate with the storage medium 530 to execute a series of instruction operations in the storage medium 530 on the fusion multimodal interaction device 500. The wired / wireless network interface 550 may include Bluetooth, Wi-Fi, 4G, or NB-IoT communication submodules. Bluetooth and Wi-Fi are usually used as basic configurations for close-range interaction with mobile terminals; 4G and NB-IoT communication modules can be optionally configured according to product version requirements to achieve remote networking or periodic communication functions. In addition, it may also include one or more input / output interfaces 560 and / or one or more operating systems 531. For example, but not limited to:
[0152] Embedded operating systems: Android (including Android Things), embedded Linux (such as Yocto and OpenWrt), RTOS (such as FreeRTOS, RT-Thread, and Zephyr), LiteOS, NuttX, and VxWorks;
[0153] General operating systems: Linux, Unix, Windows Server, Mac OS X, FreeBSD, etc.
[0154] Those skilled in the art will understand that Figure 5 The illustrated structure of the fusion multimodal interaction device does not constitute a limitation on the fusion multimodal interaction device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0155] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. The computer-readable storage medium stores instructions, which, when executed on a computer, cause the computer to execute the steps of the fused multimodal interaction method.
[0156] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0157] In addition, although adopting specific order to describe each operation, this should be understood as requiring such operation to be carried out in the specific order shown or in sequential order, or requiring that all illustrated operations should be carried out to obtain desired results. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although comprising some specific implementation details in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of separate embodiment can also be implemented in a single implementation in combination. On the contrary, the various features described in the context of a single implementation also can be implemented in a plurality of implementations individually or in the mode of any suitable subcombination.
[0158] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. A fusion multimodal interaction method, characterized in that: Applied to a fusion multimodal interaction device, the fusion multimodal interaction method includes: Collecting input signals, the input signals including: touch data, voice data, and motion data; Performing gesture recognition processing on the touch data to obtain a gesture trajectory; Comparing the gesture trajectory with a preset gesture model to obtain a gesture recognition result; Performing speech recognition processing on the speech data to obtain a speech recognition result; According to a preset motion feature model, the motion data is processed to obtain a motion recognition result; According to the preset cloud-based large model, the gesture recognition results, the voice recognition results, and the motion recognition results are fused and analyzed to obtain multimodal interactive operation recognition data.
2. The fusion multimodal interaction method according to claim 1, characterized in that: The touch data includes: the number of contact points, the contact area, and the moving direction. The gesture recognition processing is performed on the touch data to obtain the gesture trajectory, which includes: Generating continuous coordinate data according to the number of contact points, the contact area, and the moving direction; Based on the continuous coordinate data, a gesture trajectory is obtained.
3. The fusion multimodal interaction method according to claim 1, characterized in that: The performing speech recognition processing on the speech data to obtain a speech recognition result includes: Analyze the voice data based on a preset voice activity detection algorithm to obtain voice audio frames; Triggering a voice wake-up instruction according to the voice audio frame; Based on the voice wake-up instruction, voice recognition processing is performed on the voice data to obtain a voice recognition result.
4. The fusion multimodal interaction method according to claim 1, characterized in that: The identifying operation and processing of the motion data according to the preset motion feature model to obtain the motion recognition result includes: Analyzing the acceleration, angular velocity, and movement direction of the motion data to obtain change data; Performing integral calculation based on the change data to obtain the displacement of the fusion multimodal interaction device in various directions; Based on a preset behavior recognition model, the displacement of the fusion multimodal interactive device in various directions is analyzed to obtain a motion recognition result.
5. The fusion multimodal interaction method according to claim 1, characterized in that: The input signal also includes image data, and the step of collecting the input signal includes: Performing recognition processing on the image data to obtain an image recognition result; According to the preset cloud-based large model, the gesture recognition results, the voice recognition results, the motion recognition results, and the image recognition results are fused and analyzed to obtain high-end multimodal interactive operation recognition data.
6. The fusion multimodal interaction method according to claim 1, characterized in that: After fusing and analyzing the gesture recognition result, the voice recognition result, and the motion recognition result according to the preset cloud-based large model to obtain multimodal interactive operation recognition data, the method further includes: According to the preset agent behavior model, the operation identification data and the current state of the fusion multimodal interaction device are dynamically selected and processed to obtain a response strategy corresponding to the operation identification data.
7. The fusion multimodal interaction method according to any one of claims 1 to 6, characterized in that: After fusing and analyzing the gesture recognition result, the voice recognition result, and the motion recognition result according to the preset cloud-based large model to obtain multimodal interactive operation recognition data, the method further includes: Based on the long-term memory mechanism of the cloud-based large model, N interaction data are recorded, where N is a positive integer; Vectorizing the N interaction data to obtain N interaction memory matrices; Based on the similarities between the N interaction memory matrices, calling historical data related to the operation identification data; The historical data is matched with the current operation identification data and the behavioral decisions corresponding to the current interaction data to obtain an accurate state judgment result.
8. The fusion multimodal interaction method according to claim 7, characterized in that: After matching the historical data with the current operation identification data and the behavior decision corresponding to the current interaction data to obtain an accurate state judgment result, the method further includes: generating feedback data corresponding to the state judgment result by combining the recognition operation recognition data with the state judgment result; The feedback data is output by the fusion multimodal interaction device.
9. A fusion multimodal interactive device, characterized in that: The fusion multimodal interaction device includes: A display module, configured to display the feedback data; A touch control module, configured to collect the touch data, identify gestures based on the touch data, and obtain a gesture trajectory; A sound module, used for collecting the voice data and playing sound feedback; Motion sensing module, used to detect user's shaking, tapping, and wearing posture motion data; Communication module, used for remote communication or memory content update; The main control module is connected to the display module, the touch module, the sound module, the motion sensing module, and the communication module respectively, and is used to fuse and analyze the gesture recognition results, the voice recognition results, and the motion recognition results to obtain multimodal interactive operation recognition data.
10. The fusion multimodal interaction device according to claim 9, characterized in that: The fusion multimodal interaction device also includes a camera module, which is connected to the main control module and is used to collect image data and perform user expression recognition, image recognition analysis or image interaction processing.
11. The fusion multimodal interaction device according to claim 9, characterized in that: The communication module is a 4G communication module or an NB-IoT communication module, which is used to achieve remote networking, content update or data interaction with cloud services in a non-Wi-Fi environment.
12. The fusion multimodal interaction device according to any one of claims 9 to 11, characterized in that: The fusion multimodal interaction device further includes a vibration module, which is connected to the main control module and is used to provide tactile feedback including simulated emotional expression or execution event reminder.
13. The fusion multimodal interaction device according to claim 9, characterized in that: The communication module includes a Bluetooth module and a Wi-Fi module. The Bluetooth module and the Wi-Fi module are respectively connected to the main control module. The Bluetooth communication module is used for device pairing, mobile terminal interaction, data synchronization and firmware upgrade, and the Wi-Fi module is used for device networking and cloud data communication.
14. The fusion multimodal interaction device according to claim 9, characterized in that: The fusion multimodal interaction device further includes: a memory and at least one processor, wherein the memory stores instructions, and the memory and the at least one processor are interconnected via a line; The at least one processor calls the instruction in the memory to enable the fusion multimodal interaction device to execute the fusion multimodal interaction method according to any one of claims 1 to 8.
15. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the fusion multimodal interaction method according to any one of claims 1 to 8 is implemented.
Citation Information
Cited By
Interaction control method, wearable device and computer readable storage medium
CN121455330A
Distributed screen multi-modal input fusion method and system based on open source gap
CN121541818A
Distributed screen multi-modal input fusion method and system based on open source honkong
CN121541818B