Intelligent multi-modal interaction method, system, and electronic device
By using a matrix acquisition array and time series model to analyze action trends in interactive devices, predictive compensation perception of user behavior is achieved, solving the problem of interaction latency and improving the responsiveness and sensitivity of interactive devices.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- DREAM INTELLIGENT TECHNOLOGY (DONGGUAN) CO LTD
- Filing Date
- 2026-03-13
- Publication Date
- 2026-06-19
AI Technical Summary
Existing interactive devices suffer from interaction latency, which affects the user experience.
A matrix-type acquisition array conforming to human body characteristics is used to generate continuous motion signals, analyze the dynamic change trend of motion using a time series model, predict user behavior, and implement a multimodal interaction method based on the prediction results to compensate for perception.
Reduce interaction latency, improve the responsiveness and sensitivity of interactive devices, and enhance the user experience.
Smart Images

Figure CN122239931A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of human-computer interaction technology, and in particular to an intelligent multimodal interaction method, system, and electronic device. Background Technology
[0002] Currently, interactive devices are input / output mechanisms that establish connections and transmit information between humans and electronic devices. Their core function is to achieve bidirectional transmission of commands and information feedback through direct interaction with human senses or motor organs. Based on functional differences, they can be divided into input devices and output devices. They play a crucial role in data transmission and feedback processing within electronic device systems, serving as input / output devices for establishing connections and exchanging information between humans and electronic devices. However, existing interactive devices suffer from a certain degree of interaction latency, affecting the user experience. Summary of the Invention
[0003] The purpose of this invention is to provide an intelligent multimodal interaction method, system, and electronic device to solve the technical problem that the interaction latency of interactive devices affects the user's interactive experience.
[0004] In a first aspect, this application provides an intelligent multimodal interaction method applied to an interactive device, wherein the sensor used for data acquisition in the interactive device is correspondingly pre-set with a matrix acquisition array conforming to human body feature structure; the method includes: In response to an action execution operation on the interactive device where the complete action corresponding to the action execution operation has not been completed, initial action feedback data corresponding to the action execution operation is determined, and a continuous action signal is generated based on the action execution operation using the matrix acquisition array; wherein, the continuous action signal contains a continuous action parameter sequence; Based on the continuous action parameter sequence, the dynamic change trend of the action is driven and analyzed by a time series model, and user behavior is predicted in advance based on the dynamic change trend of the action to obtain the behavior prediction result. Based on the behavior prediction results, the initial action feedback data is compensated using a prediction compensation perception method to obtain the behavior multimodal perception results of the interactive device in response to the action execution operation; wherein, the behavior multimodal perception results are used to represent the feedback of the interactive device to the user's behavioral intention corresponding to the action execution operation. The interactive device performs a response operation based on the behavioral multimodal perception results in response to the action.
[0005] In one possible implementation, the matrix acquisition array is a variable-configuration matrix sensing array made of a flexible material with a sensitivity greater than a specified value, and the matrix sensing array is used for layout detection and motion sensing of two-dimensional or three-dimensional posture; wherein the flexible material comprises conductive polymers and / or flexible strain films.
[0006] In one possible implementation, the matrix sensing array includes a variable number of flexible strain elements N×N; the action execution operation based on the action, utilizing the matrix acquisition array to generate a continuous action signal, includes: Based on the action execution operation, the flexible strain unit N×N is used to detect the user's action pressure distribution signal, motion posture change signal, sound signal and image signal in real time. The user's human body displacement change signal is detected by converting the trigger change data obtained from different sensing points in the matrix sensing area corresponding to the matrix sensing array into a simulated action sequence. Multiple signals detected by the N×N flexible strain unit are sampled by a multi-channel ADC in the form of a time-series matrix and transmitted to the MCU. The N×N matrix is mapped to a two-dimensional attitude data matrix corresponding to a two-dimensional or three-dimensional attitude model, and pressure and displacement heat maps are established and key area motion curves are extracted to obtain the attitude construction result. Among them, N is set based on the application scenario, and N is greater than or equal to 1.
[0007] In one possible implementation, the step of predicting user behavior in advance based on the dynamic change trend of the action to obtain the behavior prediction result includes: Based on the user's movement speed, direction, and dynamic change trend of the action, a sequence prediction algorithm is used to predict the user's movement trend and posture trajectory at the next moment, thereby obtaining a behavior prediction result. Based on the behavior prediction result, the next frame target image of the virtual object is generated in advance, so that the virtual object can perform a synchronous response with reduced latency. The sequence prediction algorithm is LSTM and / or TCN.
[0008] In one possible implementation, the step of driving and analyzing the dynamic change trend of the action based on the continuous action parameter sequence through a time series model includes: Based on the continuous motion parameter sequence, the following calculation formula is used to drive and analyze the dynamic change trend of the motion through a time series model:
[0009] in, Indicates the dynamic trend of action changes; Represents a sequence of continuous action parameters; t represents the time sampling interval; t represents the current sampling time.
[0010] In one possible implementation, the step of predicting user behavior in advance based on the dynamic change trend of the action to obtain the behavior prediction result includes: Based on the dynamic change trend of the actions, user behavior is predicted in advance using the following calculation formula to obtain the behavior prediction result:
[0011] in, This indicates the outcome of predicting behavior in advance; This indicates the current position / state of the action in the continuous sequence of action parameters; Indicates the current speed of the action; Indicates the position / state of the action at the next moment; The position / state of the action at the previous moment; Indicates the time span predicted in advance; Indicates the time sampling interval.
[0012] In one possible implementation, the interactive device includes an AR device; prior to determining the initial action feedback data corresponding to the action execution operation in response to an action execution operation performed on the interactive device and before the complete action corresponding to the action execution operation has been completed, the method further includes: In response to the AR device detecting that the real-world brightness of the user's location is lower than a specified brightness, the AR device emits invisible infrared light towards the user's eyes, and the AR device models the user's eye visual data based on the light reflected by the user in response to the invisible infrared light, thereby obtaining an eye visual modeling result, and the user's gaze direction is identified based on the eye visual modeling result. The AR device emits an invisible laser into the real scene in front of the user, and uses the invisible laser to measure the size and distance of objects in the real scene in front of the user to obtain real scene object data. Based on the real scene object data, the objects in the real scene are modeled to obtain a virtual three-dimensional real scene. The virtual 3D scene is displayed on the interface corresponding to the AR device according to the user's gaze direction, so that the user can perform the action operation on the virtual 3D scene.
[0013] Secondly, this application provides an intelligent multimodal interaction system applied to an interactive device, wherein the sensors used for data acquisition in the interactive device are correspondingly pre-set with a matrix acquisition array conforming to human body feature structures; the system includes: A generation module is configured to, in response to an action execution operation performed on the interactive device and the complete action corresponding to the action execution operation not being completed, determine initial action feedback data corresponding to the action execution operation, and generate a continuous action signal based on the action execution operation using the matrix acquisition array; wherein, the continuous action signal contains a continuous action parameter sequence; The prediction module is used to drive and analyze the dynamic change trend of the action based on the continuous action parameter sequence through a time series model, and to predict user behavior in advance based on the dynamic change trend of the action to obtain the behavior prediction result in advance. The compensation module is used to compensate the initial action feedback data based on the behavior prediction result using a prediction compensation perception method, so as to obtain the behavior multimodal perception result of the interactive device in response to the action execution operation; wherein, the behavior multimodal perception result is used to represent the feedback of the interactive device to the user behavior intention of the user performing the action corresponding to the action execution operation. The response module is used to perform the response process of the interactive device in response to the action based on the behavior multimodal perception results.
[0014] Thirdly, this application also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program that can run on the processor, and the processor executes the computer program to implement the method described in the first aspect above.
[0015] Fourthly, this application also provides a computer-readable storage medium storing computer-executable instructions that, when invoked and executed by a processor, cause the processor to perform the method described in the first aspect above.
[0016] This application brings the following beneficial effects: This application provides an intelligent multimodal interaction method, system, and electronic device, applied to an interactive device. The sensor used for data collection in this interactive device is pre-configured with a matrix acquisition array conforming to human body characteristics. The method responds to an action execution operation on the interactive device where the complete action corresponding to the action execution operation is not yet completed. It determines the initial action feedback data corresponding to the action execution operation and generates a continuous action signal based on the matrix acquisition array. The continuous action signal contains a continuous action parameter sequence. Based on the continuous action parameter sequence, a time series model is used to drive and analyze the dynamic change trend of the action, and user behavior is predicted in advance based on the dynamic change trend, obtaining a behavior prediction result. Based on the behavior prediction result, a prediction compensation perception method is used to compensate for the initial action feedback data, obtaining a multimodal perception result of the interactive device's behavior in response to the action execution operation. In this scheme, the behavior multimodal perception result is used to represent the feedback of the interactive device to the user's behavioral intent corresponding to the action execution operation; the interactive device responds to the action execution operation based on the behavior multimodal perception result; unlike the prior art that converts the action into discrete events such as start or end, in this scheme, the action corresponding to the action execution operation is identified as a continuous signal, the action is converted into a continuous parameter sequence, and then the continuous action parameters are used to drive dynamic changes. The action recognition in this scheme is based on the trend of change rather than absolute value. Moreover, unlike the response in the prior art which is a one-way trigger structure triggered by the completed action, the response in this scheme is based on the prediction result. It uses a closed-loop interaction structure of prediction compensation perception through action execution feedback and then execution perception, thereby improving the timeliness and sensitivity of the interactive device's response, reducing interaction latency, improving the user experience of the interactive device, and solving the technical problem that the interaction latency of the interactive device affects the user's interaction experience.
[0017] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the specific embodiments of this application or the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0019] Figure 1 A flowchart illustrating the intelligent multimodal interaction method provided in an embodiment of this application; Figure 2 Another flowchart illustrating the intelligent multimodal interaction method provided in this application embodiment; Figure 3 This is a schematic diagram of the structure of an intelligent multimodal interaction device provided in an embodiment of this application; Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of this application is shown. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0021] The terms "comprising" and "having," and any variations thereof, used in the embodiments of this application, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the steps or units listed, but may optionally include other steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices.
[0022] Currently, the interaction latency of interactive devices affects the user's interactive experience. Based on this, embodiments of this application provide an intelligent multimodal interaction method, system, and electronic device, which can solve the technical problem of the interaction latency of interactive devices affecting the user's interactive experience.
[0023] The embodiments of the present invention will be further described below with reference to the accompanying drawings.
[0024] Figure 1 This is a flowchart illustrating an intelligent multimodal interaction method provided in an embodiment of this application. The method is applied to an interactive device, where sensors for data collection are pre-configured with a matrix-type data acquisition array conforming to human body characteristics. Figure 1 As shown, the method includes: Step S110: In response to an action execution operation for the interactive device and the complete action corresponding to the action execution operation not being completed, determine the initial action feedback data corresponding to the action execution operation, and generate a continuous action signal based on the action execution operation using a matrix acquisition array.
[0025] Among them, the continuous action signal contains a continuous sequence of action parameters.
[0026] As one possible implementation, the matrix acquisition array is a variable-configuration matrix sensing array made of a flexible material with a sensitivity greater than a specified value. The matrix sensing array is used for layout detection and motion sensing of two-dimensional or three-dimensional postures. The flexible material comprises a conductive polymer and / or a flexible strain film. For example, a flexible matrix sensing array made of a high-sensitivity flexible material can achieve layout detection and motion sensing of two-dimensional or three-dimensional postures.
[0027] In practical applications, matrix acquisition arrays can be fitted to the human body structure that needs to be acquired, without defining a specific shape, but using a matrix array that conforms to human characteristics, such as the fingers of the hand, a piece of skin, a joint, a region, or the outer area that touches the human body.
[0028] In this application embodiment, the action corresponding to the action execution operation is regarded as a continuous signal process rather than a "start / end" event. Furthermore, this application embodiment allows for the presence of local noise, contact instability, or wearing deviation. Moreover, the response of this application solution is based on "predicted results" rather than "completed actions," using prediction to compensate for sensing, computation, and transmission delays. The action modeling method of this solution is completely different from existing solutions, which convert actions into discrete events, while this solution converts actions into a continuous parameter sequence.
[0029] In one alternative implementation, the matrix sensing array includes a variable number of flexible strain elements N×N; such as Figure 2 As shown, the above-mentioned action-based operation utilizes a matrix acquisition array to generate continuous action signals, which may specifically include the following steps: Step S210: Based on the action execution operation, the user's action pressure distribution signal, motion posture change signal, sound signal and image signal are detected in real time by the flexible strain unit N×N. The user's human body displacement change signal is detected by converting the trigger change data obtained from different sensing points in the matrix sensing area corresponding to the matrix sensing array into a simulated action sequence. Step S220: Multiple signals detected by the N×N flexible strain unit are sampled by a multi-channel ADC in the form of a time-series matrix and transmitted to the MCU. The N×N matrix is mapped to a two-dimensional attitude data matrix corresponding to a two-dimensional or three-dimensional attitude model, and pressure and displacement heat maps are established and key area motion curves are extracted to obtain the attitude construction result.
[0030] Where N is set based on the application scenario, and N is greater than or equal to 1. It should be noted that the matrix sensor array design and signal acquisition (mainly the wide threshold value of the algorithm and how the data is applied to improve the stability of the physical structure) are exemplary. In this embodiment, an N×N variable configuration flexible sensor unit array is used. Pressure sensing, physically adaptive structure, data noise filtering in the algorithm, and adaptive application range are all included. N can be set according to the application scenario, such as 1×1, 3×5, 8×10 or higher arrays. Each unit uses a highly sensitive flexible material (such as conductive polymer, flexible strain film). It should be noted that the array can detect: displacement changes (trigger changes obtained at different sensing points in the matrix sensing area are converted into simulated action sequences), pressure distribution, contact action, and action frequency; the array samples through multiple ADCs and transmits the data to the MCU to form a two-dimensional attitude data matrix. Therefore, the matrix flexible sensor array consists of a variable number of flexible strain units (N×N) for real-time acquisition of signals such as pressure changes, human movement, motion posture, heart rate, blood oxygen, sound, and images.
[0031] If combined with an IMU and photoelectric vital signs module, it is possible to simultaneously acquire: temperature, respiration, sound and speech, and image signals (via a camera); these multimodal signals are ultimately formed into a fused input data packet and coefficientd, which is then provided to the AI module for computation.
[0032] For the implementation flow of the data processing module, for example, the basic data processing module runs under a real-time operating system (RTOS) and mainly includes: raw signal preprocessing: dynamic filtering, low-pass / high-pass processing, noise suppression, and array normalization; attitude construction: mapping an N×N matrix to a two-dimensional / three-dimensional attitude model, establishing a pressure or displacement heatmap, and extracting motion curves in key areas; multimodal data fusion: aligning image, sound, vital signs, and touch data, unifying them to a timestamp system, and outputting them to the AI module for unified decision-making. Through the above processing, the system can complete the fusion of motion-vital signs-audio-visual data in milliseconds.
[0033] The overall architecture of the system corresponding to this solution mainly consists of the following modules: matrix sensor array (N×N), multi-source signal acquisition module (image, etc.), data processing module, AI interaction and adaptive learning module, control feedback module, audio and video output interface module, communication module, and terminal devices (mobile phone / computer / VR, etc.). Specifically: An intelligent multimodal interaction system based on matrix sensor array acquisition and AI multimodal presentation and control interaction, characterized by comprising: a matrix sensor array composed of N×N flexible sensing units for acquiring human motion displacement, pressure, etc., where N is greater than or equal to 1; images and sounds; a data processing module for sampling, filtering, feature extraction, and posture construction of the sensor array signals; an AI interaction and adaptive learning module for behavior recognition, trend prediction, and response control based on actions, vital signs, and historical data; a control feedback module for generating control signals to drive electromechanical actuators or virtual feedback units to achieve synchronous response; an audio-visual output interface module for transmitting audio, image, or video control signals generated by AI interaction to terminal devices, whereby the terminal devices complete the audio-visual content output; and a communication module for enabling data transmission and remote interactive control between multiple terminals.
[0034] For the specific data processing flow, for example, the basic data processing module runs under a real-time operating system (RTOS) and mainly includes: raw signal preprocessing: dynamic filtering, low-pass / high-pass processing, noise suppression, and array normalization; attitude construction: mapping an N×N matrix into a two-dimensional / three-dimensional attitude model, pressure or displacement heatmap, and extracting motion curves for key areas; multimodal data fusion: aligning image, sound, vital signs, and touch data, unifying them to a timestamp system, and outputting them to the AI module for unified decision-making. Through the above processing, the system can complete the fusion of motion-vital signs-audio-visual data in milliseconds.
[0035] In practical applications, users can wear sensors to analyze and predict local user actions and transmit the data to remote devices. These remote devices can be virtual characters, robotic arms, smart toys, VR / AR virtual avatars, etc., and the interaction methods can be interactive or mirror simulation.
[0036] This intelligent multimodal interaction system, based on matrix sensor array data acquisition, AI multimodal presentation, and control interaction, belongs to the fields of smart wearables, artificial intelligence, and human-computer interaction. The system includes: a matrix sensor array, image, audio and voice, terminal device screen touch control, text acquisition, data processing modules, AI interaction and adaptive learning modules, control feedback modules, audio and video output interface modules, and communication modules.
[0037] Step S120: Based on the continuous action parameter sequence, the dynamic change trend of the action is driven and analyzed by the time series model, and the user behavior is predicted in advance based on the dynamic change trend of the action to obtain the behavior prediction result.
[0038] In this application embodiment, unlike existing action-triggered feedback schemes which typically possess one or a combination of the following characteristics: the action is identified as a discrete event (such as triggering, clicking, exceeding a threshold); the action identification in this application scheme is based on "change trend" rather than "absolute value". Furthermore, this application scheme differs from existing control logic: existing schemes trigger behavior through events; this scheme uses continuous action parameters to drive dynamic changes. Moreover, the timing relationship differs: existing schemes respond after the action is completed; this scheme generates responses in advance based on action trend prediction.
[0039] As an optional implementation, the above-mentioned method of driving and analyzing the dynamic change trend of actions based on continuous action parameter sequences through a time series model may specifically include the following steps: Based on a continuous sequence of motion parameters, the following calculation formula is used to drive and analyze the dynamic change trend of motion through a time series model:
[0040] in, Indicates the dynamic trend of action changes; Represents a sequence of continuous action parameters; t represents the time sampling interval; t represents the current sampling time.
[0041] It should be noted that, T ( t ): The dynamic trend of the action. This is a vector representing the trend of the action at time t. t At any given moment, the direction (vector direction) and intensity (vector magnitude) of the action. P ( t ): A sequence of continuous motion parameters. This is raw data acquired from a sensor array, typically including position coordinates, angles, pressure values, etc. t The instantaneous state at a given moment. Δ t Time sampling interval: The fixed time step of the matrix acquisition array. Differential operator: Driving logic. Velocity is derived by differentiating position, or acceleration is derived by differentiating velocity, thus describing the trend. This is a typical non-weighted, calculus-based calculation formula.
[0042] In one possible implementation, the above-mentioned prediction of user behavior based on the dynamic change trend of actions to obtain the behavior prediction result may specifically include the following steps: Based on the user's movement speed, direction, and dynamic change trend, a sequence prediction algorithm is used to predict the user's movement trend and posture trajectory at the next moment, thus obtaining the behavior prediction result. The next frame target image of the virtual object is generated in advance based on the behavior prediction result, so that the virtual object can perform a synchronous response with reduced latency. The sequence prediction algorithm is LSTM and / or TCN.
[0043] For the implementation of the AI interaction and adaptive learning module, the AI module, for example, includes three sub-units: 1. Action recognition sub-module: Based on deep neural networks (such as CNN + LSTM), it takes temporal matrix data from the sensor array as input and can recognize: stress patterns, user vital signs, gestures / limb movements, directional movements, posture changes, and user behavior classification. 2. Behavior prediction sub-module: Uses sequence prediction algorithms (such as LSTM, Transformer, TCN) to predict the following at the next moment: motion trend and posture trajectory.
[0044] In practical applications, sensors capture users' movements and subtle posture changes in real time, and predictively generate the target action for the next frame in advance through predictive algorithms, enabling virtual digital humans (such as virtual anchors, virtual idols, or virtual companions) to respond synchronously with extremely low latency. By combining parameters such as voice and force changes to generate emotional states, the digital human's facial expressions, voice, and interactions can be driven to engage in natural conversational or companion-like emotional interactions with the user.
[0045] As an example, in interactive toy entertainment scenarios, users trigger sensors to drive virtual characters or multimedia content, predicting the next action based on the user's speed and direction of movement. The smart toy or digital content then provides feedback based on the predicted parameters, including sound, light, or physical actions, creating a complete closed-loop experience.
[0046] As an optional implementation, the above-mentioned prediction of user behavior based on the dynamic change trend of actions to obtain the prediction result may specifically include the following steps: Based on the dynamic changes in action trends, user behavior is predicted in advance using the following calculation formula to obtain the prediction results:
[0047] in, This indicates the outcome of predicting behavior in advance; This indicates the current position / state of the action in the continuous sequence of action parameters; Indicates the current speed of the action; Indicates the position / state of the action at the next moment; The position / state of the action at the previous moment; Indicates the time span predicted in advance; Indicates the time sampling interval.
[0048] In practical applications, terminal devices include, but are not limited to, smartphones, tablets, personal computers, VR / AR devices, or other multimedia playback devices. Furthermore, the AI interaction and adaptive learning module includes: a motion recognition submodule for recognizing user posture and movement patterns; a behavior prediction submodule for predicting action trends based on time-series models; and a text-audio-visual decision-making submodule for generating multimodal control signals to drive the terminal device to perform text dialogue or audio / video output. The text-audio-visual output interface module provides standardized communication interfaces or data protocols, including but not limited to Bluetooth, USB, Wi-Fi, or internet interfaces, for data interaction with the terminal device.
[0049] Step S130: Based on the behavior prediction results, the initial action feedback data is compensated and perceived using a prediction compensation perception method to obtain the behavior multimodal perception results of the interactive device for the action execution operation.
[0050] Among them, the behavior multimodal perception result is used to represent the feedback of the interactive device to the user's behavioral intent in response to the action execution operation.
[0051] In this embodiment, the response is based on "predicted results" rather than "completed actions," and the solution utilizes prediction to compensate for perception, computation, and transmission delays. Furthermore, the system runs a real-time operating system (RTOS) through a microcontroller unit (MCU) to achieve high-frequency data sampling and low-latency feedback control. The AI interaction and adaptive learning module supports cloud-based model training and local model fusion to achieve personalized user behavior modeling and response optimization. Moreover, the system in this solution features an audio-visual synchronization algorithm, ensuring that the audio and video signals output by the terminal device maintain timing consistency with the user's actions, with a response latency in the millisecond range, imperceptible to human senses. The communication module supports point-to-point or multi-terminal interaction, enabling remote synchronous control and action sharing. The system is suitable for scenarios such as intelligent sports training, remote rehabilitation, virtual interaction, emotional companionship, and multimedia experiences.
[0052] Existing solutions employ a one-way triggering structure, while this application's solution utilizes a closed-loop interactive structure where action execution feedback precedes perception. By integrating sensor data acquisition and AI algorithms, it achieves human motion recognition and multimodal interactive control, exhibiting characteristics such as high sensitivity, high synchronization, and immersive experience. It is suitable for scenarios such as intelligent rehabilitation, virtual training, remote collaboration, and emotional companionship.
[0053] Step S140: Based on the behavioral multimodal perception results, the interactive device responds to the action execution operation.
[0054] In one possible implementation, the AI interaction module performs calculations, recognition, prediction, and response control based on action signals and user behavior models; the media output interface module transmits the multimodal feedback signals generated by the AI interaction to terminal devices (including but not limited to smart hardware, mobile phones, computers, VR devices, etc.), and the terminal completes the presentation and output of text, audio, and video content; the control feedback module generates execution signals to drive external human-computer interaction devices such as smart hardware and smart wearables. For the text / audio-visual decision submodule, based on the recognition and prediction results, it generates: dynamic audio change parameters, video rendering instructions, subtitle feedback content, virtual scene change control, and execution commands for external hardware. The AI module can work through cloud training combined with local import execution, achieving personalized behavioral interaction without consuming excessive cloud computing power.
[0055] For the implementation of control feedback, for example, the control feedback module converts the action commands output by the AI into: drive signals for electromechanical actuators, intensity adjustment of the flexible haptic feedback device, and motion rendering control for the virtual scene. This module communicates with peripherals through interfaces such as PWM, DAC, I²C, SPI, or UART.
[0056] For the implementation of the audio and video output interface, for example, the audio and video commands generated by the AI module are transmitted to the terminal device through this module, supporting the following output methods: Bluetooth (BLE), Wi-Fi (dual-band), WebRTC, and 2.4G radio frequency. All sound, images, video, subtitles, and other content are presented by the terminal device, ensuring clear system boundaries. The terminal may include: mobile app, PC software, web browser, VR / AR glasses, and smart hardware display devices. The audio and video output response latency can be controlled at the millisecond level, and a synchronization mechanism ensures that the user perceives no significant delay.
[0057] In practical applications, the communication module synchronizes with multiple terminals and supports: point-to-point (P2P) transmission, dual-terminal interactive synchronization, remote action reproduction, encrypted channels (AES / SSL / TLS), and synchronized time base (NTP / built-in local clock); it can be used to build scenarios such as remote collaboration, remote rehabilitation, and virtual interaction.
[0058] Regarding the application scenarios of this application, the solution can be applied to: 1. Intelligent sports training scenarios: hand / limb motion capture, pressure distribution analysis, and AI motion correction; 2. Remote rehabilitation assistance: posture assessment, motion amplitude detection, and synchronized motion training; 3. Virtual interaction: motion-driven virtual avatars and VR immersive interaction; 4. Emotional companionship: multimodal interactive feedback based on body posture, sound, and text, with AI-generated audio-visual content automatically changing according to movements; 5. Multimedia experience: AI-driven adaptive changes in video / audio, and real-time motion-driven animation or scene rendering. The system possesses the characteristics of high sensitivity, low latency, and immersive experience.
[0059] Unlike existing technologies that convert actions into discrete events such as start or end, in this embodiment, the action corresponding to the action execution operation is identified as a continuous signal. The action is converted into a continuous parameter sequence and then the continuous action parameters are used to drive dynamic changes. The action recognition in this solution is based on the trend of change rather than absolute values. Moreover, unlike the unidirectional triggering structure in existing technologies where the response is triggered by the completed action, the response in this solution is based on the prediction result. It uses a closed-loop interaction structure of prediction compensation perception, action execution feedback, and then execution perception, thereby improving the responsiveness and sensitivity of the interactive device, reducing interaction latency, and improving the user experience of the interactive device.
[0060] In some embodiments, the interactive device includes an AR device; before determining the initial action feedback data corresponding to the action execution operation in response to an action execution operation on the interactive device and before the complete action corresponding to the action execution operation has been completed, the method may further include the following steps: In response to the AR device detecting that the brightness of the real scene where the user is located is lower than the specified brightness, the AR device emits invisible infrared light towards the user's eyes, and the user's eye visual data is modeled based on the light reflected by the user in response to the invisible infrared light detected by the AR device, to obtain the eye visual modeling result, and the user's gaze direction is identified based on the eye visual modeling result. An AR device emits an invisible laser into the real-world scene in front of the user. The size and distance of objects in the real-world scene in front of the user are detected by the invisible laser ranging, and the real-world object data is obtained. Based on the real-world object data, the objects in the real-world scene are modeled to obtain a virtual 3D real-world scene. The virtual 3D real-world scene is displayed on the corresponding interface of the AR device according to the user's line of sight, so that the user can perform actions on the virtual 3D real-world scene.
[0061] By utilizing active invisible light detection technology, the limitations imposed by ambient light on the core elements of AR interaction (eye tracking and environmental perception) are eliminated. This allows for a consistent, accurate, and natural virtual-real fusion interactive experience for users under any lighting conditions, particularly ensuring the usability and robustness of complex actions based on gaze and gestures in low-light scenarios. Thus, even in low-light environments, it can still accurately and without interference identify the user's gaze direction and construct a precise virtual 3D scene. This ensures that users can perform actions (such as gesture control and object grasping) on the virtual scene naturally, smoothly, and accurately within the AR interface, solving the problems of eye tracking failure and inaccurate environmental modeling caused by insufficient visible light in traditional AR devices under low-light conditions.
[0062] Figure 3 A schematic diagram of an intelligent multimodal interaction device is provided. This device can be applied to interactive devices, where the sensors used for data collection correspond to a pre-set matrix acquisition array conforming to human body characteristics. Figure 3 As shown, the intelligent multimodal interaction device 300 includes: The generation module 301 is configured to, in response to an action execution operation performed on the interactive device and the complete action corresponding to the action execution operation not being completed, determine initial action feedback data corresponding to the action execution operation, and generate a continuous action signal based on the action execution operation using the matrix acquisition array; wherein, the continuous action signal contains a continuous action parameter sequence; The prediction module 302 is used to drive and analyze the dynamic change trend of the action based on the continuous action parameter sequence through a time series model, and to predict user behavior in advance based on the dynamic change trend of the action to obtain the behavior prediction result in advance. The compensation module 303 is used to compensate the initial action feedback data based on the behavior prediction result using a prediction compensation perception method, so as to obtain the behavior multimodal perception result of the interactive device in response to the action execution operation; wherein, the behavior multimodal perception result is used to represent the feedback of the interactive device to the user behavior intention of the user performing the action corresponding to the action execution operation. The response module 304 is used to perform the response process of the interactive device to the action based on the behavior multimodal perception results.
[0063] The intelligent multimodal interaction device provided in this application embodiment has the same technical features as the intelligent multimodal interaction method provided in the above embodiments, so it can also solve the same technical problems and achieve the same technical effects.
[0064] An electronic device provided in this application embodiment, such as Figure 4As shown, the electronic device 400 includes a processor 402 and a memory 401. The memory stores a computer program that can run on the processor. When the processor executes the computer program, it implements the steps of the method provided in the above embodiments.
[0065] See Figure 4 The electronic device also includes a bus 403 and a communication interface 404. The processor 402, the communication interface 404 and the memory 401 are connected via the bus 403. The processor 402 is used to execute executable modules, such as computer programs, stored in the memory 401.
[0066] The memory 401 may include high-speed random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Communication between this system network element and at least one other network element is achieved through at least one communication interface 404 (which can be wired or wireless), such as the Internet, wide area network, local area network, or metropolitan area network.
[0067] Bus 403 can be an ISA bus, PCI bus, or EISA bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 4 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.
[0068] The memory 401 is used to store programs. After receiving an execution instruction, the processor 402 executes the program. The method executed by the apparatus defined by the process disclosed in any of the preceding embodiments of this application can be applied to the processor 402 or implemented by the processor 402.
[0069] Processor 402 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of processor 402 or by instructions in software form. The processor 402 can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly manifested as execution by a hardware decoding processor, or execution by a combination of hardware and software modules in the decoding processor. The software module can reside in a mature storage medium in the field, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory 401, and processor 402 reads the information from memory 401 and, in conjunction with its hardware, completes the steps of the above method.
[0070] Corresponding to the above-described intelligent multimodal interaction method, this application embodiment also provides a computer-readable storage medium storing computer-executable instructions. When the computer-executable instructions are invoked and executed by a processor, the computer-executable instructions cause the processor to perform the steps of the above-described intelligent multimodal interaction method.
[0071] The intelligent multimodal interaction device provided in this application embodiment can be specific hardware on a device or software or firmware installed on the device. The implementation principle and technical effects of the device provided in this application embodiment are the same as those in the foregoing method embodiments. For the sake of brevity, any parts not mentioned in the device embodiment can be referred to the corresponding content in the foregoing method embodiments. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can all be referred to the corresponding processes in the above method embodiments, and will not be repeated here.
[0072] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some communication interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.
[0073] For example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0074] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0075] In addition, the functional units in the embodiments provided in this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0076] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the intelligent multimodal interaction method described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0077] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In addition, the terms "first", "second", "third", etc. are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0078] Finally, it should be noted that the above-described embodiments are merely specific implementations of this application, used to illustrate the technical solutions of this application, and not to limit them. The protection scope of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this application; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application. All should be covered within the protection scope of this application. Therefore, the protection scope of this application should be determined by the protection scope of the claims.
Claims
1. An intelligent multimodal interaction method, characterized in that, The method is applied to an interactive device, wherein the sensor used for data acquisition in the interactive device is correspondingly equipped with a matrix acquisition array that conforms to the human body's characteristic structure; the method includes: In response to an action execution operation on the interactive device where the complete action corresponding to the action execution operation has not been completed, initial action feedback data corresponding to the action execution operation is determined, and a continuous action signal is generated based on the action execution operation using the matrix acquisition array; wherein, the continuous action signal contains a continuous action parameter sequence; Based on the continuous action parameter sequence, the dynamic change trend of the action is driven and analyzed by a time series model, and user behavior is predicted in advance based on the dynamic change trend of the action to obtain the behavior prediction result. Based on the behavior prediction results, the initial action feedback data is compensated using a prediction compensation perception method to obtain the behavior multimodal perception results of the interactive device in response to the action execution operation; wherein, the behavior multimodal perception results are used to represent the feedback of the interactive device to the user's behavioral intention corresponding to the action execution operation. The interactive device performs a response operation based on the behavioral multimodal perception results in response to the action.
2. The method according to claim 1, characterized in that, The matrix acquisition array is a variable-configuration matrix sensing array, which is made of a flexible material with a sensitivity greater than a specified value. The matrix sensing array is used for layout detection and motion sensing of two-dimensional or three-dimensional posture. The flexible material includes conductive polymers and / or flexible strain films.
3. The method according to claim 2, characterized in that, The matrix sensing array includes a variable number of flexible strain elements N×N; the action execution operation based on the action utilizes the matrix acquisition array to generate a continuous action signal, including: Based on the action execution operation, the flexible strain unit N×N is used to detect the user's action pressure distribution signal, motion posture change signal, sound signal and image signal in real time. The user's human body displacement change signal is detected by converting the trigger change data obtained from different sensing points in the matrix sensing area corresponding to the matrix sensing array into a simulated action sequence. Multiple signals detected by the N×N flexible strain unit are sampled by a multi-channel ADC in the form of a time-series matrix and transmitted to the MCU. The N×N matrix is mapped to a two-dimensional attitude data matrix corresponding to a two-dimensional or three-dimensional attitude model, and pressure and displacement heat maps are established and key area motion curves are extracted to obtain the attitude construction result. Among them, N is set based on the application scenario, and N is greater than or equal to 1.
4. The method according to claim 1, characterized in that, The step of predicting user behavior in advance based on the dynamic change trend of the action to obtain the behavior prediction result includes: Based on the user's movement speed, direction, and dynamic change trend of the action, a sequence prediction algorithm is used to predict the user's movement trend and posture trajectory at the next moment, thereby obtaining a behavior prediction result. Based on the behavior prediction result, the next frame target image of the virtual object is generated in advance, so that the virtual object can perform a synchronous response with reduced latency. The sequence prediction algorithm is LSTM and / or TCN.
5. The method according to claim 1, characterized in that, The process of driving and analyzing the dynamic change trend of actions based on the continuous action parameter sequence through a time series model includes: Based on the continuous motion parameter sequence, the following calculation formula is used to drive and analyze the dynamic change trend of the motion through a time series model: in, Indicates the dynamic trend of action changes; Represents a sequence of continuous action parameters; t represents the time sampling interval; t represents the current sampling time.
6. The method according to claim 1, characterized in that, The step of predicting user behavior in advance based on the dynamic change trend of the action to obtain the behavior prediction result includes: Based on the dynamic change trend of the actions, user behavior is predicted in advance using the following calculation formula to obtain the behavior prediction result: in, This indicates the outcome of predicting behavior in advance; This indicates the current position / state of the action in the continuous sequence of action parameters; Indicates the current speed of the action; Indicates the position / state of the action at the next moment; The position / state of the action at the previous moment; Indicates the time span predicted in advance; Indicates the time sampling interval.
7. The method according to claim 1, characterized in that, The interactive device includes an AR device; before determining the initial action feedback data corresponding to the action execution operation in response to an action execution operation on the interactive device and before the complete action corresponding to the action execution operation is completed, the method further includes: In response to the AR device detecting that the real-world brightness of the user's location is lower than a specified brightness, the AR device emits invisible infrared light towards the user's eyes, and the AR device models the user's eye visual data based on the light reflected by the user in response to the invisible infrared light, thereby obtaining an eye visual modeling result, and the user's gaze direction is identified based on the eye visual modeling result. The AR device emits an invisible laser into the real scene in front of the user, and uses the invisible laser to measure the size and distance of objects in the real scene in front of the user to obtain real scene object data. Based on the real scene object data, the objects in the real scene are modeled to obtain a virtual three-dimensional real scene. The virtual 3D scene is displayed on the interface corresponding to the AR device according to the user's gaze direction, so that the user can perform the action operation on the virtual 3D scene.
8. An intelligent multimodal interaction system, characterized in that, The system is applied to interactive devices, wherein the sensors used for data collection in the interactive devices are correspondingly pre-set with a matrix-type acquisition array conforming to human body feature structures; the system includes: A generation module is configured to, in response to an action execution operation performed on the interactive device and the complete action corresponding to the action execution operation not being completed, determine initial action feedback data corresponding to the action execution operation, and generate a continuous action signal based on the action execution operation using the matrix acquisition array; wherein, the continuous action signal contains a continuous action parameter sequence; The prediction module is used to drive and analyze the dynamic change trend of the action based on the continuous action parameter sequence through a time series model, and to predict user behavior in advance based on the dynamic change trend of the action to obtain the behavior prediction result in advance. The compensation module is used to compensate the initial action feedback data based on the behavior prediction result using a prediction compensation perception method, so as to obtain the behavior multimodal perception result of the interactive device in response to the action execution operation; wherein, the behavior multimodal perception result is used to represent the feedback of the interactive device to the user behavior intention of the user performing the action corresponding to the action execution operation. The response module is used to perform the response process of the interactive device in response to the action based on the behavior multimodal perception results.
9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions that, when invoked and executed by a processor, cause the processor to perform the method according to any one of claims 1 to 7.