An intelligent terminal interaction system based on AI adaptive perception

The AI-adaptive perception-based intelligent terminal interaction system solves the problems of fixed multimodal fusion methods and insufficient user state coupling in the interaction system, realizes dynamic adjustment of interaction strategies and fault degradation protection, and improves the interaction adaptability and reliability of portable terminals.

CN122431534APending Publication Date: 2026-07-21GUANGDONG HENGXIANG SAFETY TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGDONG HENGXIANG SAFETY TECH CO LTD
Filing Date
2026-05-04
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing smart terminal interaction systems suffer from fixed multimodal fusion methods, insufficient coupling between interaction strategies and real-time user status, inadequate control over latency and power consumption during local deployment of portable terminals, difficulty in closed-loop updates of interaction feedback, and lack of degradation protection under abnormal operating conditions.

Method used

An AI-based adaptive perception-based intelligent terminal interaction system is adopted. It uses a state vector composed of feature processing, user status, and terminal status under a unified time reference, dynamic weight generation, and cross-media collaborative output. Combined with feedback updates and fault degradation, it forms a closed-loop control system, including modules for multimodal information acquisition, preprocessing, user status perception, adaptive decision-making, cross-media interaction execution, and fault self-diagnosis.

Benefits of technology

It improves interactive adaptability, real-time performance and reliability, reduces interactive latency and power consumption, and realizes fault self-diagnosis and degradation protection under abnormal operating conditions, while taking into account user privacy and model optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122431534A_ABST
    Figure CN122431534A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of intelligent terminal human-computer interaction, and discloses an intelligent terminal interaction system based on Al adaptive perception. The system comprises a multi-modal information acquisition module, a preprocessing module, a user state perception module, an adaptive decision module, a cross-medium interaction execution module, a feedback updating module and a fault self-diagnosis module. The system synchronously aligns and encodes the features of visual, thermal imaging, touch, voice, posture and device state data to generate a state vector; dynamically adjusts the weights of each mode according to the state vector and outputs an interaction strategy; cooperatively executes through at least two of display, touch and voice; and locally updates or performs degradation protection according to user feedback and device state. The system can improve the interaction adaptability, real-time performance and reliability under abnormal working conditions on a portable intelligent terminal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of human-computer interaction technology for smart terminals, specifically to a smart terminal interaction system based on AI adaptive perception, applicable to smartphones, smart wearable devices, portable smart interactive terminals, and other electronic devices with local perception and interaction functions. Background Technology

[0002] Existing smart terminals typically possess camera, touch, voice input, and display output capabilities, but most interactive systems still primarily rely on multimodal fusion methods with fixed rules or fixed weights. When ambient light changes, the user is in motion, the user's hand movements are unstable, the ambient noise is high, or the terminal is operating under high load, fixed fusion methods struggle to accurately reflect the reliability of data from each modality, easily leading to misidentification, incorrect feedback, or interaction delays.

[0003] Furthermore, existing terminal interaction systems often treat perception, decision-making, execution, and updating as independent processes. User feedback after interaction execution is difficult to promptly influence subsequent decisions, preventing the system from continuously adjusting to user habits and operating environments. While some solutions employing complex models can improve recognition capabilities, they incur significant computational, power, and storage overhead, hindering local deployment on portable terminals.

[0004] Furthermore, existing interactive systems do not adequately utilize the device's operating status. When the terminal experiences excessive temperature rise, abnormal current, abnormal vibration, or excessive processor load, the interactive system typically lacks a self-diagnostic and degradation operation mechanism for the interactive link, which may lead to slow response, abnormal haptic feedback, display lag, or user misoperation.

[0005] Therefore, it is necessary to provide an intelligent terminal interaction system that can form a closed loop of multimodal perception, user state perception, dynamic weight fusion, cross-media execution, feedback update and fault degradation control, so as to improve the interaction adaptability, real-time performance and reliability of different users and different scenarios. Summary of the Invention

[0006] Technical problems to be solved This invention aims to solve the problems in existing smart terminal interaction systems, such as fixed multimodal fusion methods, insufficient coupling between interaction strategies and real-time user status, inadequate control of latency and power consumption during local deployment of portable terminals, difficulty in closed-loop updating of interaction feedback, and lack of degradation protection under abnormal operating conditions.

[0007] Technical solution

[0008] To address the aforementioned issues, this invention provides an AI-based adaptive perception-based intelligent terminal interaction system. Instead of simply juxtaposing multiple sensors and output methods, this system uses feature processing under a unified time reference as the front end, a state vector composed of user and terminal states as intermediate constraints, dynamic weight generation and cross-media collaborative output as the core processing path, and closed-loop control through feedback updates and fault degradation.

[0009] Specifically, the system includes a multimodal information acquisition module, a preprocessing module, a user state perception module, an adaptive decision-making module, a cross-media interaction execution module, a feedback update module, and a fault self-diagnosis module. The multimodal information acquisition module collects data such as visual, thermal imaging, tactile, voice, posture, and device operating status. Visual data can come from a visible light image sensor to acquire images of the user's face, gestures, or the object being manipulated; thermal imaging data can come from an infrared thermal imaging sensor to acquire the user's local temperature or the local heat distribution of the terminal; tactile data can come from a flexible tactile sensor to acquire the pressing position, pressure, contact area, and contact duration; voice data can come from a microphone to acquire voice commands, voice intensity, and voice rhythm; posture data can come from an inertial measurement unit to characterize the terminal's movement and grip state; and device operating status data can come from temperature, current, vibration, and processor load acquisition units to describe the terminal's operating status.

[0010] The preprocessing module transforms data from different sensors into comparable and fusionable feature data. First, it assigns timestamps to each modality's data and uses tactile, speech, or other faster-responding modalities as reference time bases, resampling or interpolating to compensate for other modalities. Then, it performs denoising based on data type; for example, median or Gaussian filtering is used for visual data, moving average filtering for tactile data, bandpass filtering for speech data, and temporal smoothing for thermal imaging data. Finally, it encodes the different modalities into feature vectors of a unified dimension or those mappable to a unified dimension.

[0011] The user state awareness module extracts at least three of the following from the preprocessed features: user gaze area, hand gesture amplitude, touch pressure, voice intensity, voice emotion intensity, ambient brightness, remaining battery level, processor load, and terminal temperature. These features are then concatenated into a state vector. This state vector simultaneously represents the user's current interaction intent, user operation stability, external environmental conditions, and terminal operating pressure, enabling subsequent decisions to be based not only on a single input signal but also on a combination of scenario reliability and device capabilities.

[0012] The adaptive decision-making module comprises a modality encoding subunit, a weight generation subunit, and a policy output subunit. The modality encoding subunit maps input features such as visual, tactile, thermal imaging, speech, and device status into encoded vectors of a unified dimension. The weight generation subunit generates weights for each modality based on the state vectors and normalizes them to ensure the sum of the weights is 1. The policy output subunit outputs the interaction policy based on the weighted fusion result. The interaction policy includes at least two types of parameters: display content parameters, haptic feedback parameters, and speech output parameters. Display content parameters may include display content type, cue level, and brightness; haptic feedback parameters may include vibration frequency, amplitude, and duration; and speech output parameters may include cue content, volume, and playback order.

[0013] The cross-media interaction execution module controls at least two of the following methods—display, haptic, and voice—to perform collaboratively based on the interaction strategy. Compared to a single output method, cross-media collaborative output can select a more reliable combination of outputs in low-light conditions, noisy environments, motion states, or when the user's gaze is unstable. For example, it can reduce complex visual cues, enhance haptic cues, or increase the priority of voice output when the user cannot look at the screen.

[0014] The feedback update module acquires at least one type of feedback information, including user confirmation, cancellation, number of repeated operations, response time, satisfaction rating, and device power consumption changes, and converts this feedback information into parameter update signals. Feedback updates can be performed locally, or, upon user authorization or when preset upload conditions are met, the incremental model data, after parameter trimming and privacy perturbation, can be uploaded to the cloud aggregation unit. The system does not require uploading original user images, original voice, or original haptic data, thus reducing privacy risks.

[0015] The fault self-diagnosis module performs joint analysis on at least two of the following signals: temperature, current, vibration, and processor load, to generate an anomaly score. When the anomaly score exceeds a preset threshold, the system triggers degraded operation or safety protection control. Degraded operation may include disabling high-power display effects, reducing haptic feedback intensity, reducing model inference frequency, retaining basic voice or communication capabilities, recording anomaly logs, and uploading fault information. Thus, the system can maintain basic interactive capabilities and prevent fault escalation even under abnormal terminal operating conditions.

[0016] In a preferred embodiment, the system's main sensing, weight generation, and policy output are all executed locally on the terminal, with the cloud only used for optional model aggregation and parameter feedback. This ensures both low-latency interaction and continuous model optimization across multiple terminal scenarios.

[0017] Beneficial effects

[0018] Compared with existing technologies, this invention provides an AI-based adaptive perception-based intelligent terminal interaction system, which has the following beneficial effects: It can dynamically adjust the weight of different modalities in interaction decisions based on user status and terminal status, thereby improving interaction adaptability; It enables user and device feedback to influence subsequent strategy generation, forming an interactive closed loop. It can make major decisions locally on the terminal, reducing interaction latency and dependence on the cloud; It can improve system availability under abnormal operating conditions through fault self-diagnosis and degradation operation mechanisms; The optional privacy-preserving incremental upload method allows for iterative optimization of the model while protecting user privacy. Attached Figure Description

[0019] Figure 1 This is a schematic diagram of the overall system structure of the present invention.

[0020] Figure 2 This is a schematic diagram of the multimodal information acquisition and preprocessing module of the present invention.

[0021] Figure 3 This is a schematic diagram of the user state perception and adaptive decision-making module of the present invention.

[0022] Figure 4 This is a schematic diagram of the cross-media interactive execution module of the present invention.

[0023] Figure 5 This is a flowchart illustrating the feedback update and cloud aggregation process of this invention.

[0024] Figure 6 This is a flowchart of the fault self-diagnosis and degradation operation of the present invention. Detailed Implementation

[0025] Example 1: System Overall Structure refer to Figure 1 The intelligent terminal interaction system provided in this embodiment includes a multimodal information acquisition module, a preprocessing module, a user state perception module, an adaptive decision-making module, a cross-media interaction execution module, a feedback update module, a fault self-diagnosis module, and an optional cloud aggregation unit. Each module can be integrated into the same portable terminal, or the cloud aggregation unit can be located on a server.

[0026] The multimodal information acquisition module collects data related to user operation, user status, usage environment, and device operation during terminal interaction. The preprocessing module converts data from different sources into feature data under a unified time base. The user status perception module generates a state vector based on the multimodal features. The adaptive decision-making module generates modal weights based on the state vector and outputs an interaction strategy. The cross-media interaction execution module controls at least two outputs from display, touch, and voice according to the interaction strategy. The feedback update module converts user feedback and device feedback into parameter update signals. The fault self-diagnosis module monitors the device's operating status and triggers degraded operation when an anomaly occurs.

[0027] Example 2: Multimodal Acquisition and Preprocessing refer to Figure 2 The multimodal information acquisition module may include at least three of the following: a visible light image sensor, an infrared thermal imaging sensor, a flexible tactile sensor, a voice / posture sensor, and a device status sensor. The visible light image sensor is used to acquire images of the user's face, gestures, or the object being manipulated; the infrared thermal imaging sensor is used to acquire the user's local temperature or the local heat distribution of the terminal; the flexible tactile sensor is used to acquire the pressing position, pressure, contact area, and duration; the voice / posture sensor is used to acquire voice intensity, voice rhythm, terminal posture, and amplitude of movement; and the device status sensor is used to acquire the terminal's operating status, such as temperature, current, vibration, and processor load.

[0028] The preprocessing module includes a synchronization alignment unit, a noise suppression unit, and a feature encoding unit. The synchronization alignment unit appends a timestamp to each modality data point and performs resampling or interpolation compensation based on the reference sampling time. Let the sampling time of the i-th modality data be t_i, and the reference sampling time be t_r, then the time compensation amount is: Δt_i=t_i-t_r Where Δt_i represents the time compensation amount of the i-th mode relative to the reference sampling time; t_i represents the sampling time of the i-th mode data; t_r represents the sampling time of the reference mode; and i represents the mode number. When Δt_i is not zero, interpolation compensation or resampling is performed on the mode data to ensure that all mode data enter subsequent processing under a unified time base.

[0029] The noise suppression unit can perform median filtering or Gaussian filtering on visual data, moving average filtering on tactile data, bandpass filtering on speech data, and temporal smoothing on thermal imaging data. The feature encoding unit generates at least two of the following: visual features F_v, tactile features F_t, thermal imaging features F_h, speech features F_a, and device status features F_d. Here, F_v represents the visual feature vector, F_t represents the tactile feature vector, F_h represents the thermal imaging feature vector, F_a represents the speech feature vector, and F_d represents the device status feature vector.

[0030] Example 3: User State Awareness and Adaptive Decision Making refer to Figure 3 The user state awareness module constructs a state vector S based on the preprocessed multimodal features. The state vector S can be represented as: S=[g,m,p,e,l,b,c] Where S represents the state vector; g represents the user's gaze area or gaze point offset; m represents the amplitude of hand movements or terminal motion; p represents the pressure intensity or contact duration; e represents the intensity of voice emotion or voice strength level; l represents ambient brightness; b represents battery level; and c represents processor load or chip temperature. This state vector is used to characterize the user's current interaction intent, environmental conditions, and device operating status.

[0031] The adaptive decision-making module comprises a modality encoding subunit, a weight generation subunit, and a policy output subunit. Let the n input modal features be F_1, F_2, and F_n. The modality encoding subunit maps the i-th modal feature F_i to a uniform-dimensional encoding vector E_i(F_i). Here, n represents the number of modalities involved in the fusion; F_i represents the input feature of the i-th modality; and E_i(·) represents the encoding mapping function corresponding to the i-th modality.

[0032] The weight generation subunit generates the weight score g_i for each mode based on the state vector S, and obtains the weight α_i of the i-th mode through normalization: α_i=exp(g_i) / Σ_{j=1}^{n}exp(g_j) Where α_i represents the weight of the i-th mode; g_i represents the weight score of the i-th mode generated based on the state vector S; j represents the summation index; n represents the total number of modes; and exp(·) represents the exponential operation. The above normalization process makes the sum of the weights of all modes equal to 1.

[0033] The weighted fusion feature F can be expressed as: F=Σ_{i=1}^{n}α_iE_i(F_i) Where F represents the fused features; α_i represents the weight of the i-th modality; E_i(F_i) represents the encoding vector of the i-th modality; Σ represents the summation over all participating modalities. The policy output subunit outputs the interaction policy A based on the fused features F, the state vector S, and the historical feedback information H: A=π(F,S,H) Where A represents the interaction strategy; π(·) represents the strategy output function; and H represents historical feedback information. The interaction strategy A includes at least two types of parameters from the following categories: display content parameters, haptic feedback parameters, and voice output parameters.

[0034] Example 4: Cross-media interactive execution refer to Figure 4 The cross-media interaction execution module includes a display output unit, a haptic feedback unit, and a voice output unit. The display output unit displays text, icons, prompts, navigation information, or status images to the user. The haptic feedback unit can employ piezoelectric haptic feedback devices, linear motors, or eccentric rotor motors to output haptic signals of different frequencies, amplitudes, and durations. The voice output unit plays prompts, confirmations, alarms, or navigation voices.

[0035] In one embodiment, interaction strategy A includes display parameters A_v, haptic parameters A_t, and voice parameters A_a. Here, A_v represents the display output parameter, A_t represents the haptic feedback parameter, and A_a represents the voice output parameter. The system can output visual cues at time T0, provide haptic feedback within the range of T0 to T0+20ms, and provide voice cues within the range of T0 to T0+100ms to achieve consistency in multi-channel perception. T0 represents the reference time at which the current interaction strategy begins execution.

[0036] Example 5: Feedback Updates and Cloud Aggregation refer to Figure 5 The feedback update module collects at least one type of feedback information from users, including confirmations, cancellations, number of repeated operations, response wait times, satisfaction ratings, and changes in device power consumption. The system converts this feedback into a feedback evaluation value R. R=λ_1u-λ_2d-λ_3r Where R represents the feedback evaluation value; u represents the interaction success rate; d represents the normalized value of the interaction delay; r represents the normalized value of the number of repeated operations; λ_1, λ_2 and λ_3 represent the weight coefficients corresponding to the interaction success rate, interaction delay and number of repeated operations, respectively.

[0037] In optional cloud-based collaborative scenarios, the terminal uploads the local model increment Δθ to the cloud aggregation unit when the upload conditions are met. The model increment Δθ can be expressed as: Δθ = θ_local - θ_global Where Δθ represents the parameter increment of the local model relative to the global model; θ_local represents the updated model parameters locally; and θ_global represents the current global model parameters. Before uploading, parameter pruning and privacy perturbation are performed on Δθ to avoid uploading raw user data.

[0038] Example 6: Fault Self-Diagnosis and Degraded Operation refer to Figure 6 The fault self-diagnosis module collects at least two of the following signals: temperature, current, vibration, and processor load, and calculates an anomaly score Q: Q = μ₁T + μ₂I + μ₃V + μ₄L Where Q represents the anomaly score; T represents the degree of temperature anomaly; I represents the degree of current anomaly; V represents the degree of vibration anomaly; L represents the degree of processor load anomaly; and μ_1, μ_2, μ_3 and μ_4 represent the weight coefficients of the corresponding anomaly items.

[0039] When Q exceeds the anomaly scoring threshold Q_th, the fault self-diagnosis module triggers degraded operation. Here, Q_th represents the preset anomaly scoring threshold. Degraded operation may include at least one of the following: disabling high-power display effects, reducing haptic feedback intensity, reducing model inference frequency, retaining basic voice or communication capabilities, recording anomaly logs, and uploading fault information. Through these methods, the system can maintain basic interactive capabilities and prevent the fault from escalating under abnormal terminal operating conditions.

[0040] Example 7: Implementation of the Method and Electronic Device This invention also provides an AI-based adaptive perception-based intelligent terminal interaction method. The method includes: collecting at least two types of multimodal data; performing time alignment, denoising, normalization, and feature encoding on the collected data; and generating a state vector based on the multimodal features. Modal weights are generated from the state vector and weighted and fused; an interaction strategy is output based on the fused features; at least two of the display, haptic, and voice modes are controlled to perform collaboratively; user feedback and device feedback are collected; and model parameters are updated or degraded operation is triggered based on the feedback.

[0041] The present invention also provides an electronic device, comprising a processor, a memory, and a program stored in the memory and executable by the processor. When the program is executed by the processor, it implements the aforementioned smart terminal interaction method. The processor may be a general-purpose processor, an embedded processor, a neural network accelerator, a digital signal processor, or a combination thereof; the memory may be volatile memory, non-volatile memory, or a combination thereof.

[0042] The above embodiments are merely illustrative of preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. All equivalent substitutions, modifications, or combinations made within the scope of the inventive concept and claims should be included within the scope of protection of the present invention.

Claims

1. An intelligent terminal interaction system based on AI adaptive perception, characterized in that, include: The module includes a multimodal information acquisition module, a preprocessing module, a user status perception module, an adaptive decision-making module, a cross-media interactive execution module, a feedback update module, and a fault self-diagnosis module. The multimodal information acquisition module is used to acquire at least two types of data from visual data, thermal imaging data, tactile data, voice data, posture data, and device operating status data. The preprocessing module is used to perform timestamp alignment, noise reduction, normalization, and feature encoding on the collected data to obtain multimodal feature data; The user state perception module is used to extract user state features and terminal state features based on the multimodal feature data, and generate a state vector; The adaptive decision module is used to generate modal weights based on the state vector, dynamically weight and fuse different modal features, and output an interaction strategy. The cross-media interaction execution module is used to control at least two of the following methods—display, touch, and voice—to perform collaboratively according to the interaction strategy. The feedback update module is used to adjust the parameters of the adaptive decision module based on user feedback and device feedback; The fault self-diagnosis module is used to generate an anomaly score based on the equipment operating status data, and trigger degraded operation or safety protection control when the anomaly score exceeds a preset threshold.

2. The intelligent terminal interaction system based on AI adaptive perception according to claim 1, characterized in that, The multimodal information acquisition module includes at least three of the following: a visible light image sensor, an infrared thermal imaging sensor, a flexible tactile sensor, a microphone, an inertial measurement unit, a temperature sensor, a current sensor, and a vibration sensor; wherein, the flexible tactile sensor is disposed adjacent to the terminal housing, the touch panel, or the display module, and is used to acquire the pressing position, the contact area, the pressing force, and the contact duration.

3. The intelligent terminal interaction system based on AI adaptive perception according to claim 1, characterized in that, The preprocessing module includes a synchronization alignment unit, a noise suppression unit, and a feature encoding unit; the synchronization alignment unit is used to resample or interpolate other modal data using the sampling time of the selected modality as a reference time; the noise suppression unit is used to filter and denoise at least one of visual data, tactile data, speech data, and thermal imaging data. The feature encoding unit is used to generate at least two of the following: visual features, tactile features, thermal imaging features, voice features, and device status features.

4. The intelligent terminal interaction system based on AI adaptive perception according to claim 1, characterized in that, The user state perception module is used to extract at least three of the following: user gaze area, hand movement amplitude, touch pressure intensity, voice intensity, voice emotion intensity, ambient brightness, remaining battery level, processor load, and terminal temperature, and concatenate the extracted features into a state vector.

5. The intelligent terminal interaction system based on AI adaptive perception according to claim 1, characterized in that, The adaptive decision-making module includes a modality encoding subunit, a weight generation subunit, and a policy output subunit; the modality encoding subunit is used to map each modality feature into an encoding vector of a uniform dimension; the weight generation subunit is used to generate weights corresponding to each modality based on the state vector; The strategy output subunit is used to output at least two types of parameters from display content parameters, haptic feedback parameters, and voice output parameters based on the weighted fusion result.

6. The intelligent terminal interaction system based on AI adaptive perception according to claim 5, characterized in that, The weight generation subunit adopts a normalized weight generation method so that the sum of the weights of each modality is 1; the weighted fusion result is used to adjust the priority of display, haptic and voice output in low light, noisy environment, motion state or processor high load state.

7. The intelligent terminal interaction system based on AI adaptive perception according to claim 1, characterized in that, The feedback update module is used to collect at least one type of feedback information from user confirmation, cancellation, number of repeated operations, response time, satisfaction score, and device power consumption changes, and adjust the weight generation parameters of the adaptive decision module according to the feedback information; when the upload conditions are met, the feedback update module performs parameter pruning and privacy perturbation on the local model incrementally before uploading it to the cloud aggregation unit.

8. The intelligent terminal interaction system based on AI adaptive perception according to claim 1, characterized in that, The fault self-diagnosis module is used to perform joint analysis on at least two of the temperature signal, current signal, vibration signal and processor load signal; when the abnormal score exceeds the preset threshold, the fault self-diagnosis module triggers at least one of the following: turning off unnecessary display output, reducing tactile feedback intensity, retaining basic voice or communication capabilities, recording abnormal logs and uploading fault information.

9. A smart terminal interaction method based on AI adaptive perception, characterized in that, Includes the following steps: S1, collect at least two types of data from visual data, thermal imaging data, tactile data, voice data, posture data, and device operating status data; S2 performs time alignment, denoising, normalization, and feature encoding on the collected data to obtain multimodal feature data; S3, extract user state features and terminal state features based on multimodal feature data, and generate state vectors; S4, generate modal weights based on the state vector, and perform weighted fusion of features from different modalities; S5 outputs an interaction strategy based on the weighted fusion result and controls at least two of the display, haptic, and voice methods to be executed collaboratively. S6 collects user and device feedback and updates model parameters or triggers degraded operation based on the feedback.

10. An electronic device, characterized in that, It includes a processor, a memory, and a program stored in the memory and executable by the processor, wherein the program, when executed, implements the AI-adaptive perception-based intelligent terminal interaction method as described in claim 9.