Non-contact robot man-machine interaction system based on multi-sensor fusion

By using multi-sensor fusion and deep learning algorithms, the problem of insufficient data acquisition and response in non-contact robot interaction systems has been solved, achieving high-precision understanding of user intent and personalized response, improving system stability and interaction fluency, and adapting to the needs of complex scenarios.

CN122008269APending Publication Date: 2026-05-12BEIJING HAIBAICHUAN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING HAIBAICHUAN TECH CO LTD
Filing Date
2026-01-28
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing non-contact robot interaction systems suffer from limitations such as single data acquisition dimensions, insufficient preprocessing accuracy, and poor multimodal data synchronization. This results in limited context awareness, difficulty in accurately understanding complex interaction intentions, a lack of flexibility and personalization in response methods, insufficient system stability and reliability, high interaction latency, and difficulty in adapting to dynamically changing scenario requirements.

Method used

A non-contact robot human-computer interaction system employing multi-sensor fusion includes a multimodal sensor module, a data preprocessing module, a data synchronization module, a feature extraction module, a multivariate comprehensive evaluation module, a context perception and creative response generation module, and a response execution and feedback module. Through the collaborative effect of multimodal sensors, combined with data preprocessing and synchronization technologies, high-quality data input is achieved, providing a foundation for context perception and creative response generation. Heterogeneous graph neural networks and deep learning algorithms are used for user intent prediction and personalized response generation.

Benefits of technology

It improves the perception accuracy and adaptability of contactless interaction, reduces interaction latency, enhances the smoothness of human-computer collaboration and the personalization of user experience, improves the stability and reliability of the system, and ensures timely and flexible response.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122008269A_ABST
    Figure CN122008269A_ABST
Patent Text Reader

Abstract

The invention discloses a non-contact robot man-machine interaction system based on multi-sensor fusion. Through the synergistic effect of the multi-modal sensor module and the data synchronization module, multi-dimensional data such as physiological signals, environmental parameters and interaction states are comprehensively captured, and a high-quality input basis is provided for the context awareness and creative response generation module in combination with precise cleaning and standardized processing of the data preprocessing module. The situation perception module can deeply understand user intentions and environment changes by means of multi-dimensional information after feature extraction, the perception precision of non-contact interaction is effectively improved, and the adaptability of the system to complex scenes is enhanced. By means of the design, the robot can still accurately capture user requirements under the condition that physical contact is not needed, interaction delay is reduced, and the smoothness of man-machine cooperation is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of robot interaction technology, specifically a non-contact robot human-computer interaction system based on multi-sensor fusion. Background Technology

[0002] Non-contact robot interaction is a technological model that enables human-machine communication without physical contact. It primarily relies on voice recognition, gesture control, eye tracking, facial expression recognition, and remote control, allowing users to input commands and obtain information without directly touching the device. This interaction method is widely used in medical, service, public display, and smart home scenarios, and has significant advantages, especially in environments with high hygiene requirements or accessibility. Combining artificial intelligence and multimodal perception technology, non-contact robot interaction can analyze user intentions in real time, provide accurate feedback, effectively reduce the risk of cross-infection, and improve the convenience and safety of interaction. As the technology continues to mature, non-contact interaction is gradually becoming an important direction for the intelligent development of robots.

[0003] However, existing technologies suffer from problems such as limited data acquisition dimensions, insufficient preprocessing accuracy, and poor synchronization of multimodal data. This results in limited context awareness, difficulty in accurately understanding complex interaction intentions, lack of flexibility and personalization in response methods, insufficient system stability and reliability, high interaction latency, and difficulty in adapting to dynamically changing scenario requirements. Summary of the Invention

[0004] The purpose of this invention is to provide a non-contact robot human-computer interaction system based on multi-sensor fusion in order to solve the problems mentioned above.

[0005] The technical solution adopted in this invention is as follows: a non-contact robot human-computer interaction system based on multi-sensor fusion, the system comprising: a multimodal sensor module, a data preprocessing module, a data synchronization module, a feature extraction module, a multivariate comprehensive evaluation module, a context perception and creative response generation module, and a response execution and feedback module; The context awareness and creative response generation module includes: a context modeling submodule, an intent prediction submodule, and a creative response generation submodule. The output of the multimodal sensor module is connected to the input of the data synchronization module. The output of the data synchronization module is connected to the input of the data preprocessing module. The output of the data preprocessing module is connected to the input of the feature extraction module. The output of the feature extraction module is connected to the input of the context awareness and creative response generation module and the input of the multivariate comprehensive evaluation module, respectively. The output of the context awareness and creative response generation module is connected to the input of the response execution and feedback module. The output of the response execution and feedback module is connected to the input of the multivariate comprehensive evaluation module.

[0006] In a preferred embodiment, the multimodal sensor module internally includes: a physiological signal acquisition unit, an environmental perception unit, an interactive state sensor, and a sensor control center; the physiological signal acquisition unit integrates a photoelectric heart rate sensor, an eye-tracking camera, and a facial electromyography sensor; the environmental perception unit includes a lidar RGB-D camera and a temperature and humidity sensor; the interactive state sensor is equipped with a touch pressure sensor and a voice microphone array; and the sensor control center enables multi-device collaboration via an SPI bus, with a built-in FPGA chip for raw data packaging and dynamic power consumption management.

[0007] In a preferred embodiment, the data preprocessing module internally comprises: a raw data cleaning sublayer, a spatiotemporal standardization sublayer, a feature dimensionality reduction sublayer, and a temporal segmentation sublayer; the raw data cleaning sublayer uses a multi-level filtering mechanism to process physiological signal environmental point clouds and speech data; the spatiotemporal standardization sublayer performs data scale unification, including physiological feature standardization, image interpolation scaling, and coordinate transformation; the feature dimensionality reduction sublayer processes high-dimensional data through principal component analysis and the t-SNE algorithm; and the temporal segmentation sublayer uses a sliding window technique to divide data segments and attach timestamps and data quality labels.

[0008] In a preferred embodiment, the data synchronization module internally comprises: a hardware-triggered synchronization subsystem, a timestamp calibration algorithm layer, a cross-modal delay compensation layer, and a synchronization quality monitoring layer; the hardware-triggered synchronization subsystem achieves multi-sensor clock alignment through GPIO trigger signals; the timestamp calibration algorithm layer runs a distributed time synchronization protocol using GPS timing and NTP calibration; the cross-modal delay compensation layer establishes a device delay model and aligns data through a pre-compensation algorithm; and the synchronization quality monitoring layer calculates timestamp consistency, data integrity, and event synchronization rate indicators, triggering a reset mechanism when these indicators are abnormal.

[0009] In a preferred embodiment, the feature extraction module internally comprises: a time-domain feature calculation unit, a frequency-domain feature conversion unit, a spatial-domain feature encoding unit, and a semantic feature parsing unit; the time-domain feature calculation unit extracts statistical features for physiological and interactive signals, the frequency-domain feature conversion unit processes periodic signals through Fourier transform and wavelet transform, the spatial-domain feature encoding unit processes visual and spatial data, and the semantic feature parsing unit performs deep encoding on text and speech content.

[0010] In a preferred embodiment, the multivariate comprehensive evaluation module is internally configured with: an evaluation index system construction unit, a dynamic weight allocation unit, a real-time evaluation algorithm unit, and a historical data comparison unit; the evaluation index system construction unit defines four types of core indicators, the dynamic weight allocation unit calculates the indicator weights using an improved entropy weight method, the real-time evaluation algorithm unit runs a hierarchical evaluation model, and the historical data comparison unit maintains a time-series database of evaluation results.

[0011] In a preferred embodiment, the context modeling submodule, through real-time integration and dynamic representation of multi-source heterogeneous data, comprises: Multi-source data input layer: Simultaneously receives three types of basic data—user status data (such as heart rate variability sequence HRV, eye movement trajectory coordinates (x,y,t), facial expression feature vector F), and environmental perception data (scene semantic labels S). env Illumination intensity L, obstacle coordinate set O), historical interaction data (intention vector sequence I1...I from the past 30 rounds of dialogue). n The response satisfaction score (R) is preprocessed by timestamp alignment (error ≤ 5ms) and outlier filtering (based on the 3σ criterion).

[0012] Dynamic Feature Fusion Layer: Employs an adaptive weighted algorithm to fuse input features, where the user state weight w u Dynamically adjusted by the physiological signal-to-noise ratio (e.g., automatically reducing w when HRV signal-to-noise ratio < 0.6). u Environmental feature weights w e Introduce a scene complexity factor (e.g., in a kitchen scene, the complexity factor is increased due to the high density of targets). e Up to 0.4), historical data weight w h Decayed via exponential moving average (EMA) (half-life set to 5 minutes).

[0013] Contextual State Representation Layer: A dynamic contextual graph is constructed based on Heterogeneous Graph Neural Network (HGNN). Nodes include users (attributes: emotional polarity, attention entropy), environmental entities (attributes: spatial location, interaction priority), and interaction events (attributes: timestamp, intent completion). Edge weights are dynamically updated through an attention mechanism (e.g., the "user-object" edge weight is enhanced when the user looks at an object).

[0014] Real-time update engine: An improved LSTM network is used to realize the temporal evolution of the context state, outputting the current context vector S every 100ms. t Meanwhile, it uses Kalman filtering to predict the situational trends (such as the user's attention shift trajectory) for the next 2 seconds, providing forward input for the intent prediction submodule.

[0015] The formula for dynamically updating the situation state is: ; In the formula: S t The context state vector at time t contains a fusion representation of the user, environment, and interaction. w u (t) represents the user state weight at time t, which is calculated from the signal-to-noise ratio SNR_u(t) of the HRV signal: wu(t) = 0.5·SNRu(t) + 0.2; U t This represents the user's state vector at time t (dimension 256×1), which is composed of emotional features (such as valence-arousal 2D values) and attention features (such as the percentage of dwell time in the region of interest ROI). w e (t) represents the environmental feature weight at time t (range [0.1, 0.5]), which is positively correlated with the scene complexity factor C(t): we(t) = 0.1 + 0.4·C(t) / Cmax, where Cmax is the preset maximum complexity). E t This represents the environmental feature vector at time t (dimension 128×1), which includes scene semantic encoding (e.g., the unique heat vector corresponding to "kitchen") and dynamic obstacle density; w h (t) represents the historical state weight at time t (range [0.1, 0.3]), satisfying w u (t)+w e (t)+w h (t)=1; S t−1 This represents the situational state vector at time t-1; η(t) represents the situational entropy gradient coefficient (range [0.01, 0.05]), which characterizes the impact of situational uncertainty on state updates; ∇H(S t−1 H(S) represents the situational entropy at time t-1. t−1 The gradient vector of ) has a higher entropy value (the more ambiguous the context), the greater the correction of the gradient term to St.

[0016] In a preferred embodiment, the intent prediction submodule achieves accurate prediction of user explicit commands and implicit needs by fusing temporal interaction data and knowledge reasoning. Its core components include: the input fusion layer receiving the real-time context vector S output by the context modeling submodule. t Historical interaction intent sequence I1... t−1The knowledge graph is embedded and weighted by a multi-head attention mechanism; the temporal intent modeling layer uses a bidirectional LSTM network, with a forward LSTM capturing recent intent trends and a backward LSTM mining long-term habitual patterns, outputting a temporal intent feature vector Ht; the knowledge-enhanced reasoning layer processes the knowledge graph based on a graph attention network, dynamically calculates entity association weights, and generates a knowledge feature vector Kt; the intent decoding layer fuses Ht through a gating mechanism. t With K t The softmax function outputs an explicit intent probability distribution, while the variational autoencoder generates an implicit demand vector, ultimately outputting a structured intent prediction result (including intent category, confidence level, and execution priority). The formula for dynamic fusion of intent probabilities is: ; In the formula: P(I t ): t time intention I t The overall probability distribution (dimension equal to the number of intent categories, such as 10 types of interactive intents). λ(t): Time-series feature weights (range [0.3, 0.8]), dynamically adjusted by the historical intent prediction accuracy A(t). λ(t) = 0.3 + 0.5·A(t) / Amax, where Amax is the preset maximum accuracy; Softmax(H t H represents the temporal intent probability distribution of the LSTM output. t The temporal feature vector at time t (dimension 256×1); GAT(K t ,S t Knowledge graph reasoning probability: It integrates knowledge features Kt and context vector St through the GAT network to output the probability of knowledge-guided intent. ϵ(t): Intention evolution coefficient (value range [0.05, 0.2]), which is positively correlated with the temporal continuity of the intention sequence (e.g., ϵ(t) increases when the same intention is repeated). ΔP(I t−1 ): The change in the probability of intent at time t-1, reflecting the trend of intent (e.g., if the probability of "query" intent increases continuously, then ΔP is positive).

[0017] In a preferred embodiment, the creative response generation submodule takes the intent prediction result and context vector as input, and outputs a personalized interactive response through a multimodal generation and dynamic optimization mechanism. Its core components include a response task decomposition layer, a multimodal generation engine, and a real-time optimization unit.

[0018] The response task decomposition layer first parses the explicit instructions and implicit requirements in the intent prediction results, breaking down complex tasks into a sequence of executable sub-tasks (e.g., "recommended recipes" is broken down into "ingredient recognition - recipe matching - step generation - visual guidance"), and assigns sub-task priorities based on the user status in the context vector (e.g., child / elderly / professional user). The multimodal generation engine contains three parallel generators: the visual generator uses dynamic image generation technology to generate interactive interface elements (e.g., guidance animations, operation buttons) based on user preferences (e.g., cartoon / realistic style) and scene requirements (e.g., AR annotation / static prompts); the language generator generates natural language feedback based on an intelligent language model, adjusting the tone in conjunction with emotional states (e.g., using soothing short sentences when anxious, simplifying instructions when focused); and the physical action generator generates trajectory sequences for robotic arms / chassis through motion planning algorithms (e.g., gentle grasping actions adapted to fragile items), and simultaneously outputs multimodal collaborative signals (e.g., aligning voice and action timestamps).

[0019] The real-time optimization unit dynamically adjusts its response content through a closed-loop feedback mechanism: on the one hand, it collects micro-feedback during user interaction (such as eye-tracking duration and voice interruption frequency) and updates the generation strategy through reinforcement learning algorithms (e.g., automatically increasing the proportion of visual animation when users frequently ignore text prompts); on the other hand, it combines dynamic environmental changes (e.g., switching to high-contrast visual output when the light dims, and increasing voice volume when noise increases) to ensure the robustness and adaptability of the response. Furthermore, this unit maintains a user creative preference library, recording highly satisfactory response patterns from historical interactions (e.g., visualizing specific user preferences in charts rather than text lists), and quickly adapts to new scenarios through transfer learning to achieve subsequent accurately matched personalized experiences.

[0020] In a preferred embodiment, the response execution and feedback module is internally configured with: an actuator control layer, a multimodal output scheduling layer, a feedback signal acquisition layer, and an execution status monitoring layer; the actuator control layer includes a three-level control architecture, the multimodal output scheduling layer manages cross-modal collaborative output, the feedback signal acquisition layer deploys closed-loop feedback sensors, and the execution status monitoring layer diagnoses three types of states in real time.

[0021] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are: 1. In this invention, the multimodal sensor module and data synchronization module work together to comprehensively capture multidimensional data such as physiological signals, environmental parameters, and interaction states. Combined with the precise cleaning and standardization of the data preprocessing module, a high-quality input foundation is provided for the context perception and creative response generation modules. The context perception module, leveraging the multidimensional information extracted from features, can more deeply understand user intentions and environmental changes, effectively improving the perception accuracy of non-contact interaction and enhancing the system's adaptability to complex scenarios. This design allows the robot to accurately capture user needs without physical contact, reducing interaction latency and enhancing the smoothness of human-machine collaboration.

[0022] 2. In this invention, the response execution and feedback module and the multi-dimensional comprehensive evaluation module form a closed-loop mechanism, continuously optimizing the interaction process by combining the dynamic strategies output by the context awareness and creative response generation modules. The system can not only adjust its response based on real-time feedback but also provide diverse interaction solutions through the creative response generation module, enhancing the personalization and flexibility of the user experience. Simultaneously, the layered design of each module ensures efficient data processing and timely responses, further strengthening the stability and reliability of the system's operation, making contactless human-computer interaction more practical and scalable in real-world applications. Attached Figure Description

[0023] Figure 1 This is an overall system block diagram of the present invention; Figure 2 This is a system block diagram of the context awareness and creative response generation module in this invention. Detailed Implementation

[0024] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0025] Example: Reference Figure 1-2 A non-contact robot human-computer interaction system based on multi-sensor fusion, the system comprising: a multimodal sensor module, a data preprocessing module, a data synchronization module, a feature extraction module, a multivariate comprehensive evaluation module, a context perception and creative response generation module, and a response execution and feedback module; The context awareness and creative response generation module includes: a context modeling submodule, an intent prediction submodule, and a creative response generation submodule. The output of the multimodal sensor module is connected to the input of the data synchronization module. The output of the data synchronization module is connected to the input of the data preprocessing module. The output of the data preprocessing module is connected to the input of the feature extraction module. The output of the feature extraction module is connected to the input of the context awareness and creative response generation module and the input of the multivariate comprehensive evaluation module, respectively. The output of the context awareness and creative response generation module is connected to the input of the response execution and feedback module. The output of the response execution and feedback module is connected to the input of the multivariate comprehensive evaluation module.

[0026] The multimodal sensor module internally includes: a physiological signal acquisition unit, an environmental perception unit, an interactive status sensor, and a sensor control center. The physiological signal acquisition unit integrates a photoelectric heart rate sensor, an eye-tracking camera, and a facial electromyography sensor. The environmental perception unit includes a lidar RGB-D camera and a temperature and humidity sensor. The interactive status sensor is equipped with a touch pressure sensor and a voice microphone array. The sensor control center enables multi-device collaboration via an SPI bus and uses a built-in FPGA chip to package raw data and support dynamic power consumption management.

[0027] The data preprocessing module is internally configured with: a raw data cleaning sublayer, a spatiotemporal standardization sublayer, a feature dimensionality reduction sublayer, and a temporal segmentation sublayer. The raw data cleaning sublayer uses a multi-level filtering mechanism to process physiological signal environmental point clouds and speech data. The spatiotemporal standardization sublayer performs data scaling, including physiological feature standardization, image interpolation and scaling, and coordinate transformation. The feature dimensionality reduction sublayer processes high-dimensional data through principal component analysis and the t-SNE algorithm. The temporal segmentation sublayer uses a sliding window technique to divide data segments and adds timestamps and data quality labels.

[0028] The data synchronization module internally comprises: a hardware-triggered synchronization subsystem, a timestamp calibration algorithm layer, a cross-modal delay compensation layer, and a synchronization quality monitoring layer. The hardware-triggered synchronization subsystem achieves multi-sensor clock alignment through GPIO trigger signals. The timestamp calibration algorithm layer runs a distributed time synchronization protocol using GPS timing and NTP calibration. The cross-modal delay compensation layer establishes a device delay model and aligns data through a pre-compensation algorithm. The synchronization quality monitoring layer calculates timestamp consistency, data integrity, and event synchronization rate indicators and triggers a reset mechanism when these indicators are abnormal.

[0029] The feature extraction module is internally configured with: a time-domain feature calculation unit, a frequency-domain feature conversion unit, a spatial-domain feature encoding unit, and a semantic feature parsing unit; the time-domain feature calculation unit extracts statistical features for physiological and interactive signals, the frequency-domain feature conversion unit processes periodic signals through Fourier transform and wavelet transform, the spatial-domain feature encoding unit processes visual and spatial data, and the semantic feature parsing unit performs deep encoding on text and speech content.

[0030] The multi-dimensional comprehensive evaluation module is internally configured with: an evaluation index system construction unit, a dynamic weight allocation unit, a real-time evaluation algorithm unit, and a historical data comparison unit. The evaluation index system construction unit defines four types of core indicators, the dynamic weight allocation unit uses an improved entropy weight method to calculate the indicator weights, the real-time evaluation algorithm unit runs a hierarchical evaluation model, and the historical data comparison unit maintains a time-series database of evaluation results.

[0031] The scenario modeling submodule, through real-time integration and dynamic representation of multi-source heterogeneous data, comprises: Multi-source data input layer: Simultaneously receives three types of basic data—user status data (such as heart rate variability sequence HRV, eye movement trajectory coordinates (x,y,t), facial expression feature vector F), and environmental perception data (scene semantic labels S). env Illumination intensity L, obstacle coordinate set O), historical interaction data (intention vector sequence I1...I from the past 30 rounds of dialogue). n The response satisfaction score (R) is preprocessed by timestamp alignment (error ≤ 5ms) and outlier filtering (based on the 3σ criterion).

[0032] Dynamic Feature Fusion Layer: Employs an adaptive weighted algorithm to fuse input features, where the user state weight w u Dynamically adjusted by the physiological signal-to-noise ratio (e.g., automatically reducing w when HRV signal-to-noise ratio < 0.6). u Environmental feature weights w e Introduce a scene complexity factor (e.g., in a kitchen scene, the complexity factor is increased due to the high density of targets). e Up to 0.4), historical data weight w h Decayed via exponential moving average (EMA) (half-life set to 5 minutes).

[0033] Contextual State Representation Layer: A dynamic contextual graph is constructed based on Heterogeneous Graph Neural Network (HGNN). Nodes include users (attributes: emotional polarity, attention entropy), environmental entities (attributes: spatial location, interaction priority), and interaction events (attributes: timestamp, intent completion). Edge weights are dynamically updated through an attention mechanism (e.g., the "user-object" edge weight is enhanced when the user looks at an object).

[0034] Real-time update engine: An improved LSTM network is used to realize the temporal evolution of the context state, outputting the current context vector S every 100ms. t Meanwhile, it uses Kalman filtering to predict the situational trends (such as the user's attention shift trajectory) for the next 2 seconds, providing forward input for the intent prediction submodule.

[0035] The formula for dynamically updating the situation state is: ; In the formula: S t The context state vector at time t contains a fusion representation of the user, environment, and interaction. w u (t) represents the user state weight at time t, which is calculated from the signal-to-noise ratio SNR_u(t) of the HRV signal: wu(t) = 0.5·SNRu(t) + 0.2; U t This represents the user's state vector at time t (dimension 256×1), which is composed of emotional features (such as valence-arousal 2D values) and attention features (such as the percentage of dwell time in the region of interest ROI). w e (t) represents the environmental feature weight at time t (range [0.1, 0.5]), which is positively correlated with the scene complexity factor C(t): we(t) = 0.1 + 0.4·C(t) / Cmax, where Cmax is the preset maximum complexity). E t This represents the environmental feature vector at time t (dimension 128×1), which includes scene semantic encoding (e.g., the unique heat vector corresponding to "kitchen") and dynamic obstacle density; w h (t) represents the historical state weight at time t (range [0.1, 0.3]), satisfying w u (t)+w e (t)+w h (t)=1; S t−1 This represents the situational state vector at time t-1; η(t) represents the situational entropy gradient coefficient (range [0.01, 0.05]), which characterizes the impact of situational uncertainty on state updates; ∇H(S t−1 H(S) represents the situational entropy at time t-1. t−1 The gradient vector of ) has a higher entropy value (the more ambiguous the context), the greater the correction of the gradient term to St.

[0036] The intent prediction submodule achieves accurate prediction of user explicit commands and implicit needs by fusing temporal interaction data and knowledge reasoning. Its core components include: the input fusion layer receiving the real-time context vector S output by the context modeling submodule. t Historical interaction intent sequence I1... t−1The knowledge graph is embedded and weighted by a multi-head attention mechanism; the temporal intent modeling layer uses a bidirectional LSTM network, with a forward LSTM capturing recent intent trends and a backward LSTM mining long-term habitual patterns, outputting a temporal intent feature vector Ht; the knowledge-enhanced reasoning layer processes the knowledge graph based on a graph attention network, dynamically calculates entity association weights, and generates a knowledge feature vector Kt; the intent decoding layer fuses Ht through a gating mechanism. t With K t The softmax function outputs an explicit intent probability distribution, while the variational autoencoder generates an implicit demand vector, ultimately outputting a structured intent prediction result (including intent category, confidence level, and execution priority). The formula for dynamic fusion of intent probabilities is: ; In the formula: P(I t ): t time intention I t The overall probability distribution (dimension equal to the number of intent categories, such as 10 types of interactive intents). λ(t): Time-series feature weights (range [0.3, 0.8]), dynamically adjusted by the historical intent prediction accuracy A(t). λ(t) = 0.3 + 0.5·A(t) / Amax, where Amax is the preset maximum accuracy; Softmax(H t H represents the temporal intent probability distribution of the LSTM output. t The temporal feature vector at time t (dimension 256×1); GAT(K t ,S t Knowledge graph reasoning probability: It integrates knowledge features Kt and context vector St through the GAT network to output the probability of knowledge-guided intent. ϵ(t): Intention evolution coefficient (value range [0.05, 0.2]), which is positively correlated with the temporal continuity of the intention sequence (e.g., ϵ(t) increases when the same intention is repeated). ΔP(I t−1 ): The change in the probability of intent at time t-1, reflecting the trend of intent (e.g., if the probability of "query" intent increases continuously, then ΔP is positive).

[0037] The creative response generation submodule takes the intent prediction result and context vector as input, and outputs a personalized interactive response through a multimodal generation and dynamic optimization mechanism. Its core components include a response task decomposition layer, a multimodal generation engine, and a real-time optimization unit.

[0038] The response task decomposition layer first parses the explicit instructions and implicit requirements in the intent prediction results, breaking down complex tasks into a sequence of executable sub-tasks (e.g., "recommended recipes" is broken down into "ingredient recognition - recipe matching - step generation - visual guidance"), and assigns sub-task priorities based on the user status in the context vector (e.g., child / elderly / professional user). The multimodal generation engine contains three parallel generators: the visual generator uses dynamic image generation technology to generate interactive interface elements (e.g., guidance animations, operation buttons) based on user preferences (e.g., cartoon / realistic style) and scene requirements (e.g., AR annotation / static prompts); the language generator generates natural language feedback based on an intelligent language model, adjusting the tone in conjunction with emotional states (e.g., using soothing short sentences when anxious, simplifying instructions when focused); and the physical action generator generates trajectory sequences for robotic arms / chassis through motion planning algorithms (e.g., gentle grasping actions adapted to fragile items), and simultaneously outputs multimodal collaborative signals (e.g., aligning voice and action timestamps).

[0039] The real-time optimization unit dynamically adjusts its response content through a closed-loop feedback mechanism: on the one hand, it collects micro-feedback during user interaction (such as eye-tracking duration and voice interruption frequency) and updates the generation strategy through reinforcement learning algorithms (e.g., automatically increasing the proportion of visual animation when users frequently ignore text prompts); on the other hand, it combines dynamic environmental changes (e.g., switching to high-contrast visual output when the light dims, and increasing voice volume when noise increases) to ensure the robustness and adaptability of the response. Furthermore, this unit maintains a user creative preference library, recording highly satisfactory response patterns from historical interactions (e.g., visualizing specific user preferences in charts rather than text lists), and quickly adapts to new scenarios through transfer learning to achieve subsequent accurately matched personalized experiences.

[0040] The response execution and feedback module is internally configured with: an actuator control layer, a multimodal output scheduling layer, a feedback signal acquisition layer, and an execution status monitoring layer. The actuator control layer contains a three-level control architecture, the multimodal output scheduling layer manages cross-modal collaborative output, the feedback signal acquisition layer deploys closed-loop feedback sensors, and the execution status monitoring layer diagnoses three types of states in real time.

[0041] As described above, this invention, through the synergistic effect of the multimodal sensor module and the data synchronization module, comprehensively captures multidimensional data such as physiological signals, environmental parameters, and interaction states. Combined with the precise cleaning and standardization processing of the data preprocessing module, this provides a high-quality input foundation for the context perception and creative response generation modules. The context perception module, leveraging the multidimensional information extracted from features, can more deeply understand user intentions and environmental changes, effectively improving the perception accuracy of non-contact interaction and enhancing the system's adaptability to complex scenarios. This design allows the robot to accurately capture user needs without physical contact, reducing interaction latency and enhancing the smoothness of human-machine collaboration.

[0042] In this invention, the response execution and feedback module and the multi-dimensional comprehensive evaluation module form a closed-loop mechanism, continuously optimizing the interaction process by combining the dynamic strategies output by the context awareness and creative response generation modules. The system can not only adjust its response based on real-time feedback but also provide diverse interaction solutions through the creative response generation module, enhancing the personalization and flexibility of the user experience. Simultaneously, the layered design of each module ensures efficient data processing and timely responses, further strengthening the stability and reliability of the system's operation, making contactless human-computer interaction more practical and scalable in real-world applications.

[0043] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0044] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A non-contact robot human-computer interaction system based on multi-sensor fusion, characterized in that: The system includes: a multimodal sensor module, a data preprocessing module, a data synchronization module, a feature extraction module, a multivariate comprehensive evaluation module, a context perception and creative response generation module, and a response execution and feedback module; The context awareness and creative response generation module includes: a context modeling submodule, an intent prediction submodule, and a creative response generation submodule. The output of the multimodal sensor module is connected to the input of the data synchronization module. The output of the data synchronization module is connected to the input of the data preprocessing module. The output of the data preprocessing module is connected to the input of the feature extraction module. The output of the feature extraction module is connected to the input of the context awareness and creative response generation module and the input of the multivariate comprehensive evaluation module, respectively. The output of the context awareness and creative response generation module is connected to the input of the response execution and feedback module. The output of the response execution and feedback module is connected to the input of the multivariate comprehensive evaluation module.

2. The non-contact robot human-computer interaction system based on multi-sensor fusion as described in claim 1, characterized in that: The multimodal sensor module internally includes: a physiological signal acquisition unit, an environmental perception unit, an interactive status sensor, and a sensor control center. The physiological signal acquisition unit integrates a photoelectric heart rate sensor, an eye-tracking camera, and a facial electromyography sensor. The environmental perception unit includes a lidar RGB-D camera and a temperature and humidity sensor. The interactive status sensor is equipped with a touch pressure sensor and a voice microphone array. The sensor control center enables multi-device collaboration via an SPI bus and uses a built-in FPGA chip to package raw data and support dynamic power consumption management.

3. The non-contact robot human-computer interaction system based on multi-sensor fusion as described in claim 1, characterized in that: The data preprocessing module is internally configured with: a raw data cleaning sublayer, a spatiotemporal standardization sublayer, a feature dimensionality reduction sublayer, and a temporal segmentation sublayer. The raw data cleaning sublayer uses a multi-level filtering mechanism to process physiological signal environmental point clouds and speech data. The spatiotemporal standardization sublayer performs data scaling, including physiological feature standardization, image interpolation and scaling, and coordinate transformation. The feature dimensionality reduction sublayer processes high-dimensional data through principal component analysis and the t-SNE algorithm. The temporal segmentation sublayer uses a sliding window technique to divide data segments and adds timestamps and data quality labels.

4. The non-contact robot human-computer interaction system based on multi-sensor fusion as described in claim 1, characterized in that: The data synchronization module internally comprises: a hardware-triggered synchronization subsystem, a timestamp calibration algorithm layer, a cross-modal delay compensation layer, and a synchronization quality monitoring layer. The hardware-triggered synchronization subsystem achieves multi-sensor clock alignment through GPIO trigger signals. The timestamp calibration algorithm layer runs a distributed time synchronization protocol using GPS timing and NTP calibration. The cross-modal delay compensation layer establishes a device delay model and aligns data through a pre-compensation algorithm. The synchronization quality monitoring layer calculates timestamp consistency, data integrity, and event synchronization rate indicators and triggers a reset mechanism when these indicators are abnormal.

5. A non-contact robot human-computer interaction system based on multi-sensor fusion as described in claim 1, characterized in that: The feature extraction module is internally configured with: a time-domain feature calculation unit, a frequency-domain feature transformation unit, a spatial-domain feature encoding unit, and a semantic feature parsing unit; The temporal feature calculation unit extracts statistical features for physiological and interactive signals, the frequency domain feature transformation unit processes periodic signals through Fourier transform and wavelet transform, the spatial domain feature encoding unit processes visual and spatial data, and the semantic feature parsing unit performs deep encoding on text and speech content.

6. The non-contact robot human-computer interaction system based on multi-sensor fusion as described in claim 1, characterized in that: The multi-dimensional comprehensive evaluation module is internally configured with: an evaluation index system construction unit, a dynamic weight allocation unit, a real-time evaluation algorithm unit, and a historical data comparison unit. The evaluation index system construction unit defines four types of core indicators, the dynamic weight allocation unit uses an improved entropy weight method to calculate the indicator weights, the real-time evaluation algorithm unit runs a hierarchical evaluation model, and the historical data comparison unit maintains a time-series database of evaluation results.

7. A non-contact robot human-computer interaction system based on multi-sensor fusion as described in claim 1, characterized in that: The scenario modeling submodule, through real-time integration and dynamic representation of multi-source heterogeneous data, comprises: Multi-source data input layer: synchronously receives three types of basic data—user status data, environmental perception data, and historical interaction data, and preprocesses them through timestamp alignment and outlier filtering; Dynamic Feature Fusion Layer: Employs an adaptive weighted algorithm to fuse input features, where the user state weight w u The signal-to-noise ratio of physiological signals is dynamically adjusted, and the environmental feature weights w e Introducing a scenario complexity factor and historical data weight w h Decaying via exponential moving average; Contextual State Representation Layer: A dynamic contextual graph is constructed based on a heterogeneous graph neural network. Nodes include users, environmental entities, and interaction events, and edge weights are dynamically updated through an attention mechanism. Real-time update engine: An improved LSTM network is used to realize the temporal evolution of the context state, outputting the current context vector S every 100ms. t Meanwhile, the Kalman filter is used to predict the situational trend in the next 2 seconds, providing look-ahead input for the intent prediction submodule; The formula for dynamically updating the situation state is: ; In the formula: S t The context state vector at time t contains a fusion representation of the user, environment, and interaction. w u (t) represents the user state weight at time t, which is calculated from the signal-to-noise ratio SNR_u(t) of the HRV signal: wu(t) = 0.5·SNRu(t) + 0.2; U t This represents the user's state vector at time t, which is concatenated with emotion features and attention features. w e (t) represents the environmental feature weight at time t, which is positively correlated with the scene complexity factor C(t): we(t) = 0.1 + 0.4·C(t) / Cmax, where Cmax is the preset maximum complexity). E t This represents the environmental feature vector at time t, which includes scene semantic encoding and dynamic obstacle density; w h (t) represents the historical state weight at time t, satisfying w u (t)+w e (t)+w h (t)=1; S t−1 This represents the situational state vector at time t-1; η(t) represents the situational entropy gradient coefficient, which characterizes the impact of situational uncertainty on state updates; ∇H(S t−1 H(S) represents the situational entropy at time t-1. t−1 The gradient vector of ) has a higher entropy value, and the greater the correction of the gradient term to St.

8. A non-contact robot human-computer interaction system based on multi-sensor fusion as described in claim 1, characterized in that: The intent prediction submodule achieves accurate prediction of user explicit commands and implicit needs by fusing temporal interaction data and knowledge reasoning. Its core components include: the input fusion layer receiving the real-time context vector S output by the context modeling submodule. t Historical interaction intent sequence I1... t−1 The knowledge graph is embedded and weighted by a multi-head attention mechanism; the temporal intent modeling layer uses a bidirectional LSTM network, with a forward LSTM capturing recent intent trends and a backward LSTM mining long-term habitual patterns, outputting a temporal intent feature vector Ht; the knowledge-enhanced reasoning layer processes the knowledge graph based on a graph attention network, dynamically calculates entity association weights, and generates a knowledge feature vector Kt; the intent decoding layer fuses Ht through a gating mechanism. t With K t The softmax function outputs an explicit intent probability distribution, while the variational autoencoder generates an implicit demand vector, ultimately outputting a structured intent prediction result (including intent category, confidence level, and execution priority). The formula for dynamic fusion of intent probabilities is: ; In the formula: P(I t ): t time intention I t The comprehensive probability distribution; λ(t): Time series feature weights; GAT(K t ,S t Knowledge graph reasoning probability: It integrates knowledge features Kt and context vector St through the GAT network to output the probability of knowledge-guided intent. ϵ(t): Intention evolution coefficient, which is positively correlated with the temporal continuity of the intention sequence; ΔP(I t−1 ): The change in the probability of intent at time t-1, reflecting the trend of intent.

9. A non-contact robot human-computer interaction system based on multi-sensor fusion as described in claim 1, characterized in that: The creative response generation submodule consists of a response task decomposition layer, a multimodal generation engine, and a real-time optimization unit.

10. A non-contact robot human-computer interaction system based on multi-sensor fusion as described in claim 1, characterized in that: The response execution and feedback module is internally configured with: an actuator control layer, a multimodal output scheduling layer, a feedback signal acquisition layer, and an execution status monitoring layer; the actuator control layer contains a three-level control architecture.