Interactive method and interactive terminal fusing multi-modal perception and naked-eye 3D
By combining multimodal perception and naked-eye 3D interaction methods with active spatial detection, passive environmental perception and visual information acquisition, a three-level verification mechanism and hardware redundancy wake-up architecture are constructed to solve the problems of false triggering and inaccurate response of existing AI interaction devices, and achieve an efficient and stable user interaction experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Filing Date
- 2026-03-25
- Publication Date
- 2026-07-10
AI Technical Summary
Existing AI interactive device wake-up mechanisms are susceptible to environmental noise interference, have a high false trigger rate, lack multi-level verification mechanisms, have difficulty accurately distinguishing user intentions, struggle to balance response speed and accuracy, and lack proactive perception capabilities, thus affecting user experience and the depth of device application.
Employing multimodal perception and naked-eye 3D interaction methods, combining active spatial detection, passive environmental perception, and visual information acquisition, and through a three-level progressive anti-mistouch logic and hardware redundancy wake-up architecture, a wake-up mechanism with high accuracy and low false trigger rate is achieved, and multimodal interaction is performed through an AI core engine.
Significantly reduces system false trigger rate, improves wake-up accuracy and response speed, provides diversified interaction methods, enhances user experience and device adaptability, and ensures stability and covert deployment in complex environments.
Smart Images

Figure CN122363500A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of human-computer interaction technology, and more specifically, to an interaction method and terminal that integrates multimodal perception and naked-eye 3D. Background Technology
[0002] With the rapid development of artificial intelligence technology, human-computer interaction devices are evolving from single-function devices to intelligent and scenario-based devices. However, existing AI interaction devices still have significant technical deficiencies in their wake-up mechanisms.
[0003] On the one hand, existing devices' wake-up mechanisms mostly rely on a single sensor or simple preset rules for triggering. For example, some devices only use pyroelectric infrared (PIR) sensors for human detection. This single-modal sensing method is easily affected by environmental noise, such as pet movement, drastic changes in light, and heat source interference, resulting in a high false trigger rate and seriously affecting the user experience. At the same time, single-modal sensing lacks a multi-level verification mechanism, making it difficult to accurately distinguish between genuine user intent and environmental interference in complex dynamic environments, easily leading to false wake-ups or missed wake-ups.
[0004] On the other hand, traditional interactive devices typically only passively respond to explicit user commands, lacking the ability to proactively sense user approach or behavioral intentions. Existing solutions struggle to continuously perceive the user's presence and behavioral intentions, leading to frequent service interruptions and failing to deliver a human-like proactive service experience. In scenarios requiring highly predictive services, such as smart shopping guides and elderly care, this passive response mode severely limits the depth of device application.
[0005] Furthermore, existing wake-up mechanisms struggle to balance accuracy and response speed. To reduce false triggers, some solutions employ higher wake-up thresholds, but this leads to sluggish responses to the user's true intent; conversely, to improve response speed, some solutions lower the judgment criteria, resulting in frequent false triggers. There is a lack of a multi-level verification mechanism that can guarantee both high accuracy and rapid response.
[0006] Therefore, there is an urgent need to build a new human-computer interaction wake-up mechanism that combines high accuracy, low false trigger rate and active perception capability to overcome the limitations of existing technologies. Summary of the Invention
[0007] The purpose of this application is to provide an interactive method and terminal that integrates multimodal perception and naked-eye 3D to solve the above-mentioned technical problems.
[0008] In a first aspect, the present invention provides an interactive method integrating multimodal perception and naked-eye 3D. The method is applied to an interactive terminal, which includes at least three types of heterogeneous sensors, a naked-eye 3D display module, and an interactive system. The first type of sensor is an active spatial detection unit, used to continuously output signals reflecting the activity state of targets in space; the second type of sensor is a passive environmental change perception unit, used to output signals reflecting transient environmental changes; and the third type of sensor is a visual information acquisition and processing unit, used to output target recognition results based on visual features. The method includes: Determine whether the signal detected by the first type of sensor meets the main triggering condition, wherein the main triggering condition includes detecting a stable state signal that meets the time window requirement; When the main triggering condition is met, determine whether the transient change characteristics output by the second type of sensor meet the auxiliary verification condition; When the auxiliary verification conditions are met, it is determined whether the target recognition result based on visual features output by the third type of sensor meets the third verification condition. When the third verification condition is met, a wake-up command is generated, the interactive system is activated by the wake-up command, the virtual image is presented by the naked-eye 3D display module, and multimodal interaction is performed with the user.
[0009] In an optional implementation, the interactive terminal further includes an independent triggering link not controlled by the main operating system. This independent triggering link is used to perform logical operations on the output signals of the first type of sensor and the second type of sensor, and then directly connect them to the hardware wake-up pin of the main control processing unit. The method further includes: When an anomaly is detected in the main control software link or when it is in emergency mode, the independent trigger link not controlled by the main operating system is connected by a switch to wake it up. In normal mode, the main control processing unit executes multi-level progressive anti-accidental touch verification logic, and automatically reverts to normal mode after the fault is recovered.
[0010] In an optional implementation, the first type of sensor is a microwave radar sensor, which is installed behind a non-metallic shield. The antenna structure is optimized by dielectric loading to match the dielectric constant of the non-metallic shield, or an electromagnetic wave focusing structure is provided inside the non-metallic shield; the second type of sensor is a passive infrared sensor. The method further includes: A unified synchronous trigger signal is sent to each sensor via a hard-wired trigger bus to control each sensor to start sampling within the same clock cycle and perform microsecond-level time alignment. The microwave radar sensor operates in the ISM band, the passive infrared sensor is made of pyroelectric crystal material, and the independent triggering link includes a monostable multivibrator and logic gate circuits.
[0011] In an optional implementation, before presenting the virtual image via the glasses-free 3D display module, the following steps are also included: AI companion avatars can be created using at least one of the following methods: face reconstruction based on single or multiple frames of images; pose estimation and 3D model reconstruction based on video sequences; direct import of 3D model files; semantic parsing and generative modeling based on natural language text or speech descriptions. Specifically, for voice description methods, acoustic emotional features in the voice are extracted simultaneously, and the initial facial expression style or body posture of the generated image is automatically adjusted based on these emotional features. The video sequence-based pose estimation and 3D model reconstruction includes: using a pose estimation network to extract the coordinates of two-dimensional joints in the video frame, and reconstructing them into a parametric 3D human body model through a 3D pose enhancement network; The generative modeling based on natural language text description or speech description includes: parsing semantic tags using a large language model and calling a text-based 3D model algorithm to generate a 3D avatar with a neural radiation field representation or a Gaussian splash representation.
[0012] In an optional implementation, the multimodal interaction with the user includes: By integrating and analyzing micro-expressions captured by the camera, emotional features of speech extracted by the microphone, and physiological rhythm signals detected by radar, a label for the user's current emotional state is generated. Based on the combination of the emotional state label and semantic intent, the optimal action sequence and voice style are matched from the preset behavior rule base to drive the virtual character to make an empathetic response; The physiological rhythm signals include heart rate, respiratory rhythm, or heart rate variability indicators obtained through microwave radar phase change analysis. The method further includes: dividing user regions using spatial location-aware data and maintaining multi-path concurrent session state management to enable continuous tracking and switching of independent user contexts.
[0013] In an optional implementation, the interactive system includes an AI core engine that supports a cloud-edge collaborative deployment architecture. In local resource-constrained mode, a lightweight inference engine optimized by model quantization or pruning is run, handling only basic interactive tasks; In cloud-enhanced mode, complex inference tasks and high-performance rendering tasks are offloaded to cloud servers, and rendering results are received and displayed via streaming media protocols. The AI core engine is also equipped with a model routing middleware, which is used to dynamically select the optimal large language model instance for inference based on task type, response latency requirements and service availability score. The AI core engine adopts an edge computing architecture. In the absence of events, the main processor enters a deep low-power sleep state, and only the low-power coprocessor maintains environmental monitoring. All biometric data is extracted and compared locally, and only irreversible feature encoding is uploaded to the cloud.
[0014] In an optional implementation, it further includes: The working strategy is dynamically adjusted according to the device deployment scenario. The working strategy includes the configuration of perception sensitivity, wake-up threshold, interaction content and service interface. The deployment scenarios include high-dynamic pedestrian flow scenarios, low-dynamic private scenarios, and regular activity scenarios. Different scenarios correspond to different detection time windows and sensitivity threshold configurations. The method also includes: dynamically adjusting the system wake-up sensitivity through scene adaptation adjustment, user habit self-learning based on historical interaction data, and emotion linkage adjustment combined with multimodal sentiment analysis results; The method further includes: capturing the user's micro-expressions and gestures within a preset time after wake-up, using a sequence prediction model to predict the user's intentions and preload relevant information.
[0015] In an optional implementation, the naked-eye 3D display module integrates a viewpoint tracking unit, and the method further includes: The system detects the user's eye position in real time and dynamically adjusts the grating parameters or liquid crystal molecule orientation of the electronically controlled grating based on the detection results, so that the optimal viewing area follows the user's movement. The naked-eye 3D display module uses instantiation rendering technology to share material and mesh resources for multiple virtual characters, and combines frustum culling and occlusion culling to optimize rendering performance; The method further includes: converting the neural radiation field representation into a Gaussian splash representation through knowledge distillation, dynamically selecting key viewpoints for high-resolution rendering based on eye-tracking data, and using low-resolution rendering for adjacent viewpoints combined with super-resolution reconstruction for restoration.
[0016] In an optional implementation, it further includes: In highly privacy-sensitive scenarios, a purely local mode is adopted, in which the entire process of collecting, extracting, comparing and storing biometric data is completed in a local trusted execution environment; When cloud computing power is required, a cloud-edge collaborative mode is adopted to perform irreversible feature encoding processing on sensitive biometric data before uploading it to the cloud. It has a built-in automatic data cleanup mechanism, supports user-defined data retention periods, and automatically destroys sensitive intermediate data temporarily cached after the session ends; When the interactive system is deployed on a local terminal device, the AI core engine obtains the hierarchical structure of the application interface by calling the standardized auxiliary function interface provided by the operating system, thereby realizing the automatic recognition and interactive operation of interface elements.
[0017] Secondly, the present invention provides an interactive terminal that integrates multimodal perception and naked-eye 3D, comprising: At least three types of heterogeneous sensors are used. The first type of sensor is an active space detection unit, which is used to continuously output signals reflecting the activity status of targets in space. The second type of sensor is a passive environmental change sensing unit, which is used to output signals reflecting transient environmental changes. The third type of sensor is a visual information acquisition and processing unit, which is used to output target recognition results based on visual features. A glasses-free 3D display module is used to present virtual images; Interactive systems are used for multimodal interaction with users; The controller is connected to the at least three types of heterogeneous sensors, the naked-eye 3D display module, and the interactive system, respectively, and the controller is configured to perform the method described in any one of the foregoing embodiments.
[0018] This invention constructs a human-like "perception" system with context awareness by combining a multi-source fusion verification mechanism of active spatial detection and passive environmental perception with visual feature-based target existence verification. This significantly reduces the system's false trigger rate and greatly improves wake-up accuracy. A closed-loop processing architecture for sensitive data based on local edge computing ensures that sensitive information such as biometrics is parsed entirely on the edge, with only irreversible feature encoding synchronized to the cloud for identity verification, effectively blocking the risk of sensitive information leakage. A high-precision hardware clock synchronization protocol is used to achieve system-level time base alignment, coordinating with lightweight preprocessing logic on the edge to construct a low-latency wake-up link, compressing end-to-end response latency to an extremely short interval imperceptible to the user. A multimodal avatar construction interface is provided, supporting image sequences, 3D data, natural language, and acoustic signals, allowing users to quickly create their own digital avatars through diverse input methods, significantly lowering the barrier to entry. Based on the dynamic switching of behavior mapping strategies driven by the multimodal emotional state perception layer, the wake-up sensitivity and digital avatar expressiveness are adjusted in real time according to the user's multidimensional emotional state representation, significantly improving the emotional resonance of human-computer interaction. Software-defined cross-scenario adaptive configuration capabilities enable the same hardware platform to flexibly adapt to various application scenarios such as business, home, and office, significantly reducing operation and maintenance complexity and costs. Employing an independent trigger control circuit based on pure hardware logic gates, coupled with high-gain directional signal focusing technology, it achieves concealed deployment in complex, concealed environments and ensures the availability of basic functions in the absence of software intervention or system failure. Through a redundant wake-up architecture with primary and backup collaboration, the system can automatically switch to a hardware-level redundant bypass wake-up path in the event of a primary control unit malfunction or firmware failure. While ensuring the availability of basic functions, it records fault logs and automatically reverts to a high-security mode after recovery, significantly improving the long-term operational stability of the device under complex operating conditions. Attached Figure Description
[0019] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 A schematic diagram of an interaction method integrating multimodal perception and naked-eye 3D provided in an embodiment of this application; Figure 2 This is a schematic diagram of the overall system architecture provided in an embodiment of the present invention; Figure 3 This is a detailed schematic diagram of the primary / standby collaborative wake-up architecture provided in an embodiment of the present invention; Figure 4This is a multimodal wake-up process logic diagram provided in an embodiment of the present invention; Figure 5 This is a schematic diagram of the 3D virtual character generation and behavior mapping process provided in an embodiment of the present invention; Figure 6 This is a diagram illustrating the communication structure between the large language model AI and hardware provided in this embodiment of the invention. Figure 7 This is a schematic diagram of a commercial application scenario of the present invention in a shopping mall environment; Figure 8 This is a schematic diagram illustrating the application of an embodiment of the present invention in a home environment. Detailed Implementation
[0021] The technical solutions in the embodiments of this application will now be described with reference to the accompanying drawings.
[0022] This embodiment provides an interactive method integrating multimodal perception and naked-eye 3D, which is applied to an interactive terminal. The interactive terminal includes at least three types of heterogeneous sensors, a naked-eye 3D display module, and an interactive system. The first type of sensor is an active spatial detection unit, used to continuously output signals reflecting the activity state of targets within space; the second type of sensor is a passive environmental change perception unit, used to output signals reflecting transient environmental changes; and the third type of sensor is a visual information acquisition and processing unit, used to output target recognition results based on visual features.
[0023] Specifically, this embodiment employs a layered heterogeneous hardware architecture to implement the aforementioned sensing functions. The active space detection unit is preferably a microwave radar sensor, operating in the ISM band, capable of penetrating non-metallic obstructions for detection, and continuously monitoring the Doppler frequency shift signal within the detection area. The passive environmental change sensing unit is preferably a pyroelectric infrared sensor (PIR), made based on pyroelectric crystal materials, which outputs a transient jump signal only when it detects changes in infrared radiation caused by human movement. The visual information acquisition and processing unit is preferably a high-definition camera module, working in conjunction with an embedded image processing chip to run a lightweight target detection algorithm. These three types of sensors differ significantly in their detection principles, signal characteristics, and power consumption characteristics, complementing each other and forming the hardware foundation for multimodal sensing.
[0024] like Figure 1 As shown, the method specifically includes the following steps: Step S100: Determine whether the signal detected by the first type of sensor meets the main triggering condition. The main triggering condition includes detecting a stable state signal that meets the time window requirement.
[0025] In practice, the microwave radar sensor, as the primary detection device, is continuously operational. When a moving target appears within the detection area, the radar echo signal generates a Doppler frequency shift. This frequency shift signal is continuously monitored; the subsequent process is not triggered solely based on a single transient signal, but rather within a defined time window (e.g., 1 to 2 seconds). Only when a valid frequency shift signal is detected continuously or frequently within this time window—identified as a "steady-state signal"—is the primary trigger condition met. This design effectively filters out occasional transient interference in the environment, such as small animals quickly crossing or transient pulses of external electromagnetic noise, thus establishing the first line of defense against accidental triggering.
[0026] Step S200: When the main triggering condition is met, determine whether the transient change characteristics of the second type of sensor output meet the auxiliary verification condition.
[0027] Specifically, once the microwave radar detects a suspected target, it immediately reads the status of the PIR sensor. Because the PIR sensor is highly sensitive to the infrared radiation characteristics of the human body and its output signal exhibits typical transient jump characteristics, this is used as an auxiliary verification condition. It is determined whether the PIR sensor synchronously outputs a valid jump signal within the time window during which the radar detects stable activity. If the radar detects activity but the PIR sensor consistently outputs no signal, it may be due to a non-heat source object (such as a curtain being blown by the wind), and the auxiliary verification condition will not be met, thus terminating the process. This fusion logic of "active detection + passive verification" further narrows down the range of suspected targets and improves the accuracy of wake-up determination.
[0028] Step S300: When the auxiliary verification conditions are met, determine whether the target recognition result based on visual features output by the third type of sensor meets the third verification conditions.
[0029] After passing the first two verifications, the visual information acquisition and processing unit is activated, i.e., the camera is started to acquire images. At this time, a lightweight human detection algorithm (such as the YOLO series or SSD algorithm) is run to analyze the acquired images in real time. It identifies whether there are contours in the image that conform to human morphological characteristics, such as specific head-to-shoulder ratios and torso contours. Only when the visual algorithm confirms that the target is a "human body" is the third verification condition satisfied. This step is the final "identity verification," which completely eliminates false triggers caused by other heat sources or moving objects, ensuring high reliability of wake-up.
[0030] In step S400, when the third verification condition is met, a wake-up command is generated, the interactive system is activated by the wake-up command, the virtual image is presented through the naked-eye 3D display module, and multimodal interaction is carried out with the user.
[0031] Specifically, once all three levels of verification pass, a hardware interrupt signal or software wake-up command is generated, switching the main control processing unit from a low-power sleep state to a full-function operating state. The interactive system is activated, loading user preference settings and the AI core engine. Simultaneously, the glasses-free 3D display module powers on, rendering and presenting a 3D virtual image based on preset or user-customized content. Subsequently, the system enters a multimodal interaction mode, preparing to receive user input such as voice commands and gestures.
[0032] Through the above scheme, this embodiment constructs a three-level progressive anti-accidental touch logic of "main trigger - auxiliary verification - third verification". It uses a "stable state within a time window" as the main trigger condition, rather than a single transient signal, establishing a defense depth and effectively solving the problem of high false trigger rate of traditional single sensors. Simultaneously, the high-power vision module is only activated after passing the first two verifications, achieving a balance between low-power standby and high-accuracy wake-up, significantly improving the user experience.
[0033] In some embodiments, the verification threshold of the second-level PIR can be dynamically adjusted based on the Doppler frequency shift characteristics of the first-level radar; or, the activation timing of the third-level vision can be determined based on the phase alignment results of the first two-level signals.
[0034] A verification threshold is a critical value used to determine whether an event or state meets specific conditions. In this scenario, the second-level PIR might be used to further verify the target information detected by the first-level radar, determining whether a valid target has actually been detected by setting a verification threshold.
[0035] Doppler shift refers to the change in the frequency of the wave received by an observer when there is relative motion between the wave source and the observer. In radar systems, Doppler shift characteristics can reflect the target's motion state, such as speed and direction.
[0036] Targets in different motion states may have varying effects on PIR detection. For example, fast-moving targets may produce different signal characteristics in PIR detection, which could lead to false positives or false negatives if a fixed verification threshold is used. Dynamically adjusting the PIR verification threshold based on the radar's Doppler frequency shift characteristics can make PIR detection more accurate and flexible, adapting to targets in different motion states.
[0037] A functional relationship can be established, taking the radar's Doppler frequency shift characteristics as input and outputting an appropriate PIR verification threshold. For example, when the radar detects a target moving at a high speed, the PIR verification threshold can be appropriately increased to reduce false positives caused by interference signals generated by the target's rapid movement; when the target moving at a low speed, the verification threshold can be decreased to increase detection sensitivity.
[0038] The first two signals of the third-level vision system come from the radar of the first level and the PIR of the second level, respectively. The radar signal and the PIR signal may have a certain phase relationship in time and space.
[0039] Phase alignment refers to adjusting the phases of two or more signals to achieve synchronization or a specific phase relationship in time or space. The phase alignment results of the first two stages of signals can reflect the temporal and spatial consistency of target information detected by radar and PIR.
[0040] Third-level vision is typically used to provide more detailed target information, such as the target's appearance and shape. Determining the timing of vision activation based on the phase alignment results of the first two levels of signals ensures that the vision system activates at the appropriate time to acquire accurate target information.
[0041] For example, if the phase alignment results of the first two signal stages indicate a high degree of consistency between the target information in time and space, it suggests that a valid target may have indeed been detected. In this case, the vision system can be activated for further detection and analysis. If the phase alignment results are inconsistent, it may mean that there is interference or misjudgment in the detected signal. In this case, the vision system can be deactivated to avoid unnecessary waste of resources. For example, when the radar and PIR detect the target simultaneously, and the phases of the signals are basically aligned, it indicates that the target does exist and its position is relatively fixed. Activating the vision system at this time can more accurately obtain detailed information about the target.
[0042] Figure 2 This invention demonstrates the overall system architecture of an intelligent interactive terminal integrating multimodal perception and naked-eye 3D, as provided in an embodiment of the present invention. The system adopts a layered heterogeneous hardware architecture, with a master-slave collaborative wake-up mechanism at its core. It integrates a multimodal perception module, an AI core engine, a naked-eye 3D display module, and a scene adaptation module, constructing a complete interactive link from perception to processing to presentation.
[0043] Among them, the multimodal sensing module integrates a variety of heterogeneous sensors and is responsible for collecting user and environmental data, serving as the system's sensing entry point.
[0044]
[0045] The main control processing unit adopts a primary-backup collaborative wake-up architecture, which includes two wake-up paths: Sensor signals from the main wake-up path (normal mode) are aggregated to the STM32F4 series MCU via the I2C / SPI / UART interface.
[0046] The MCU implements dual-mode wake-up timing control and triple verification mechanism: First layer: Microwave radar continuously detects activity signals (meeting time window requirements); Second layer: PIR sensor outputs transient jump signal; The third step: Activate the OV5640 camera to quickly detect human body contours; After successful verification, the MCU sends a wake-up command to the Rockchip RK3566 main control chip; The RK3566 runs on the Linux operating system and carries the core AI engine, performing natural language understanding, task planning, and multimodal data fusion processing.
[0047] The microwave radar and PIR sensor signals of the backup wake-up path (emergency mode) are directly connected to the RK3566 main control chip through a hardware logic AND gate: When an MCU fault, power failure, or high-speed response mode is detected, the system automatically switches to this path. This enables physical layer forced wake-up that bypasses the software protocol stack, ensuring high system availability under extreme failure scenarios.
[0048] Multi-sensor synchronous trigger bus (PTP): The MCU sends a unified synchronous trigger signal to each sensor via a dedicated bus. Each sensor initiates sampling or outputs data tags within the same clock cycle; Achieve microsecond-level time alignment, eliminate time drift between multimodal data, and improve the accuracy of data fusion.
[0049] The naked-eye 3D display module consists of three components: an electronically controlled grating projection screen, a viewpoint tracking unit, and a stereoscopic image output. The working principle and main functions of each component are clearly explained.
[0050]
[0051] The core AI engine and interaction unit include: The image-behavior mapping unit is used to map the semantic intent and emotional state output by the AI core engine into the facial expressions, body movements and speech synthesis commands of the virtual character; The scene adaptation module is used to dynamically adjust the working strategy according to the device deployment scenario (commercial / home / office), including perception sensitivity, wake-up threshold, interaction content and service interface.
[0052] like Figure 2 As shown, the WT4102 microwave radar, HC-SR501 PIR sensor, OV5640 camera, and INMP441 microphone array collect multi-source data; each sensor aggregates data to the STM32F4 MCU via I2C / SPI / UART interfaces for wake-up determination. Normal mode: The MCU executes triple verification logic, and wakes up the RK3566 after successful verification; Emergency mode: The hardware logic AND gate directly triggers the RK3566 wake-up pin.
[0053] The RK3566 main control chip runs the AI core engine, completing natural language understanding, task planning, and multimodal data fusion.
[0054] Interactive output: The image-behavior mapping unit drives the virtual character; the naked-eye 3D display module outputs stereoscopic images; the viewpoint tracking unit realizes dynamic view zone following; and the scene adaptation module dynamically adjusts system parameters according to the deployment environment.
[0055] Figure 4 The flowchart illustrates the complete logic of the multimodal wake-up process according to an embodiment of the present invention. It begins in the device's deep sleep state and uses a three-level progressive judgment mechanism to determine whether to wake the system. The core anti-accidental touch logic is "continuous detection + transition verification + contour detection". This process ensures low power consumption while keeping the wake-up latency within an extremely short range imperceptible to the user.
[0056] Deep sleep (standby mode): Device status: The system is in deep sleep mode with extremely low power consumption, less than 10mW.
[0057] Sensors that remain powered: Only microwave radar and PIR sensors remain powered.
[0058] Main control status: The STM32F4 MCU is in low-power standby mode, and the RK3566 main control chip is in complete sleep mode.
[0059] Screen status: Naked-eye 3D display module is off.
[0060] Trigger condition detection: Microwave radar: detects frequency offset signal (f_d), which is a continuously detected signal based on the Doppler effect (f_d=2v / λ·cosθ, where v is the target velocity, λ is the wavelength, and θ is the angle).
[0061] PIR sensor: detects changes in infrared radiation and outputs a GPIO interrupt signal. The principle is that human movement causes changes in the voltage on the surface of the pyroelectric crystal.
[0062] After detecting a frequency offset signal, the microwave radar outputs an activity detection signal.
[0063] After the PIR sensor detects a change in infrared radiation, it outputs a transition signal that triggers a GPIO interrupt.
[0064] The signals from both sensors trigger an MCU interrupt, waking up the MCU to perform a fusion judgment.
[0065] Fusion Judgment (MCU Execution): Determines whether the continuous microwave detection activity exceeds 1 second. The judgment criterion is that the continuous detection time is greater than the preset time window (1 second). Determines whether there is a PIR transition during this period. The judgment criterion is that there is at least one valid transition signal within the time window.
[0066] The system will only proceed to the next stage if both conditions are met simultaneously; if either condition is not met, the system will ignore the signal and return to a deep sleep state.
[0067] Start the OV5640 camera and acquire image data. Use a lightweight object detection model (such as YOLO series / SSD) to quickly detect human contours within a limited time.
[0068] The system determines whether the human body conforms to preset morphological features (such as head-to-toe ratio and key point distribution). If the human body outline is recognized, the system enters the wake-up phase. If the human body outline is not recognized, the process terminates and the system returns to a deep sleep state.
[0069] System wake-up and interaction preparation: Activate the RK3566 main control chip and start the glasses-free 3D display module. Load user-personalized configurations (such as virtual avatar and interaction preferences). Start the AI core engine to bring the system into full-function interactive mode.
[0070] Example 2: This embodiment, based on Embodiment 1, further optimizes the reliability architecture and sensor performance of the interactive terminal. Specifically, as follows: Figure 3 As shown, the interactive terminal also includes an independent triggering link not controlled by the main operating system. This independent triggering link can be constructed from pure hardware logic circuits, serving as a hardware-level redundant backup path. It is used to perform logical operations on the output signals of the first and second types of sensors and then directly connect them to the hardware wake-up pin of the main control processing unit. Furthermore, this independent triggering link can also use an ultra-low-power coprocessor specifically for monitoring. If the main processor crashes, this coprocessor will wake it up.
[0071] like Figure 3 As shown, in normal operating mode, the multi-level progressive anti-mistouch verification logic described in Embodiment 1 is executed by the main control processing unit (such as an MCU) to ensure the accuracy of wake-up. However, to cope with extreme conditions such as main control software link abnormalities, firmware crashes, or system crashes caused by extreme electromagnetic interference, this embodiment designs an emergency mode. When a main control software link abnormality is detected or the system is in emergency mode, the system connects an independent trigger link not controlled by the main operating system via a switch (such as a relay or analog switch) to perform wake-up, which can be a physical wake-up.
[0072] Specifically, this pure hardware logic circuit includes monostable multivibrators and logic gates. Taking the signals output by a microwave radar sensor and a passive infrared sensor as examples, the microwave radar signal is first fed into a monostable multivibrator (e.g., based on a 555 timer) to shape the irregular pulsating signal into a fixed-width pulse signal. Subsequently, this pulse signal and the transition signal output by the PIR sensor are input together to a logic gate (e.g., an AND gate). Only when both signals are valid simultaneously does the logic gate output a high level, directly triggering the hardware wake-up pin of the main control processing unit. This process completely bypasses the MCU's software protocol stack, achieving physical layer forced wake-up with zero software intervention. The advantage of this design is that even if the main control program crashes or enters an infinite loop, the system can still maintain basic wake-up response capabilities, greatly improving the system's survivability and fault tolerance.
[0073] Furthermore, after the fault is recovered, a self-test program will be automatically executed. Once the software link is confirmed to be back to normal, the hardware direct connection will be automatically cut off, and the system will revert to normal mode, restoring the highly secure multi-level verification logic, thereby achieving a dynamic balance between security and reliability.
[0074] Regarding the specific implementation of the first type of sensor, in this embodiment, the first type of sensor is preferably a microwave radar sensor, which operates in the ISM band; the second type of sensor is preferably a passive infrared sensor, which is made based on pyroelectric crystal material. Considering that in practical applications, the equipment often needs to be concealed behind non-metallic shielding objects (such as glass counters or wooden partitions), this embodiment optimizes the antenna structure of the microwave radar sensor by using dielectric loading to match the dielectric constant of the non-metallic shielding object.
[0075] Specifically, the presence of non-metallic obstructions alters the effective electrical length of the antenna, leading to impedance mismatch and signal reflection. By optimizing the dielectric loading and adjusting the physical dimensions and dielectric constant distribution of the antenna's microstrip lines, the antenna's radiation characteristics behind the obstruction can be made equivalent to those in free space, significantly reducing penetration loss and improving detection sensitivity. As another alternative implementation, an electromagnetic wave focusing structure, such as an artificial electromagnetic surface (AMS) or a Fresnel-like lens structure, can be placed inside the non-metallic obstruction to focus the emitted microwave energy onto a specific detection area, enhancing the signal-to-noise ratio of the target area and enabling high-precision detection of the human body at long distances or under weak signal conditions.
[0076] Furthermore, to address the time alignment issue in multimodal data fusion, this embodiment introduces a hard-wired trigger bus mechanism. A unified synchronization trigger signal is sent to each sensor via the hard-wired trigger bus, controlling each sensor to start sampling within the same clock cycle for microsecond-level time alignment.
[0077] Specifically, the main control unit, acting as the master node, sends synchronization pulses to sensors such as microwave radar, PIR, and cameras via a dedicated synchronization signal line. Each sensor integrates synchronization trigger logic, initiating sampling or marking data timestamps the instant a pulse is received. This hardware-level synchronization mechanism eliminates time deviations caused by sampling clock drift between sensors, improving time alignment accuracy to the microsecond level. This is crucial for subsequent data fusion; for example, when determining whether microwave and PIR signals are caused by the same target, high-precision time synchronization can effectively eliminate environmental noise interference, further enhancing the robustness of multimodal sensing.
[0078] Figure 3 A detailed schematic diagram of the primary and backup collaborative wake-up architecture of this invention is presented. Employing a dual-path redundancy design, it covers normal mode (primary wake-up path) and emergency mode (backup wake-up path), supporting intelligent dynamic switching and automatic fault rollback, ensuring stable wake-up capability of the system under different operating conditions.
[0079] The normal mode (main wake-up path) path consists of the WT4102 microwave radar, HC-SR501 PIR sensor, STM32F4 MCU, OV5640 camera and RK3566 main control chip connected in sequence.
[0080] Workflow: The WT4102 microwave radar actively probes the space, detecting frequency shift signals and using Doppler frequency shift to detect moving objects. The HC-SR501 PIR sensor performs passive infrared detection, outputting a transition signal to detect changes in infrared radiation caused by human movement. The STM32F4 MCU receives signals from the microwave radar and PIR sensor, executing dual-mode wake-up timing control and a triple verification mechanism. Only when continuous microwave detection and PIR transition meet the conditions does the OV5640 camera activate for human contour detection. After successful triple verification, the STM32F4 MCU sends a wake-up command to the RK3566. The RK3566 main control chip receives the wake-up command and starts, running the AI core engine and entering full-function interactive mode.
[0081] Key mechanisms: The triple verification logic consists of continuous microwave detection (with time window requirements), PIR transition verification, and visual contour detection.
[0082] In normal mode, the relay is disconnected, and the hardware direct connection is isolated.
[0083] The emergency mode path consists of a WT4102 microwave radar, an HC-SR501 PIR sensor, a Timer555 timer, hardware logic AND gates, and an RK3566 main control chip connected in sequence.
[0084] Triggering conditions: Fault detection finds that the STM32F4 MCU is unresponsive for a long time (such as watchdog timeout), abnormal power supply fluctuations, or the user has configured it to "fast response mode".
[0085] Workflow: A dedicated low-power coprocessor or power management chip monitors the system's health status and detects an MCU fault. A switching control signal closes a relay, activating the hardware direct connection. The WT4102 microwave radar continuously detects activity signals and outputs a frequency offset signal. A Timer555 converts the microwave radar signal into a fixed-width pulse signal. An HC-SR501 PIR sensor detects changes in infrared radiation and outputs a transition signal. A hardware logic AND gate performs a logical AND operation between the shaped pulse signal and the PIR transition signal. The logic gate output directly triggers the RK3566 hardware wake-up pin, bypassing the software protocol stack for direct wake-up.
[0086] Key Mechanism: Zero software intervention, completely bypassing the STM32F4 MCU and software protocol stack. Physical layer forced wake-up, with pure hardware circuitry directly triggering the RK3566 wake-up pin. Ensuring basic functional availability under extreme fault scenarios, achieving extremely fast response.
[0087] The switching logic of the intelligent switching and fallback mechanism is as follows: When the MCU is working normally, the relay is open, and the system is in normal mode (main wake-up path). When an MCU fault, power failure, or user configuration to ultra-fast response mode is detected, the relay is closed, and the system switches to emergency mode (backup wake-up path). Once the MCU recovers and passes the self-test, the system automatically cuts off the hardware direct connection, the relay is open, the system switches back to normal mode, and restores the high-security strategy (triple verification path).
[0088] Example 3: This embodiment, based on the above embodiments, provides a detailed description of the creation process of the AI companion avatar. Before presenting the virtual avatar through the naked-eye 3D display module, the AI companion avatar is created based on at least one of the following methods: face reconstruction based on single or multiple frames of images; pose estimation and 3D model reconstruction based on video sequences; direct import of 3D model files; semantic parsing and generative modeling based on natural language text descriptions or speech descriptions.
[0089] Specifically, such as Figure 5As shown, this embodiment provides multiple parallel image generation paths to adapt to different user habits and data sources. For face reconstruction based on single-frame or multi-frame images, the system calls the camera to capture the user's facial image or receives image files uploaded by the user. Subsequently, using generative adversarial networks (such as the StyleGAN2 series algorithms) or 3D reconstruction algorithms, the 3D geometric structure and texture information are inferred from the 2D image to generate a basic face mesh with the user's facial features. This method is simple to operate, requiring only one photo to complete the modeling, greatly reducing the threshold for image creation.
[0090] For video sequence-based pose estimation and 3D model reconstruction, a pose estimation network is used to extract 2D joint coordinates from video frames, and then a 3D pose enhancement network is used to reconstruct a parametric 3D human model. Specifically, firstly, pose estimation networks such as OpenPose or HRNet are used to extract 2D joint coordinates frame by frame from user-uploaded video clips; then, these 2D coordinate sequences are input into a 3D pose enhancement network such as VIBE or SMPLify to reconstruct a parametric 3D human model (such as the SMPL-X model) that includes skeleton binding and skinning weights. This approach can capture the user's dynamic movement features, enabling the generated virtual avatar to have more natural body language expression capabilities.
[0091] For direct import of 3D model files, it supports direct parsing of standard 3D format files, such as OBJ, FBX, or GLTF. It automatically reads the geometric meshes, material maps, and skeletal binding information contained in the file and converts them into its internally compatible rendering format. This method is suitable for users with professional modeling capabilities or third-party content providers, greatly expanding the sources of image resources.
[0092] For generative modeling based on natural language text or speech descriptions, a large language model (LLM) is used to parse semantic tags, and a text-based 3D model algorithm is invoked to generate a 3D avatar with a neural radiation field (NeRF) representation or a Gaussian splatting representation. Specifically, users can describe the desired image by inputting text (such as "a cartoon bear wearing a spacesuit") or speech. The semantics of the input are parsed using a large language model (LLM) to extract key feature tags (such as style, clothing, species, etc.). Subsequently, text-based 3D model algorithms such as RODIN or Make-A-Character are invoked to generate the corresponding 3D avatar with a neural radiation field (NeRF) representation or a Gaussian splatting representation. Among them, the neural radiation field (NeRF) representation can model complex lighting and geometric details with fine detail, while the Gaussian splatting representation has a significant advantage in rendering efficiency, making it particularly suitable for real-time rendering on embedded terminals. By converting the NeRF representation to a Gaussian splatting representation through knowledge distillation, the visual fidelity is maintained while significantly reducing the demand for hardware computing power, achieving smooth display of high-quality images at the edge.
[0093] Furthermore, for voice description methods, acoustic emotional features are extracted simultaneously from the speech, and the initial facial expression style or body posture of the generated avatar is automatically adjusted based on these emotional features. Specifically, when a user describes an avatar via voice, not only is the speech content recognized, but acoustic features such as pitch, speech rate, energy, and pause rhythm are extracted in parallel. These acoustic features are mapped to emotional tags (such as "cheerful," "serious," and "gentle"). Based on these emotional tags, the parameters in the generation algorithm are automatically adjusted. For example, when a user's tone is detected to be light and energetic, a smiling expression and lively body posture are automatically assigned to the generated avatar; when a flat and low tone is detected, a calm expression and dignified posture are assigned to the avatar. This "expressive" generation method makes the avatar creation process itself full of emotional interaction, and the generated avatar is more in line with the user's subconscious psychological expectations.
[0094] By setting up the above-mentioned various creation methods, this embodiment achieves full coverage from "zero-threshold" rapid generation to "professional" fine customization, meeting the needs of different user groups and significantly improving the personalized expression capabilities and user stickiness of interactive terminals.
[0095] Example 4: This embodiment, based on the above embodiments, provides a detailed explanation of the emotion perception and response mechanism in the multimodal interaction process. Specifically, multimodal interaction with the user includes: fusing and analyzing micro-expressions captured by the camera, voice emotion features extracted by the microphone, and physiological rhythm signals detected by radar to generate a label for the user's current emotional state.
[0096] like Figure 5As shown, during the interaction, multimodal data is continuously collected to perceive the user's emotional state. For the visual dimension, a camera captures the user's facial image, and a trained convolutional neural network (CNN) is used to analyze the image in real time, identifying minute muscle movements in key areas such as the corners of the eyes and mouth, thereby extracting micro-expression features. For the auditory dimension, a microphone array collects the user's speech, extracting acoustic features such as Mel-frequency cepstral coefficients (MFCC), fundamental frequency, and short-time energy, and inputting them into a sequence model (such as LSTM or Transformer) for emotion classification, identifying vocal emotion features such as joy, anger, and sadness. For the physiological dimension, electromagnetic wave signals emitted by a microwave radar sensor are used to analyze the echo phase changes caused by the rise and fall of the human chest cavity. A demodulation algorithm is used to extract micro-motion signals caused by heartbeat and respiration, and then physiological rhythm signals such as heart rate, respiratory rhythm, and heart rate variability (HRV) are calculated. These physiological indicators can objectively reflect the user's stress level and emotional arousal, compensating for the shortcomings of visual and auditory modalities, which are easily masked by subjective performance.
[0097] The three heterogeneous features mentioned above are fused to generate a user's current emotional state label. Specifically, a weighted fusion algorithm can be used to dynamically adjust the weights of each modality based on the current ambient lighting and noise levels. For example, in noisy environments, the weight of speech emotion features is reduced, while the weights of physiological signals and visual features are increased, thereby ensuring the robustness of emotion recognition. The generated emotional state labels (such as "anxiety," "focus," and "excitement") will serve as key input parameters for subsequent behavioral decisions.
[0098] Furthermore, based on the combination of emotional state tags and semantic intent, the optimal action sequence and voice style are matched from a preset behavioral rule base to drive the virtual character to make an empathetic response.
[0099] Specifically, the behavior rule base pre-stores various response strategies corresponding to different emotion-intent combinations. For example, when the semantic intent of a user is identified as "checking bills" and the emotional state label is "anxious," a "soothing" response strategy is matched from the rule base. At this time, the virtual avatar's action sequence is adjusted to "leaning forward, making gentle eye contact, and using gestures to guide," and the speech style generated by the text-to-speech (TTS) module is also adjusted to "slowing down the speech rate and softening the tone." This emotion-driven empathetic response makes the interaction process no longer a mechanical execution of instructions, but rather imbued with human-like emotional warmth, effectively alleviating the user's negative emotions and enhancing the friendliness of human-computer interaction.
[0100] Furthermore, for complex scenarios involving multiple concurrent users, such as in homes or public places, this embodiment also provides a multi-user interaction management mechanism. It uses spatial location-aware data to divide user areas and maintains multi-path concurrent session state management to continuously track and switch independent user contexts.
[0101] Specifically, the system utilizes the ranging and angle-measuring capabilities of microwave radar, combined with visual positioning information from cameras, to divide the interaction space into multiple independent user areas. To accurately identify the current interaction object, a multi-user identity confidence fusion formula is used for identity determination. ; in, This represents the confidence level of user i's identity; , , , These represent similarity scores based on face recognition, voiceprint recognition, spatial location association, and behavioral pattern matching, respectively. , , , These are dynamic weighting coefficients, and their sum is 1. The confidence level of each candidate user is calculated using this formula, and the user with the highest confidence level is determined as the current interaction subject.
[0102] After identifying the interaction subjects, independent multi-channel concurrent session state management is maintained. Each user has an independent context stack for storing historical conversation records, preference settings, and current emotional state. When users move between different areas or multiple users take turns interacting, the context can be quickly switched based on the identity recognition result, achieving a seamless interactive experience. For example, when user A pauses interaction to handle other tasks and then returns to continue interacting, user A's previous session state can be automatically restored without the user having to repeat background information. This mechanism effectively avoids identity confusion and context crosstalk in multi-user scenarios, ensuring the continuity and personalization of the interaction.
[0103] Figure 5 The flowchart for 3D virtual avatar generation and behavior mapping is divided into two branches. The upper branch shows the 3D virtual avatar generation path, demonstrating the process of users creating personalized AI companion avatars using various input methods. The lower branch shows the behavior mapping and emotional response path, demonstrating how the system drives the virtual avatar to make empathetic responses based on the user's emotional state and semantic intent. This attached diagram illustrates the core technical advantages of this invention: "zero-threshold avatar customization" and "emotion-driven interaction."
[0104] Upper branch: 3D virtual image generation path.
[0105]
[0106] Detailed explanation of the processing flow for each input method: Algorithms for image input (real-time image capture / recognition of existing images): StyleGAN2 / DeepFaceLab.
[0107] The image is first preprocessed, then a generative adversarial network is used to generate a textured 3D face mesh. The output is a high-quality 3D face model.
[0108] Algorithm for video input (uploaded video clips): OpenPose+VIBE→SMPL-X.
[0109] OpenPose extracts 2D key coordinates from video frames; VIBE (Video Inference for Body Expression) reconstructs the 2D sequence into a parametric 3D human body model; SMPL-X generates a complete human body model with facial expressions and gestures. The output is a parametric 3D human body model (including skeletal rigging).
[0110] 3D model import formats: OBJ / FBX / GLTF. The system directly parses the skeleton, materials, and texture information of the model file. The output is a 3D model that can be directly used for rendering.
[0111] Algorithms for natural language description: LLM → RODIN / Make-A-Character → NeRF.
[0112] The large language model parses semantic tags (such as "long blonde hair, blue eyes, cartoon style"); RODIN or Make-A-Character calls the Wensheng 3D model algorithm; and generates 3D avatars represented by NeRF (Neural Radiation Field). The output is the 3D avatar represented by NeRF.
[0113] The speech description processing flow is as follows: ASR (Automatic Speech Recognition) transcribes speech into text; simultaneously, acoustic emotional features (pitch, speech rate, energy) are extracted from the speech; the text is fed into LLM for semantic parsing; the emotional features are used to automatically adjust the initial facial expression style of the generated avatar. The output is a 3D avatar with emotional adaptation.
[0114] 3D models generated from various input methods are uniformly converted into the standard FBX format (Unity compatible format) and sent to the "Image-Behavior Mapping Unit".
[0115] Sub-branch behavior mapping and emotional response path: It can collect multi-source user state data in parallel and perform fusion analysis. Specific information is as follows: The camera primarily captures facial expressions and eye movements. Facial expressions reflect a user's current emotions, attitudes, and other psychological states; eye movements reveal the user's focus and attention. Relevant features are extracted through micro-expression recognition and gaze direction analysis. Micro-expression recognition can capture subtle, instantaneous changes in a user's facial expressions, which often reflect their true emotions; gaze direction analysis can pinpoint the location of a user's attention within a specific scene.
[0116] A Convolutional Neural Network (CNN) can be used. CNNs have powerful image feature extraction capabilities and can automatically learn feature patterns in images, thereby accurately recognizing the information contained in facial expressions and eye movements.
[0117] The microphone captures speech signals, which contain information such as the user's speech content and tone. Emotional features are extracted, such as MFCC (Mel-frequency cepstral coefficients), pitch, speech rate, and energy. MFCC effectively describes the spectral characteristics of speech; pitch reflects changes in volume; speech rate indicates the speed of speaking; and energy represents the intensity of the speech signal. These features help determine the user's emotional state and intended expression.
[0118] LSTM / Transformer sequence models can be used. LSTM can process sequence data and capture long-term dependencies in speech signals; Transformer models, on the other hand, are efficient in processing sequence data and have powerful parallel computing capabilities, enabling them to better extract features from speech signals.
[0119] Microwave radar collects physiological signals closely related to human physiological activities. It extracts heart rate and respiratory rhythm characteristics. Heart rate reflects the frequency of the heartbeat, and respiratory rhythm reflects the frequency and depth of breathing; these are important indicators of human physiological state. Phase change analysis of the microwave radar is also employed. The microwave signals emitted by the radar are reflected when they encounter the human body. The phase changes of the reflected signals are related to human physiological activities (such as heartbeat and respiration). By analyzing these phase changes, heart rate and respiratory rhythm information can be accurately extracted.
[0120] PIR sensors collect body temperature distribution data, enabling the detection of temperature distribution around the human body. They then extract characteristics of body temperature changes. These changes can reflect information such as a person's health status and activity level.
[0121] Infrared radiation detection technology can be used. The human body emits infrared radiation, and PIR sensors obtain information on body temperature distribution and changes by detecting the intensity and distribution of infrared radiation.
[0122] It integrates multi-dimensional emotional features and outputs the user's current emotional state label, such as anxiety, focus, excitement, frustration, confusion, happiness, sadness, etc.
[0123] The core AI engine is deployed on a local RK3566 platform and integrates natural language understanding capabilities. It outputs semantic intent commands, such as welcome, reminder, sadness, inquiry, recommendation, and comfort.
[0124] The matching results drive the 3D model to make coordinated changes in facial expressions and postures. For example, TTS (text-to-speech) synthesis produces speech output that matches the emotional style. The naked-eye 3D display module presents a virtual image with emotional resonance.
[0125] Example 5: This embodiment, based on the above embodiments, provides a detailed description of the core computing architecture of the interactive system. Specifically, the interactive system includes an AI core engine that supports a cloud-edge collaborative deployment architecture.
[0126] like Figure 6 As shown, the AI core engine, acting as the "brain" of the system, is not deployed in a static manner but dynamically adjusted based on the current computing resources and task requirements. In local resource-constrained mode, a lightweight inference engine optimized by model quantization or pruning runs, handling only basic interactive tasks. Specifically, when the device is offline or network bandwidth is insufficient, it automatically switches to local mode. In this mode, the AI core engine loads lightweight models that have undergone INT8 quantization or structured pruning. Although these models have a small number of parameters, they are sufficient to handle basic tasks such as voice wake-up, simple command recognition, and preset action triggering. This design ensures that the device still has basic interactive capabilities in environments without or with weak network coverage, avoiding the embarrassing situation of "paralysis upon network outage" and significantly improving system availability.
[0127] In cloud-enhanced mode, complex inference tasks and high-performance rendering tasks are offloaded to cloud servers, and rendering results are received and displayed via streaming media protocols. Specifically, when users initiate complex semantic understanding requests (such as multi-turn dialogue inference or code generation) or when high-fidelity 3D scenes need to be rendered, local computing power may be insufficient to meet real-time requirements. In this case, raw, non-sensitive data (such as text commands or low-resolution preview images) is uploaded to a high-performance cloud server via an encrypted channel. The cloud server utilizes a powerful GPU cluster for inference operations or ray tracing rendering, and the generated audio or video streams are pushed back to the terminal in real time via low-latency streaming media protocols such as WebRTC and RTMP. This "thin client + powerful cloud" model enables terminal devices to provide an interactive experience comparable to high-performance workstations while maintaining a slim and low-power design.
[0128] To achieve seamless switching and optimal resource allocation between the two modes, the AI core engine is also equipped with a model routing middleware. This middleware dynamically selects the optimal large language model instance for inference based on task type, response latency requirements, and service availability scores. Specifically, the model routing middleware maintains a dynamic service quality score table. When a task is received, the middleware first parses the task's type label (such as "casual conversation," "professional knowledge question answering," or "emergency control") and evaluates the response latency and service availability scores of each candidate model instance (including local small models, cloud-based general large models, and cloud-based vertical domain-specific models). For example, for "emergency control" tasks, the middleware prioritizes the local model with the lowest latency; for "professional knowledge question answering," it prioritizes the cloud-based dedicated model with higher accuracy. This intelligent scheduling mechanism maximizes the use of cloud computing resources while ensuring response speed, achieving an optimal balance between cost and performance.
[0129] Furthermore, the AI core engine in this embodiment adopts an edge computing architecture. In the absence of events, the main processor enters a deep low-power sleep state, with only a low-power coprocessor maintaining environmental monitoring. All biometric data undergoes feature extraction and comparison locally, with only irreversible feature encoding uploaded to the cloud. Specifically, in standby mode, the high-performance main application processor (such as the RK3566) is powered off or enters deep sleep mode, reducing power consumption to the milliwatt level. At this time, only a low-power coprocessor (such as an MCU) works with sensors to perform environmental monitoring. Once a wake-up event is detected, the main processor is quickly awakened. During interaction, the flow of sensitive biometric data such as facial images and voiceprint recordings is strictly restricted. All feature extraction and comparison processes are completed within the trusted execution environment (TEE) or secure sandbox of the local chip, ensuring that the raw data never leaves the device. Only the extracted irreversible feature encoding (such as a feature vector digest processed by a hash algorithm) is uploaded to the cloud for cross-device identity synchronization or historical record matching. Because the hashing process is one-way, even if the cloud obtains the feature encoding, it cannot deduce the original facial image or voice, thus blocking the path of privacy leakage at the physical level and completely eliminating users' concerns about privacy and security.
[0130] By combining the aforementioned cloud-edge collaborative architecture with privacy protection mechanisms, this embodiment achieves both a low-latency, highly intelligent interactive experience and a robust data security defense, solving the dual pain points of high latency and high privacy leakage risk caused by cloud dependence in existing technologies.
[0131] Example 6: This embodiment, based on the above embodiments, provides a detailed description of the system-scene adaptive strategy. Specifically, this embodiment also includes dynamically adjusting the working strategy according to the device deployment scenario. This working strategy includes the configuration of perception sensitivity, wake-up threshold, interaction content, and service interfaces.
[0132] In practical applications, the deployment environments of interactive terminals vary greatly, and fixed parameter configurations cannot meet the needs of all scenarios. Therefore, this embodiment divides the deployment scenarios into high-dynamic pedestrian flow scenarios, low-dynamic private scenarios, and regular activity scenarios, with different detection time windows and sensitivity threshold configurations corresponding to different scenarios.
[0133] Specifically, in high-dynamic pedestrian traffic scenarios, such as shopping mall entrances and transportation hubs, where pedestrian flow is frequent and background noise is complex, conventional high-sensitivity configurations are prone to frequent false triggers by passersby, leading to resource waste and a degraded user experience. To address these scenarios, the system automatically implements a "noise-resistant, fast-response" strategy: shortening the main trigger time window of the microwave radar to 0.5 to 1 second, requiring the target to exhibit more obvious activity characteristics within a short period before triggering the main judgment; simultaneously, appropriately lowering the sensitivity threshold of the passive infrared sensor, for example, increasing the trigger threshold to 120% to 150% of its original value, to suppress high-frequency environmental noise. While this configuration sacrifices some ability to capture weak signals, it significantly reduces the false trigger rate and ensures stable operation in noisy environments.
[0134] Conversely, in low-dynamic, private settings such as family bedrooms and nursing home wards, user activity is often subtle and slow, with extremely low tolerance for missed detections. For these scenarios, an automatic switch to a "high-sensitivity, deep-detection" strategy is implemented: the main trigger window of the microwave radar is extended to 2-3 seconds to capture slow body movements or subtle breathing signals; simultaneously, the sensitivity of the passive infrared sensor and microwave radar is increased, for example, by boosting the microwave radar's signal gain by 20%-30%, ensuring timely response to subtle user activities, such as getting up at night, and providing prompt care services.
[0135] For routine scenarios such as offices and exhibition halls, the system employs a balanced detection threshold configuration with a time window set between 1 and 2 seconds, maintaining the factory calibration sensitivity to balance response speed and accuracy. It should be understood that the above numerical ranges are merely illustrative of a preferred embodiment, and in actual applications, adaptive fine-tuning can be performed based on the specific environmental noise level.
[0136] Furthermore, to achieve more intelligent and personalized services, this embodiment also introduces a multi-dimensional dynamic adjustment mechanism. This method further includes: dynamically adjusting the system's wake-up sensitivity through scene adaptation adjustment, user habit self-learning based on historical interaction data, and emotion-linked adjustment combining multimodal sentiment analysis results.
[0137] Specifically, the user habit self-learning module continuously records the user's interaction history, including high-frequency interaction periods, common interaction distances, and typical approach speeds. Using a local incremental learning algorithm, a user-specific wake-up response mapping relationship is constructed. For example, if it learns that the user habitually engages in entertainment interactions between 8:00 PM and 10:00 PM each night, the wake-up threshold is automatically lowered during this period to achieve a "zero-wait" response; while during late-night sleep, the threshold is automatically raised to avoid accidental wake-ups that disturb the user's rest. Furthermore, the emotional linkage adjustment mechanism can dynamically adjust its strategy based on the user's emotional state identified in Example 4. When the system detects that the user is anxious or agitated, it automatically lowers the wake-up threshold and preloads reassuring service interfaces to ensure the user receives immediate response and support; when the system detects that the user is focused on work, it appropriately raises the wake-up threshold to avoid frequently interrupting the user's train of thought.
[0138] Furthermore, to further reduce interaction latency, this embodiment also provides an intent prediction mechanism. The method further includes: capturing the user's micro-expressions and gestures within a preset time after wake-up, using a sequence prediction model to predict the user's intent and preload relevant information.
[0139] Specifically, within a very short time after the system is activated (e.g., 100 to 300 milliseconds), the system does not wait for the user to issue a complete voice command. Instead, it uses the camera to capture the user's facial micro-expressions (such as eye focus direction and lip movements) and hand gestures (such as raising a hand or pointing). This temporal data is input into a pre-trained sequence prediction model, such as a Long Short-Term Memory (LSTM) network or a Gated Recurrent Unit (GRU). Based on the currently captured behavioral sequence features, the model predicts the user's possible intention category. For example, when the model detects that the user's eyes are focused on the bottom of the screen and the body is leaning forward, it predicts that the user's intention is "to view details"; when the model detects that the user raises a hand and the lips are moving slightly, it predicts that the user's intention is "voice input". Based on this, the system preloads relevant UI interfaces, data content, or API interfaces from a local database or cloud server in advance, so that the system can achieve a "second-level" response the moment the user actually issues a command, greatly improving the smoothness and intelligence of the interaction.
[0140] Example 7: This embodiment, based on the above embodiments, provides a detailed description of the optimization of the display effect and the improvement of the rendering performance of the naked-eye 3D display module. Specifically, the naked-eye 3D display module integrates a viewpoint tracking unit, and the method further includes: real-time detection of the user's eye position, and dynamic adjustment of the grating parameters of the electronically controlled grating or the orientation of the liquid crystal molecules based on the detection results, so that the optimal viewing area follows the user's movement.
[0141] Specifically, the viewpoint tracking unit preferably employs a miniature infrared eye-tracking module, or reuses the camera module from the aforementioned embodiments, combined with iris recognition algorithms such as MediaPipe Iris, to capture the spatial coordinates of the user's pupil center in real time. This coordinate data is transmitted to the grating control driver via a high-speed serial interface. As an active optical element, the electro-optical grating's period, slit width, or the deflection angle of the liquid crystal molecules can be electrically adjusted. An internal mapping model between the viewpoint position and grating parameters is pre-stored. When left-right displacement of the user's head is detected, the driver calculates and updates the grating parameters in real time, dynamically adjusting the grating's diffraction angle so that the separation interface of the left and right eye views is always aligned with the user's current pupil position. This dynamic parallax correction mechanism effectively solves the "viewpoint dead zone" problem present in traditional naked-eye 3D display technology, allowing users to move freely within a certain range without image ghosting or reversal, significantly improving the continuity and immersion of stereoscopic vision.
[0142] In terms of rendering performance optimization, the naked-eye 3D display module uses instantiation rendering technology to share material and mesh resources for multiple virtual characters, and combines frustum culling and occlusion culling to optimize rendering performance.
[0143] Specifically, in multi-user concurrent interaction or complex scene rendering, if there are multiple virtual characters with similar appearances but different postures (such as multiple digital clones of different users) in the scene, traditional rendering methods require loading a separate set of mesh and texture resources for each character, resulting in a huge waste of video memory resources and congestion of rendering bandwidth. The instantiation rendering technology used in this embodiment allows the graphics processing unit (GPU) to render multiple objects with the same geometric structure in a single draw call. Only one set of mesh and texture resources needs to be loaded as a "master". By passing different transformation matrices (position, rotation, scaling) and a few differential parameters to each instance, multiple characters with different appearances can be generated. This technology merges the original dozens of draw calls into one, significantly reducing the communication overhead between the CPU and GPU, and significantly improving rendering efficiency.
[0144] Simultaneously, by combining view frustum culling and occlusion culling techniques, invalid rendering load is further eliminated. View frustum culling determines whether objects in the scene are within the field of view based on the camera's view frustum range; if an object is completely outside the view frustum, its rendering process is skipped. Occlusion culling determines whether an object is completely occluded by other objects in front; if occluded, it is also not rendered. The combination of these two culling techniques renders only the pixels that the user can actually see, reducing the rendering load by 30% to 50% in complex scenes, effectively alleviating the bottleneck of computational strain on embedded platforms.
[0145] To address the issue of low rendering frame rates for high-fidelity virtual avatar models (such as models generated based on NeRF neural radiation fields) on edge devices, this embodiment further includes: converting the neural radiation field representation into a Gaussian splash representation through knowledge distillation, dynamically selecting key viewpoints for high-resolution rendering based on eye-tracking data, and using low-resolution rendering for adjacent viewpoints combined with super-resolution reconstruction for restoration.
[0146] Specifically, while Neural Radiation Fields (NeRF) can meticulously model complex lighting and geometric details, its rendering process requires extensive sampling and integration calculations for each ray, resulting in massive computational demands and hindering real-time interaction on mobile devices. Gaussian Splatting, on the other hand, uses an explicit set of 3D Gaussian ellipsoids to represent the scene, and during rendering, it performs projection and blending through a rasterization pipeline, achieving speeds several orders of magnitude faster than NeRF. This embodiment utilizes knowledge distillation, using a pre-trained NeRF model as the "teacher network" to guide the training of a lightweight Gaussian Splatting model as the "student network." During distillation, by constraining pixel consistency, structural similarity (SSIM), and perceptual loss in the rendered images from the same viewpoint, the Gaussian Splatting model inherits the high-fidelity visual features of the NeRF model while maintaining extremely high rendering efficiency, making it highly suitable for deployment on resource-constrained interactive terminals.
[0147] Furthermore, to further reduce computational consumption while maintaining visual clarity, a gaze-based rendering strategy was introduced. Based on eye-tracking data, the central region (key viewpoint) currently being gazed upon by the user is identified. This region is rendered at full resolution or even through oversampling to ensure the sharpness of the core visual content. For the peripheral regions (adjacent viewpoints) that the user's peripheral vision focuses on, a lower resolution is used for coarse rendering. Subsequently, a super-resolution reconstruction algorithm (such as the deep learning-based ESRGAN algorithm) is used to upscale the low-resolution peripheral region image to the target resolution and merge it with the central region. Since the human eye is not sensitive to peripheral visual blur, this hierarchical rendering strategy significantly reduces the number of pixels requiring real-time computation while maintaining a virtually unaffected subjective visual experience for the user, making it possible to smoothly run high-fidelity naked-eye 3D interaction on embedded platforms.
[0148] Through the comprehensive implementation of the above-mentioned viewpoint tracking linkage, rendering pipeline optimization and model lightweight conversion, this embodiment successfully breaks through the application bottleneck of naked-eye 3D technology on consumer terminal devices, and achieves a stereoscopic display effect with high image quality, low latency and low power consumption, significantly improving the immersiveness and comfort of user interaction.
[0149] Example 8: This embodiment, based on the above embodiments, provides a detailed description of the system's privacy protection mechanism and automated expansion capabilities. With the increasing prevalence of smart interactive devices in sensitive scenarios such as homes and healthcare, the privacy and security of user data has become a core concern. Therefore, this embodiment constructs a tiered privacy protection system to adapt to the security needs of different scenarios.
[0150] Specifically, in highly privacy-sensitive scenarios, a purely local mode is adopted, where the entire process of collecting, extracting, comparing, and storing biometric data is completed in a local trusted execution environment.
[0151] The Trusted Execution Environment (TEE) described here refers to a secure area isolated within the main control processor (such as the RK3566 chip). It possesses independent memory space and encryption keys, and is logically isolated from the general-purpose operating system (such as Android or Linux) at the hardware level. In practice, when a user registers their face or voiceprint, data processing is not limited to the application layer. Instead, a secure monitoring call command directly transmits the raw data collected by the camera or microphone to the TEE. Within the TEE, the system runs a digitally signed trusted application, executing feature extraction algorithms (such as facial feature vector extraction). The extracted feature vectors are encrypted and stored in a secure storage partition accessible only to the TEE. This process ensures that even if the general-purpose operating system is rooted or subjected to malware attacks, attackers cannot read the raw biometric data, thus building a robust data security defense at the hardware level.
[0152] When cloud computing power is required, a cloud-edge collaborative mode is adopted to perform irreversible feature encoding processing on sensitive biometric data before uploading it to the cloud.
[0153] Specifically, when large cloud-based models are needed for identity synchronization or personalized model fine-tuning, data must be uploaded. Uploading raw images or audio is strictly prohibited. Instead, the extracted feature vectors are irreversibly encoded locally. For example, a one-way hash algorithm (such as SHA-256) is used to calculate a fixed-length hash digest; or homomorphic encryption is used to encrypt the feature vectors. Due to the one-way nature of hash algorithms, the cloud server only receives a meaningless string of codes and cannot deduce the user's facial image or voice content. The cloud verifies identity or updates the model by comparing these codes with records in the database, thus ensuring functionality while completely blocking the risk of privacy leaks.
[0154] In addition, to further protect users' data sovereignty, this embodiment also has a built-in automatic data cleanup mechanism, supports users to customize the data retention period, and automatically destroys sensitive intermediate data temporarily cached after the session ends.
[0155] Specifically, the storage module is configured with a data lifecycle management policy. Users can select the data retention period in the settings interface, such as "session-only," "24 hours," or "permanent retention." For data selected as "session-only," a secure erase command will be automatically triggered before the interaction ends and the device enters sleep mode. This command not only deletes the file index but also overwrites the original data blocks on the storage medium multiple times, ensuring that the data cannot be recovered. This mechanism effectively avoids the potential privacy risks of historical data leakage when the device is lost or lent out.
[0156] In addition to privacy protection, this embodiment further expands the application boundaries of the interactive system. When the interactive system is deployed on a local terminal device, the AI core engine obtains the hierarchical structure of the application interface by calling the standardized auxiliary function interface provided by the operating system, thereby achieving automated recognition and interactive operation of interface elements.
[0157] Specifically, modern operating systems such as Android and iOS provide accessibility service frameworks, originally designed to assist people with disabilities in using devices. This embodiment creatively utilizes this mechanism to give the AI core engine the ability to operate across applications. When a user issues a command such as "Help me open WeChat and send a message to [user's name]", the AI core engine first analyzes the user's intent and identifies the target application as "WeChat". Subsequently, the engine requests permissions from the system through Accessibility APIs to obtain a snapshot of the current screen's window content. The system returns a tree-structured data containing attribute information for all UI elements (such as buttons, text boxes, and icons), i.e., the interface hierarchy structure.
[0158] The AI core engine traverses and performs semantic analysis on the tree structure, locating the coordinates of the "WeChat" icon and simulating a click to launch the application. Once inside the application, the engine retrieves the new interface hierarchy, identifying elements such as the search box, contact list, and message input box. Through simulated gesture clicks and text input, the engine automatically completes a series of operations: finding contacts, entering message content, and sending. This process requires no modification to third-party applications or API integration, achieving true "cross-application automation." This design significantly expands the functional boundaries of smart interactive terminals, upgrading them from simple information query devices to intelligent assistants capable of operating the phone on behalf of the user, significantly improving user interaction efficiency and experience in complex task scenarios.
[0159] Through the implementation of the above-mentioned tiered privacy protection and cross-application automated operation, this embodiment maximizes the potential of the AI core engine while ensuring the security and controllability of user data, achieving a perfect balance between security and functionality.
[0160] Example 9: like Figure 2 As shown, this embodiment provides an interactive terminal that integrates multimodal perception and naked-eye 3D. This interactive terminal is the physical carrier for implementing the interactive methods in the aforementioned embodiments. The interactive terminal includes at least three types of heterogeneous sensors, a naked-eye 3D display module, an interactive system, and a controller.
[0161] Specifically, at least three types of heterogeneous sensors constitute the terminal's perception layer. The first type is an active space detection unit, continuously outputting signals reflecting the activity status of targets within space. The second type is a passive environmental change sensing unit, outputting signals reflecting transient environmental changes. The third type is a visual information acquisition and processing unit, outputting target recognition results based on visual features. In terms of hardware connectivity, the first type of sensor is preferably a microwave radar module, connected to the controller via a UART serial port or SPI interface, continuously transmitting Doppler frequency shift data. The second type of sensor is preferably a pyroelectric infrared sensor (PIR), whose output pin is directly connected to the controller's GPIO interrupt pin to achieve microsecond-level transient signal triggering. The third type of sensor is preferably a high-definition camera module, connected to the controller via a MIPI CSI-2 or DVP interface, transmitting high-bandwidth video stream data. This combination of heterogeneous sensors enables the terminal to simultaneously acquire spatial moving target information, infrared heat source change information, and visual image information, providing a physical basis for subsequent multimodal fusion perception.
[0162] The naked-eye 3D display module is used to present virtual avatars. This module integrates a high-resolution LCD screen, an electrically controlled grating or lenticular lens optical path assembly, and a viewpoint tracking unit. The screen connects to the controller's display output via a video signal cable (such as LVDS, MIPI, or HDMI interface) to receive rendered image data. The viewpoint tracking unit transmits the captured user's eye coordinates back to the controller in real time via USB or I2C interface for dynamic adjustment of grating parameters. The interaction system is used for multimodal interaction with the user. At the hardware level, the interaction system includes audio acquisition and playback components (such as microphone arrays and speakers), a network communication module (such as Wi-Fi and 5G modules), and a storage unit. The microphone array transmits audio data via I2S or PDM interfaces, and the speakers receive synthesized speech signals via I2S or analog audio interfaces. The network communication module is responsible for establishing a cloud-edge collaborative data channel, while the storage unit is used to cache local models, user preference data, and temporary interaction logs.
[0163] The controller is the core hub of the entire interactive terminal, connecting to at least three types of heterogeneous sensors, a naked-eye 3D display module, and an interactive system. The controller is configured to execute the methods described in any of the above embodiments. To support the complex real-time sensing, AI inference, and rendering tasks in the foregoing embodiments, the controller preferably employs a heterogeneous multi-core architecture. For example, the controller may include a low-power microcontroller unit (MCU) and a high-performance application processor (SoC). The MCU is responsible for maintaining the basic operation of the sensors, executing low-power standby monitoring, and hardware-level redundant wake-up logic; the SoC is responsible for running the operating system, AI core engine, and graphics rendering pipeline. The controller's internal integrated storage medium stores computer program instructions. When these instructions are executed by the processor, steps such as determining the primary trigger condition, auxiliary verification condition, third-level verification condition, generating wake-up instructions, and driving virtual avatar interaction are implemented. It should be understood that the controller can also be a single high-performance processor, or a distributed controller composed of a cloud server and local edge computing nodes, as long as it can implement the logical control functions of the above methods.
[0164] Through the implementation of the above hardware architecture, this embodiment deeply integrates multimodal perception, edge computing, and naked-eye 3D display technology into a single terminal device. The hardware connection of heterogeneous sensors ensures the real-time performance and accuracy of multi-source data acquisition; the close cooperation between the controller and each module ensures a low-latency link from perception to response; and the modular hardware design provides flexible physical support for system function expansion and scene adaptation, thus fully realizing the technical concept of this invention at the product level.
[0165] Example 10: To verify the effectiveness of the technical solution of the present invention in practical applications, this embodiment combines... Figure 7 and Figure 8 Taking shopping mall guide scenarios and family companionship scenarios as examples, the multimodal perception, concealed installation, multi-user interaction and scenario adaptation strategies in the above embodiments are explained in detail.
[0166] like Figure 7As shown, in a shopping mall scenario, the interactive terminal is concealed inside a glass counter. Traditional devices placed inside counters often suffer signal attenuation due to metal or glass obstructions, or fail because infrared sensors cannot penetrate them. This embodiment utilizes a microwave radar sensor as the first type of sensor. Its antenna structure is optimized through dielectric loading to match the dielectric constant of the glass, thus enabling detection through non-metallic obstructions. When a customer lingers in front of the counter, the microwave radar continuously outputs a signal reflecting the target's activity status in the space. The system determines whether this signal meets the stable state signal requirements of the time window, serving as the primary trigger condition. Simultaneously, a passive infrared sensor (PIR) serves as the second type of sensor, assisting in verifying the customer's intention to linger. Due to the dense crowds and complex noise in shopping mall environments, the system automatically switches to a high-dynamic crowd flow scenario strategy, shortening the detection time window and lowering the sensitivity threshold to suppress high-frequency false touches. When both the primary trigger condition and the auxiliary verification condition are met, the camera is activated for visual recognition. After confirming the target is a human body, a wake-up command is generated. At this time, the naked-eye 3D display module presents a virtual sales assistant image, proactively greeting the customer. Furthermore, by integrating with the mall's database, personalized product information can be pushed based on customers' browsing history. This process transforms the service from a "passive response" to a "proactive service," and the entire wake-up determination is completed locally at the edge, eliminating the need to upload data to the cloud. This effectively protects customer privacy and significantly reduces response latency.
[0167] like Figure 8 As shown, in a family-oriented scenario, the interactive terminal is placed on the living room TV cabinet. This scenario involves the complex situation of multiple members interacting concurrently. When a child stands in front of the device, a facial image is captured by a camera, combined with voiceprint features extracted by a microphone and spatial location data detected by radar, and the confidence score of each candidate user is calculated using a multi-user identity confidence fusion formula.
[0168] During interactions between children and parents, the system monitors the level of activity in real time. For example, if a child wants to play a game while a parent wants to check health alerts, the system uses a multi-user intent priority arbitration function to determine the priority. Because the urgency score for the "health alert" intent is higher than that for the "entertainment" intent, the system prioritizes responding to the parent's request. Simultaneously, the system dynamically adjusts the virtual avatar's behavior strategy based on a cross-generational interaction strategy adjustment function. ; in, The overall interaction strategy index determines the liveliness / seriousness of the virtual avatar.
[0169] This is the dynamic adjustment amplitude coefficient.
[0170] The switching cycle refers to the period during which interaction strategies switch between generational preferences.
[0171] The interaction weights for children are weighted by the duration of their speech, the duration of their gaze, and the activity level of their gestures.
[0172] The interaction weight is assigned to parents, and the activity level of parent interactions is weighted.
[0173] As a benchmark for children's preferences, the images that children prefer are usually more lively.
[0174] Based on parents' preferences, the images favored by parents typically have a lower liveliness level.
[0175] The comprehensive interaction strategy index calculated based on this formula controls the virtual avatar to dynamically fluctuate between "lively" and "calm," balancing children's enjoyment with parents' practical needs. Furthermore, when the system detects that a parent is in an anxious state, it automatically adjusts the avatar's expression to a soothing style, achieving emotional resonance.
[0176] Through the implementation of the two typical application scenarios mentioned above, this invention fully demonstrates its highly reliable wake-up under concealed installation, accurate identity recognition and policy adjustment under multi-user concurrency, and proactive service capabilities with low latency and high privacy protection. It effectively solves the pain points existing in the background technology and significantly improves the intelligence level of human-computer interaction and user experience.
[0177] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. An interactive method integrating multimodal perception and naked-eye 3D, characterized in that, The method is applied to an interactive terminal, which includes at least three types of heterogeneous sensors, a naked-eye 3D display module, and an interactive system. The first type of sensor is an active spatial detection unit, which is used to continuously output signals reflecting the activity status of targets in space; the second type of sensor is a passive environmental change sensing unit, which is used to output signals reflecting transient environmental changes; and the third type of sensor is a visual information acquisition and processing unit, which is used to output target recognition results based on visual features. The method includes: Determine whether the signal detected by the first type of sensor meets the main triggering condition, wherein the main triggering condition includes detecting a stable state signal that meets the time window requirement; When the main triggering condition is met, determine whether the transient change characteristics output by the second type of sensor meet the auxiliary verification condition; When the auxiliary verification conditions are met, it is determined whether the target recognition result based on visual features output by the third type of sensor meets the third verification condition. When the third verification condition is met, a wake-up command is generated, the interactive system is activated by the wake-up command, the virtual image is presented by the naked-eye 3D display module, and multimodal interaction is performed with the user.
2. The method according to claim 1, characterized in that, The interactive terminal also includes an independent triggering link not controlled by the main operating system. This independent triggering link is used to perform logical operations on the output signals of the first type of sensor and the second type of sensor and then directly connect them to the hardware wake-up pin of the main control processing unit. The method also includes: When an anomaly is detected in the main control software link or when it is in emergency mode, the independent trigger link not controlled by the main operating system is connected by a switch to wake it up. In normal mode, the main control processing unit executes multi-level progressive anti-accidental touch verification logic, and automatically reverts to normal mode after the fault is recovered.
3. The method according to claim 2, characterized in that, The first type of sensor is a microwave radar sensor, which is installed behind a non-metallic shield. The antenna structure is optimized by dielectric loading to match the dielectric constant of the non-metallic shield, or an electromagnetic wave focusing structure is provided inside the non-metallic shield. The second type of sensor is a passive infrared sensor; The method further includes: A unified synchronous trigger signal is sent to each sensor via a hard-wired trigger bus to control each sensor to start sampling within the same clock cycle and perform microsecond-level time alignment. The microwave radar sensor operates in the ISM band, the passive infrared sensor is made of pyroelectric crystal material, and the independent triggering link includes a monostable multivibrator and logic gate circuits.
4. The method according to claim 1, characterized in that, Before presenting the virtual image through the naked-eye 3D display module, it also includes: AI companion avatars can be created using at least one of the following methods: face reconstruction based on single or multiple frames of images; pose estimation and 3D model reconstruction based on video sequences; direct import of 3D model files; semantic parsing and generative modeling based on natural language text or speech descriptions. Specifically, for voice description methods, acoustic emotional features in the voice are extracted simultaneously, and the initial facial expression style or body posture of the generated image is automatically adjusted based on these emotional features. The video sequence-based pose estimation and 3D model reconstruction includes: using a pose estimation network to extract the coordinates of two-dimensional joints in the video frame, and reconstructing them into a parametric 3D human body model through a 3D pose enhancement network; The generative modeling based on natural language text description or speech description includes: parsing semantic tags using a large language model and calling a text-based 3D model algorithm to generate a 3D avatar with a neural radiation field representation or a Gaussian splash representation.
5. The method according to claim 1, characterized in that, The multimodal interaction with the user includes: By integrating and analyzing micro-expressions captured by the camera, emotional features of speech extracted by the microphone, and physiological rhythm signals detected by radar, a label for the user's current emotional state is generated. Based on the combination of the emotional state label and semantic intent, the optimal action sequence and voice style are matched from the preset behavior rule base to drive the virtual character to make an empathetic response; The physiological rhythm signals include heart rate, respiratory rhythm, or heart rate variability indicators obtained through microwave radar phase change analysis. The method further includes: dividing user regions using spatial location-aware data and maintaining multi-path concurrent session state management to enable continuous tracking and switching of independent user contexts.
6. The method according to claim 1, characterized in that, The interactive system includes an AI core engine, which supports a cloud-edge collaborative deployment architecture. In local resource-constrained mode, a lightweight inference engine optimized by model quantization or pruning is run, handling only basic interactive tasks; In cloud-enhanced mode, complex inference tasks and high-performance rendering tasks are offloaded to cloud servers, and rendering results are received and displayed via streaming media protocols. The AI core engine is also equipped with a model routing middleware, which is used to dynamically select the optimal large language model instance for inference based on task type, response latency requirements and service availability score. The AI core engine adopts an edge computing architecture. In the absence of events, the main processor enters a deep low-power sleep state, and only the low-power coprocessor maintains environmental monitoring. All biometric data is extracted and compared locally, and only irreversible feature encoding is uploaded to the cloud.
7. The method according to claim 1, characterized in that, Also includes: The working strategy is dynamically adjusted according to the device deployment scenario. The working strategy includes the configuration of perception sensitivity, wake-up threshold, interaction content and service interface. The deployment scenarios include high-dynamic pedestrian flow scenarios, low-dynamic private scenarios, and regular activity scenarios. Different scenarios correspond to different detection time windows and sensitivity threshold configurations. The method also includes: dynamically adjusting the system wake-up sensitivity through scene adaptation adjustment, user habit self-learning based on historical interaction data, and emotion linkage adjustment combined with multimodal sentiment analysis results; The method further includes: capturing the user's micro-expressions and gestures within a preset time after wake-up, using a sequence prediction model to predict the user's intentions and preload relevant information.
8. The method according to claim 1, characterized in that, The naked-eye 3D display module integrates a viewpoint tracking unit, and the method further includes: The system detects the user's eye position in real time and dynamically adjusts the grating parameters or liquid crystal molecule orientation of the electronically controlled grating based on the detection results, so that the optimal viewing area follows the user's movement. The naked-eye 3D display module uses instantiation rendering technology to share material and mesh resources for multiple virtual characters, and combines frustum culling and occlusion culling to optimize rendering performance; The method further includes: converting the neural radiation field representation into a Gaussian splash representation through knowledge distillation, dynamically selecting key viewpoints for high-resolution rendering based on eye-tracking data, and using low-resolution rendering for adjacent viewpoints combined with super-resolution reconstruction for restoration.
9. The method according to claim 6, characterized in that, Also includes: In highly privacy-sensitive scenarios, a purely local mode is adopted, in which the entire process of collecting, extracting, comparing and storing biometric data is completed in a local trusted execution environment; When cloud computing power is required, a cloud-edge collaborative mode is adopted to perform irreversible feature encoding processing on sensitive biometric data before uploading it to the cloud. It has a built-in automatic data cleanup mechanism, supports user-defined data retention periods, and automatically destroys sensitive intermediate data temporarily cached after the session ends; When the interactive system is deployed on a local terminal device, the AI core engine obtains the hierarchical structure of the application interface by calling the standardized auxiliary function interface provided by the operating system, thereby realizing the automatic recognition and interactive operation of interface elements.
10. An interactive terminal integrating multimodal perception and naked-eye 3D, characterized in that, include: At least three types of heterogeneous sensors, the first type of which is an active space detection unit, used to continuously output signals reflecting the activity status of targets in space; The second type of sensor is a passive environmental change sensing unit, which is used to output signals that reflect transient changes in the environment; the third type of sensor is a visual information acquisition and processing unit, which is used to output target recognition results based on visual features. A glasses-free 3D display module is used to present virtual images; Interactive systems are used for multimodal interaction with users; The controller is connected to the at least three types of heterogeneous sensors, the naked-eye 3D display module, and the interactive system, respectively, and the controller is configured to perform the method of any one of claims 1 to 9.